
Maintenance
Part of Making sense of quality management, 2027 edition
Quality management metrics: definitions, targets and traps
Quality management measurement done honestly: escapes against defects, testing whether your checking works, severity over counts, and finding rework in the accounts.
A defect count is the easiest quality number to produce and the least informative one you can put on a page. It goes down when things improve and it also goes down when people stop writing things up, and nothing in the number distinguishes those two.
This page is about the measures that do carry information, and about the two questions to ask before adopting any of them.
What to take away
- Split every failure into caught and escaped. The ratio between them says more about your operation than the total ever will.
- Measure your checking as well as your output. Nobody knows how much their inspection misses because nobody tests it.
- Weight by consequence rather than counting. Twenty trivial errors and one serious one are not twenty-one of anything.
- Most of the cost of poor quality is rework, and rework is almost never recorded as rework.
Caught and escaped are two different measures
One number, split in two, changes what you can see.
Caught vs Escaped Failures
Caught
- What it tells you
- Process
- Who found it
- You
- Countable with no measurement
- No
- Rising share means
- Detection up
Escaped
- What it tells you
- Detection
- Who found it
- Customer
- Countable with no measurement
- Yes
- Rising share means
- Worse situation
Failures you found before the customer did tell you about your process. Failures the customer found tell you about your detection. The total tells you about neither, and it is what almost everyone reports.
Some consequences of holding them apart:
Caught and escaped
- A falling total with a rising escaped share is a worse situation than a flat total, and looks better on a chart.
- Improving detection makes your caught count go up, which looks like the process got worse. Anyone reporting a detection improvement should say so in the same sentence, or they will be answering for it for a year.
- Escapes are countable even in operations with no measurement at all, because the customer tells you. Start there if you are starting from nothing.
The awkward part: you never see the escapes nobody reported. That undercount is real and it is not fixable by trying harder. What you can do is calibrate it occasionally by asking a sample of customers directly, which gives you a sense of the multiple between complaints and actual problems. The bias in complaint data is treated in quality management.
Test your checking, because it is a process too
Inspection has a defect rate of its own and nobody measures it. Two cheap methods exist.
Testing Your Inspection Process
- Seed a documented defect into the flow
- See whether the check catches it
- Tell people afterwards, not before
- Recheck items that already passed
- Count anything the first check missed
- Estimate a rough detection rate
Seed a known defect. Put a deliberate, documented fault into the flow and see whether the check catches it. Do it rarely and tell people afterwards, since the point is to test the system and not to trap anyone. If your check misses the seeded defect, the useful question is whether it was looking for that kind of fault at all.
Recheck what passed. Take a set of items that were checked and passed, and check them again properly. Anything found is something the first check missed. This is uncomfortable and it is the single most informative quality measurement most organizations have never done.
Both give you a rough detection rate, which turns a vague sense that inspection is imperfect into a number you can reason with. If your check finds a fraction of what is there, then your caught count understates the true defect rate by roughly the inverse of that fraction, and the arithmetic changes what improvement is worth.
Count consequences, not events
A defect rate treats a typo and a wrong shipment as the same event. Nobody experiences them that way.
Severity Bands Done Honestly
- Use three or four bands only
- Define bands by customer consequence
- Write definitions before categorizing
- Report each band separately
- Avoid weighted scores that hide serious failures
- Name who can reclassify the number
Use a small number of severity bands, defined in advance by what happens to the customer, not what happened internally. Three or four bands is plenty.
Write definitions down before anybody categorizes, because assigning severity after the fact drifts toward whatever makes the month look reasonable. Severity bands are a strategic planning example of a definition written before the fact.
Then report the bands separately. Combining them into a weighted score is tempting and loses the thing you wanted: a serious failure hidden inside a favorable weighted average is exactly the case the severity split exists to prevent. Two lines on a page beat one clever index.
The related caution about a measure changing behavior once it becomes the target is the substance of Goodhart's law, and severity classification is unusually exposed to it, because the classifier and the person being measured are frequently the same.
The cost is mostly rework, and rework is invisible
Ask an organization what poor quality costs and you get either nothing or a number somebody constructed for a slide.
Where Rework Hides
Where it hides
- Support and service time
- Contacts caused by errors
- Expedited work
- Rushed to recover internal slip
- Extra checking
- Steps added after incidents
- Waiting
- Work paused for queries
- Write-offs and credits
- Concessions to keep customers
What to look for
- Support and service time
- Expedited work
- Extra checking
- Waiting
- Write-offs and credits
You do not need a costing model. You need to notice that rework is being done and is being recorded as normal work. Some places to look:
Where rework hides
| Where it hides | What to look for |
|---|---|
| Support and service time | Contacts that exist only because something was wrong |
| Expedited work | Anything rushed to recover a slip caused internally |
| Extra checking | Steps added after an incident and never removed |
| Waiting | Work paused while a query is resolved |
| Write-offs and credits | Concessions given to keep a customer |
Sizing this once, roughly, over a fortnight, is more persuasive than any published estimate, because the numbers are yours and nobody can dispute the source. Keep it as an order of magnitude and say so. The method for counting when nothing records the thing you want is set out in operating processes metrics.
Two questions before adopting any quality measure
Would a bad result change what we do? If the honest answer is that it would produce an explanation rather than an action, you are building a reporting obligation and not a measure.
Before Adopting a Quality Measure
Would a bad result change what we do?
Adopt the measure
You are building a reporting obligation
Who can make this number look better without improving anything? Every measure has such a route. Name it out loud when you adopt the measure, so that later movement in that direction is discussed rather than celebrated. Reclassification, definition changes and stopping recording are the usual three.
Neither question needs data to answer, and both are best asked before collection starts. The wider version of this discipline, applied to plans rather than to operations, is in strategic planning metrics.
Where the measure has to come from outside
Some quality measurements come from a customer contract or regulator. Then the definition is not yours, and the source's current text governs. Sampling schemes are included: the arithmetic for accepting or rejecting a batch is well documented. The NIST engineering handbook covers lot acceptance sampling and shows the reasoning.
Where a scheme is specified for you, follow it rather than a general principle. Take advice on anything with contractual weight.
For internal use, the ordinary caution applies: a sample tells you about the population it came from, and a convenient sample tells you about convenience. Where the work sits alongside recurring checks rather than continuous measurement, quality management checklist covers the review shape.
Common questions
How many quality measures should we have?
Few enough that someone can hold them in mind. A caught count, an escaped count, and a severity split will occupy most organizations for a year. Add a measure when a specific decision needs it, and retire one at the same time.
Our defect numbers went to zero. Is that good?
Check whether reporting became harder or riskier in the same period, and check the escaped count. Genuine zeroes exist and are rare, and a zero arriving without a matching change in the process is worth investigating with the same energy you would give a spike.
Should individuals' defect rates be tracked?
Rarely, and only where the work is genuinely comparable between people. The moment a defect count is attached to a person, the count starts describing what people are willing to report. Team and process level measurement keeps the information flowing, and the difference between the two matters more here than almost anywhere else in management.
What if our volumes are too small for any of this?
Then read the cases individually rather than summarizing them. At low volume, knowing every failure in detail is both possible and better than a rate that swings wildly on one event. Percentages calculated from a handful of cases mislead more than they inform.







