Quality management metrics: definitions, targets and traps. Quality management metrics: definitions, targets and traps
Image: Operations Process Control

Maintenance

Part of Making sense of quality management, 2027 edition

Quality management metrics: definitions, targets and traps

Quality management measurement done honestly: escapes against defects, testing whether your checking works, severity over counts, and finding rework in the accounts.

A defect count is the easiest quality number to produce and the least informative one you can put on a page. It goes down when things improve and it also goes down when people stop writing things up, and nothing in the number distinguishes those two.

This page is about the measures that do carry information, and about the two questions to ask before adopting any of them.

What to take away

  • Split every failure into caught and escaped. The ratio between them says more about your operation than the total ever will.
  • Measure your checking as well as your output. Nobody knows how much their inspection misses because nobody tests it.
  • Weight by consequence rather than counting. Twenty trivial errors and one serious one are not twenty-one of anything.
  • Most of the cost of poor quality is rework, and rework is almost never recorded as rework.

Caught and escaped are two different measures

One number, split in two, changes what you can see.

Caught vs Escaped Failures

Caught

What it tells you
Process
Who found it
You
Countable with no measurement
No
Rising share means
Detection up

Escaped

What it tells you
Detection
Who found it
Customer
Countable with no measurement
Yes
Rising share means
Worse situation

Failures you found before the customer did tell you about your process. Failures the customer found tell you about your detection. The total tells you about neither, and it is what almost everyone reports.

Some consequences of holding them apart:

Caught and escaped

  • A falling total with a rising escaped share is a worse situation than a flat total, and looks better on a chart.
  • Improving detection makes your caught count go up, which looks like the process got worse. Anyone reporting a detection improvement should say so in the same sentence, or they will be answering for it for a year.
  • Escapes are countable even in operations with no measurement at all, because the customer tells you. Start there if you are starting from nothing.

The awkward part: you never see the escapes nobody reported. That undercount is real and it is not fixable by trying harder. What you can do is calibrate it occasionally by asking a sample of customers directly, which gives you a sense of the multiple between complaints and actual problems. The bias in complaint data is treated in quality management.

Test your checking, because it is a process too

Inspection has a defect rate of its own and nobody measures it. Two cheap methods exist.

Testing Your Inspection Process

  1. Seed a documented defect into the flow
  2. See whether the check catches it
  3. Tell people afterwards, not before
  4. Recheck items that already passed
  5. Count anything the first check missed
  6. Estimate a rough detection rate

Seed a known defect. Put a deliberate, documented fault into the flow and see whether the check catches it. Do it rarely and tell people afterwards, since the point is to test the system and not to trap anyone. If your check misses the seeded defect, the useful question is whether it was looking for that kind of fault at all.

Recheck what passed. Take a set of items that were checked and passed, and check them again properly. Anything found is something the first check missed. This is uncomfortable and it is the single most informative quality measurement most organizations have never done.

Both give you a rough detection rate, which turns a vague sense that inspection is imperfect into a number you can reason with. If your check finds a fraction of what is there, then your caught count understates the true defect rate by roughly the inverse of that fraction, and the arithmetic changes what improvement is worth.

Count consequences, not events

A defect rate treats a typo and a wrong shipment as the same event. Nobody experiences them that way.

Severity Bands Done Honestly

  • Use three or four bands only
  • Define bands by customer consequence
  • Write definitions before categorizing
  • Report each band separately
  • Avoid weighted scores that hide serious failures
  • Name who can reclassify the number

Use a small number of severity bands, defined in advance by what happens to the customer, not what happened internally. Three or four bands is plenty.

Write definitions down before anybody categorizes, because assigning severity after the fact drifts toward whatever makes the month look reasonable. Severity bands are a strategic planning example of a definition written before the fact.

Then report the bands separately. Combining them into a weighted score is tempting and loses the thing you wanted: a serious failure hidden inside a favorable weighted average is exactly the case the severity split exists to prevent. Two lines on a page beat one clever index.

The related caution about a measure changing behavior once it becomes the target is the substance of Goodhart's law, and severity classification is unusually exposed to it, because the classifier and the person being measured are frequently the same.

The cost is mostly rework, and rework is invisible

Ask an organization what poor quality costs and you get either nothing or a number somebody constructed for a slide.

Where Rework Hides

Where it hides

Support and service time
Contacts caused by errors
Expedited work
Rushed to recover internal slip
Extra checking
Steps added after incidents
Waiting
Work paused for queries
Write-offs and credits
Concessions to keep customers

What to look for

Support and service time
Expedited work
Extra checking
Waiting
Write-offs and credits

You do not need a costing model. You need to notice that rework is being done and is being recorded as normal work. Some places to look:

Where rework hides

Where it hidesWhat to look for
Support and service timeContacts that exist only because something was wrong
Expedited workAnything rushed to recover a slip caused internally
Extra checkingSteps added after an incident and never removed
WaitingWork paused while a query is resolved
Write-offs and creditsConcessions given to keep a customer

Sizing this once, roughly, over a fortnight, is more persuasive than any published estimate, because the numbers are yours and nobody can dispute the source. Keep it as an order of magnitude and say so. The method for counting when nothing records the thing you want is set out in operating processes metrics.

Two questions before adopting any quality measure

Would a bad result change what we do? If the honest answer is that it would produce an explanation rather than an action, you are building a reporting obligation and not a measure.

Before Adopting a Quality Measure

Would a bad result change what we do?

Yes

Adopt the measure

No

You are building a reporting obligation

Who can make this number look better without improving anything? Every measure has such a route. Name it out loud when you adopt the measure, so that later movement in that direction is discussed rather than celebrated. Reclassification, definition changes and stopping recording are the usual three.

Neither question needs data to answer, and both are best asked before collection starts. The wider version of this discipline, applied to plans rather than to operations, is in strategic planning metrics.

Where the measure has to come from outside

Some quality measurements come from a customer contract or regulator. Then the definition is not yours, and the source's current text governs. Sampling schemes are included: the arithmetic for accepting or rejecting a batch is well documented. The NIST engineering handbook covers lot acceptance sampling and shows the reasoning.

Where a scheme is specified for you, follow it rather than a general principle. Take advice on anything with contractual weight.

For internal use, the ordinary caution applies: a sample tells you about the population it came from, and a convenient sample tells you about convenience. Where the work sits alongside recurring checks rather than continuous measurement, quality management checklist covers the review shape.

Common questions

How many quality measures should we have?

Few enough that someone can hold them in mind. A caught count, an escaped count, and a severity split will occupy most organizations for a year. Add a measure when a specific decision needs it, and retire one at the same time.

Our defect numbers went to zero. Is that good?

Check whether reporting became harder or riskier in the same period, and check the escaped count. Genuine zeroes exist and are rare, and a zero arriving without a matching change in the process is worth investigating with the same energy you would give a spike.

Should individuals' defect rates be tracked?

Rarely, and only where the work is genuinely comparable between people. The moment a defect count is attached to a person, the count starts describing what people are willing to report. Team and process level measurement keeps the information flowing, and the difference between the two matters more here than almost anywhere else in management.

What if our volumes are too small for any of this?

Then read the cases individually rather than summarizing them. At low volume, knowing every failure in detail is both possible and better than a rate that swings wildly on one event. Percentages calculated from a handful of cases mislead more than they inform.

More in Maintenance

Latest from Review Desk