
Guides
Part of Operating processes: a clear guide with practical examples
Operating processes metrics explained with examples
Operating processes measured honestly: five decisions inside any figure, why the average describes nobody, queue length as a leading count, and low-volume work.
Most arguments about operational numbers are not about the numbers. They are about what was counted, and they happen after the figure is on a slide, when nobody wants to reopen the definition.
This page is about the mechanics: where the clock starts, why the average is the wrong summary for almost everything you care about, and how to count when you have no system that records the thing you need.
What to take away
- Write the definition before you produce the number. Five decisions sit inside any operational measure, and if you do not make them deliberately, the tool made them for you.
- The average is the number least likely to describe anyone's experience. Look at the worst tenth of cases, because that is where the complaints and the cost live.
- Queue length is free to collect, updates immediately, and tells you more about where you stand than throughput does.
Five decisions inside every measure
Before anyone can dispute your figure, they will dispute one of these. Settle them in writing first.
What counts as one item. A request, a batch, a customer, a document. Different choices produce different answers from identical work, and the choice should follow the unit the person waiting experiences. There is more on picking it in operating processes framework.
When the clock starts. The moment the work formally entered your system is convenient and usually wrong. The honest start is when the person began waiting, which is often earlier and in somebody else's inbox.
When it stops. At handover, at acceptance, or when nothing comes back. These can differ by days. Stopping at handover flatters you and hides rework.
What is excluded. Canceled items, duplicates, test records, the enormous one-off. Every exclusion is defensible and every exclusion moves the number, so list them next to the figure rather than in a note nobody opens.
Which cases quietly disappear. This is the one people miss. Items that were never logged, that were handled informally as a favor, or that the requester gave up on, are absent from your data and are disproportionately the bad ones. If a route exists for work to be done without being recorded, your measure describes the recorded subset.
The average is the wrong summary
Operational times are not symmetric. Most cases go through quickly and a tail of them take much longer, so the mean sits above what most people experience and below what the unhappy ones do. It describes nobody.
Three summaries are more useful and none of them need a statistics package.
The median. Half of cases are worse than this. It is the honest answer to "how long does this normally take."
The worst tenth. Sort your cases by duration and look at the slowest tenth. This is where complaints, escalations and extra cost come from, and it is the number that moves when something is genuinely broken.
The count over a threshold. How many took longer than the promise you made. This is the only version a customer would recognize, and it is the one worth reporting outward.
A worked habit: whenever someone shows you an average, ask for the spread and for the slowest few. If the answer is that the data is not held that way, the average was computed from individual records that still exist, and the request is reasonable.
Queue length is underrated
Throughput tells you what happened last week. Queue length tells you what next week will feel like, and you can collect it by counting, today, with no system at all.
Three things it gives you cheaply. It is a leading indicator: a queue that grew this week produces delays next week, before any completion measure moves. It localizes the constraint, because the step with work piled in front of it is the one that is limiting, whatever the complaints say. And it converts arguments about capacity into arithmetic, since a queue holding two weeks of work at the current finishing rate will take two weeks to clear even if nothing new arrives. The relationship between how much you are carrying, how fast work leaves, and how long each item takes is fixed, and it is stated as Little's law.
Count it at the same time on the same day each week. Consistency matters more than precision, and a rough count taken every Monday beats an exact one taken whenever somebody remembers.
Counting when nothing records it
Frequently the measure you want is not in any system, and building the collection would take longer than the decision can wait. The alternative is not to guess.
Tally for a fixed period. Two weeks of somebody marking a sheet each time a defined event happens answers most operational questions. Define the event tightly enough that two people would mark the same cases, and agree the definition before the fortnight starts, not during it.
Sample rather than census. Take a set of recent items chosen without picking them, for instance every case that arrived on three particular days, and follow each one fully. A small number of cases examined properly beats a large number summarized badly.
Read the timestamps you already have. Emails, tickets, file dates and system logs contain start and end points that nobody set up for measurement. They are messy and usually sufficient to size a problem.
Be straight about what a fortnight of tallying can support. It can tell you whether something is a large problem or a small one, and where the time goes. It cannot support a claim about a small change over a long period, and it will be quoted for years unless you write the collection period next to the number.
When the volume is too low to compute anything
Below a certain volume, percentages are theater. A rate calculated from a handful of cases moves wildly with one event, and a month in which two things went wrong looks like a crisis next to a month in which none did.
The right response is to stop summarizing and read the cases. Every one of them, with the story attached. At low volume you can afford to know each case individually, which is far more informative than any statistic derived from them, and it does not pretend to precision you cannot have.
If someone insists on a trend, show counts rather than rates, over a long enough period that the pattern is visible, and say plainly that a single month means nothing. The general caution about reading too much into one result is the same one that applies to reading anybody else's write-up in operations case studies.
Comparing two periods without fooling yourself
Before concluding that something improved, check three things.
Did the mix change? If easier cases became a larger share of the work, every average improves without anything getting better. This is the most common false improvement in operations reporting and it is invisible unless you split by case type. A total that moves in one direction while every subgroup inside it moves the other way is common enough to have a name, Simpson's paradox.
Did the definition change? A new system, a reconfigured field, a different default for when an item is marked complete. Definition changes almost always coincide with the improvement projects that motivated them, which makes them hard to separate.
Did the denominator change? Rates fall when the bottom grows. A dropping defect rate with a rising defect count is a real and different situation from a dropping count, and the two get reported identically.
Where you cannot rule these out, say so beside the figure. An improvement reported with a caveat survives scrutiny. One reported cleanly and then dismantled in a meeting costs you the next three numbers as well.
What to publish and what to keep
Publish few numbers, with their definitions attached, and keep the underlying cases where you can get at them. Most of the value in operational measurement comes from being able to go back to the individual items when a figure looks odd, and most reporting throws those away in favor of a monthly summary.
Two habits worth holding. Never publish a measure without the date range and the definition in the same place, because both get separated from the number within a week. And keep a short note of what changed in the measurement itself, since the most confusing charts in any organization are the ones where a definition moved and nobody wrote it down.
Where these numbers feed a plan rather than a daily operation, they need different properties, and that is treated in strategic planning metrics. Where they concern whether output is acceptable rather than fast, the quality management view applies. The three measures worth watching in combination, and why they resist gaming better as a set, are in operating processes.
Common questions
Should measures be visible to the team doing the work?
Yes, if they are used to run the work and not to rank people. Teams fix things they can see. The moment the same number is used to compare individuals, it stops describing the work and starts describing what people believe is being watched.
How often should we look?
At the cadence at which you could actually act. Daily numbers on a process you review monthly produce noise and anxiety. If the answer is that you would not do anything differently on Tuesday, look weekly.
What about a measure that everyone disputes?
The dispute is usually about one of the five definitional decisions. Ask which one, specifically, and settle it in a sentence. A number nobody accepts is worse than no number, because the meeting becomes about the measurement rather than about the work.



