
Industry
Part of The operations case studies worth studying twice
Operations case study framework: turn one claim into a two-week trial
Operations case studies are worth acting on only after a small test: how to turn a claim into a two-week trial with a unit, a comparison, and a stop rule.
Reading a write-up well is a defensive skill. It stops you copying something that will not work. It does not tell you whether the thing works here, and no amount of careful reading will.
The only thing that answers that is a test you run yourself, small enough to fit inside two weeks and designed so that it can come out negative. This page is about designing that test.
What to take away
- A claim from someone else's operation is a hypothesis about yours, and the cheapest response to a hypothesis is a small test rather than a debate.
- The design decisions that matter are the unit, the comparison, the duration, and the stop rule. Get those written before anyone starts, because after the results arrive they cannot be settled honestly.
- Most tests fail because nothing was measured before. One week of baseline is worth more than four weeks of the new thing.
Four decisions, made in writing
The unit. What gets the new treatment: a person, a team, a shift, a case type, a day of the week. This is the first decision and it constrains everything else. Choose the smallest unit where the change is meaningful, and be honest about whether that unit is really independent of the others. Two people sitting together are not two units, because they will talk.
Four decisions before testing
- Unitsmallest meaningful, truly independent
- Comparisonbefore-after, side-by-side, alternating
- Durationcover full work cycle, set end date
- Stop rulenumber or condition to drop
The comparison. Against what. Three options and they are not equal. Before and after on the same unit is easiest and weakest, because everything else also changed. Side by side with a similar unit is stronger and requires that the units really are similar. Alternating periods, where the same unit switches back and forth, handles slow drift and is often the most practical in a small operation.
The duration. Long enough to cover the natural cycle of the work. If your volume varies by week of the month, a two-week test that lands in the quiet half tells you about the quiet half. Set the end date in advance, because a test that runs until the result looks good is not a test.
The stop rule. What result would make you drop the idea. Write it as a number or a described condition. This is the item most often skipped and it is the one that makes the whole thing worth doing, because without it every outcome becomes an argument for continuing.
Why the baseline is the part that gets skipped
Enthusiasm arrives before the measurement. Somebody reads a write-up on Thursday and wants to start on Monday, and taking a week to measure the current state feels like delay.
Build a usable baseline
- Choose one measure
- Count same way as after
- Cover one full work cycle
- If no system, tally by hand for a week
It is not delay, it is the difference between a test and an anecdote. Without a baseline you will compare the new thing against a remembered version of the old one. Memory of an operation's performance is systematically flattering to whatever came before a change people disliked and unflattering to whatever came before one they wanted.
A usable baseline is small. One measure, counted the same way you will count it afterward, over one full cycle of the work. If no measure exists, tally by hand for a week. The mechanics of counting something when no system records it are in operating processes metrics.
The things that will confound you, and what to do about each
Attention. A group that knows it is being studied behaves differently, and the effect is well enough documented to be named after the plant where it was first argued about: the Hawthorne effect. You cannot remove it in a small operation. You can reduce it by running the comparison group with the same visibility as the test group, so that both are being watched.
Confounders and countermeasures
Confounder
- Attention
- Hawthorne effect
- Selection
- Volunteers not representative
- Regression
- Extreme readings normalize
- Mix
- Work type changes
Countermeasure
- Attention
- Watch both groups equally
- Selection
- Avoid volunteer units
- Regression
- Don't trigger off bad month
- Mix
- Check composition start/end
Selection. If the unit volunteered, it is not representative. Volunteers are better at the thing, more motivated, or both. This is the single most common reason a pilot succeeds and the rollout does not.
Regression. If you started the test because a number was unusually bad, some of the improvement will happen without you. Extreme readings tend to be followed by less extreme ones for reasons that have nothing to do with any intervention, an effect known as regression toward the mean, and it is why triggering a test off a bad month is a trap.
Mix. If the type of work changed during the test, the measure moves and the change did nothing. Check the composition of the work at the start and end, and split the result by type if it moved.
What a small test can and cannot establish
It can tell you whether an effect is large. Big effects are visible in small tests and that is most of what you need operationally, because a change too small to see in two weeks with your volumes is a change too small to be worth the disruption.
What a small test can do
Can establish
- Effect size
- Large effects
- Scale
- One team works
- Decision
- Worth disrupting
Cannot establish
- Effect size
- Small effects
- Scale
- Forty people consistent
- Decision
- Modest improvement
It cannot establish a small effect, and it should not be asked to. Two weeks of an operation with ordinary variation cannot distinguish a modest improvement from a good fortnight. Saying so plainly at the end is more useful than a claim the data will not carry.
It also cannot tell you what will happen at scale. A practice that works with one team and its most experienced member fails when it needs forty people to do it consistently. The gap between a pilot and a rollout is mostly about consistency, and that is a different problem, treated in change management.
The general logic of designing a comparison so that it can answer a question, including why the assignment matters more than the sample size, is set out in the material on design of experiments in the NIST engineering statistics handbook.
Writing the test up so it is worth something later
Four lines, written at the end, on the same page as the four decisions written at the start.
Four lines to write up
- What we did, repeatable detail
- What we measured, with definition
- What happened, including misfits
- What we decided, with reason
The fourth line is the one people leave out, and it is why an organization runs the same test three times in five years. Recording the decision, including a decision to stop, makes the record cumulative rather than a set of unconnected experiments.
The wider discipline of reading and writing operational accounts honestly is in operations case studies, and a worked trace of how one internal result was checked before anybody believed it is in business development ops case study.
Common questions
We cannot run a comparison group. Is a before and after worth anything?
Yes, with caveats stated. Take a longer baseline, look for whether the change in the measure coincides with the intervention rather than drifting toward it, and name the other things that changed in the same period. It is weaker evidence and it is not nothing.
How small is too small?
If the thing you are measuring happens fewer than about ten times in the test period, you are reading individual cases rather than a pattern. That is fine, and the honest response is to read the cases properly rather than to compute a rate from them.
What if the result is neutral?
Then you have learned that the effect is not large here, which is a real result and saves you a rollout. Neutral results are the most underreported and the most useful, because they are the ones nobody else publishes.
Who should design the test?
Not the person who wants the change to work, or at least not alone. The stop rule in particular should be written by someone with no stake in the outcome, because it is the one decision that costs something to honor.







