Why Small Consistent Misses Cost More Than Big Rare Ones
Why this matters
Every shop can name the job that went wrong last quarter. Everybody was there, everybody talked about it, somebody probably lost sleep. Almost no shop can name the job type that runs 4% over on every single instance, because 4% of a three-hour call is about seven minutes and nobody has ever called a meeting about seven minutes.
The arithmetic does not care which one you remember. Exposure is miss size times frequency, and a small number times a large number routinely beats a large number times a small one. The blowout gets the post-mortem because it is visible. The chronic miss gets nothing, and it is usually the larger of the two.
Exposure is size times frequency, and attention is neither
Write both misses in the same unit - hours per quarter - and the comparison stops being a matter of opinion.
Exposure = instances per quarter x estimated hours per instance x variance as a fraction of estimate.
That is the only ranking that matters, and it is not the ranking anybody's memory produces. Memory ranks by salience: how surprising it was, how loud the customer got, how recent it is. Salience correlates with size and correlates negatively with frequency, because a thing that happens on every job stops registering as an event at all. So the shop's attention lands almost exactly where the exposure is not.
This is not a claim that the blowout does not matter. It is a claim that you cannot rank the two by feeling, and that the feeling is systematically wrong in one direction.
The small miss is invisible by construction, not by neglect
There is a reason nobody catches a 4% bias job by job, and it is not laziness.
Estimate variance on a typical residential service job type runs somewhere in the range of 15% to 25% absolute. That is the ordinary spread - access, customer talk, a fitting that fights you. A 4% bias sits completely inside that noise. On any single job you cannot distinguish a 4% systematic underbid from an ordinary bad afternoon, and neither can the tech.
A useful rule of thumb: you can start to see a bias in the data once it is roughly as large as the spread divided by the square root of the number of instances. At 18% spread, seeing a 4% bias needs the square root of the instance count to be around 4.5, so roughly 20 instances or more. That is a quarter of volume on a common job type, and it is why the small consistent miss is only ever visible in aggregate. Waiting for it to show up on a job is waiting for something that cannot happen.
The blowout is the opposite. It is visible on one instance because it is far outside the spread. It announces itself. That is exactly why it gets the attention it does, and exactly why the attention is not evidence that it is the bigger problem.
The arithmetic, worked
A shop runs two problems side by side for a year.
The chronic miss. A high-volume service call type: 120 instances a quarter, estimated at 3.0 hours each, running a signed variance of +4% - consistently a little over, never dramatically.
- Estimated hours per quarter: 120 x 3.0 = 360.0 hours
- Excess at 4%: 14.4 hours per quarter
- Annualized: 57.6 hours per year
Per instance that is 4% of 3.0 hours, or 0.12 hours - about seven minutes. No tech reports it. No office notices it. It is not a story anyone tells.
The rare blowout. A project type where 3 jobs a quarter run badly over: estimated at 8.0 hours each, coming in 40% over.
- Estimated hours per quarter across the three: 3 x 8.0 = 24.0 hours
- Excess at 40%: 9.6 hours per quarter
- Annualized: 38.4 hours per year
Per instance that is 40% of 8.0 hours, or 3.2 hours - a full afternoon lost, visible to everyone, and worth a conversation on its own merits.
The comparison. 14.4 hours per quarter against 9.6 hours per quarter. The chronic miss is larger by 4.8 hours per quarter, a ratio of 1.5x. Over a year, 57.6 hours against 38.4 hours - 19.2 hours more, from the problem nobody in the building can name.
And the attention split runs the other way entirely. Each of the three blowouts gets a post-mortem, so that problem gets three investigations a quarter. The 120-instance problem gets zero, because there is nothing to investigate on any individual instance.
The second cost: a bias corrupts the schedule, not just the margin
Margin is the obvious cost and it is not the whole cost. Every downstream system that consumes an estimate inherits its bias.
Take the same 120 jobs at 3.0 estimated hours. The shop schedules 360.0 hours of work and performs 374.4 hours - the 4% excess again, 14.4 hours it never planned for. That overflow does not evaporate. It lands as one of three things:
- Overtime, which costs more per hour than the hours you priced.
- A pushed appointment, which costs a customer relationship.
- A rushed last call of the day, which is where callbacks come from.
Spread across a 13-week quarter, 14.4 hours is a little over 1.1 hours a week of schedule overflow. On a small crew that reads as a late finish roughly once a week, forever, with no identifiable cause. The shop experiences it as "we are always a bit behind" and treats it as a dispatch problem, hires for it, or tightens the schedule further - which makes it worse, because the estimate the schedule is built on is still 4% light.
The blowout does not do this. It is one bad day that everyone can see and route around.
Where the intuition is right
Three cases where prioritizing the rare miss is correct, and they are worth naming precisely so this card does not get over-applied:
When the rare miss is catastrophic rather than merely large. A job that ends in an injury, a claim, a lost commercial account, or a licence problem is not on the same axis as hours. Exposure arithmetic does not apply to outcomes you cannot absorb at any frequency.
When the rare miss is getting more frequent. Three a quarter that were one a quarter last year is a trend, and a trend gets priced by where it is heading, not by where it is. Check the count over the last four quarters before dismissing a low-frequency problem.
When the rare miss is cheap to fix and the chronic one is not. A blowout traceable to a single missing intake question is a half-hour fix. Correcting a 4% chronic bias means moving a price on your highest-volume job type, which touches your win rate and needs to be done carefully. Cost of the fix belongs in the ranking alongside size of the problem, and sometimes it flips the order legitimately. That is a different reason than "the blowout felt bigger," which is the reason shops usually give.
How to find your consistent misses
They will not appear as outliers, so do not look for outliers.
Use the sign test first, before any averaging. For one job type, count how many instances came in over the estimate at all, regardless of size. If your estimating on that type were unbiased you would expect roughly half over and half under. Twenty-four of thirty over is not chance, and it tells you the type is biased before you compute a single percentage. This is the cheapest bias detector there is and most shops have never run it.
Read the median, not the mean. One blowout drags a mean and leaves a median alone. If the median instance of a type is over, the type is biased. If the mean is over and the median is not, you have a tail problem, which is a different diagnosis with a different cure.
Rank job types by exposure in hours, not by variance percentage. A type at 11% bias on four large instances can carry nearly the same exposure as a type at 19% bias on twelve smaller ones. The percentage ranking will hide it.
Look hardest at your highest-volume type, even when it looks healthy. Volume multiplies everything, including a bias too small to be interesting on its own. The type you run most is the one where a small number does the most damage, and it is usually the one nobody reviews because it never causes trouble.
How to verify you are reading this right
Take your two candidate problems, write both as exposure in hours per quarter, and check that the instance counts are from the same window and the variance percentages share the same base - both as a fraction of the estimate, not one as a fraction of estimate and the other as a fraction of actual. Mixing those two bases is the most common arithmetic error in this comparison and it moves the answer by enough to reverse it on a close call.
Then check the count you used for the rare miss against the whole trailing year rather than one quarter. Three blowouts in one quarter and none in the other three is one blowout a quarter annualized, not three, and treating a single bad quarter as the run rate overstates the rare side of the comparison by a factor of four.
Finally, once you have found a chronic bias and corrected it, verify the exposure actually closed rather than moved. A labor multiplier that fixes the bias while the spread stays wide has fixed the average and not the predictability, and the schedule overflow that motivated half of this card will still be there.
References
- See related:
The Estimate Accuracy Metrics Worth Trackingfor the bias, spread and exposure definitions used here. - See related:
The Quarterly Costing Review SOPfor the meeting where exposure ranking drives what gets fixed. - See related:
The Costing Mistakes That Compound Quietlyfor the mechanisms that let a chronic miss survive a review. - See related:
The Real Cost of an Inefficient Schedulefor the downstream scheduling effects touched on here.