Why Estimates Miss and Which Misses Matter
Why this matters
Every estimate misses. The question worth answering is not whether you were wrong on a job, it is whether the same wrongness will show up again next month with a different crew on a different street. A shop that treats every miss as a failure burns its energy on noise and its crew's trust on interrogations. A shop that treats every miss as bad luck repeats the same underbid four hundred times a year.
The sorting rule is simple to state and hard to apply: a miss matters when it has a direction that survives repetition, and it matters in proportion to how often you run that job, not how badly any single one went. The rest of this card is how to tell those apart.
The two different things a miss can be
Borrow the distinction from measurement. A scale can be off in two independent ways, and the fix is different for each.
Bias is a consistent lean in one direction. Your estimates for a job type land under actual most of the time, by a similar amount. Bias is an assumption error: a number in your head or your template is wrong, and it is wrong the same way every time. Bias is correctable by changing the number, and correcting it is the highest-return work in estimating because it pays out on every future job of that type.
Scatter is variability without a lean. Some run over, some run under, the middle sits near zero, the range is wide. Scatter is an information error: the job varies more than your estimate can see from where you are standing when you price it. You do not fix scatter by moving the template number, because the middle is already right. You fix it by gathering more information before you commit, by pricing a range, or by adding a conditions clause.
Confusing the two produces the two classic wrong moves. Correcting scatter as if it were bias pads a template that was already centered, and you lose bids on the jobs that would have run clean. Ignoring bias because "jobs vary" leaves a permanent leak running under the floor.
Where misses actually come from
Each source leaves a different signature in your data. Learn the signature and you can usually name the cause before anyone tells you the story.
| Source | What it is | Signature in the numbers |
|---|---|---|
| Stale time assumption | The template hours were built for a method, a tool, or a crew you no longer use | Consistent one-direction labor bias on one job type, stable over months |
| Takeoff omission | Small parts, fittings, consumables, fasteners that never make the list | Material multiple consistently above 1.0, labor near estimate |
| Unread site conditions | Access, occupancy, working height, what was behind the wall | Wide labor scatter inside one job type, correlated with building age or type |
| Undocumented scope drift | Work added on site that never became a change order | Both labor and material over, no change order on the record |
| Best-case optimism | The estimate describes the job when everything goes right | Estimate matches your fastest jobs, misses the middle; distribution leans over |
| Supplier price drift | Cost figures in the price book aged out | Material multiple creeping up slowly across all job types at once |
| Non-productive time | Staging, cleanup, paperwork, moving between stops, waiting | Small consistent labor over on every job type, not one |
| Crew and learning curve | Who ran it, not what it was | Variance sorts by name, not by job type |
Two of these are commonly misread. Non-productive time shows up as a small overrun on everything and gets dismissed as background noise, when it is actually the most systematic bias a shop can carry: an hour of it per tech per day is a full day per tech every two weeks that nobody bid. And crew gets blamed for what is really an unread-conditions problem, because the newer tech happens to draw the harder houses.
The sign test
The first question about any pattern is not how big it is. It is how many jobs went the same direction.
If misses were genuinely random, roughly half your jobs of a type would land over and half under. When 8 of 10 land over, the estimate is not noisy, it is off-center, and the size of the median miss tells you by how much. Count first, size second. A shop that looks only at averages can be fooled by one catastrophic job dragging a mean that is otherwise fine.
Use the median, not the mean, for the size of a typical miss. One job where the customer's dog got out and the crew spent two hours on it will move a mean and will not move a median, and you want the number that describes the typical job, because that is the job you are about to bid again.
The frequency filter: volume times bias
This is the sorting rule that decides where your attention goes, and it is counterintuitive enough that most shops get it backwards. The worst miss is rarely the most expensive one.
A job type with a modest bias that you run constantly costs far more than a rare job type with a dramatic miss. The math is just multiplication, but nobody runs it, because the dramatic miss is memorable and the modest one is invisible.
Rank your job types by estimated hours per job, times the median percentage bias, times the count in the period. The top of that list is your work queue for template corrections, regardless of which job you remember being angry about.
The misses that do not matter
Three categories, and letting them go is what buys you the attention to fix the ones that do:
- The genuine one-off. A frozen valve, a customer who changed their mind twice, a supplier who shipped wrong. It happened, it is not repeatable, it does not go in a template. Log the cause line and move on.
- Small misses inside your noise band. Every shop has a range where variance is measurement error and normal job variation rather than signal. Investigating inside that band teaches the crew that the log is a trap and produces nothing.
- The tails, in both directions. A single job at 60% over and a single job at 40% under, on a job type whose middle sits at 2%, are the shape of a scattered distribution, not two separate problems. Chasing the tails on a centered distribution is how a review meeting fills an hour without changing anything.
A worked read
A quarter's closed jobs, two job types:
Type A: 36 jobs, average estimate 6.0 labor hours each. Median labor variance 9% over. Of the 36, 29 landed over estimate, which is 81%, well past the roughly half you would expect if the misses were random.
Type B: 8 jobs, average estimate 20.0 labor hours each. Variance runs from 35% under to 50% over. Median 2% over. 4 of the 8 landed over, an even split.
The owner's instinct is to work on Type B, because Type B contains the job that ran 50% over and everybody remembers it. Run the multiplication instead.
- Type A: 9% of 6.0 hours is 0.54 hours given away on a typical job. Across 36 jobs, that is 19.4 hours in the quarter.
- Type B: 2% of 20.0 hours is 0.4 hours on a typical job. Across 8 jobs, that is 3.2 hours in the quarter.
Type A is costing about six times the hours Type B is, and Type A is the one with no memorable disaster in it. That is the whole lesson of the frequency filter in one comparison.
The diagnoses also differ. Type A shows a clear sign (81% over) with a tight, modest miss: that is bias, and the fix is a number. Raising the template from 6.0 to 6.5 hours is an 8.3% correction, which sits just under the 9% median miss deliberately, because you would rather approach the right answer from below and re-measure than overshoot and start losing bids.
Type B shows an even split and a wide range: that is scatter, and moving the template number would be wrong, because the middle is already close. The fix for Type B is upstream of the estimate. Look at what separates the 35%-under jobs from the 50%-over ones. If it is building age, access, or what was found behind the wall, the answer is a better pre-quote site question or a conditions clause, not a bigger number.
Verification, one quarter later: the Type A median should move toward zero and the over-count should move toward half. If Type A's median lands at 1% over with 19 of 34 jobs over, the correction took and you leave it alone. If it lands at 7% over with 27 of 34 over, the bias was bigger than the correction and you take another step, again from below.
What changes the answer
- Small samples. Under about eight jobs of a type, the sign test is weak. Six of eight over is not a pattern you should reprice on. Hold the job type on a watch list, keep logging, and decide when you have enough of them. Repricing on four jobs is how a shop ends up with a template that fits four unusual houses.
- Seasonality. A job type whose median variance moves with the season is not biased, it is two different jobs sharing one template. Heavy-season work with rushed staging and unfamiliar helpers genuinely costs more hours than the same work in a slow month. Split the template or accept the seasonal spread as a known band.
- A method change mid-period. New tooling, a new install standard, or a new supplier resets the baseline. Variance measured across the change date is two populations mixed together, and the median means nothing. Date your changes so you can cut the data at them.
- A price book that quotes flat rate. When the customer price comes from a published book, labor variance measures whether the book fits your crew and your market. The correction lives in your book adjustment factor, not in coaching an estimator who did not make the number.
How to verify you got this right
Before acting on a pattern, three checks:
- Count the sign, out loud, with the base. "29 of 36 Type A jobs landed over estimate, 81%." If you cannot say it in that form, you are looking at an average and calling it a pattern.
- Re-run the pattern with the biggest single miss removed. If the pattern disappears, you had one bad job and a story, not a bias.
- Check the confound. Sort the same jobs by who ran them and by customer. If the pattern is really a crew pattern or a single-customer pattern wearing a job-type costume, correcting the template will overprice everyone else's version of the same work.
The failure mode that survives all three checks is the one to watch for: a correction that gets made and never verified. A template raised 8% in March that nobody re-measures in June is indistinguishable, from the outside, from a template that was never touched. The correction is not finished when the number changes. It is finished when the next batch of jobs shows the miss moved.
References
- U.S. Small Business Administration: cost estimating and job costing guidance for small contractors.
- Standard field-service and construction practice on estimate-to-actual variance analysis.
- See related: The Variance Threshold Worth Investigating.
- See related: How to Find the Job Types You Consistently Underbid.
- See related: The Estimating Feedback Loop Explained.