How to Improve an Estimate After the Third Miss
Why this matters
The third miss is where most shops act, and it is the worst-understood moment in the whole feedback loop. Two misses feel like bad luck. Three feels like a pattern, so somebody raises the number - usually by the full amount of the worst overrun, usually across the whole job type, usually without writing down why. Six months later the shop is losing bids it used to win, nobody remembers the change was made, and the original cause is still there waiting for the jobs where it actually applies.
Three data points genuinely tell you something. They tell you far less than the size of the change most people make. This article is about what a sample of three licenses, what it does not, and how to act on it without overcorrecting into a different problem.
Step 1: Verify the three misses are the same miss
Before anything else, confirm you have three instances of one thing rather than three unrelated events that happen to share a job type.
Check all four:
- Same direction. Two over and one under is not a pattern, it is spread. Spread and bias are different diseases with different fixes.
- Same bucket. Three labor overruns is a pattern. A labor overrun, a material overrun, and a subcontract overrun are three separate investigations wearing one job-type label.
- Same phase, if you have phase data. Three jobs over by similar amounts in three different phases is a scoping problem, not a template problem.
- Not one cause with three names. Three "jobs ran long" entries that all trace to the same crew, the same month, or the same customer site are one observation, not three.
Fail any of these and you do not have a third miss. You have one miss and some noise, and the correct action is to keep watching.
Step 2: Understand exactly what three same-direction misses license
Here is the honest arithmetic. If your estimate for a job type were genuinely unbiased - as likely to run over as under, with each job independent of the last - then three consecutive misses in one named direction happens by chance about 1 time in 8. Three in the same direction, either direction, happens about 1 time in 4.
That is suggestive. It is not proof, and the independence assumption is the weak part: three jobs run by the same crew in the same six weeks are not three independent draws, so the real chance of a coincidental run is higher than 1 in 8. Four consecutive in one direction drops to about 1 in 16, and five to about 1 in 32, which is where most people stop arguing.
What follows from that is the operating rule for this whole article: at three, you can believe the sign. You cannot believe the size. The direction of your error is probably real. The magnitude you observed is one draw from a distribution and is as likely to be at the wide end as the middle. Act on the sign confidently. Act on the size cautiously.
Step 3: Look for the assumption before you touch the number
A number is the symptom. Somewhere in the estimate is a sentence, spoken or unspoken, that turned out not to be true. Find it and you can fix the cause instead of padding the effect.
Pull the three job records side by side and hunt for what they share that the estimate did not anticipate. Look at the phase where the hours landed, the notes, the part list, and the visit count. Ask the tech who ran two of them a single open question: "what did you have to do on this one that the quote did not account for?" That question gets a specific answer far more reliably than "why did it run long," which gets you a defense.
You are looking for a named condition. "The existing mount had to be relocated." "The access required a second person for the lift." "The customer's shutoff was seized and needed replacing before we could start." A named condition is worth ten multipliers, because it comes with a trigger: you can ask about it at the sale.
If you find one, the fix is conditional, not blanket. Add a scoping question to the estimating checklist and a conditional line to the template that fires when the answer is yes. Now jobs with the condition get priced for it and jobs without it stay competitive.
Step 4: If no assumption surfaces, correct part of the observed gap
Sometimes the three jobs share nothing you can name. The template is just light. Now you correct the number, and the size of the correction is the decision.
Use the median gap, not the mean and not the worst. Then apply about half of it.
Half feels timid and it is the right call for a specific reason. Extreme observations tend to be followed by less extreme ones, so the median of three is more likely to overstate the true bias than to understate it. A half correction moves you in the correct direction on the first cycle without buying an overprice you will not detect for two quarters. If the bias is real and larger than half, the next window shows the same sign and you take the rest. If it was partly noise, you never overshot.
The exception, stated with it: when all three misses are large and tightly clustered - say all three between 40% and 50% over their own quoted hours - the spread is small enough that the median is a better estimate of the truth, and you can take two thirds to three quarters of the gap rather than half.
Step 5: Write the trigger, not just the multiplier
Whatever you change, the change has to carry three things or it will not survive contact with the next estimator: what changed, when it changed, and what evidence caused it. One line on the template is enough. "Base labor raised from 6.0 to 7.25 hours, based on three consecutive overruns, revisit after 5 jobs."
Without the date and the reason, the next person to look at this template sees a number with no story, cannot tell whether it is a considered correction or a typo, and either leaves a bad number in place forever or removes a good one.
Step 6: Set the confirmation window and the stopping rule before you deploy
Decide now, while you are unattached to the outcome, what would tell you the correction worked and what would tell you it went too far.
A workable default for a job type running one to two per week: the next 5 completed jobs of that type. Fewer than 5 and you are back to reading noise. More than about 10 and you have left a possibly-wrong number in the field for a full quarter.
Two stopping rules, both written down before the first job lands:
- Done: at least 3 of the next 5 land within 10% of the new quoted hours, in either direction. Stop, leave the number, move to another job type.
- Take the rest: 4 or 5 of the next 5 still run over in the same direction. The bias was real and larger than half. Apply the remaining half of the original gap.
And one counter-signal you must watch alongside both: win rate on that job type. A correction that fixes your margin and quietly costs you a third of your bids has not improved anything. If you cannot measure win rate, at minimum count quotes issued versus jobs sold for that type over the same window.
Step 7: Treat the fourth miss as a different question
If you corrected and the misses continue, do not simply raise the number again. Four consecutive same-direction misses after a correction says something more specific than "still light." It says you probably corrected the wrong layer.
Go back to step 3 with more determination. Check whether the overrun moved phases after the correction - if the hours used to land in install and now land in commissioning, you fixed one thing and exposed another. Check whether the job type itself needs splitting, because a category that mixes two genuinely different jobs will never estimate cleanly no matter what multiplier you put on it. That split is usually the real answer, and it is the one people avoid because it means rebuilding a template rather than editing a field.
A worked case
A shop quotes a particular equipment-swap job type at 6.0 labor hours. Three in a row:
| Job | Quoted hours | Actual hours | Gap | Over its own 6.0 quote |
|---|---|---|---|---|
| 1 | 6.0 | 8.5 | 2.5 | 41.7% |
| 2 | 6.0 | 7.5 | 1.5 | 25.0% |
| 3 | 6.0 | 9.0 | 3.0 | 50.0% |
Step 1. All three over, all three in labor, all three with the overage landing in the same phase, and three different techs across two months. It is one pattern.
Step 2. Same direction three times. Sign is credible: this template is light. Size is not - the three gaps run from 1.5 to 3.0 hours, a 2x spread across just three jobs, which is exactly the situation where the observed median is a shaky estimate of the truth.
Step 3. Records pulled side by side. All three jobs required relocating the existing disconnect, which the template assumed was reusable in place. The tech confirms it: roughly 2 to 3 hours of extra work when the existing mount will not serve, and it is visible on a walkthrough if anyone thinks to look.
That is a named condition, so the fix is conditional. The base stays at 6.0 hours. A scoping question goes on the estimating checklist: "existing disconnect reusable in place - yes or no." A conditional line adds 2.5 hours when the answer is no. The 2.5 hours is the median gap of the three, which here is defensible because all three jobs had the condition, so the median gap is measuring the condition rather than a general bias.
Step 6. Confirmation window is the next 5 jobs of this type, recording the answer to the new question on every one.
The alternative path, if step 3 had found nothing. Median gap is 2.5 hours, half of which is 1.25 hours, so the base would move from 6.0 to 7.25 hours - a 20.8% increase over the old 6.0. Not the 50% that the worst job would have justified, and not the 41.7% the median gap would have justified. Half, on purpose, with a review at 5 jobs.
Notice how differently the two paths behave over the next quarter. The conditional fix prices the jobs that carry the condition correctly and leaves everything else alone. The blanket 7.25-hour fix raises every quote by 20.8%, including the jobs that were fine at 6.0, and pays for the fix with lost bids on the clean work. Both are better than doing nothing. Only one of them is right.
What changes the answer
Job frequency. A type you run twice a week supports a 5-job confirmation window inside a month. A type you run twice a year does not, and the correct response to three misses spread over eighteen months is a conditional line plus a scoping question - never a blanket number change you will not be able to evaluate before the next estimator inherits it.
Whether the three jobs shared a crew. Three misses from the same two-person crew is at least as likely to be an execution pattern as an estimating one, and raising the estimate would bake that crew's pace into every future quote. Separate the two before correcting.
Direction. Three consecutive jobs coming in under quote deserves the same investigation and almost never gets it. Persistent underruns mean you are leaving margin on the table or, worse, losing bids on a padded number for a job type you are actually efficient at.
Where the miss sits. A material miss is usually a count or a spec, and it is often fixable exactly rather than statistically - go find what was left off the list. Save the half-correction approach for labor, where the true value really is a distribution.
How to verify you got this right
At the end of the confirmation window, check three things in order.
First, did the misses stop, and in which direction? Overruns that turn into consistent underruns mean you overcorrected, and the fix is to give back part of the change, not to leave it and enjoy the margin.
Second, did win rate hold? A correction that improved job margin while cutting sold volume on that type may still be right, but that is a pricing decision you should make on purpose rather than discover.
Third, and most often skipped: is the change still documented and still attributed? Open the template and confirm the line from step 5 is there with its date. A correction nobody can trace back to evidence gets removed by the next person who thinks the number looks high, and you will run the whole loop again from job one.
References
- U.S. Small Business Administration (SBA), pricing and cost analysis guidance for small business
- Trade-standard practice for estimating template maintenance and version control
- See related: How to Adjust an Estimate Template From Real Data, How to Find the Job Types You Consistently Underbid, How to Separate Estimating Error From Execution Error