The Difference Between a Bad Estimate and a Bad Job

Why this matters

A job came in 38 percent over on labor. That single fact supports two completely opposite responses: raise the price on that job type, or fix how the work runs. Pick wrong and you make things worse in a way that is hard to see. Raise the price when the real problem was a coordination failure and you have priced in your own inefficiency permanently, lost the bids you should have won, and left the coordination failure running. Retrain a tech when the template was 2 hours light and you have told a competent person they are slow, which is the fastest way to lose them.

Variance tells you something is wrong. Attribution tells you what to change. Most shops compute the first and skip the second, then wonder why their corrections never seem to make the numbers behave.

The two failure classes

A bad estimate is a systematic error in what you told yourself the job would take. It repeats. It shows up regardless of who runs the job. It is a defect in the template, the scope reading, or the load factor. The correction lives in the estimating system.

A bad job is a one-time execution or coordination failure on an otherwise correctly-estimated piece of work. It does not repeat in a pattern. It is caused by something that happened on that specific job: a part that was not staged, another trade blocking the work, a dispatch that sent the wrong crew size, a diagnosis that went down a dead end. The correction lives in operations.

There is a third thing that looks like both and is neither: a scope failure, where the job that was performed was not the job that was estimated. That is not an estimating error and not an execution error. It is a change that was never priced, and its correction is a change-order discipline problem. It is worth naming separately because attributing it to either of the first two classes produces a fix that cannot work.

The attribution test

Four questions, in this order. The first one that produces a clear answer usually ends the investigation.

1. Does it repeat across jobs of the same type? Pull the last 5 to 10 closed jobs of the type. If the median sits well over the template and most of them ran over, the estimate is wrong and this job is just one instance of it. If the median sits close to the template and this job is an outlier, the estimate is fine.

This question comes first because it is the cheapest to answer and the most decisive. Everything downstream assumes you already know whether the type behaves.

2. Does it repeat across technicians? If job after job of the type runs over regardless of crew, the estimate is wrong. If the overruns concentrate on one crew while others hit the template, you have a capability or method difference. Be careful here: check the assignment pattern first, because a shop that routinely sends its strongest tech to its hardest sites will produce data that makes the strongest tech look slowest.

3. Which bucket moved? Labor, materials, trips, and subcontract point in different directions. A labor overrun with materials on target usually means either time was lost or the task hours are light. Materials over with labor on target usually means a quantity or waste-factor error in the template. Both over together, with a return trip, usually means something was missing from the material list and the crew waited or drove for it - which is one defect, not two.

4. Does the timeline of the record explain it? Read the time entries in order. A block of hours with no corresponding progress is a wait, not slow work. Two techs logged where the task needs one is a dispatch decision. Hours logged on a day the site was supposedly unavailable is a data-quality problem you need to fix before you attribute anything.

Signal table

What you observe Points to The correction that follows
Median of the job type is well over template, most jobs over Bad estimate Correct the template layer
One job far out, type median near template Bad job Cause sentence and an operational fix
Overruns cluster on one crew, others hit template Capability or method Coaching, shadowing, or a method standard
Overruns cluster on one property class or access type Job type needs splitting Two templates, not one correction
Material over and labor over with a return trip Missing material line Fix the list, then re-measure
Big labor overrun, no material movement, no return trip Time lost on site Read the timeline for the wait
Actual work does not match the estimated scope Scope failure Change-order discipline, not pricing
Overruns start on a specific date across all types Something changed shop-wide Look at staffing, supplier, or process change

Why the wrong attribution is expensive in both directions

Calling a bad job a bad estimate is the more common error, because raising a template feels like a decisive fix and requires nobody to have an uncomfortable conversation. The cost is compounding. The padding rides on every future bid of that type, you lose a slice of the work you would have won, and the actual cause - the wait, the missing part, the crew-size call - stays live and produces the same overrun again on top of the now-higher price.

Calling a bad estimate a bad job is less common and more corrosive. It puts a measurement problem on a person. The tech knows the template is light, cannot say so credibly because the shop has decided the issue is speed, and either starts padding their own time entries to make the numbers safe or leaves. Both outcomes destroy the data you need to ever find the real problem.

Worked example: one 38 percent overrun, attributed

The record. A job type estimated at 16.0 labor hours closes at 22.0 hours. That is 6.0 hours over on a 16.0-hour estimate, which is 38 percent over. It clears every flag threshold, so it goes to review.

Question 1, the job type. This type has closed 7 times in the year including this one. Sorted actual hours: 15.5, 16.0, 16.0, 16.5, 17.0, 17.5, 22.0. The median is the 4th value, 16.5 hours, against a 16.0-hour template. That is 0.5 hours over the median, about 3 percent over the 16.0-hour template. The type behaves. This job is the outlier, not the type.

Read that carefully, because it is where a shop under time pressure goes wrong. The mean of those 7 values is 17.2 hours, which is about 8 percent over the 16.0-hour template, and someone looking only at the mean would start a template correction on the strength of a single bad job pulling the average. The median says the template is right.

Question 2, the crew. The lead on this job ran 4 of the 7 instances. Their other three came in at 16.0, 16.5, and 17.0 hours, all within about an hour of the 16.0-hour template. Nothing here suggests a capability gap. If those three had all sat near 20 hours, this becomes a different conversation entirely.

Question 3, the buckets. Materials landed at 1.02x the allowance, effectively on target. Subcontract, none. Trips, 2 against an assumed 1. So the overrun is almost purely labor, with one extra trip attached.

Question 4, the timeline. The time entries show a continuous 3.5-hour block on day one where both techs were logged on site but the work log shows no task progress. The lead's note reads that the area was still occupied by another trade and they could not get access. That is 3.5 labor hours of waiting, not slow work.

The second trip accounts for another 1.5 hours between drive time and the pickup. Cause noted: a component that the material list carried but the truck did not, a staging miss rather than a list error.

Totalling the explanation. 3.5 hours of access waiting plus 1.5 hours on the return trip is 5.0 hours of the 6.0-hour overrun. The remaining 1.0 hour sits inside normal job-to-job scatter for a 16.0-hour job and does not need a cause.

The verdict. Bad job, not bad estimate. The template stays at 16.0 hours. Two operational corrections come out of it: a pre-start access confirmation on any job sharing a site with another trade, and a truck-stock check against the material list before the crew rolls. Neither one is a price change, and pricing this job type higher would have fixed nothing.

What would have flipped it. If the sorted list had read 19.0, 20.0, 20.5, 21.0, 21.5, 22.0, 22.0, the median would be 21.0 hours against a 16.0-hour template, 31 percent over, with 7 of 7 jobs over. Same 22.0-hour job in front of you, opposite conclusion: the template is 5 hours light and the correction is a template correction. The single job never decides this. The distribution does.

The mixed case, which is most of them

Real jobs usually carry both. The template is a bit light and the job also ran badly. Attribute the portion you can explain from the record, and leave the rest to the type-level rollup rather than forcing every hour into a story.

The practical rule: explain what you can with named causes, and only what remains unexplained after that gets folded into the estimating question. In the example above, folding all 6.0 hours into the template question would have produced a template correction of 38 percent when the type's real gap was about 3 percent.

What changes the answer

Time-and-materials billing. The customer absorbed the overrun, so there is no margin loss to attribute. It still matters, because the quoted range you gave was wrong by 38 percent and that damages trust, but the correction is to your range-setting, not your cost template.

A job type with almost no history. Question 1 cannot be answered, so the attribution rests on the timeline read alone. Be more conservative: do not correct a template off a type with fewer than 5 closed jobs unless the miss is both large and repeated.

A shop-wide date break. If overruns across several unrelated job types all began in the same month, stop attributing job by job. Look for the shared cause: a new supplier with longer counter waits, a dispatch software change, a lost lead, a fuel or route change. Per-job attribution will find seven different stories for one cause.

Self-reported hours with no verification. If time entries are reconstructed at week end, the timeline read in question 4 is unreliable and you cannot separate a wait from slow work. Fix the capture before you invest in attribution.

How to verify the attribution held

The test is whether the correction worked. An estimate correction should show up as the type's median moving to within about 10 percent of the new template over the next 5 to 8 jobs. An operational correction should show up as the specific cause disappearing from cause sentences, and as the return-trip count on that job type dropping.

If you made an operational correction and the type's median keeps climbing, you attributed wrong. If you made a template correction and the same cause sentence keeps appearing, you attributed wrong in the other direction. Either way you learn it within a quarter, which is the reason to write the attribution down at the time rather than carrying it in your head.

References

  • U.S. Small Business Administration (SBA), cost control and operational review for small contractors
  • Standard root-cause practice, distinguishing systemic from special-cause variation
  • See related: The Job Costing SOP; Estimate vs Actual - Cost Variance Review; Why Your Average Job Is Lying to You