How to Separate Estimating Error From Execution Error

Why this matters

A job came in over. You know that much. What you do next depends entirely on which of two completely different things went wrong, and the two have opposite cures. If the bid was light, you fix the template and every future job of that type prices correctly. If the bid was right and the work ran slow, changing the template does nothing except make you more expensive than the shop down the road, and you have permanently priced a training problem into your customer's invoice.

Shops that never separate these two drift in a predictable direction: every overrun gets absorbed into the estimate, the estimate keeps climbing, the win rate slides, and nobody can explain why the numbers went up. The variance told them something was wrong. It did not tell them what, and they guessed.

The two errors respond to opposite actions

An estimating error is a defect in the prediction. The template said 6 hours because whoever built it never counted the access time, the haul-out, or the second trip for a fitting nobody stocks. The work was done competently and it still took longer. The cure lives in the template.

An execution error is a defect in the performance. The template was honest, and this particular run of the job absorbed extra hours because of who did it, how they did it, what was on the truck, or how the day went around them. The cure lives in scheduling, stocking, training, or method - never in the price.

The reason this matters more than it sounds: an estimating error is systematic and shows up on every instance of that job type. An execution error is variable and shows up on some instances. That difference is the entire diagnostic, and you can test it with data you already have.

Step 1: Normalize the scope before you attribute anything

You cannot attribute a variance you have not cleaned. If the customer added work mid-job and it was approved, the estimate you compare against is the estimate plus the approved change, not the original. If it was not approved, it is neither an estimating nor an execution error - it is unbilled scope, which is a third category with its own cure (a change-order discipline), and folding it into either of the other two corrupts both.

Do this first because everything downstream is arithmetic on the wrong baseline otherwise. A shop that routinely absorbs small unapproved adds will read its own data as a chronic estimating miss and raise prices to cover work it should have been billing.

Step 2: Split the variance by who ran it, before you split it by anything else

Pull the last 8 to 12 completed jobs of one type. Fewer than 8 and you are reading noise. Group the actual hours by lead tech or crew. You are looking for two numbers: the floor, meaning the best performer's average, and the spread, meaning the gap between best and worst.

The floor is the honest cost of the job when it is done well. If the floor sits above the estimate, that excess is estimating error, full stop, because it happened even in the hands of the person who runs the job best. No amount of coaching removes it.

The spread is execution capacity. It is the difference between your best and your worst, and it belongs to training, method, or assignment - not to the bid.

Step 3: Apply the repeatability test

Ask one question of the variance: did it show up regardless of who ran it?

  • Present on essentially every instance, across multiple techs, in a similar size - estimating error.
  • Present on some instances and absent on others, and the split lines up with a person, a truck, a neighbourhood, a season, or a shift - execution error.
  • Present on exactly one instance, at several times the size of anything else in the set - neither. That is an event, and it gets investigated as an event, not folded into an average that it will distort.

The reason to test repeatability rather than reading the average: the average of a set containing one blowout says the job type is unprofitable when eleven of twelve instances were fine. Averages hide the shape. Sort the list and look at it before you compute anything.

Step 4: Check the three execution conditions on the jobs that ran long

For the instances the repeatability test flagged as execution, check three things in order, because they account for most of it and they are cheap to check:

  1. Crew composition. Was it the normal pairing? A lead running with a helper who has not done this job type before produces a predictable overrun that has nothing to do with either person's ability.
  2. Truck stock. Did the job include a supply run? A single unplanned parts trip is often the whole variance by itself, and it is a stocking decision, not a skill problem.
  3. Method drift. Was the job done in the sequence the template assumes? A tech who does the demolition before the isolation, or who sets up twice because they did not stage, is not slower - they are running a different job.

If none of the three explain it, then you are looking at genuine speed difference, which is a coaching conversation and needs a separate track from anything in this article.

A worked example, carried through

A shop tracks a standard equipment changeout. The template calls for 6.0 hours of labor. Twelve completed instances over a quarter, actual labor hours, grouped by lead:

Lead Instances Actual hours each Total Average
Tech A 6 7.5, 8.0, 7.0, 7.5, 8.5, 7.5 46.0 7.67
Tech B 6 6.5, 7.0, 6.5, 7.0, 6.5, 6.5 40.0 6.67
All 12 86.0 7.17

The naive read: 7.17 actual against 6.0 estimated is 1.17 hours over, which is 19.4% over the 6.0-hour estimate. Raise the template to 7.2 hours and move on.

That read is wrong, and here is the cost of it. Split by lead:

  • Tech B's floor is 6.67 hours, which is 0.67 hours over the 6.0-hour estimate, or 11.1% over the 6.0-hour estimate. Tech B runs the job well and it still takes longer than the template says. That 0.67 hours is estimating error. It is in every instance in the set.
  • Tech A averages 7.67 hours, 1.67 hours over the estimate, or 27.8% over the 6.0-hour estimate. But 0.67 of those hours is the same estimating error Tech B has. The remaining 1.0 hour is the gap between A and B, and it is execution.

Now check the spread inside each group before trusting either average. Tech B's six values run 6.5 to 7.0, a range of 0.5 hours - tight, consistent, believable as a floor. Tech A's run 7.0 to 8.5, a range of 1.5 hours, three times B's spread. That widening is itself a signal: it is what a process still being learned looks like, versus one that is settled.

Checking Tech A's three longest instances against the execution conditions: two of the three included a mid-job supply run for a fitting that is not on the truck stock list. That is one clean cause, and it is a stocking fix, not a pricing fix.

The correction that follows. Move the template from 6.0 to 6.7 hours, about a 1.12x multiplier on the old number. Add the missing fitting to standard truck stock. Put Tech A's remaining gap on the coaching track with a specific target of the 6.5 to 7.0 band, not a vague "work faster."

What the naive read would have cost. Had the template gone to 7.2 hours instead, every job Tech B runs would come in about 0.5 hours under the bid, and the shop would be carrying a permanent price premium of roughly half an hour per changeout to fund a stocking gap and a training gap. On a job type running twelve instances a quarter, that is 6 hours a quarter of price you are charging the market for nothing you are delivering. In a competitive bid situation that is the difference between winning and being the second-cheapest number on the sheet, and you would never trace the loss back to this decision.

What changes the answer

Sample size below 8. With four instances you cannot distinguish a floor from a lucky run. Widen the window to two quarters rather than acting on thin data, or hold the correction and keep measuring.

One tech runs the job type. You have no floor to read, because performance and prediction are confounded. Two options: have a second lead run the next two instances specifically to establish a comparison, or treat the current average as a provisional ceiling and revisit when the sample widens. Do not assume the single tech is the floor, and do not assume they are the problem.

The job type is not really one type. If "changeout" silently includes both attic and ground-level work, the spread you are reading is a scope difference, not an execution difference. Split the type and re-run. This is the most common false positive in the whole procedure, and the tell is that the long instances share a site characteristic rather than a person.

A new person joined mid-window. Their learning curve is real and temporary. Exclude their first few instances of a job type from the floor calculation and track them separately, or the ramp gets written into your price permanently.

The variance is in materials, not hours. This whole procedure is built for labor, where a floor exists because a person can only go so fast. Material variance rarely has an execution component in the same sense - it is usually a quantity error in the takeoff, a price change from the supplier, or waste. Attribute those three separately rather than forcing them into this frame.

How to verify you got this right

Run the next 6 instances after the correction and check three things:

  1. Did the floor move to the new estimate? Tech B's next instances should land near 6.7 hours with a small variance, not consistently under it. If they now come in at 6.0 again, you corrected a temporary condition, not a structural one.
  2. Did the spread close where you intervened? If the stocking fix worked, Tech A's supply-run instances disappear and A's range narrows toward B's. If the range stays at 1.5 hours, the supply run was correlation, not cause, and you should re-open the execution check.
  3. Did the coaching gap close on a schedule? A gap that is still 1.0 hour after six more instances is not closing on its own, and continuing to leave it out of the price is a choice you should make deliberately rather than by default.

The check that catches the most self-deception: pull the corrected template's next quarter of instances and compute the floor again from scratch, ignoring what you decided last time. If the floor has moved again in the same direction, you are chasing a drifting number and something upstream is changing - scope creep on the job type, a supplier change, an aging customer base of equipment. Chasing it with template edits every quarter treats the symptom.

References

  • See related: Estimate vs Actual: Cost Variance Review and How to Compare Estimated Against Actual on Every Job for computing the variance this article attributes.
  • See related: The Job Costing SOP for the per-job capture procedure that produces the actuals used here.
  • See related: Tracking Margin by Job Type to Find the Leaks for the roll-up view above the single job type.
  • SBA small-business financial management guidance on cost tracking and variance review as a management routine.