How to Adjust an Estimate Template From Real Data

Why this matters

Finding out that a job type runs 23 percent over is the easy half. The hard half is changing the template without wrecking it: correcting the wrong layer, correcting off one bad job, or stacking a new correction on top of last quarter's correction that nobody documented. A template that has absorbed four undated adjustments from three different people is worse than an unadjusted one, because it now carries padding nobody can find and nobody dares remove.

The write-back is a procedure with its own rules. Done well, your template gets tighter every quarter and your bids get more confident. Done casually, it inflates in one direction forever, you slowly price yourself out of a whole job type, and you never know which of the four changes did it.

Know which layer you are correcting

An estimate template is not one number. It is a stack, and the same overrun can live in any layer. Correcting the wrong one gives you a template that is right on average and wrong on every individual job.

Layer What it holds The miss that belongs here
Task hours The wrench time for each defined task The task genuinely takes longer than the template says
Named non-task steps Setup, protection, staging, cleanup, closeout paperwork A real, repeatable step nobody wrote down
Load factor The multiplier on task hours covering travel, interruption, and site friction Task times are right but the day never delivers the assumed productive hours
Material allowance Quantity and waste factor Consistent short counts or an unlisted consumable
Contingency The buffer for the identified risk on this job type The job type has a genuine long tail you are choosing to price for

The diagnostic question: does the overrun show up on every job of this type, or only on the ones where a specific condition was present? Every job means task hours or a missing named step. Only under a condition means contingency, or the job type needs splitting.

Step 1: Confirm the miss is systematic before you touch anything

Do not open the template until you have a sample and a direction. The default gate: at least 5 closed jobs of that type since the last correction, a median more than 10 percent off the template, and more than half the jobs missing in the same direction.

The last condition is the one people skip and it is the important one. A median 15 percent over, built from 5 jobs where 2 ran way over and 3 ran slightly under, is a bimodal job type wearing a template's clothing. Splitting it into two job types will fix more than any correction factor will.

Skipping this step is how a shop ends up correcting for one memorable disaster. The job you remember is by definition not the typical job.

Step 2: Target the median, not the mean and not the worst case

Set the new template to the median of the closed actuals for that job type. The mean gets dragged by the long tail that service work always has, and bidding at the mean prices you above the majority of instances you will actually run. The worst case is worse still: it converts an occasional bad job into a permanent price increase on every good one.

If you need protection against the tail, that is what the contingency layer is for, and contingency is a conscious, visible, removable number. Padding buried in task hours is none of those things.

Step 3: Decide between a correction factor and a rewrite

A correction factor multiplies the existing template - useful when the miss is proportional, meaning bigger instances of the job miss by proportionally more. A rewrite replaces specific line values - correct when the miss is a fixed amount that does not scale, like a setup step that costs the same whether the job is small or large.

Test it with two sizes. Pull your small instances and your large instances of the job type and compute each group's overrun as a percentage of its own estimate. If the small jobs run 40 percent over and the large ones 8 percent over, the miss is a fixed quantity, not a proportional one, and a single multiplier will overprice the large jobs while still underpricing the small ones. Add a fixed line instead.

If both groups run within a few points of the same percentage, a multiplier is honest.

Step 4: Change one layer per cycle

Change the task hours or the load factor, not both. If you move two layers and the next quarter's variance closes, you learned nothing about which one was wrong, and you have no way to unwind the half of the change that overshot.

The exception is when a cause sentence explicitly names two layers, for example a missing setup step plus a material line that was never on the list. Those are two distinct defects with two distinct fixes, and both can go in.

Step 5: Date-stamp and annotate every change

One line per change on the template itself: date, old value, new value, sample size it was based on, and the one-sentence cause. Without it, the next estimator cannot tell padding from measurement, so they either trust a stale correction forever or strip out a correction that was earned.

This is also what stops double-correction. A shop that corrects a job type in March for a missing setup step, then corrects it again in June for the same step because nobody recorded the first change, has now built two setup steps into one job.

Step 6: Watch two counter-signals for a full quarter

Post-change variance. Pull the next 5 to 8 closed jobs of that type. The median should now land within about 10 percent of the new template. If it does not, the correction went to the wrong layer.

Win rate on that job type. Track bids sent against jobs won, before and after. A correction that raises a price will cost you some work, and you need to know whether it is a little or a lot.

Read the win-rate signal carefully. Bid counts are small and noisy, and the number of bids it takes to distinguish a real drop from chance is larger than people expect. Do not reverse a template correction on a difference of one or two closed bids. Ask for at least 20 bids after the change, and prefer 40 before treating a swing of under 15 percentage points as real.

Worked example: one template, one correction, one confirmation

The starting point. A job type with a template of 6.0 labor hours. Eleven jobs of that type have closed since the last correction. The median actual is 7.4 hours.

The miss is 1.4 hours over on a 6.0-hour template, which is 23 percent over. Eight of the 11 jobs ran over, about 73 percent of them. That clears the step 1 gate: 11 jobs is above the 5-job minimum, 23 percent is above the 10 percent bar, and the direction is consistent.

Choosing the layer. The cause sentences on the flagged jobs are read together. Seven of the eight overruns name the same thing: covering and protecting the work area before starting, and reversing it at the end, which nobody ever put in the template. The techs have been doing it on every job for years. Estimated from the time logs where it was separately coded, it runs about 0.9 hours.

So this is a named non-task step, not a task-hours miss. That distinction matters: raising the task hours from 6.0 to 7.4 would hide a real, nameable step inside the wrench time, where the next estimator would read it as slow work and eventually "correct" it back out.

Making the change. Add a line: site protection and reinstatement, 0.9 hours. New template total is 6.9 hours.

The residual. Median actual 7.4 hours against a 6.9-hour template leaves 0.5 hours unexplained, which is 7 percent of the 6.9-hour template. That is below the 10 percent bar, so nothing else changes this cycle. It is logged and watched. Chasing a 7 percent residual with a second adjustment in the same cycle is how templates start carrying padding.

Proportional or fixed? Checked before committing: the small instances of this job type ran a median 0.8 hours over their estimates and the large ones ran 1.0 hours over theirs. Close to the same absolute quantity across both sizes, which confirms a fixed line rather than a multiplier. A multiplier that produced 1.4 extra hours on a 6.0-hour job would have produced far more on the large instances and overpriced them.

Confirmation, one quarter later. Eight more jobs of this type close. Median actual is 7.1 hours against the 6.9-hour template, so 0.2 hours over, about 3 percent over the 6.9-hour template. Inside tolerance. The correction went to the right layer.

Win rate. Before the change, the shop closed 12 of 20 bids on this job type, 60 percent. After, it closed 10 of 20, 50 percent. That is a 10 percentage point drop built on a difference of 2 closed bids out of 20, which is well inside what chance produces at that sample size. The correct action is to keep the correction and keep counting, not to reverse it. Revisit after another 20 bids, and if the post-change rate is still sitting near 50 percent across 40 bids, the question becomes whether this job type is worth running at a price that reflects its real hours - which is a strategy decision, not an estimating one.

What changes the answer

A supplier or code change rather than a measurement finding. If a code update adds a required step, correct the template immediately on a sample of zero. You are not measuring a trend, you are recording a known fact. The sample-size gate exists for statistical inference, not for facts you already have.

A crew that changed. If your overruns started when a lead left, you have a capability gap, not a template gap. Correcting the template locks the slower pace into your prices permanently and hides the training issue. Check the timeline before you touch the layer stack.

A job type you bid competitively against a known market. When you already know roughly what the work sells for, a template correction that pushes you above that number does not mean the template is wrong. It means the job type may not clear at your cost structure. Fix the template anyway, so the number is true, and take the pricing question separately.

Very low volume. Below about 5 jobs a year of a type, you will never accumulate a sample. Correct on the cause sentences: three jobs, three different surprises, means the scope work is thin rather than the hours.

How to verify you got it right

  • The template's change log has a dated entry with a sample size and a cause sentence, and an estimator who was not in the review can read it and understand why.
  • The next 5 to 8 closed jobs land within roughly 10 percent of the new template, and the direction of the residual is mixed rather than all one way.
  • Nothing was corrected twice. Search the log for the same cause appearing in two entries.
  • The corrections over the last year are not all upward. A template that only ever goes up is absorbing padding, not measurement.

References

  • U.S. Small Business Administration (SBA), pricing and cost control practices for small contractors
  • Standard construction estimating practice, production rates and labor factoring
  • See related: The Job Costing SOP; Estimate vs Actual - Cost Variance Review; How to Build a Labor Hour Database From Your Own Jobs