The Signs Your Estimating Is Drifting

Why this matters

An estimating system that was never right announces itself early. You lose money on the first few jobs, somebody notices, you fix it. An estimating system that used to be right and is slowly stopping is the dangerous one, because it degrades under your own confidence. Nothing looks broken. Each quarter is a little worse than the last by an amount too small to argue about, and by the time the margin number is bad enough to call a meeting, you have been bidding wrong for a year and a half of backlog.

Drift is also the failure mode most likely to be misdiagnosed, because when it finally surfaces it surfaces as a margin problem, and margin problems get blamed on prices, crews, and suppliers before anyone suspects the template. The signals below move earlier than margin does and point at the template directly.

Drift is not bias, and the two need different fixes

Bias is a constant offset: this job type has always run about 15 percent over, in every quarter, since you started measuring. It is a wrong number, it is correctable in one move, and it does not get worse on its own.

Drift is a trend: this job type ran near even two years ago, a little over last year, and clearly over now. The number was right when it was set. Something underneath it changed and the template did not follow.

The distinction determines the fix entirely. Correcting bias means changing a number once. Correcting drift means finding the underlying thing that moved, changing the number to match it, and then deciding how you will notice next time, because whatever moved once will move again. A shop that treats drift as bias applies a correction, feels good for two quarters, and drifts past the new number too.

The test is trivial and almost nobody runs it: plot the per-type median ratio of actual to estimated hours by quarter. Bias is a flat line above 1.0. Drift is a sloped line.

The lagging signal everybody uses, and why it is late

Gross margin per job is the signal shops actually watch, and it is the last one to move. It sits downstream of an averaging effect (good jobs offsetting bad ones), a mix effect (a shifting share of job types), and a backlog delay (jobs bid four months ago closing now). Any of the three can hide a real trend for two or three quarters.

Use margin as confirmation, never as detection. If margin is where you first noticed drift, you are roughly a year late, and every unclosed bid in your backlog carries the error.

Six signals that move earlier

1. Change-order rate rising on a stable job mix. Change orders are scope arriving after the bid. A rising rate on work that has not changed means your scoping is capturing less of the job than it used to, which is drift in the questions you ask, not in the numbers you use. This is often the earliest of all the signals because it fires on jobs that are still open.

2. Quoted hours clustering on round numbers. Look at the last thirty estimates for a type. If the hour figures are increasingly 8, 12, 16, 24 and increasingly identical across jobs that differ visibly in size, the estimator has stopped running the model and started recalling a number. Recalled numbers freeze at the moment the estimator last checked them, so the estimate stops tracking reality from that day forward, silently.

3. Time from scope to quote falling with no process change. Faster quoting is usually good and is sometimes the same clustering problem in a different costume. Check what got faster. If the site visit shortened or the take-off is now being skipped on familiar types, the speed came out of the measurement.

4. Win rate moving with no deliberate price change. Rising win rate at an unchanged price means the market has moved up and you have not, which is the exact situation where your cost inputs have gone stale. Falling win rate at an unchanged price means the opposite. Either direction is information, and a shop that does not track win rate by job type cannot see it.

5. Supply runs per job rising. Count mid-job trips to a supplier. A rising count means take-off quantities or truck stock assumptions no longer match the work, and the labor cost of those trips lands in the labor bucket where it reads as slow crews. This one is cheap to count and almost nobody counts it.

6. Template age relative to change events. Keep a date on every template and a short log of events that invalidate inputs: a payroll or insurance change, a supplier changing stock or package sizes, a crew composition change, a service-area change, a method or tooling change. A template that has not been touched since before two logged events is a drift candidate regardless of what its numbers currently look like.

Two false alarms

A single bad quarter is not drift. Small shops have small samples, and one large job with a genuine site surprise will move a quarterly median all by itself. Require a consistent direction across at least three consecutive reads before calling drift, and check whether removing the single largest job flattens the trend. If it does, you had one job, not a trend.

A mix shift inside a job type is not drift either. If "service call" quietly came to include a heavier class of work than it did two years ago, your ratio will slope upward while every input remains correct. The fix there is splitting the type, not correcting the number. Check this before anything else: compare the median job size within the type across the same quarters. If job size is trending too, you have a taxonomy problem masquerading as an estimating problem.

Read a series, not a reading

One quarter's ratio is a number. Four quarters is a direction. Six to eight is a slope you can act on with confidence.

Keep the series per job type, in one unit, with the sample count next to every point. A quarter built on three closed jobs sitting next to a quarter built on twenty is not a comparison, and a slope drawn through them is decoration. This is the same sample-count discipline that governs review cadence; a type that only produces three closes a quarter cannot support quarterly drift detection at all and should be read semiannually or annually.

A worked read across four quarters

A job type quoted at 16.5 crew-hours by the shop's model. Median actual-to-estimate ratio by quarter, with sample counts: Q1 1.02 (18 jobs), Q2 1.06 (21), Q3 1.11 (19), Q4 1.18 (22).

Is it drift? Four consecutive quarters, same direction every time, samples between 18 and 22 throughout. Removing the single largest job from each quarter moves each point by less than a full point of ratio and does not flatten the slope. That is drift, not noise.

How big. Q1 at 1.02 against 16.5 quoted is about 16.8 actual crew-hours. Q4 at 1.18 is about 19.5. So the same job type is taking 2.7 more crew-hours than it did a year ago, a 16 point move in the ratio over three quarters, which is roughly 5 points of ratio per quarter.

Mix check first. Median job size within the type is flat across the four quarters, and the share of the type's jobs that are the heavier variant did not move. So this is not a taxonomy problem.

Now decompose the actual hours by model input, comparing Q4 medians to Q1 medians:

Input Q1 actual median Q4 actual median Change
Productive hours 13.9 15.5 +1.6
On-site non-productive 1.8 2.2 +0.4
Travel 1.1 1.8 +0.7
Total 16.8 19.5 +2.7

The three input changes sum to the 2.7 hour total move, and each column sums to the quarter's median actual, which is the confirmation that nothing is missing from the decomposition.

This is the lesson. No single input moved enough to trigger anything. Travel went up 0.7 hours, which nobody notices on a two-day job. On-site non-productive went up 0.4 hours, inside anyone's tolerance. Productive hours went up 1.6, the largest piece, and even that is about 12 percent over the 13.9 hour Q1 median, which most shops would call normal spread on a single job. Drift is usually additive across several small movements, and that is precisely why a total-only reader misses it until margin moves.

Causes, one per input. Travel: the service area widened over the year and the model's travel line was set when the radius was smaller. On-site non-productive: a supplier package-size change meant more staging and more handling per job. Productive hours: two experienced techs left and the typical crew skill mix on this type shifted, which is a real cost change and not a performance complaint.

Which of the three earlier signals had fired. Supply runs per job had been rising for two quarters, consistent with the staging change. Win rate on this type went from 41 percent to 52 percent across the year with no price change, which was the market moving up while the shop held still. Change-order rate was flat, correctly, because scoping had not degraded.

What gets corrected, and separately. Travel line updated to the current service radius, and a rule that it gets re-derived whenever the service area changes rather than annually. On-site non-productive re-measured against the new package sizes. Productive hours are the awkward one: the honest answer is that the input was right for the crew it was priced at, so the correction is a crew-mix multiplier in the model rather than a permanent increase to the base hours, so the number does not have to be walked back when the crew mix recovers.

What they did not do. They did not add 2.7 hours to the template as a lump. That correction would have been right on average and wrong on every job, and it would have hidden all three causes, guaranteeing the same drift resumed from the new baseline.

What changes the answer

Very low volume. Under about eight closes a quarter for a type, quarterly points are too noisy to slope. Move to semiannual points, and lean harder on the leading signals, which do not need a per-type sample: supply runs, change-order rate and win rate are all countable across the whole shop.

A deliberate price change in the window. Win rate is uninterpretable across a quarter in which you changed prices on purpose. Mark the change on the series and read the two sides separately rather than through it.

High inflation in materials or wages. Input cost drift and estimating drift look identical in the margin number and are completely different problems. Read hours-based ratios rather than money-based ones to separate them: hours do not inflate. If your hour ratios are flat and margin is falling, the estimate is fine and your rates are stale.

A newly hired estimator. Expect a step change, not a slope, and expect it in both directions across types. Judge them on spread first and bias second; a new estimator whose numbers are consistently 10 percent light is far easier to fix than one whose numbers are centered but scattered.

How to verify a drift call before acting

  • Confirm the slope survives removing the largest job in each period. If it does not, you have one job, not a trend.
  • Confirm the sample count is stable across periods. A ratio that rises while the sample count halves is more likely a reporting problem than a cost problem, because the jobs that stop getting actuals recorded are not a random subset.
  • Decompose to inputs and check the pieces sum to the total move. If they do not, an input is missing from your model and you are about to correct the wrong one.
  • Look for the physical change. Real drift has a cause you can name and date: a supplier change, a service area change, a crew change, a method change. If you cannot name one, hold the correction and keep measuring for another period, because an unexplained slope is as likely to be a taxonomy or reporting artifact as a real one.

References

  • Trade-standard practice for cost trend analysis and project cost control
  • U.S. Small Business Administration (SBA), small business financial performance monitoring guidance
  • See related: The Estimate Accuracy Metrics Worth Tracking, Why Estimates Miss and Which Misses Matter, How to Set an Estimate Review Cadence, How to Adjust an Estimate Template From Real Data