The Estimate Accuracy Metrics Worth Tracking

Why this matters

Most shops that start job costing track exactly one number: how far off the last job was. That number cannot tell you anything, because a single job's variance is mostly luck. Two shops can both report "about 20% over on average" and have completely different problems - one is systematically underbidding and needs a multiplier, the other bids the average correctly and cannot predict any individual job, which needs a scope definition, not a price change. Same headline number, opposite cures.

The metrics below are the smallest set that distinguishes those cases. Track them per job type, never blended across the whole company, because a blend of a type you overbid and a type you underbid reads as perfect accuracy while both are broken.

The core six

Metric What it answers How to compute Healthy default
Signed variance (bias) Am I systematically light or heavy? Mean of (actual - estimate) / estimate, keeping the sign Within +/- 5% of the estimate
Absolute variance (spread) How far off is a typical job, direction ignored? Mean of the absolute value of the same ratio Under 15% of the estimate
Hit rate What share land close enough to price on? Share of instances within +/- 15% of estimate 70% or better
Tail share How much of my total miss is a few jobs? Total absolute miss from the worst 10% of instances, over total absolute miss Under 25%
Coverage What share of closed jobs produced usable actuals? Jobs with complete actuals over jobs closed 95% or better
Correction latency How long from signal to fixed template? Days job-close to costing-close, plus quarters signal to template change 5 business days, one quarter

Those defaults are starting points for a residential service shop with a mixed job book. A shop running mostly repeat maintenance on known equipment should hold itself tighter, say 10% spread and an 80% hit rate. A shop doing renovation or diagnostic-heavy work on unknown systems will not reach 15% spread on its rough-in work and should not pretend to - set the target from your own best-performing job type and work the others toward it.

Bias and spread are different diseases

This is the single distinction the rest of the card hangs on, so it is worth stating precisely.

Signed variance keeps direction. Two jobs, one 20% over and one 20% under, average to zero signed variance. That shop bids the type correctly on average.

Absolute variance throws direction away. The same two jobs average to 20% absolute variance. That shop cannot predict any single instance of the type.

When bias is close in size to spread, the misses are nearly all in one direction. That is the easy case: the estimate is wrong by a consistent amount, and a multiplier on the template fixes it. This is the case that job costing is famous for catching.

When bias is small and spread is large, the average is right and every individual job is a coin flip. A multiplier does nothing here except shift the whole distribution - you will start overbidding the easy instances and still lose money on the hard ones, and your win rate will drop on exactly the jobs you were making money on. The real cause is almost always that the "job type" contains two or more genuinely different jobs, or that a site condition you are not asking about on the phone drives the hours. The cure is splitting the type or adding an intake question, not repricing.

When spread is small and bias is small, the type is under control and you should leave it alone and spend your review time elsewhere.

Why tail share earns its place

An average is a bad summary of a skewed distribution, and estimate variance is always skewed - a job can run three times its estimate but it cannot run less than zero hours. Tail share tells you whether your average is describing the population or being dragged by a handful of instances.

A high tail share (say a third or more of your total miss coming from the worst tenth of jobs) means the body of the work is fine and you have a distinct sub-type hiding inside the type. Those jobs almost always share something - a building age, an access condition, a customer who is present and talking, a system nobody in the shop has trained on. Find the shared trait, pull those instances into their own type with their own estimate, and both types get more accurate at once.

A low tail share with a high spread is worse news, because it means the unpredictability is evenly distributed and there is no sub-type to extract. That points at intake or at method, and it is a longer fix.

Coverage gates every other number

Coverage is the least interesting metric on the card and the one that invalidates all the others when it slips, because the jobs missing their actuals are not a random sample. The job that ran clean gets closed out the same afternoon. The job that went sideways, where the tech is embarrassed about the hours and the office is chasing a change order that never got signed, is the one still sitting open six weeks later.

That means low coverage biases every accuracy metric optimistic, in the exact direction that stops you from finding your problems. A shop reporting 8% bias at 70% coverage may well be running 15% bias in reality. Get coverage above 95% before you act on any of the other five, and treat a coverage drop as an alarm in itself rather than a data-entry annoyance.

What is not worth tracking

A single blended accuracy number for the whole shop. It nets an overbid type against an underbid type and reports health. It is the metric most likely to be on a dashboard and least likely to be useful.

Share of jobs that were profitable. Too coarse to act on. A job at 1% margin and a job at 40% margin both count as profitable, and the metric moves only when something is already badly wrong.

Variance measured in total across labor, materials and subs together. Labor over and materials under net to a clean-looking total while both are wrong. Compute variance per bucket and keep them separate all the way through, because the corrections are different: labor variance goes to the time template, material variance goes to the takeoff quantities or the supplier price file.

Estimator-level accuracy scoring in the first year. Tempting and destructive. It converts the review into a performance evaluation, and the reliable response is that estimators start padding, which improves their score and destroys your win rate. Track by job type until the data is trusted.

A worked quarter

A shop closes 51 jobs in a quarter. 44 have complete actuals, so coverage is 44 of 51, about 86%, below the 95% target and worth flagging before reading anything else. Of the 44, three job types account for the bulk:

Job type Instances Signed variance Absolute variance Hit rate within +/- 15%
A - standard service call 30 +4% 22% 47%
B - equipment changeout 12 +19% 20% 25%
C - small install 9 -2% 6% 89%

Type C is healthy. Near-zero bias, tight spread, high hit rate. It gets no attention this quarter.

Type B is the easy fix. Signed variance of +19% against absolute variance of 20% means essentially every instance ran over, by similar amounts. That is a one-directional, structural underbid. Apply a multiplier of about 1.19x to the labor line of the template and re-measure next quarter. The hit rate of 25% is low precisely because the whole distribution sits to the right of the estimate - shifting it should move most instances back inside the band without touching the spread at all.

Type A is the interesting one and the one most shops get wrong. Signed variance of +4% is inside the healthy band, so the estimate is right on average. But absolute variance of 22% and a hit rate of 47% say fewer than half of these jobs land anywhere near their number. Raising the price here would be a mistake: you would start losing the easy calls to a competitor and still eat the hard ones.

Dig into the tail. Type A is estimated at 2.0 hours per instance, so 30 instances carry 60.0 estimated hours. At 22% mean absolute variance, the mean absolute miss is 0.44 hours per instance, so total absolute miss across the type is 13.2 hours. The three worst instances ran 4.5, 3.8 and 3.0 actual hours against the 2.0-hour estimate, misses of 2.5, 1.8 and 1.0 hours, or 5.3 hours combined.

That is 5.3 of 13.2 hours, about 40% of the type's total absolute miss, carried by 3 of 30 instances

  • a tail share well over the 25% default and the clearest signal in the quarter. Strip those three out and the remaining 27 instances carry 7.9 hours of absolute miss, a mean of 0.29 hours, which against the 2.0-hour estimate is about 15% absolute variance rather than 22%. The body of the work is borderline acceptable. The tail is the whole problem.

Check what the three share. In this case all three were in buildings where the equipment was in a crawl space rather than a utility room. That is a scope trait, it is knowable at booking with one question, and it belongs in its own job type with its own hours. Split it, and both the new type and the remaining Type A become predictable without anyone changing how they work.

Note what the naive read would have concluded: "Type A averages only 4% over, we are fine there, let us focus on Type B." Type B is a 12-instance type with a clean, easy correction. Type A is a 30-instance type hiding a sub-type, and it is the larger exposure by volume even though its headline bias looks better.

How to verify you are reading these right

Re-derive one metric by hand from raw rows. Pull the raw actuals for one job type and compute signed and absolute variance yourself rather than trusting the report. The most common data defect is that "actual hours" quietly includes or excludes drive time depending on who entered it, which moves every number by a consistent amount and is invisible in the summary.

Confirm your bias number is not a coverage artifact. Take the jobs excluded for missing actuals, force-close five of them with the best figures you can reconstruct, and see whether the bias moves. If it jumps, coverage is your first project and everything else waits.

Check that the correction you applied last quarter actually landed. A template edited but never deployed to the field - the estimator still working from a saved copy, the price book not republished is common enough that it should be checked explicitly. The tell is a bias number that does not move at all between quarters despite a correction being agreed.

Sanity-check a percentage against its own base out loud. Every accuracy number here is a ratio to the estimate, in hours, over one job type, over one quarter. A number quoted without all four of those qualifiers is the point where these reviews usually go wrong.

References

  • See related: Estimating Confidence From Tracking Actuals for getting the capture habit started.
  • See related: Tracking Margin by Job Type to Find the Leaks for the margin view alongside this accuracy view.
  • See related: How to Read Your Own Job Costing Data for sample-size gates and distribution reading.
  • SBA guidance on small-business performance measurement and management reporting cadence.