How to Run a Monthly Estimate Accuracy Review
Why this matters
Reviewing individual jobs tells you about individual jobs. It will never tell you whether your estimating, as a system, is getting better or worse. That is a different question with different numbers, and a shop can run diligent per-job reviews for a year without ever answering it.
The monthly review answers three things nobody else in the shop is asking: is our estimating leaning in one direction, is it getting tighter, and did the corrections we made two months ago actually work. Forty-five minutes, same slot every month, four numbers on a page. It is the cheapest management meeting in the business and the one most often skipped, because nothing is on fire.
Step 1: Set the cadence and protect the slot
Monthly, 45 minutes, the same day each month, and it happens whether the month was good or bad. Attendance is whoever estimates, whoever owns the templates, and whoever runs operations. In a small shop that is two people, and the meeting still happens with two people.
Monthly is deliberate. Weekly is too fast to see a trend in shop-level numbers and turns into a per-job review, which you already have. Quarterly is too slow: a correction made in January and reviewed in April has spent three months either working or quietly not working, and the difference is a quarter of your volume.
The single most common failure is the meeting getting cancelled in the busy season. That is exactly when estimating drifts most, because peak-season jobs cost more hours and nobody is watching. Protect it hardest when you least want it.
Step 2: Prepare four numbers before anyone sits down
All four cover the month's closed jobs. Prepared in advance, on one page, no arithmetic in the room.
- Bias. The median labor variance across all closed jobs, as a percentage. Median, not mean, so one runaway job cannot move it. This is the lean.
- Hit rate. The share of jobs landing within your noise band. Use plus or minus 10% of estimate as the default band and tune it to your job sizes. This is the tightness.
- Over-share. The share of jobs that landed over estimate at all, regardless of size. A calibrated system sits near half.
- Tail count. The number of jobs that landed more than 25% over. These are the ones that will hurt margin regardless of how the median looks.
Set targets and commit to them rather than leaving the metrics as a scoreboard nobody scores. Workable defaults for a shop with the loop running: bias inside plus or minus 3%, hit rate at or above 60%, over-share between 45% and 55%, tail count under 10% of closed jobs. Tune them once you know your own spread. A shop doing large project work will run a wider band; a shop doing standardized service calls should beat these.
Step 3: Open with verification, not with new findings
First agenda item is always the corrections made in previous months, not what happened this month. Everyone wants to start with the new numbers, and starting there is how corrections go unverified for a year.
For each open correction: the change, the date it took effect, how many jobs of that type have closed since, and the median variance on those jobs only. Three outcomes:
- Verified. The bias moved inside the band. Close it and stop tracking it.
- Not enough jobs yet. Fewer than about eight since the change date. Carry it forward, do not read anything into the number yet, and resist the urge to comment on the direction of three data points.
- Failed. Enough jobs, bias barely moved. The correction was too small or the diagnosis was wrong. Send it back for re-diagnosis rather than simply doubling the number, because a wrong cause corrected harder is still wrong.
Verifications first also sets the tone of the meeting: this is a place where things get closed, not just a place where problems get listed.
Step 4: Read the four shop-level numbers as a set
Read them together, because each one alone can mislead:
- Bias near zero with a low hit rate means your estimates are centered but scattered. That is an information problem at quote time, not a template problem. Correcting template numbers will not help and may hurt.
- Bias positive with over-share well above half means a genuine systemic lean. Something is under-counted across many job types: non-productive time, consumables, travel between stops. Look for the cause that is common to everything rather than tuning individual templates.
- Bias near zero with a high tail count means most jobs are fine and a small number are blowing up. That is a scoping or site-conditions problem on a subset. Go find what those jobs share.
- Hit rate rising while close rate falls is the one to watch for. It means you corrected upward and are now losing bids. A lost bid produces no cost data, so the estimating numbers keep improving while the business gets worse. Always look at close rate in the same meeting for exactly this reason.
Step 5: Look at the movers, not the whole list
The per-job-type table goes in the pack, but the meeting discusses only movers: any job type whose bias shifted by more than about 8 percentage points from the prior month, in either direction.
A stable job type running steadily 4% over needs no monthly conversation. It is on the correction list or it is inside the band, and either way it is handled. A job type that was at 4% last month and 19% this month has had something change, and finding out what within a month of it happening is the entire point of a monthly cadence.
Sudden movers usually have one of four causes: a supplier price change, a new person on that work, a seasonal shift, or a template that got edited by someone outside the process. Check the last one first. It takes thirty seconds and it explains more movers than people expect.
Step 6: Decide, assign, date, and cap
Every item that gets discussed ends in one of three states before the room moves on: corrected (a specific change, an owner, a date), watched (named, with what evidence would trigger action), or closed (explicitly not a problem, with the reason).
Cap the corrections at three per month. That is not a productivity target, it is a verification constraint: each correction has to be measurable next month, and a batch of nine simultaneous changes across overlapping job types cannot be attributed when the numbers move. Three changes you can verify beat nine you cannot.
Step 7: Record the month's four numbers so a trend can exist
Log bias, hit rate, over-share, tail count, and closed-job count in a running sheet. One row per month.
This is the step that gets skipped because it produces nothing this month. It is also the only reason the meeting is worth anything in month six, when the question is whether the estimating system is improving and the answer is a trend line rather than an opinion. Twelve rows of four numbers is the entire artifact, and it will outlast every person who attended.
A worked monthly review
Three consecutive months at one shop:
| Month | Closed jobs | Bias (median) | Hit rate (within 10%) | Over-share | Tail (over 25%) |
|---|---|---|---|---|---|
| 1 | 58 | +7% | 24 of 58 (41%) | 42 of 58 (72%) | 7 |
| 2 | 61 | +5% | 29 of 61 (48%) | 40 of 61 (66%) | 5 |
| 3 | 55 | +3% | 31 of 55 (56%) | 33 of 55 (60%) | 4 |
Read the set. Bias is falling toward zero. Hit rate is climbing. Over-share is walking down toward half. Tail count is shrinking. All four are moving the same direction, which is what a calibrating system looks like, and it is more convincing than any one of them alone would be.
Against the targets, month 3 has bias at +3%, right at the edge of the plus or minus 3% target, and hit rate at 56%, still short of the 60% target. Over-share at 60% is above the 45 to 55 range, so there is still a real lean left. The system is close but not yet calibrated, and the honest read is: keep going, do not declare victory.
Close rate over the same three months held roughly steady. That matters. If close rate had dropped noticeably while these four numbers improved, the improvement would be partly an artifact of pricing yourself out of the jobs you used to underbid, and the correct response would be to stop correcting upward.
Verification segment, month 3. Three corrections were made at the end of month 1:
- Correction A, a template raised 8% on a common job type. 11 jobs have closed since the change date, median variance +2%, down from +11%. Verified. Closed.
- Correction B, a consumables allowance raised on a second job type. 3 jobs since the change date. Too few. Carry forward, say nothing about the direction.
- Correction C, a template raised 6% on a third job type. 12 jobs since, median variance +14%, essentially unchanged from the +15% before the change. Failed. The 6% went in and the bias did not move, which means the cause was not template hours at all. Back to diagnosis, and the first place to look is whether those jobs share a site condition or a crew rather than a job type.
Movers segment, month 3. One job type moved from +4% to +19%, a 15 point jump. The template-edit check comes first: the price book was updated three weeks ago for a supplier increase, and the labor hours on that template were reset to a default in the process. That is the whole explanation, it took one minute to find, and the fix is restoring the hours and adding the template to the dated change log so it cannot happen invisibly again.
Output of the meeting: one correction closed, one carried, one sent back to diagnosis, one template restored. Four items, all with owners and dates, and four numbers added to the running sheet.
What changes the answer
- Low volume. Under about 25 closed jobs a month, the shop-level metrics are too noisy to read monthly. Run the meeting monthly for verification and movers, but read bias and hit rate on a rolling three-month window so the numbers mean something.
- A mixed book of work. If a third of your volume is time and materials, split the metrics. T and M jobs measure estimating accuracy but not margin exposure, and blending them into one bias number tells you neither cleanly.
- Right after a method or pricing change. The month a new price book or install standard lands, the numbers describe a transition, not a system. Note the change date on the running sheet and expect two noisy months.
- Seasonal peaks. Compare a peak month to last year's peak month, not to last month. A shop that reads July against June every year will conclude its estimating collapses annually, when the truth is that peak work costs more hours and always did.
How to verify you got this right
- Every percentage in the pack carries its count and its base. "31 of 55 closed jobs landed within 10% of estimate, 56%." A bare percentage in a monthly pack is the easiest number in the business to read against the wrong denominator six months later.
- The buckets reconcile to the job count. Jobs within the band, jobs over the band, and jobs under the band must sum to closed jobs. If they do not, some jobs have no frozen estimate to compare against, and those are worth finding: they are usually the rushed ones.
- The running sheet has an unbroken row per month. A gap is not a missing row, it is a month where nobody looked, and drift in a gap month is invisible forever.
- Corrections closed roughly matches corrections opened, over a few months. A list that only grows means the meeting detects and never confirms, which is half a loop wearing the clothes of a whole one.
References
- U.S. Small Business Administration: job costing, pricing, and financial review guidance for small contractors.
- Standard construction and field-service practice on estimating database maintenance and periodic accuracy review.
- See related: The Estimating Feedback Loop Explained.
- See related: The Post-Job Cost Review SOP.
- See related: How to Find the Job Types You Consistently Underbid.