How to Set an Estimate Review Cadence
Why this matters
Most shops that decide to "start reviewing estimates" pick monthly because a month is a familiar unit, then quit inside two quarters. The reason is not discipline. It is that one interval is being asked to answer four different questions at once, and it can only answer one of them well. A month is far too slow to catch a job that went sideways while anyone still remembers it, and far too fast to prove that a job type you run twice a month is systematically underbid.
Cadence is a sampling problem, not a calendar problem. Every question you want to ask needs a certain number of closed jobs before the answer means anything, and the interval that answers it is however long your shop takes to accumulate them. Set it that way and the review stops feeling arbitrary, because each tier fires exactly when it has something new to say.
This article is about designing the schedule. The mechanics of the monthly meeting itself are a separate sibling; see the references.
Step 1: Write the four questions down before you pick any interval
There are only four, and they sit at four different altitudes:
- Did this job land where we said? Scope: one job. Needs a sample of one. Answerable immediately.
- Is anything on fire right now? Scope: the last handful of closes. Needs enough jobs to spot an outlier, which is a handful.
- Which direction is the shop as a whole leaning? Scope: all job types together. Needs enough total closes to see a median move, typically 20 to 30.
- Is this specific job type systematically wrong? Scope: one job type. Needs enough closes of that one type, which is the slowest counter in the building.
Write these four on the wall in this order. The single most common cadence failure is trying to answer question 4 in a monthly meeting, getting three jobs of the type, arguing about them, and either changing the template on noise or concluding the whole exercise is useless.
Step 2: Compute your accumulation rate per job type
Pull twelve months of closed jobs and count them by job type. You want two numbers per type: closes per month, and how many months it takes to reach twelve closes.
Twelve is the working floor for believing a per-type median. Below about eight closes a median is easily swung by one job, and between eight and twelve you can read direction but should not set a correction factor from it. If you can only get eight, you can conclude "this type leans over" and go looking; you cannot conclude "add 15 percent."
Now sort your types by that months-to-twelve number. What comes out is almost always a steep curve: one or two types hit twelve inside two months, a middle group takes four to six, and a long tail takes a year or never gets there. That curve, not the calendar, is your cadence.
The types in the tail are not a failure of measurement. They are types where your correct policy is to bid them from the model rather than from history, and to review them individually rather than statistically.
Step 3: Per-job tier, at close, on a threshold
Fire a single-job look at every close where variance crosses a stated line. A workable default: labor hours more than 25 percent over the scope-corrected estimate, or any job closing below zero margin. Fifteen minutes, one person, one sentence of cause written on the job record.
This tier exists for recall, not for statistics. Fourteen days after a job, nobody can tell you which morning the access problem appeared, and the sentence you write becomes speculation. Run it at close and the sentence is worth something. It is that written sentence, accumulating on job records, that makes every slower tier possible: without it the quarterly review is reading numbers with no attached causes and can only guess.
Set the threshold as a number and apply it to everyone's jobs including the owner's. A discretionary trigger is read as an accusation and poisons the reporting the whole system runs on.
Step 4: Weekly tier, triage only, hard capped
Fifteen minutes, once a week, on the jobs that closed that week. The only questions: did any job cross the per-job threshold and not get its sentence written, and is any open job already tracking past its estimate.
Cap it at fifteen minutes and cap it hard. The failure mode here is a weekly meeting that grows into an analysis session, at which point people start finding patterns in four data points. Four jobs is not a pattern, and a shop that changes an estimate on four jobs will change it back six weeks later, which teaches everyone that the numbers do not mean anything.
The second job of the weekly tier is catching an overrun while the job is still open, which is the only moment a change-order conversation is still available.
Step 5: Monthly tier, shop-level only
Once a month, on the aggregate across all job types: median variance, spread, the share of jobs in the tail, and coverage (what fraction of closed jobs actually have usable actuals). Nothing per-type.
The discipline here is a subtraction: the monthly meeting is explicitly forbidden from changing a job-type template. Its output is direction and a watch list. A shop closing on the order of 20 jobs a month has enough total volume in a month to see the shop-level median move, and does not have enough of any single type to justify a correction.
Coverage belongs in this tier because it gates everything else. If only 6 of 20 closed jobs carried usable actuals last month, your quarterly per-type numbers are built on a self-selected third of the work, and the missing two thirds are not random. Coverage below about 80 percent is the finding, and fixing it outranks any variance you think you see.
Step 6: Quarterly tier, the only tier that changes a template
Per job type, and only for types that have accumulated twelve or more closes since the last change to that template. This is the tier with the authority to move a number, and the twelve-close gate is what gives it that authority.
Change one layer per cycle per type: either the base hours, or an allowance, or a condition question, not all three. If you move three inputs at once and the next quarter comes back better, you have learned nothing about which one mattered, and you cannot reverse the one that was wrong.
Date-stamp every change on the template and reset that type's counter to zero. A type whose template changed in month two of the quarter has not accumulated twelve closes against the new number no matter what the calendar says, and reading its next-quarter variance as a test of the change is a mistake shops make constantly.
Step 7: Annual tier, the rates underneath everything
Once a year, rebuild the inputs that sit below every template: the burdened labor rate, the overhead allocation, the travel load, and the standing waste allowances. These change slowly and for reasons outside any job, so reading them quarterly produces noise and an unnecessary argument.
Pull this one forward off-cycle for exactly three events: a payroll or insurance change large enough to move burden, a supplier changing standard package or stock sizes, and a service-area change that moves typical drive time. Those three invalidate an input immediately rather than gradually, so waiting for the anniversary bakes a known-wrong number into every bid in between.
A worked cadence
A shop closes about 19 jobs a month, roughly 57 a quarter, across six job types. Twelve months of history gives closes per month by type: A 7, B 4, C 3, D 2, E 2, F 1.
Months to twelve closes. A reaches twelve in under 2 months. B in 3 months. C in 4. D and E in 6. F in 12.
What that produces. Type A accumulates about 21 closes a quarter, so it gets a real per-type correction every quarter with a comfortable sample. B accumulates 12 a quarter, right at the floor, so it also reviews quarterly but any correction gets treated as provisional and confirmed the following quarter. C at 9 a quarter is under the floor, so C moves to a semiannual review at about 18 closes. D and E at 6 a quarter go semiannual as well, landing near 12. F at 3 a quarter goes annual, and even then it only just reaches 12.
The reallocation this forces. The shop's old habit was one monthly meeting that walked all six types. That meeting was reading 7 jobs of type A (worth reading) and 1 job of type F (worth nothing) with the same seriousness, and the type F discussion took the longest because it was the most surprising. The new cadence reads A and B quarterly, C through E twice a year, and takes F out of the statistical process entirely.
What happens to F. One job a month is not a sample, so F gets a per-job review at every close, with the sentence of cause written every time, and it is bid from the cost model rather than from a historical average. Over a year the accumulated sentences on those 12 jobs are worth more than the median of 12 numbers would have been, because the sample is too small for the median to be stable but the causes repeat visibly.
The check they built in. At the end of the first year the shop compares each type's variance spread before and after. Type A, on 21 closes a quarter and four correction cycles, should show the largest narrowing. Type F should show none, and that is the expected result, not a failure. If F narrows and A does not, something is wrong with the correction process, not with the cadence.
What changes the answer
A shop closing far fewer jobs. At 5 or 6 closes a month, no job type reaches twelve inside a quarter and the quarterly tier has nothing to do. Collapse to three tiers: per-job at close, monthly shop-level, and an annual per-type review. Do not compensate by lowering the twelve-close floor. A correction set on four jobs is a coin flip that costs you a template.
A shop with one dominant job type. If 70 percent of your volume is one type, that type can support a monthly correction cycle on its own while everything else stays quarterly or slower. Cadence is per type, not per shop, and there is no rule that says all types review together.
Seasonal work. A type that only runs five months a year accumulates in bursts. Run its review at the end of its season rather than on the calendar quarter, because a quarter that contains no jobs of that type produces a meeting with nothing in it and trains people to skip meetings.
A template you just changed. The counter resets. Two quarters after a change on a slow type is not stalling, it is the sample arriving.
A brand new job type. No history, no cadence. Bid it from the model, review every single one at close, and let it enter the tiered schedule once it has eight closes and you can see a direction.
How to verify the cadence is right
- Each tier produces its own kind of output. Per-job produces sentences, weekly produces catches, monthly produces direction, quarterly produces template changes, annual produces rate changes. If your quarterly meeting is producing sentences of cause, the per-job tier is not running and the quarterly tier is doing its work badly and late.
- Count corrections per type against closes per type. A type that got three template changes in a year on 12 closes is being tuned on noise. A type with 84 closes a year and zero changes is not being read at all.
- Check the reversal rate. If more than about one in five of your corrections gets reversed or re-corrected in the opposite direction within two cycles, your sample floor is too low, not your judgment. Raise the floor before you change anything else.
- Check attendance decay. A tier people stop showing up to is almost always a tier that fires more often than it has news. That is diagnostic information about the interval, so lengthen it rather than mandating attendance.
References
- U.S. Small Business Administration (SBA), small business financial review and performance monitoring guidance
- Trade-standard practice for project cost control and periodic variance reporting
- See related: How to Run a Monthly Estimate Accuracy Review, How to Compare Estimated Against Actual on Every Job, How to Adjust an Estimate Template From Real Data, The Estimating Feedback Loop Explained