The Quarterly Costing Review SOP
Purpose
To convert a quarter of job-costing data into a small number of specific, deployed corrections to the estimating templates, and to verify that the previous quarter's corrections actually worked.
This is the only meeting in the costing cycle whose output is a change to how you price. Per-job costing tells you what happened on one job. Weekly triage catches the outliers. The monthly financial review reads the profit and loss. None of those edit a template, and a shop can run all three for a year while its estimates stay exactly as wrong as they were.
Scope
In scope: completed jobs with closed cost records in the quarter just ended, grouped by job type; the estimating templates and price-book entries those job types draw from; verification of corrections approved in the previous quarter.
Out of scope: individual job disputes and change orders (handled at close, not here), technician performance conversations (a separate track, see step 8), overhead and rate-setting (annual, and driven by the profit and loss rather than by job variance), and any job type with fewer than 8 completed instances in the trailing two quarters, which is below the sample size where a correction is defensible.
Run it within three weeks of quarter close, timeboxed to 90 minutes. Later than three weeks and the field memory behind the flagged jobs is gone. Longer than 90 minutes and it becomes a discussion rather than a decision meeting.
Roles and responsibilities
| Role | Responsibility |
|---|---|
| Owner or general manager | Chairs. Approves or rejects each proposed template change. Owns the decision record. |
| Office or admin lead | Produces the pack one week ahead. Closes the coverage gap. Deploys approved edits to the live price book and confirms deployment. |
| Lead estimator | Classifies each flagged job type as bias or spread. Drafts the specific edit, with the number. |
| Senior field lead | Supplies cause on flagged instances. Confirms a proposed hours figure is achievable in the field rather than aspirational. |
The field lead's veto matters more than it looks. An hours figure agreed in an office with nobody who runs the work in the room produces a template the field ignores, and then the actuals diverge from the estimate for a reason that never appears in the data.
Procedure
1. Close the coverage gap (two weeks before the meeting)
Pull every job completed in the quarter and count how many have a closed cost record with labor, materials and any sub cost captured. Force-close the open ones with the best reconstruction available, flagged as reconstructed.
Do this first because the jobs still open are not a random sample - the ones that went badly are the ones whose paperwork never got finished. Reviewing an 80%-coverage quarter produces an optimistic picture of exactly the job types you most need to fix. Target 95% coverage before the pack is built, and if you cannot get there, note the coverage figure at the top of the pack so every number underneath is read with it in mind.
2. Build the pack (one week before)
One row per job type with at least 8 instances in the trailing two quarters. Columns:
- Instances this quarter
- Estimated hours per instance
- Signed variance (bias), as a percentage of the estimate
- Absolute variance (spread), as a percentage of the estimate
- Hit rate within plus or minus 15%
- Bias exposure: instances times estimated hours times signed variance, in hours per quarter
- Spread exposure: instances times estimated hours times absolute variance, in hours per quarter
Plus a second short table: every correction approved last quarter, what it changed, and what the metric did since.
The two exposure columns are the reason the pack is worth building rather than reading variance percentages off a report. A percentage tells you how wrong a job type is. Exposure tells you how much it costs you, and those rank differently.
3. Verify last quarter's corrections before discussing anything new
Open with the second table. For each correction: did it deploy, and did the metric move in the predicted direction and roughly the predicted size?
Three outcomes and three responses:
- Deployed and moved as predicted. Close it. Say so out loud; this is the evidence that makes the next quarter's corrections credible to the people who have to work to them.
- Deployed and did not move. The diagnosis was wrong, not the number. Re-classify the job type from scratch in step 5 rather than applying a larger multiplier in the same direction.
- Did not deploy. The most common outcome in the first year, and the most damaging, because it looks identical in the data to "deployed and did not move." Find out where it stopped - price book not republished, estimator working from a saved copy, template edited in one place and not the other - and fix the deployment path before approving anything new this quarter.
Putting this first, before the new data, is deliberate. A meeting that starts with fresh problems never gets back to whether the last set was solved.
4. Rank by exposure, not by variance
Sort the pack by bias exposure in hours per quarter, highest first. Separately note the top three by spread exposure.
A job type at 19% bias on 12 small instances and a job type at 11% bias on 4 large instances can carry almost the same exposure. Ranking by the percentage puts your attention on the first and leaves the second alone for another year.
5. Classify each of the top three: bias or spread
For each, compare signed variance against absolute variance:
- Signed close in size to absolute means nearly every instance missed in the same direction. This is a structural underbid or overbid. It is correctable with a multiplier on the template.
- Signed much smaller than absolute means the average is right and individual jobs are unpredictable. A multiplier will not help and will cost you win rate on the easy instances. Look for a hidden sub-type inside the job type, or an intake question that is not being asked.
Where the classification is spread, check the tail before concluding: what share of the type's total absolute miss comes from the worst tenth of instances, and do those instances share a site trait. A high tail share with a shared trait is a type-splitting decision, which is a better correction than any number change.
6. Decide exactly one correction per job type
One. Not a labor multiplier and a materials change and a new intake question on the same type in the same quarter, because next quarter you will not be able to tell which one worked.
Write it as a specific edit someone could make without asking a follow-up question. "Raise labor on the changeout template from 6.0 to 6.7 hours" is an edit. "Tighten up the changeout estimate" is not.
Record alongside it the expected effect: which metric should move, in which direction, by roughly how much, by when. This is what step 3 checks next quarter, and a correction without one cannot be verified.
7. Assign an owner and a deploy date
Each approved edit gets a named person and a date, and deployment is confirmed in writing to the chair. The confirmation is not bureaucracy - step 3 exists because undeployed corrections are common enough to need their own detection.
8. Route the outliers and the people issues out of this meeting
Two categories leave the room rather than being resolved in it:
- Single blowout jobs, meaning one instance at several times the size of anything else in its type. Investigate as an event, in its own session, with the people who were there. Do not let it into the averages, where it will misprice the type for everyone.
- Performance gaps between techs on the same job type. Real, and not a template problem. Send it to the coaching track. Letting it become a pricing discussion here is how a training gap gets permanently written into the customer's invoice.
9. Publish the pack and set the next date
The pack goes where the estimator and the field leads can read it, not into a folder. A field lead who sees the changeout template moved because their own jobs said it should moves from complying with the number to defending it.
A worked quarter
A shop finishes the quarter at 44 of 51 jobs with closed cost records, so coverage is about 86%, which is flagged at the top of the pack. Three job types clear the 8-instance gate:
| Job type | Instances | Est hours each | Bias | Spread | Bias exposure (hrs/qtr) | Spread exposure (hrs/qtr) |
|---|---|---|---|---|---|---|
| A - service call | 30 | 2.0 | +4% | 22% | 2.4 | 13.2 |
| B - changeout | 12 | 6.0 | +19% | 20% | 13.7 | 14.4 |
| D - small project | 4 | 24.0 | +11% | 14% | 10.6 | 13.4 |
Ranked by bias percentage: B at 19%, D at 11%, A at 4%. That ranking says spend the quarter on B and maybe glance at D.
Ranked by bias exposure: B at 13.7 hours per quarter, D at 10.6, A at 2.4. Total bias exposure across the three types is 26.7 hours per quarter. D is now nearly as large as B on a bias percentage a little over half of B's (11% against 19%), because its instances are four times the size. On the percentage ranking D would have waited another year while costing about 10.6 hours a quarter.
Classification. Type B's bias of 19% against spread of 20% means essentially every instance ran over by a similar amount - structural. The correction is a multiplier of about 1.19x on the labor line. Type D's bias of 11% against spread of 14% is also mostly one-directional, though less cleanly, so it takes a smaller multiplier of about 1.11x and a note to re-check next quarter with a wider sample, since 4 instances is at the thin end even with the trailing-two-quarter gate satisfied.
Type A is the trap. Its bias exposure of 2.4 hours per quarter is the smallest of the three and it would drop off any bias-ranked agenda. But its spread exposure of 13.2 hours per quarter is essentially tied with B's 14.4 and D's 13.4. The average is right and no individual instance is predictable. A multiplier here would be actively harmful. The correct correction is to check the tail, find the shared trait among the worst instances, and split the type - which is a scope decision, not a price decision.
What gets decided. Three corrections, one per type: a 1.19x labor multiplier on B, a 1.11x labor multiplier on D, and a type split on A driven by the tail analysis. Each with an owner, a deploy date, and an expected effect written down: B's bias inside plus or minus 5% next quarter, D's the same with wider tolerance given the sample, A's bias unchanged while its spread falls toward 15%.
What the next quarter's step 3 will check. Whether all three deployed, and whether A's spread moved. If A's spread is unchanged at 22% after the split, the trait chosen was not the driver and the type needs re-examining rather than repricing.
How to verify the SOP is working
The honest test is not whether the meeting happened. It is whether the numbers moved.
- Corrections deployed within their date, as a share of corrections approved. Below 100% in the first two quarters is normal; below 100% in the fourth is a process failure, not a busy quarter.
- Bias inside plus or minus 5% on your top three types by exposure, trending toward it quarter over quarter rather than oscillating. Oscillation - overcorrecting one quarter and back the next - means the multipliers being applied are too large for the sample size behind them.
- Coverage at or above 95% by the second quarter of running this. Coverage is the one metric that responds immediately to attention and degrades immediately without it.
- Total exposure in hours per quarter trending down across the tracked types. This is the number that says the whole discipline is paying, and it is the one to put in front of the field.
If bias is inside band on every tracked type for two consecutive quarters, the correct response is to widen the gate and bring in the job types that were below 8 instances, not to keep polishing the ones already under control.
References
- See related:
The Job Costing SOPfor the per-job capture and weekly triage this review consumes. - See related:
The Estimate Accuracy Metrics Worth Trackingfor definitions of bias, spread, hit rate and tail share. - See related:
Monthly Financial Reviewfor the profit-and-loss cadence that sits above this. - See related:
How to Separate Estimating Error From Execution Errorfor the classification step 5 and step 8 depend on. - SBA guidance on small-business management reporting and review cadence.