The Post-Job Cost Review SOP

Purpose

To convert a job that missed its cost estimate into a specific, owned, verified correction to how the shop estimates that kind of work. The review exists to change a number in a template or a step in a process. A review that ends in a conversation and no recorded change has failed, no matter how honest the conversation was.

This procedure covers the governance around the comparison: which jobs get reviewed, who is in the room, what gets decided, and how the shop confirms months later that the correction actually took. The arithmetic of building the comparison itself is a separate procedure, referenced at the end.

Scope

Applies to every completed job that trips a review trigger, in every trade the shop runs, on both fixed-price and flat-rate work. Time and materials jobs are in scope for estimating accuracy but out of scope for margin correction, because the customer absorbed the overrun.

Review triggers. A closed job is flagged for formal review when any one of these is true. These are defaults to tune to your job sizes, not commandments:

  • Labor variance is at or beyond 15% over estimate AND at least 3.0 hours over in raw terms. Both tests, not either, so a small job does not generate paperwork over one hour.
  • Earned gross margin came in 10 or more percentage points below quoted margin.
  • Material cost landed at 1.25x the estimate or higher.
  • Any job carrying work performed but never billed, regardless of size. Unbilled work is a process failure and it always gets reviewed.
  • Any job the crew flags themselves, regardless of the numbers.

Jobs that come in favorable by the same magnitudes are reviewed on the same trigger. A job that ran 30% under estimate is either an overpriced template that is costing you close rates or a data-capture failure, and both are worth five minutes.

Out of scope: performance management. This review looks at the estimate and the process. Individual performance is a separate conversation with a separate owner, and mixing them is the fastest way to make the crew stop logging honest hours.

Roles and responsibilities

Role Who typically fills it Responsibility in this procedure
Data preparer Office or admin Pulls the frozen estimate, the actuals, and the change orders into one sheet before the meeting. Does not interpret.
Job owner Lead tech or foreman on that job Supplies what actually happened on site, in sequence. Owns the cause statement.
Estimator Whoever priced it Explains the assumptions behind the frozen estimate. Not on trial.
Chair Owner, GM, or ops manager Runs the clock, forces a decision, assigns the action.
Template owner Usually the estimator or a service manager Makes the agreed change to the template, price book, or checklist, and dates it.

In a two-person shop one person fills three of these. The roles still matter because they are separate hats: the sequence of the meeting depends on the site account coming before the estimating assumptions, and on somebody being responsible for the change actually getting made.

Procedure

1. Flag at close-out, within one business day. The trigger check runs automatically or as part of the close-out checklist. Do not wait for month-end. Detail decays fast, and a crew that is asked in week four why a job ran long will honestly reconstruct a plausible story that is not what happened.

2. Schedule the review inside 5 business days of close. Fifteen minutes, standing agenda, same slot each week for whatever flagged that week. A weekly batch of two or three reviews beats a monthly batch of twelve, because by the twelfth nobody is thinking.

3. Prepare the one-page sheet in advance. The data preparer assembles: frozen estimate by bucket, actuals by bucket, change orders with their own quoted and consumed figures, the three variance numbers, who ran the job, and any photos or notes. Nobody does arithmetic in the meeting. If the sheet is not ready, the meeting moves, it does not proceed on memory.

4. Open with the site account, not the numbers. The job owner narrates what happened in order: arrival, what was found, what changed, where time went. Two to three minutes. Numbers first invites the crew to defend hours; sequence first surfaces the actual cause. This ordering is the single most load-bearing choice in the whole procedure.

5. Split the gap into named pieces. The chair drives the room to allocate the total variance into specific chunks with a cause on each, and the chunks must sum to the total. "It just ran long" is not an allocation. If 2.0 hours cannot be explained, it gets logged as unexplained rather than attributed to something convenient.

6. Classify each piece as systematic or one-off. Systematic means it will recur on the next job of this type with a different crew on a different site. One-off means it will not. This classification decides everything downstream, and it is the step people get wrong: a genuinely one-off cause that gets baked into a template overprices every future job, and a systematic cause dismissed as bad luck repeats until it shows up in the annual numbers.

7. Decide one action per systematic piece, with an owner and a date. The action is a change to an artifact: hours on a job template, a line on the pre-job checklist, a price-book cost figure, a scope question added to the intake script, a note in the conditions clause. If the answer is "no change," record that as the decision, with the reasoning.

8. Record the decision on the job. The job record carries the cause statement and the action taken. This is what makes the multi-job pattern analysis possible later. A number without a cause line cannot be aggregated into anything useful.

9. Make the change and date it. Template owner updates the artifact within the agreed window and stamps the date. The date is what makes step 10 possible: without it, you cannot separate jobs run before the fix from jobs run after.

10. Verify on the next five jobs of that type. At the next monthly review, pull the jobs of that type that started after the change date and read the median labor variance. If the median has moved inside your noise band, the correction took. If it has not, the cause was misdiagnosed and the job goes back on the agenda. Five is a working minimum; on job types you run twice a year, verify on what you get and accept a slower loop.

A worked review, start to finish

A commercial retrofit closes. Frozen estimate: 24.0 labor hours. Actual logged: 31.5 hours. That is 7.5 hours over a 24.0 hour estimate, 31.3% over. It trips both labor tests (over 15% and over 3.0 hours), so it is flagged.

The site account, given before anyone looks at the variance, produces this sequence: the crew arrived to a locked equipment room, the site contact on the work order had left the company, and it took most of a morning to find someone with a key. Then the takeoff turned out to be short on hanger material, which cost a run. Then a section was mounted at the wrong height against the approved layout and had to come down.

Allocated, with the pieces summing to the total:

Piece Hours Systematic or one-off
Waiting on site access 4.0 Systematic (this building type, this customer's process)
Undersized hanger takeoff 2.0 Systematic (the takeoff method)
Rework on mounting height 1.5 One-off (drawing was clear, crew misread it)
Total 7.5 matches the 7.5 hour gap

Three pieces, three different answers:

The 4.0 access hours are 16.7% of the 24.0 hour estimate on their own, and they are the biggest single chunk. They are also not an estimating error. Padding the template by 4.0 hours would mean quoting every future job in that building as if the door will be locked, which loses bids. The action is a pre-job checklist line: confirm a named on-site contact with key access within 48 hours of the scheduled start, and hold the crew if it is unconfirmed.

The 2.0 hanger hours are 8.3% of the 24.0 hour estimate and they are a real template miss. That job type's standard hours go up by 2.0, roughly an 8% correction, and the takeoff checklist gets a hanger-count line. Note the discipline here: the correction is 8.3%, not the 31.3% the headline variance showed. Correcting on the headline number would have overpriced the template by about four times the actual defect.

The 1.5 rework hours get no template change. One-off, drawing was correct, and it is a coaching item for the job owner outside this meeting.

Verification comes at the next monthly review: of the retrofits started after the template change date, the median labor variance should fall inside the shop's noise band. If it still runs 15% over, the 2.0 hour correction was too small or the real cause was something the room did not name.

What changes the procedure

  • Very high job volume. A shop closing 200 jobs a month cannot review every flagged job individually. Raise the triggers until the flagged set is three to five jobs a week, and handle the rest through the monthly pattern review. The triggers are a filter, and the filter should be tuned to your capacity to act, not to some ideal of thoroughness.
  • Long multi-phase jobs. Do not wait for close. Review at phase boundaries against phase budgets, because a correction found in week two can still save the job, and one found at close only helps the next customer.
  • A brand new job type. The first three of anything will miss. Review them for information, not correction, and do not adjust the template until you have enough of them to see whether the miss has a direction.
  • A crew member is the common factor. If the same name is attached to most flagged jobs, that is a real finding, and it does not get handled in this meeting. Take it out of the room and to the person's manager, or you will convert an estimating tool into a discipline tool and lose the honest data that makes it work.

Failure modes specific to running this as a standing meeting

These are process failures the procedure above does not prevent on its own:

  • The log gets turned off quietly. If a crew learns that logged hours produce uncomfortable meetings, hours start arriving rounded to the estimate. The tell is variance collapsing toward zero across the board. The fix is upstream: never open with numbers, never review individuals here, and visibly act on process causes more often than on people.
  • Actions accumulate unowned. A backlog of undated, unowned actions makes the meeting theater. Cap open actions at what the template owner can actually close in a week and stop generating new ones until the backlog clears.
  • A closed gap reopens unnoticed. Templates get edited for other reasons, someone restores an old price book, a checklist item gets dropped in a cleanup. Keep a dated change log of estimating corrections so a reopened gap is visible as a reverted line rather than as a mysterious new pattern nine months later.
  • The review becomes an archive. Reviewing a job from six weeks ago produces reconstructed memory, not evidence. If the queue slips past two weeks, drop the oldest ones unreviewed rather than pretending. Skipping them is honest, and a stale review that assigns a wrong cause is worse than no review at all.

References

  • U.S. Small Business Administration: job costing and cost control guidance for small contractors.
  • Standard construction practice for post-project cost review and corrective action tracking.
  • See related: How to Compare Estimated Against Actual on Every Job.
  • See related: How to Run a Monthly Estimate Accuracy Review.
  • See related: Why Estimates Miss and Which Misses Matter.