The Job Costing SOP

Purpose

Produce a closed cost record for every completed job within a fixed window, compare it against the estimate that won the job, and route the misses into a queue that actually gets worked. The point is not bookkeeping. The point is that the next bid on that job type is built on measured hours instead of remembered ones. A shop without this procedure is not estimating, it is guessing repeatedly and calling the guesses experience.

The stake is specific: a job type that runs consistently over on labor keeps winning work, because you are the cheapest bidder on it, and every win makes the problem bigger. Volume growth on an underbid job type is the fastest way a busy shop loses ground.

Scope

Covers every job that carried a written estimate or a catalog price, from single-visit service calls through multi-day installs. Applies to labor hours, materials, subcontracted work, equipment rental, permits, disposal, and return trips.

Out of scope: warranty return visits with no attached revenue (those route to the callback log instead, where the cause is attributed), internal shop time, and jobs killed before any work was performed. Time-and-materials jobs are in scope but their labor variance is informational rather than a margin miss, because the customer carries the overrun.

Roles and responsibilities

Role Owns
Technician or lead Logs hours against the correct job and task code the same day, flags any return trip and its reason
Office or admin Closes the cost record inside the window, attaches supplier invoices and subcontractor bills to the job, computes variance per bucket
Estimator Freezes the baseline at sale, works the flagged queue weekly, owns the write-back to the template
Owner or manager Reviews the monthly rollup by job type, approves template corrections, chairs the escalation review

On a shop small enough that one person holds several of these, keep the roles separate in the procedure anyway. The estimator reviewing their own variance without a second reader is where a systematic miss survives for years.

Procedure

1. Freeze the baseline at the moment of sale

The instant the customer accepts, capture the estimate as it stood: labor hours by task, material allowance, expected trips, expected crew size, and the assumed access conditions. Store it separately from the working estimate so later edits cannot silently rewrite what you promised yourself.

This is the step shops skip, and skipping it destroys the whole procedure. If the estimate document is a live file that the estimator revises during the job, every variance computes to roughly zero and the record is worthless. You are not measuring against what you sold, you are measuring against what you last typed.

2. Capture actuals as the work happens, not at close

Hours logged the same day are close to right. Hours reconstructed on Friday for a Tuesday job carry a systematic bias toward the round number the tech thinks the job should have taken, which is exactly the bias you built this procedure to eliminate.

Three fields per time entry are enough: the job, the task code, and whether it was the original visit or a return. Anything longer than three fields gets filled in badly.

3. Close the cost record within 5 business days of completion

Pull labor hours, supplier invoices, subcontractor bills, permit fees, rental days, and disposal into the job. Five business days is a default worth committing to, tuned to your supplier billing lag: if your main supplier invoices weekly, stretch to 10 business days rather than closing a record you know is incomplete and then never reopening it.

A record closed on partial data reads as an under-cost job forever. Whatever arrives after the close lands in a general expense bucket and disappears from job-level truth.

4. Compute variance per bucket, never in total

Total variance nets a labor overrun against a material underrun and shows you a clean job that hides two real problems. Run it separately for labor hours, materials, subcontract, and other direct cost.

Express labor as hours and as a percentage of the estimated hours, in the same line. Express materials as a ratio to the allowance. Keep one unit per bucket for the life of the record.

5. Triage against a stated threshold

The default worth setting: flag a bucket when its variance exceeds 15 percent of that bucket's estimate AND exceeds 1.0 labor hour of effect. The percentage catches the systematic misses. The absolute floor stops a 40 percent overrun on a 0.4-hour task from filling the queue with noise. The 1.0-hour floor suits a shop running short service tickets; a shop whose typical job runs a day or more should derive its own floor from a noise-floor pass rather than adopting this one, and will usually land nearer 3.0 hours. Tune the floor, not the percentage, once your queue is running.

Anything past 40 percent on labor at close goes to the escalation review in step 8 rather than the ordinary weekly queue, because a miss that size is rarely an estimating error alone. Note this is a post-close figure: the in-flight threshold for stopping and re-pricing a job that is still running is a different and lower number, because in flight you are deciding whether to keep going rather than what to learn.

6. Investigate the flagged jobs weekly

Ten minutes per flagged job, with the tech who ran it in the room or on the phone. One question drives it: was the estimate wrong, or did the job run badly? Those two answers lead to opposite corrections, and getting the attribution backwards means you pad a price when you had a training problem, or retrain a tech when the bid was short.

Write one sentence of cause per flagged job. Not a category, a sentence. "Access was through a finished ceiling that the site walk did not open" is usable. "Labor overrun" is not.

7. Roll it up by job type monthly

Group closed jobs by job type and read the median actual hours against the template hours, plus the count of jobs that ran over. The median matters more than the mean here, because one disaster job will drag an average past the point of usefulness while the median tells you what a normal instance of this job type actually costs.

8. Write the correction back to the template

Correct a template when a job type has accumulated at least 5 closed jobs since its last correction and the median sits more than 10 percent off the template. Below 5 jobs you are reading noise. There is one narrow exception, and it is narrow on purpose: a miss exceeding 50 percent that repeats on at least 3 jobs is worth acting on now, but treat that as a provisional correction you keep watching rather than a settled template, because repricing a type off three or four jobs is the classic way a shop spends the next year wondering why that work stopped closing. If you are still discovering WHICH types are the problem rather than correcting a type you already run, hold out for 8.

Quarterly is a reasonable review rhythm for a shop under about 500 jobs a year. Faster than monthly and you will chase noise.

9. Escalate the outliers separately

Jobs past 40 percent on labor, jobs with 2 or more return trips, and any job where actual crew size differed from the assumption get a named review with the owner. These are usually scope or dispatch failures rather than estimating failures, and they get fixed in a different place.

Worked example: one flagged job through the whole procedure

A two-tech replacement job type. The frozen baseline reads 12.0 labor hours (2 techs at 6.0 hours each), a material allowance carried as 1.00x, and 1 truck trip.

The closed record reads 15.5 labor hours, materials at 1.18x the allowance, and 2 truck trips.

Labor. 15.5 actual minus 12.0 estimated is 3.5 hours over, which against the 12.0-hour estimate is 29 percent over. That clears both gates in step 5: 29 percent is above the 15 percent threshold and 3.5 hours is above the 1.0-hour floor. Flagged. It does not clear the 40 percent escalation bar, so it stays in the ordinary weekly queue.

Materials. 1.18x against a 1.00x allowance is 18 percent over, above the 15 percent threshold. Flagged.

Trips. 2 against an assumed 1. Logged for the return-trip counter, which is read as a count, not a percentage.

The weekly review with the lead produces one sentence: the second trip was to collect a fitting the material list never carried, and the crew waited on site for it. So the material overrun and part of the labor overrun are the same defect, a missing line on the material list, not two independent misses. That single observation changes the correction from "raise the labor allowance" to "fix the material list, then re-measure."

Now the rollup. Across the quarter this job type closed 9 times. The median actual is 14.0 labor hours against the 12.0-hour template, so the median runs 2.0 hours over, which is 17 percent over the 12.0-hour template. 7 of the 9 jobs ran over, about 78 percent of them, so the distribution is not one bad job dragging a clean set. It is the template.

Step 8 applies: 9 closed jobs clears the 5-job gate, and 17 percent clears the 10 percent bar. Correct by about two-thirds of the measured gap, approaching the target from below: the gap is 17 percent, two-thirds of that is roughly 11 percent, so the template moves from 12.0 to 13.3 hours, a 1.11x correction. Two-thirds rather than the full gap, because correcting at full gain overshoots and the number then oscillates for a year without settling. It also moves toward the median rather than the worst case, because bidding at the worst case prices you out of the job type entirely.

Note what did not happen. The 15.5-hour job that started this was worse than the median. If the estimator had corrected the template off that one job, the template would have gone to 15.5 hours, a 1.29x correction on a sample of one, and the shop would have lost bids it could have run profitably. The single job earns an investigation. Only the rollup earns a template change.

What changes the answer

Time-and-materials work. Labor variance is informational. The customer absorbs the overrun, so the bid was not wrong in a margin sense. It still tells you your quoted range was misleading, which costs trust and future work, but it does not trigger a price correction.

A job type you run fewer than 5 times a year. The sample will never arrive. Correct on the sentence-level causes instead: three jobs, three different access problems, means your site walk is short, not your hour count.

A crew mix change. If the 15.5 hours came from a two-apprentice crew where the template assumes a lead plus a helper, the hours are not comparable and the job should not enter the rollup at all. Tag crew composition on the record or the rollup quietly drifts toward whoever you happened to send.

Seasonal load. Jobs run in peak season carry more interruption and more rushed material staging. If your rollup blends peak and shoulder months, a template corrected in the busy season will read as padded in the slow one. Split the rollup by season once you have enough volume to.

How to verify the procedure is working

Three checks, run quarterly:

  • Close rate. What share of completed jobs got a cost record inside the window? Below about 90 percent, the rollup is being computed on the jobs that happened to be easy to close, which biases every median downward.
  • Flag resolution. What share of flagged jobs carry a cause sentence? A queue full of unresolved flags means the procedure is producing data nobody uses.
  • Drift direction. Are template corrections still mostly upward after a year? Persistent one-directional correction means something structural is missing from the estimate, most often unbilled travel, staging, or closeout time, rather than the task hours themselves.

References

  • U.S. Small Business Administration (SBA), job costing and cost control for small contractors
  • Generally accepted construction accounting practice, committed cost and cost-to-complete reporting
  • See related: Estimate vs Actual - Cost Variance Review; Job Costing - Did We Make Money on This Job?; How to Build a Labor Hour Database From Your Own Jobs