The Job Costing Data a Small Shop Actually Needs

Why this matters

Most shops fail at job costing in one of two directions, and both feel responsible while they are happening. One shop tracks a quoted total and a final invoice, discovers its margin slipped, and has no way to find out why. The other builds a thirty-field capture form, gets four weeks of beautiful data, then gets busy and gets nothing at all for the next eight months.

The useful dataset is smaller than the second shop thinks and much larger than the first. This card is the field list: what to capture, at what grain, who is on the hook for each, and specifically what you can leave out without losing the ability to correct a bid. It is about the shape of the record. Which cost categories to split apart inside it is a separate question with its own card.

The test a field has to pass

A field earns a slot only if you can name the decision it would change. Not the report it appears on, the decision.

Run it out loud. "If this field said 3 instead of 1, what would I do differently?" If the honest answer is "I would look into it," the field is not doing work - looking into it is what the job record is for. If the answer is "I would add a second visit line to that template" or "I would stop quoting that job type at a flat rate," the field earns its place.

This test kills about half of what people instinctively want to capture, and it kills the right half. Every field you drop is a field that will still be accurate in month nine.

The core record

Eleven fields. A shop that captures these can answer every question the rest of this category asks. A shop missing any one of them has a specific, predictable blind spot, named in the last column.

Field Who enters it When Blind spot if missing
Job identifier System At sale Nothing ties together
Job type code Estimator At sale Cannot group; every rollup is a shop-wide average that hides everything
Quoted labor hours Estimator Frozen at sale No baseline; variance is undefined
Quoted material and subcontract Estimator Frozen at sale Labor misses and material misses look identical
Actual labor hours, by phase Field At the boundary, same day Know the job ran long, never which part of it
Actual material issued Field At issue, off the truck Material creep invisible until year-end inventory
Actual subcontract Office On sub invoice The single largest line on some jobs floats free
Visit count System or field On each arrival Return-trip cost hides inside labor and looks like slow work
Approved-change flag and revised baseline Office At approval Successful upsells report as overruns
Rework hours flag Field At the entry Execution error and estimating error become one number
Record reliability flag System At entry Reconstructed guesses get averaged in with measured hours

The last one gets skipped most often and costs the most quietly. An hour entry typed from memory two days later is fine for billing and useless for tightening a template. If you cannot separate them, your labor database is part measurement and part folklore, blended at an unknown ratio.

Grain is a separate decision from the field list

Deciding to capture labor hours does not decide how finely. Three grains, and the right one depends on your job length:

  • Job level. One total. Adequate only for jobs under about half a day, where there is no meaningful internal structure to find.
  • Phase level. Five to seven buckets per job. This is the right default for almost every shop. It is where the information is - a job that runs 20% over on total hours is a shrug, and a job that runs 20% over with all of the overage in commissioning is a specific correction.
  • Task level. Every discrete operation. Almost always a mistake below a certain scale. The capture cost is enormous, the coding noise swamps the signal, and small shops do not run enough repetitions of any single task to get a stable number out of it.

Go one grain finer than your instinct on labor, and one grain coarser than your instinct on material. Labor is where the variance lives. Material tends to be either right or catastrophically wrong, and catastrophically wrong shows up at the job level fine.

The three fields shops add and never use

A subjective difficulty rating. A 1-to-5 "how hard was this" from the tech. It seems like it would explain variance. It never does, because the rating is assigned after the fact and is really a memory of how the day felt. It correlates with the outcome you are trying to predict, so it looks predictive in review and predicts nothing in advance.

Percent complete, captured continuously. Genuinely useful on long construction-style projects with progress billing. On a job measured in days it is a number a tech invents to answer a prompt, and it is nearly always somewhere between 60% and 80% regardless of reality.

A free-text cause field with no code behind it. Notes are valuable and you should keep them. But a free-text field cannot be counted, so a shop with 400 beautifully written notes still cannot answer "how often is it the same cause." Pair every note with a short coded reason list, four or five entries, or accept that the notes are for reading one at a time and never for rollup.

What you should never store

Store inputs. Derive everything else at read time.

Do not store variance percentage, margin percentage, effective hourly rate, or any other computed figure as a field on the job record. Two reasons, and the second is the one that bites. First, a stored derivation goes stale the moment an input is corrected, and job records get corrected constantly - a late sub invoice, a reconciled change order, a rework hour recoded. Second, the formula itself changes. The day you decide travel time belongs in job labor rather than overhead, every stored variance figure in your history becomes wrong, and there is no way to tell which ones were computed under which rule.

Derived numbers computed on read are always current and always consistent with whatever definition you hold today. Stored ones are a snapshot of a definition you may no longer remember making.

A worked comparison: two shops, forty identical jobs

Two shops each complete 40 jobs of the same type over a quarter. Both quoted 8.0 labor hours per job, so 320 quoted hours in total. Both finished at 379 actual labor hours. That is 59 hours over the 320 quoted, an 18.4% overrun against the quoted 320 hours. Identical outcome, identical aggregate.

Shop A captures six fields: job id, type, quoted total, actual total, invoice, close date. It reads an 18.4% overrun on this job type, concludes the template is light, and applies a 1.18x correction to the labor line.

Shop B captures the eleven, including visit count. It cuts the same 40 jobs by that field:

  • 31 jobs closed in a single visit. Those were quoted 248 hours in total and consumed 236, so single-visit work came in 4.8% under its own 248 quoted hours.
  • 9 jobs required a second visit. Those were quoted 72 hours in total and consumed 143, so multi-visit work ran 98.6% over its own 72 quoted hours, very nearly double.

The two subsets add back to the whole: 248 plus 72 is the 320 quoted, and 236 plus 143 is the 379 actual. Nothing has been recomputed, only cut.

Shop B's conclusion is the opposite of Shop A's. The template is close to right - it is 4.8% conservative on the 31 jobs, 77.5% of the 40, that go the way they are supposed to. The entire aggregate overrun comes from the 9 jobs, 22.5% of the 40, that turned into two visits. The correction is not a labor multiplier. It is a rule about what triggers a second visit, a line in the estimate for it, and a scoping question at the sale that catches those nine before they are quoted as one-visit work.

Now watch what Shop A's correction actually does. Single-visit jobs consumed 236 hours across 31 jobs, an average of 7.61 hours each. Shop A will now quote 8.0 times 1.18, which is 9.44 hours, for work that averages 7.61. That is 24% above what the work consumes, applied to more than three quarters of its volume, on a job type it was already winning. It will lose bids it should win, keep the nine jobs that are actually bleeding, and read the resulting drop in win rate as a competitive problem.

One field, visit count, is the difference between those two futures. That field costs nothing to capture and passes the decision test cleanly: if it says 2 instead of 1, the job goes in a different bucket and a different correction follows.

What changes the answer

Job length. Under about two hours, phase-level labor coding stops paying for itself - there is one phase. Drop to job level on labor and spend the saved capture effort on visit count and material issue instead.

Whether you subcontract. A shop that never subs can drop the subcontract field and lose nothing. A shop where subs carry a third of a typical job needs it at the same grain as its own labor, including which portion of scope the sub covered, or the largest line on the job is a single opaque number.

Flat rate versus time and material. Under flat rate the quoted-hours field is internal rather than customer-facing, and shops often stop maintaining it because nobody outside sees it. Keep it. It is the only baseline you have, and without it a flat-rate shop can measure cost but never accuracy.

Volume. Below roughly 15 to 20 completed jobs per type per quarter, most cuts of this data will not reach a count you can believe. That is not a reason to capture less; it is a reason to roll up over a longer window and to resist correcting a template off three jobs.

How to verify your dataset is working

The direct test is whether it can answer questions, not whether it is complete. Take one real question - "which job type is bleeding, and is it bias or a tail" - and try to answer it from the record alone. If you need to phone a tech to interpret any row, the field that would have told you is missing.

Three specific checks:

  • Coverage per field, not overall. A dataset that is 95% complete on labor and 40% complete on material is not 85% complete, it is unusable for any question involving material.
  • Try one cut you have never made. Group by a field you have been capturing but never used. Either it splits something interesting, in which case it earned its place, or it does not, in which case stop capturing it.
  • Check the reliability flag distribution. If more than a quarter of your labor entries are reconstructed, treat the whole labor dataset as directional until capture improves. It will still tell you sign. It will not tell you size.

References

  • U.S. Small Business Administration (SBA), financial recordkeeping for small business
  • IRS, recordkeeping requirements for business records supporting income and expense
  • Trade-standard practice for job-cost coding structures
  • See related: The Cost Categories Worth Separating, How to Log Actuals in the Field Without Friction, The Estimate Accuracy Metrics Worth Tracking