How to Read Your Own Job Costing Data

Why this matters

Most shops that start tracking actuals stall at the same place. The data exists, the report runs, and nobody can say what it means. So the owner scrolls to the worst-looking percentage, gets angry about one job, and changes nothing that matters. A costing dataset is not a scoreboard of individual jobs. It is a measurement of your estimating template, and reading it correctly is a specific skill with a specific order of operations. Done in the wrong order it will point you at the loudest job instead of the expensive one, and you will spend a quarter fixing a job type that costs you almost nothing.

Step 1: Prove the data before you read a single number

Every conclusion below rests on records that are complete. Before any analysis, count what is unusable and look at what the unusable rows have in common.

Pull the period's completed jobs and flag any that have zero logged labor hours, are still open in the system, carry a part booked at catalog cost rather than what you actually paid, or have time entries dated after the job was signed off. Set those aside. If you skip this, missing hours read as jobs that came in under estimate, and your data will tell you the template is generous when it is only under-reported.

The set-aside pile is itself a finding. If the unusable rows cluster on one tech, one crew, or one week, that is a logging problem and it gets its own fix. Do not average it away with everything else.

Step 2: Sort by exposure, not by the biggest percentage

Exposure means estimated hours committed to a job type over the period: estimated hours per job times the number of those jobs you ran. It is the size of the bet. A job type you run twice a quarter at a 60% overrun is a smaller problem than a job type you run every week at 12% over, and the percentage column will never tell you that.

Rank job types by total estimated hours first, then read variance within that ranking. The top two or three lines usually carry most of your total miss, and everything below them is a distraction until those are fixed.

Step 3: Read the median, and read the spread separately

The average variance for a job type is the number most likely to mislead you, because one catastrophic job drags it. Use the median, which is the middle value when you line the jobs up in order, so a single disaster moves it by one position instead of by its full size.

Then look at the spread on its own. The simplest usable measure is the middle half of the jobs: throw out the best quarter and the worst quarter and state the range that remains. A type whose middle half runs from 4% under to 15% over behaves very differently from one whose middle half runs from 22% under to 30% over, even if both have a median near zero.

Step 4: Separate bias from spread, because the fixes are opposite

Bias is the center being in the wrong place. The median sits well off zero and the spread is tight, meaning nearly every job of that type misses in the same direction by a similar amount. This is a template problem and it has a clean fix: move the number.

Spread is the center being right and the individual jobs being all over the place. The median sits near zero and the middle half is wide. Moving the template number here does nothing, because you would be shifting a distribution that is already centered. This is a scoping problem: the job type is really two or three different jobs wearing one name, and the estimate cannot know which one it is looking at until somebody asks a question that separates them.

Getting this backwards is the most common analysis error in the whole discipline. A shop sees a wide, centered type, multiplies the template by 1.2, and now half the jobs are badly overpriced while the other half still lose hours.

Step 5: Set a minimum sample before you act on a job type

Commit to a default of eight completed jobs of a type in the period before you change that type's template, and treat anything with fewer as watch-only. Eight is not statistically magic, it is the point where one weird job stops being able to define the median for a small shop. Tune it up to twelve if your work is highly variable, or down to five for a type that is nearly identical every time, such as a fixed-scope maintenance visit.

Below the minimum, still write the variance down. Three quarters of watch-only data becomes actionable data.

Step 6: Read the labor line and the parts line separately

Total variance hides offsetting errors. A job type that comes in 3% over in total can be 30% over on labor and 25% under on parts, which means two broken numbers, not one accurate estimate. Always split the read: estimated labor hours against actual labor hours, and estimated parts cost against actual parts cost, as two independent columns.

Labor variance in hours is the one that generalizes across your whole shop. Parts variance is usually a mix of a pricing problem, which affects every job with that part, and a quantity problem, which is scoping.

Step 7: Only now cut by tech, season, or customer type

These cuts are real and they matter, but they are step seven for a reason. If you cut by tech first you will find a person who looks slow and start a personnel conversation, when the actual cause is that this tech gets assigned the messy version of the job type. Establish the type-level picture first, then ask whether one tech's jobs of that same type sit consistently outside the middle half. If they do, you have a training or assignment finding. If they do not, the type was always the problem.

A full read, start to finish

A shop closes a quarter with 68 completed jobs. Step 1 finds six unusable: five with no time logged at all and one still open. Five of those six belong to one apprentice, which becomes a separate logging fix rather than part of this analysis. That leaves 62 usable jobs across four job types.

Estimated labor hours committed, which is the exposure ranking from step 2:

Job type Jobs Est. hours each Est. hours total Share of est. hours
Service diagnostic 19 1.5 28.5 14%
Small repair 28 2.0 56.0 27%
Equipment replacement 9 12.0 108.0 51%
Miscellaneous 6 3.0 18.0 9%

Total estimated labor for the quarter is 210.5 hours. Replacement is only 9 of the 62 usable jobs, about 15% of the job count, but 108 of 210.5 estimated hours, about 51% of the quarter's estimated labor. That single line is where the reading starts.

Now the variance read from steps 3 and 4:

  • Small repair, 28 jobs. Median 6% over on labor hours, middle half running from 4% under to 15% over. Tight and slightly biased high.
  • Service diagnostic, 19 jobs. Median 2% over, middle half running from 22% under to 30% over. Centered and very wide. The mean for this type is 9% over, pulled by one job that ran 3.2 times its estimate, which is exactly why the median is the number being used.
  • Equipment replacement, 9 jobs. Median 34% over, middle half running from 26% over to 41% over. Tight and badly biased high.
  • Miscellaneous, 6 jobs. Below the eight-job minimum from step 5, so it is recorded and not acted on.

Applying each median to that type's estimated hours gives the quarter's overrun in hours: small repair 56.0 x 0.06 = 3.4 hours, service diagnostic 28.5 x 0.02 = 0.6 hours, equipment replacement 108.0 x 0.34 = 36.7 hours. Those three total about 40.7 hours of median overrun, and equipment replacement is 36.7 of that 40.7, roughly 90% of the readable overrun in the quarter.

The calls that follow:

Equipment replacement is bias, not spread, because its middle half sits entirely above zero. The fix is to move the number. Multiply the template's labor line by 1.3 rather than the full 1.34, because you want to close most of the gap and re-measure rather than overshoot on nine jobs of evidence. That correction alone would have removed about 32 of the 36.7 overrun hours on this type, if the next quarter's jobs behave like this one.

Service diagnostic is spread, not bias, and the template stays exactly where it is. The action is a scoping question added to the intake script so the two hidden versions of this job separate before anyone quotes them. Cross-referenced work on splitting a type sits in the estimating-template articles listed below.

Small repair, at 6% over on a tight spread, is inside most shops' noise. Log it, do not touch it, and see whether it holds for a second quarter before spending a change on it.

Miscellaneous stays on the watch list. Six jobs is below the minimum, and its variance will be readable in another quarter or two.

How to verify you read it right

Two checks, and both are worth the ten minutes.

Rebuild one job by hand. Take a single job from the type you are about to change and reconstruct its actual hours and parts cost from the source records, not the report: time entries, the supplier invoice, the parts pulled. If your hand number and the report disagree by more than a rounding difference, the report has a mapping problem and every conclusion above it is suspect.

Predict, then check. Write down what the correction should do before you deploy it. In the example: replacement labor variance should move from a median of 34% over toward roughly 3% over next quarter, because a 1.3x multiplier on a type actually running 1.34x leaves about 3% uncovered. Next quarter, compare. If the median lands near the prediction, your read was sound and the loop is working. If it lands well past zero into under-estimate territory, you had bias plus a one-off, and the correction was larger than the evidence supported.

The second check is what turns costing from a monthly report into a feedback loop. A correction you never verify is a guess with better formatting.

References

  • U.S. Small Business Administration (SBA), financial management and job costing basics for small business
  • Trade-standard practice for estimating feedback and production-rate correction
  • See related: The Job Costing Data a Small Shop Actually Needs, The Variance Threshold Worth Investigating, How to Compare Estimated Against Actual on Every Job, How to Find the Job Types You Consistently Underbid, How to Log Actuals in the Field Without Friction