How to Find the Job Types You Consistently Underbid
Why this matters
Every shop has one. A category of work that feels normal, gets quoted the same way it always has, and loses hours on almost every ticket. It survives because it never announces itself: no single job is a disaster, the crew never complains, and the annual numbers just come in softer than they should.
Finding it is a specific analytical job, not a matter of paying closer attention. It requires grouping jobs, requiring a minimum count before you believe anything, ranking by total exposure rather than by worst percentage, and ruling out three confounds that will otherwise send you off to fix the wrong thing. Done properly it takes an afternoon and pays for years.
Step 1: Fix your job-type taxonomy before you look at any numbers
You cannot group jobs by a category you do not have. Most shops have one of two problems here, and both make the analysis worthless.
Too fine is the common one. Two hundred distinct job codes means every category holds three jobs and nothing reaches a sample size you can trust. You will find patterns everywhere, all of them noise.
Too coarse is the quieter one. A single bucket called "service" containing both quick diagnostics and half-day repairs averages a 20% overrun on the big ones against an underrun on the small ones and reports that everything is fine. Averaging across genuinely different work is how a leak hides in plain sight.
The workable target is 8 to 15 job types, each one holding at least eight jobs in your review period, each one describing work that shares a method and a rough size. If a category never reaches eight jobs in a quarter, merge it upward into a family and treat the specific work as a variant.
Do this before you pull data. Regrouping after you see the results is how you talk yourself into the answer you expected.
Step 2: Pull one row per closed job, with the fields that let you cut it
One row, one job, these columns at minimum:
- Job type
- Estimated labor hours (from the frozen estimate)
- Actual labor hours
- Estimated and actual material cost
- Who ran it
- Customer or site
- Close date
The last three are not optional and they are the ones people leave out. Without who ran it, you cannot rule out a crew effect. Without the site, you cannot rule out one heavy customer. Without the date, you cannot tell a seasonal pattern from a permanent one. An analysis missing those columns can find that something is wrong and can never tell you what.
Use closed jobs only, and use a period long enough to fill your categories. A quarter is the usual unit. If your volume is low, use two.
Step 3: Group, and require a minimum count before you believe anything
For each job type compute four numbers:
- Count of jobs.
- Median labor variance as a percentage of estimate. Median, not mean, so one runaway job cannot invent a pattern.
- Jobs over estimate, as a count and as a share of the type.
- Average estimated hours per job.
Then apply the floor: fewer than eight jobs, no conclusion. Put the type on a watch list and keep collecting. This rule feels overcautious right up until you reprice a template off four jobs and spend the next year wondering why that work stopped closing.
Step 4: Read direction before magnitude
The share of jobs that landed over is a better first signal than the size of the median miss, and it is the one most people skip past.
If the estimates were merely noisy, roughly half the jobs of a type would land over. When three out of four or more land over, the estimate is off-center and the median tells you by how much. A type with a big median miss but an even split is scattered, not biased, and the fix for scatter is better information at quote time, not a bigger number.
Step 5: Rank by exposure, not by the worst percentage
This is the step that decides where your afternoon goes, and the intuitive ranking is wrong.
Exposure is average estimated hours, times the median percentage bias, times the count in the period. It answers the only question that matters: how many hours did this job type give away over the period. A 22% miss on twelve short tickets is a smaller problem than an 11% miss on twenty-one big ones, and only the multiplication will tell you that.
Rank every type by exposure and work the list from the top. The type at the top is frequently one nobody was worried about.
Step 6: Rule out the three confounds before you blame the job type
Before you touch a template, re-cut the same jobs three ways. Each cut asks whether the pattern is really about the job type at all:
- By who ran it. If one crew accounts for most of the overruns, you may be looking at a training or productivity issue rather than an estimating issue. May, not certainly. See below.
- By customer or site. One heavy account with hard access, strict escort rules, or restrictive work windows can drag a whole category. That is a pricing problem for that account, not for the job type.
- By season or date. If the overruns cluster in your busy months, you have peak-season cost, not a bad template. Peak work genuinely costs more hours: rushed staging, unfamiliar helpers, longer drives between packed stops.
The confound cuts are also where the analysis most often gets it wrong in an interesting way, because confounds can hide behind each other. A crew effect and a site effect look identical when the same crew keeps drawing the same kind of site. Always run at least two cuts before concluding, and when two cuts both show a pattern, cross them.
Step 7: Split the category, correct the number, or leave it
Three outcomes, and picking the right one is the whole payoff:
- Split when the confound cut found a real subgroup with a different cost profile. Two templates that each fit beat one template that fits neither.
- Correct when the bias holds across every cut. Raise the template hours by about two-thirds of the measured bias, date the change, and verify next period.
- Leave it when the count is thin, the split is even, or the pattern vanishes under a confound cut. Watch list, keep logging.
There is a fourth case worth naming: a type running consistently under estimate with well under half the jobs over. That is not free money. It is a template quoting more hours than the work takes, which prices you out of bids you never hear about. Investigate it with the same seriousness.
A worked investigation
One quarter, 186 closed jobs, six job types.
| Job type | Count | Avg est hours | Median variance | Jobs over | Exposure (hours) |
|---|---|---|---|---|---|
| Service diagnostic | 62 | 1.5 | +6% | 41 (66%) | 5.6 |
| Standard repair | 44 | 3.0 | +3% | 24 (55%) | 4.0 |
| Residential replacement | 21 | 14.0 | +11% | 17 (81%) | 32.3 |
| Commercial retrofit | 9 | 26.0 | +18% | 7 (78%) | 42.1 |
| Maintenance visit | 38 | 1.0 | -2% | 15 (39%) | favorable |
| After-hours emergency | 12 | 2.5 | +22% | 9 (75%) | 6.6 |
Read it in order.
After-hours emergency has the worst percentage miss at 22%, and it is the one the owner brings up unprompted. Its exposure is 6.6 hours for the quarter (0.22 x 2.5 x 12). Real, worth a template bump, not where the afternoon goes.
Commercial retrofit has the highest exposure at 42.1 hours (0.18 x 26.0 x 9), and only 9 jobs. Nine clears the eight-job floor by one, so the finding is provisional. Treat it as a strong lead, correct conservatively, and re-measure next quarter before doing anything drastic.
Residential replacement is second by exposure at 32.3 hours (0.11 x 14.0 x 21), with 21 jobs behind it and 81% landing over. That is the most solid finding on the sheet: good count, strong direction, large exposure.
Maintenance visit leans the other way: median 2% under, only 15 of 38 over, which is 39%. That template is quoting more hours than the visits take. Not urgent, but it goes on the list as a possible close-rate drag.
Now the confound cuts on residential replacement, the finding with the best evidence.
Cut by crew: Crew A ran 12 of them, median 4% over. Crew B ran 9, median 19% over. That looks conclusive, and a shop that stops here goes and has a difficult and unfair conversation with Crew B.
Cut by site: of the 21 replacements, 10 were in pre-1980 housing and 11 were newer. Older housing median 21% over, newer housing median 2% over. And the cross: Crew B drew 7 of the 10 older-housing jobs.
So the crew signal was mostly the site signal wearing a crew's name. The work in older housing costs more hours because of what is behind the walls and what has to be brought up to current requirements, and Crew B happened to get assigned most of it. Correcting this as a crew problem would have damaged a good tech and left the estimate exactly as wrong as before.
The action: split residential replacement into two templates by housing vintage. The older-housing variant carries a 21% measured bias, so at two-thirds gain it goes up 14%. The newer-housing variant, running 2% over on an even-ish split, gets nothing.
Note what a single blended correction would have done. Raising the whole category by two-thirds of the blended 11%, about 7%, would have left the older-housing work still badly underbid while overpricing the eleven newer-housing jobs that were quoted correctly. The split is worth more than the correction.
Recomputed exposure after the split concentrates almost all of the category's give-away in the older-housing variant: 0.21 x 14.0 x 10 is about 29 hours of the quarter's 32.3. It does not reconcile to the blended figure exactly, and it should not, because medians of subgroups do not add up to the median of the whole. Use the recomputed subgroup numbers going forward, not the blended one.
What changes the answer
- Fewer than about 60 closed jobs in the period. Widen to two quarters or a rolling window. Segmenting 30 jobs into six types puts five jobs in each and every finding is noise.
- Flat-rate priced work. Labor variance against a published book measures whether the book fits your crew and your market. The correction is a book adjustment factor, not a template rewrite, and it applies across every job priced from that book.
- A job type in transition. New method, new tooling, new install standard. Cut the data at the change date and analyze the halves separately. Mixing them produces a median that describes no job you will ever run again.
- A single dominant customer. If one account is a third of a category's volume, run the category with and without them. Two different answers means you have an account-pricing problem sitting on top of a job-type question, and the account needs its own conversation.
How to verify you got this right
- Your counts sum to your job total. In the example, 62 plus 44 plus 21 plus 9 plus 38 plus 12 equals 186. If they do not, jobs are landing in more than one type or falling out of the taxonomy entirely, and the ones falling out are usually the odd jobs where the money went.
- State every percentage with its base and its period. "17 of 21 residential replacements in the quarter landed over estimate, 81%." A share stated bare is the easiest number in this whole analysis to misread later.
- Re-run the top finding with its single largest miss removed. If the pattern holds, it is real. If it evaporates, you had one bad job and a category built around it.
- Verify the correction next period, with the change date in hand. Pull only the jobs quoted after the change and read the median and the over-share. A correction you never re-measure is indistinguishable from one you never made.
The failure mode worth guarding against is the analysis that gets run once, produces a satisfying finding, and never runs again. The categories that were fine this quarter drift. Supplier costs move, crews change, the mix of housing you work in shifts with your marketing. This is a quarterly cut, not a project, and the second run takes a fraction of the time the first one did because the taxonomy and the query already exist.
References
- U.S. Small Business Administration: job costing and profitability analysis guidance for small contractors.
- Standard construction and field-service practice on estimate-to-actual variance analysis by work category.
- See related: Why Estimates Miss and Which Misses Matter.
- See related: How to Run a Monthly Estimate Accuracy Review.
- See related: The Cost Categories Worth Separating.