The Job-Type Margin Ranking and the Count Beside It

Why this matters

A margin ranking by job type is the most actionable page of numbers a small shop owns, and it is also the one most likely to be read wrong, because the eye goes to the extremes. The top of a ranking and the bottom of a ranking are exactly where the counts are smallest, so the two rows you most want to act on are usually the two you have the least evidence for. The column that decides whether a rank means anything is not the margin column. It is the job count sitting beside it.

The ranking as it comes out

One shop's completed work for a quarter, each type's margin being the plain average of its own jobs' margin percentages, ranked best to worst:

Rank Job type Margin Jobs completed Share of revenue
1 Specialty retrofit 71 percent 3 1.9 percent
2 Emergency after-hours 66 percent 9 2.1 percent
3 Diagnostic call 62 percent 34 3.6 percent
4 Repair 54 percent 96 26.6 percent
5 Maintenance visit 49 percent 118 17.6 percent
6 Install 27 percent 41 48.1 percent

301 jobs with revenue. Read the margin column alone and the message is "do more retrofits and more after-hours work." Read the count column and two of the six rows stop being findings at all. Read the revenue column and the bottom row turns out to be nearly half the business on 13.6 percent of the jobs.

What one job does at the top of the list and at the bottom

This is the whole argument, and it is one calculation done twice.

Specialty retrofit, 3 jobs at 71 percent. Replace one of those three with a job that went badly at 18 percent. The new average is (71 plus 71 plus 18) over 3, which is 160 over 3, or 53.3 percent. The type drops 17.7 points and falls from first place to fourth on a single job.

Maintenance visit, 118 jobs at 49 percent. Do the same thing: one job comes in at 18 percent instead of 49. The average moves by (18 minus 49) over 118, which is 0.26 of a point. From 49.0 to 48.7 percent. The rank does not move.

One bad job moves the three-job type roughly seventy times as far as it moves the 118-job type. Nothing about the shop's work differs between those two rows. The difference is entirely the denominator, and it means the top of that ranking will reshuffle next quarter whether or not anything changes.

A type at the bottom on a large count is the opposite case and it is the one worth your quarter. Install at 27 percent on 41 jobs is not going to move on its own, is not an artifact of one bad week, and is where a real decision lives.

Setting the floor

Commit to two numbers, not one, and apply them before you read the ranking. A starting point to tune:

Rank a type only when it clears 30 completed revenue jobs in the period being ranked. Act on a type only when it also carries at least 5 percent of the period's revenue. The first is a reliability floor, the second a materiality floor, and they are measured in different things because they answer different questions. Population for both: jobs with revenue, in the period being ranked, not rolling, not including the free visits. What gets measured over that population is what differs - how many of them for the first gate, how much revenue they carried for the second.

  • The 30-job gate asks whether the number is real. Below it, a single job moves the average enough to change the rank, as the calculation above shows. That is a question about how the average was built, so it is counted in jobs. This is what protects a small shop.
  • The 5 percent gate asks whether the number is worth anything. A shop completing 1,200 jobs a quarter can rank a 40-job type perfectly reliably and still have nothing to gain from improving it. That is a question about what a margin point on the type would be worth, and a point is worth exactly the revenue behind it, so this one is measured in revenue. This is what protects a large shop from precision about nothing.

Neither gate subsumes the other, which is why they are not merged into a single test. A type can be perfectly measurable and worth nothing, and a type can carry a fifth of the business on eleven jobs and still be unrankable.

Apply the reliability gate to the table. Specialty retrofit (3 jobs) and emergency after-hours (9 jobs) both fall under 30, which strikes ranks 1 and 2 outright. The rankable list is diagnostic call, repair, maintenance visit and install. The materiality gate comes back further down, and it removes one more of those four.

A struck type is not a type you ignore. It is one you accumulate. Roll it across as many periods as it takes to clear 30 jobs and label it with the window you used, so the number is never quietly compared against a single-quarter figure: "retrofit, 31 jobs across ten quarters, 64 percent." At three jobs a quarter that is two and a half years, which is itself the finding: a type this rare will never be rankable on a timescale you can act on, so either merge it into a neighbouring type or accept that you manage it job by job rather than by the numbers.

The thin, high-volume type

Install at 27 percent on 41 jobs clears both floors and sits last. Three responses, and the order matters.

Price it. The default, and the one to try first when the type's estimated margin and its delivered margin agree. If you quoted 27 percent and delivered 27 percent, the work is behaving and the price is the decision.

Scope it. If you quoted higher than you delivered, the price is not the problem and raising it will not fix it. Compare estimated against actual hours on the type, and count how many of the 41 ran over without a change order attached. That is an estimating or a scope-control problem, and repricing on top of it just moves the same gap up.

Stop quoting it. Last, and usually wrong as a first move, because a ranking cannot see what a type feeds. Installs are what put equipment on the ground that becomes the maintenance visits and repairs sitting at ranks 5 and 4 with 214 jobs between them, and they are part of what makes a service route dense enough that those visits are worth driving to. Before dropping a type, run one check: pull the customers who received a maintenance visit this year and look at what their first job with you was. If a large share of them arrived through the type you are about to drop, you are not cutting a thin type, you are cutting the top of the funnel for two of your best ones, and the ranking will look excellent for about two years.

Rankable, real, and still not worth your quarter

The reliability gate strikes two rows. The materiality gate strikes a third, and this one is easy to miss because the row survives every reliability test.

Diagnostic call sits at rank 3 with 62 percent on 34 jobs. It clears the 30-job gate, so the figure is trustworthy. It carries 3.6 percent of the revenue, so it fails the 5 percent one. This row is also why that gate cannot be denominated in jobs: on the job count diagnostic calls are 34 of 301, or 11.3 percent of the quarter, which sounds like a substantial part of the business and would clear any job-count threshold you would sensibly set. Index revenue so that a diagnostic call is 1.0 unit, and the quarter's 937.6 units break down as 34.0 on diagnostics and 451.0 on installs. A margin point is worth exactly in proportion to the revenue it applies to (see the averages card for why), so a point on diagnostic calls is worth 0.34 units and a point on installs is worth 4.51 units - a factor of about 13.

So the ranking contains three different kinds of row that all look alike until you read across:

  • Too few jobs to rank. Retrofit, after-hours. Accumulate; do not act.
  • Rankable but immaterial. Diagnostic call. Measure it, hold it, spend nothing on it.
  • Rankable and material. Repair, maintenance, install. This is the whole list of places your quarter can usefully go, and it is three rows out of six.

The same discipline applies to comparing a type against its own prior period. A type with 34 jobs this quarter and 11 last quarter cannot be compared to itself either, and the comparison will look like a trend.

When two types are the same type

A ranking is only as good as the definitions under it, and the most common defect is two types that overlap at the boundary where one converts into the other.

Take a diagnostic call that turns into a repair on the same visit. Booked as one repair job, the repair carries the diagnostic hour and whatever fee was charged for it. Booked as two jobs, the diagnostic job carries an hour of labour against a small fee, which reads thin, and the repair job carries the parts and the repair labour with the diagnosis already paid for, which reads fat. Identical work, two booking conventions, and two ranks move in opposite directions.

The tell is a type whose margin is oddly sensitive to which dispatcher wrote it up. Two tests, both quick:

  • Compare hours per job and part content across the pair. Types that are genuinely different work look different on both. Types that are labelling variants look alike.
  • Count how often the pair appears on the same day at the same address. A high rate means the conversion is happening constantly and the split is arbitrary.

Where a pair fails those, merge them for ranking purposes and decide a single convention going forward. A ranking that sorts labelling habits is worse than no ranking, because it is confident.

Working the list

An order of operations that keeps you out of the two traps above:

  1. Strike every type below the reliability gate before you look at the margins at all. Read the count column first, deliberately, because reading it second never works.
  2. Check the survivors for overlap using the two tests, and merge any pair that fails.
  3. Sort what remains by the revenue each type carries, not by its rank, apply the materiality gate down that list, and start at the top of what is left. A type's position in the margin ranking tells you how much room there is; its revenue share tells you what a point of that room is worth.
  4. Set a per-type target rather than one shop-wide number. Installs and diagnostic calls have no business being held to the same percentage, and a single blended target guarantees that half your types are being judged against a figure that was never meant for them.
  5. Re-run the same ranking next period and compare types to their own prior figures, not to each other. A type that moved against itself is a finding. A type that moved position in the ranking may only mean a different type moved.

One boundary to keep in mind while doing all of this: each type's figure is a plain average of its own jobs, so it does not weight a large job above a small one, and the column cannot be averaged back into a shop-level margin. The shop figure is a separate calculation over total margin and total revenue.

References

  • See related: Gross Margin Percent Is an Average of Averages - the weighting basis this ranking inherits, and why the column will not blend
  • See related: Job Profitability by Service Type - setting up the tracking that produces this ranking in the first place
  • See related: The Month Margin Fell and Nothing About the Work Changed - what a shift in the mix of these types does to the headline figure
  • See related: A Metric That Doesn't Change a Decision Isn't Worth Tracking