Why Your Average Job Is Lying to You

Why this matters

"Our average service call runs about two hours" is the sentence most shops build their pricing on, and it is usually true and useless at the same time. An average only describes a population that is actually one population. Blend two job types that behave differently and the average lands somewhere neither of them lives, so you overprice the easy work until you lose it and underprice the hard work until you win all of it. That combination is fatal in slow motion: your mix drifts toward the jobs you are cheapest on, the average climbs, you raise the price, and you lose more of the good work.

The fix is not more decimal places. It is refusing to average across things that are not the same thing.

What an average actually assumes

A mean is a fair summary only when the values cluster around a single center. The moment your data has two clusters, or a long tail on one side, the mean sits in the empty space between them.

Three shapes show up constantly in service work, and each one breaks the average differently.

Shape What it looks like What the mean does What to use instead
Single cluster Most jobs near the center, a few slightly off Mean is honest Mean, plus the spread
Two clusters A fast version and a slow version of the "same" call Mean lands between them and describes neither Split the population, then average each
Long right tail Most jobs normal, a handful catastrophic Mean is dragged up by the tail Median for the typical job, tail counted separately

Service work runs long-tailed almost everywhere. You cannot finish a job in negative hours, so there is a hard floor, but there is no ceiling on how bad a job can go. That asymmetry alone means the mean of your job hours sits above the typical job nearly every time.

The median is the honest center, the spread is the actual payload

The median is the middle value: half your jobs came in under it, half over. It ignores how extreme the extremes are, which is exactly what you want when you are asking "what does a normal instance of this job cost me."

But the median alone still hides the thing that decides whether you can bid the job type at a fixed price. That is the spread. Two job types with identical medians behave completely differently if one runs within a narrow band and the other swings wide.

Read the spread as the gap between the 25th and 75th percentile, or if your sample is small, just the range between the second-lowest and second-highest values. A job type where 3 hours in 4 land within about 20 percent of the median is a flat-rate candidate. One where they scatter across a 2x range is not, no matter how good the median looks.

Mix shift: the average moves without anything changing

This is the trap that catches experienced owners, because nothing in the shop got worse and the number still went up.

Say you run two job types. Type A takes a median of 1.5 hours. Type B takes a median of 4.0 hours. Neither one changes at all over a year. In the first quarter you run 60 of Type A and 20 of Type B, so 80 jobs total. In the fourth quarter you run 40 of Type A and 40 of Type B, so 80 jobs again.

First quarter blended mean hours per job: (60 x 1.5 + 20 x 4.0) / 80 = (90 + 80) / 80 = 2.125 hours. Fourth quarter: (40 x 1.5 + 40 x 4.0) / 80 = (60 + 160) / 80 = 2.75 hours.

The "average job" grew from 2.125 to 2.75 hours, an increase of 0.625 hours, which is 29 percent above the 2.125-hour starting figure. Every underlying job type is identical. The only thing that changed is that Type B went from 20 of 80 jobs, 25 percent of the quarter's volume, to 40 of 80, 50 percent of it.

An owner reading only the blended number concludes their techs slowed down and starts timing people. The correct read is that the shop's work mix shifted toward the longer job type, which is a scheduling, marketing, and pricing question, not a productivity one. Any time your headline average moves and you cannot name which job type moved, check the mix first.

Averaging the wrong thing: percentages of different bases

The second-most-common lie is averaging percentages that sit on different denominators.

Three jobs, each with a labor variance:

  • Job 1: estimated 4.0 hours, actual 5.0. Over by 1.0 hour, 25 percent over the 4.0-hour estimate.
  • Job 2: estimated 2.0 hours, actual 3.0. Over by 1.0 hour, 50 percent over the 2.0-hour estimate.
  • Job 3: estimated 20.0 hours, actual 21.0. Over by 1.0 hour, 5 percent over the 20.0-hour estimate.

Average of the three percentages: (25 + 50 + 5) / 3 = 26.7 percent.

Weighted by hours, which is what actually happened to the shop: 26.0 hours estimated across the three jobs, 29.0 actual, over by 3.0 hours, which is 11.5 percent over the 26.0 estimated hours.

The unweighted average says you are running 27 percent light on labor. The weighted figure says 12 percent. The second one is the one that describes your week, because it is the one that ties to the hours you actually had to cover. The unweighted number lets a 2-hour job carry the same influence as a 20-hour job, which is how a shop convinces itself it has a catastrophic estimating problem when what it has is one small job type that is genuinely mis-templated.

Use the unweighted percentage only when you are asking "how often am I wrong," and the weighted one when you are asking "how much is it costing me." Those are different questions and the same data answers both differently.

Sample size: when the average is not lying, just too young

An average built on 3 jobs is a rumor. A useful default: treat a job type's median as directional at 5 closed jobs, actionable at 10, and stable at 20 or more. Below 5, act on the written causes from individual job reviews rather than the number.

The exception worth naming: a miss that is both large and one-directional does not need 10 jobs. If 3 out of 3 ran more than 50 percent over, that is not noise. The probability of three coin flips all landing heads is 1 in 8, and these are not coin flips - the same cause produced all three. Correct now, keep watching, and expect to correct again.

Worked example: reading one job type properly

A shop pulls 12 closed jobs of one type. Actual labor hours, sorted:

2.0, 2.2, 2.5, 2.5, 2.8, 3.0, 3.1, 3.3, 4.5, 5.0, 7.5, 9.0

The template says 3.0 hours.

Mean: the twelve values sum to 47.4 hours, so the mean is 3.95 hours. Read alone, that says the template is 0.95 hours light, since the 3.95-hour mean sits 32 percent above the 3.0-hour template, and the obvious move is to raise the template to about 4.0 hours.

Median: with 12 values the median is the average of the 6th and 7th, which are 3.0 and 3.1, so 3.05 hours. The template is essentially correct for the typical job. It is off by 0.05 hours.

Those two readings point opposite directions, and the median is right about the typical job. Raising the template to 4.0 hours would price the shop out of 8 of these 12 jobs, the ones that came in at or under 3.3 hours, to protect against a tail the template was never going to cover anyway.

The tail is a separate problem. Four jobs came in at 4.5, 5.0, 7.5 and 9.0 hours. That is 4 of 12, a third of the sample, each running at least 1.45 hours over the 3.05-hour median. That is not a rounding issue on the template. That is a second population hiding inside this job type.

The right next move is not a math move at all. Pull the four long jobs and read the cause sentences from the job cost records. If all four share a condition - older equipment, a specific access type, a particular property class - then this was never one job type. Split it into two, template each separately, and both templates will be tight. If the four causes are unrelated, you have a genuine long tail, and the correct handling is a risk allowance on the job type or a conditions clause, not an inflated hour count on every bid.

What the split looks like if it holds. Say the four long jobs all sit in one property class. The eight short jobs run a median around 2.65 hours (the 4th and 5th values of 2.0, 2.2, 2.5, 2.5, 2.8, 3.0, 3.1, 3.3 are 2.5 and 2.8, so 2.65). The four long jobs run a median of 6.25 hours (middle two of 4.5, 5.0, 7.5, 9.0 are 5.0 and 7.5). Two templates at 2.65 and 6.25 hours describe reality. One template at 3.95 hours describes nothing that ever happened.

When an average is all you have

Sometimes you inherit a number without the data behind it: a report that only prints averages, a previous owner's figures, a supplier's published rate. You can still recover most of the shape with two cheap questions.

How many jobs came in above the average? In a single-cluster population, roughly half will. If far fewer than half exceed it - say 4 of 12 - the mean is being pulled up by a tail and the typical job sits below the number you were handed. That one count is a usable proxy for a median you cannot compute.

What is the worst instance anybody remembers? If the worst case sits within about 1.5x of the average, the population is probably tight enough to price flat. If it sits at 3x or more, the average is describing a population with a tail that will eventually land on you, and pricing at that average means the tail comes out of your margin every time it appears.

Neither question needs a spreadsheet, and both can be asked out loud in a review. Use them as a stopgap while you build the records to answer properly, not as a substitute for building them.

The reporting habit that keeps you honest

Never publish a single number where three will fit. On any job-type report, carry the count, the median, and the spread together. "This job type: 12 jobs, median 3.05 hours, range 2.0 to 9.0" takes one line and cannot mislead the way "average 3.95 hours" does.

And when someone quotes an average at you, including yourself, ask two questions before you act on it: how many jobs is that, and are they the same kind of job? Most bad pricing decisions die right there.

References

  • U.S. Small Business Administration (SBA), interpreting operating metrics for small business
  • Standard descriptive statistics practice, measures of central tendency and dispersion
  • See related: Tracking Margin by Job Type to Find the Leaks; How to Build a Labor Hour Database From Your Own Jobs; The Job Costing SOP