The Variance Threshold Worth Investigating

Why this matters

A shop that investigates every miss burns its review time on noise and teaches the crew that logging honest hours produces meetings. A shop with no threshold at all investigates whatever somebody happened to notice, which is always the memorable job and rarely the expensive one.

The threshold is the setting that decides which of those two shops you are. It is not a number you copy out of a book, and it is not one number: a single percentage cannot work across a two-hour ticket and a three-day install. Here is what to start with, and how to replace the starting defaults with numbers measured from your own work.

What a threshold is actually for

It rations attention. You can genuinely act on some number of reviews a week, and that number is small. Two or three is realistic for most shops once you account for pulling the data, getting the job owner in a room, deciding, and making the change. The threshold's job is to make the flagged set match that capacity while capturing as much of your total give-away as possible.

That framing matters because it kills the instinct to set the threshold at whatever feels rigorous. A threshold that flags forty jobs a quarter in a shop that can review twenty-five does not produce more learning. It produces a backlog, then stale reviews on jobs nobody remembers, then a process everyone quietly stops running.

A false alarm has a real cost, and it is not just the time. Investigating a job that ran 8% over on a genuinely normal day tells the crew that ordinary variation is treated as a defect. The next month's hours start arriving suspiciously close to the estimate, and you have traded a small amount of review time for the honest measurement the entire system runs on.

Why a single percentage always fails

Percentages and raw gaps disagree in opposite directions at the two ends of your job-size range, and the disagreement is large.

Job estimate 1.0 hour over reads as 15% over equals
2.0 hours 50% 0.3 hours
6.0 hours 17% 0.9 hours
24.0 hours 4% 3.6 hours

A bare percentage threshold set anywhere useful for big jobs will flag half your short tickets over gaps of twenty minutes. Set anywhere sane for short tickets, it will miss a three-hour overrun on a multi-day job because three hours is only 4% of it.

A bare absolute threshold fails the mirror way: 3.0 hours over is a serious problem on a 6.0 hour job and a rounding error on a 60.0 hour job, so an absolute-only rule will ignore a big job quietly bleeding 8% while flagging every short job that hit traffic.

The dual test, and the defaults to start from

Flag a job when both tests trip. Both, not either, is the whole point:

  • Labor: at or beyond 15% over estimate, and at least 3.0 hours over in raw terms. Treat 15% as the starting default and expect to tune it: the worked calibration below runs the exercise on a real distribution, finds 15% too loose against that shop's mix, and lands on 10%. Which number is right for you comes out of the calibration, not out of this line.
  • Material: at or beyond 1.25x the estimated material cost, and enough of the bid to matter (a workable gate is material being at least 20% of estimated job cost).
  • Margin: earned gross margin 10 or more percentage points below quoted margin. This one stands alone with no second gate, because it is already scale-adjusted and it is the number that reaches your bank account.

Start there and tune. These are defaults for a shop doing a mix of service and small project work, not laws. A shop running mostly multi-day projects will find 3.0 hours far too tight and should raise it. A shop running mostly one-hour tickets has the opposite problem at both ends: 15% of a one-hour ticket is nine minutes, so the percentage test alone flags nearly everything, while a 3.0 hour absolute gate no ticket can ever reach blocks every flag and the list comes back empty. That shop lowers the absolute gate until it does real filtering, and expects the gate, not the percentage, to be the test that decides.

Two triggers stand outside the dual test and flag regardless of size:

  • Any work performed and never billed. That is a process failure, not a variance, and a small one predicts a large one.
  • Any job the crew flags themselves. A tech who says something was wrong with that job is giving you free diagnosis. Never let a threshold overrule it.

How to measure your own noise floor

The defaults above are a starting point. Your actual threshold should come from your own distribution, and measuring it takes one pass through a quarter of closed jobs.

  1. Take every closed job in the quarter with a frozen estimate. Compute the absolute labor variance percentage for each, ignoring direction.
  2. Sort them and find where the middle of the pack sits. The share landing within 10%, within 20%, and beyond 20% is enough resolution.
  3. Count how many jobs a threshold at each level would flag, per week.
  4. Pick the level whose flagged count matches what you can genuinely review, then add the absolute-hours gate to strip out the short-ticket false alarms.

The order matters. Set the percentage from your capacity, then use the absolute gate to improve the quality of what got through. Doing it the other way, picking a hours figure first, tends to produce a flagged set dominated by big jobs that were within a normal percentage.

A worked calibration

A shop closes 80 jobs in a quarter. Its honest review capacity is about two a week, roughly 26 a quarter.

The distribution of absolute labor variance:

Band Jobs Share
Within 10% either direction 44 55%
Between 10% and 20% 21 26%
Beyond 20% 15 19%
Total 80 100%

A 10% threshold alone flags everything outside the first band: 21 plus 15, or 36 jobs a quarter, about 2.8 a week. That is over capacity, and inspection shows that 14 of the 36 are short tickets where the raw gap was under 3.0 hours, mostly one-hour misses on two and three hour work.

Adding the 3.0 hour gate drops those 14 and leaves 22 flagged, about 1.7 a week, comfortably inside the 26-a-quarter capacity.

Now check what the gate threw away, because a filter you do not audit is a filter you do not understand. The 14 dropped jobs averaged about 1.4 hours over each, so together roughly 19.6 hours. The 22 that stayed averaged about 6.5 hours over each, roughly 143 hours. Total give-away in the flagged-or-dropped set is about 162.6 hours, and the dropped jobs are 19.6 of that 162.6, about 12%.

So the absolute gate removed 39% of the flagged jobs (14 of 36) while giving up 12% of the hours. That is the trade, stated plainly, and it is a good one.

The 12% does not get ignored, it gets handled differently. Those short tickets go into the monthly pattern review as an aggregate: if diagnostic calls are running consistently 30 to 50 minutes over across dozens of tickets, that is a template correction worth making, and it is visible in the aggregate even though no single ticket was ever worth a meeting. Per-job investigation and pattern analysis have different thresholds because they answer different questions.

If this shop wanted more headroom, moving the percentage to 15% while keeping the 3.0 hour gate would flag roughly 16 a quarter, about 1.2 a week, at the cost of giving up some of the 10-to-20% band. That is the right move if reviews are consistently slipping past the two-week freshness window, and the wrong move if capacity is genuinely there.

Favorable variance needs a threshold too

Set the same thresholds for jobs coming in under estimate, and actually run them. Most shops never do, and it costs them in two ways.

A job that lands 30% under estimate is one of three things: a template quoting more hours than the work needs, which is losing you bids you never hear about; a job where corners got cut, which becomes a callback later; or hours that were never logged, which means your measurement is broken and every other number on the sheet is suspect.

All three are worth fifteen minutes. The third one in particular is the highest-value flag in the entire system, because unlogged hours corrupt the data that everything else depends on, and a big favorable variance is often the only place it becomes visible.

What changes the answer

  • Job-size range. A shop whose jobs run from 1.0 to 8.0 hours needs a much tighter absolute gate than one whose jobs run from 4.0 to 200.0 hours. If your range spans more than about ten to one, consider two threshold pairs, one for service work and one for project work, rather than one compromise that fits neither.
  • Time and materials work. The customer absorbed the overrun, so the margin threshold does not apply. Keep the labor threshold, because a T and M job running well past the verbal ballpark is a customer-relationship event that needs a call before the invoice goes out.
  • A brand new job type. Suspend the thresholds and review the first several regardless. You are gathering baseline information, not detecting deviation from a baseline you do not have yet.
  • After a big correction. The month following a template change, temporarily tighten the threshold on that job type so you see the effect early rather than waiting for the aggregate to show it.
  • Peak season. If your busy months genuinely cost more hours, a fixed threshold will flood the flagged list in July and tell you what you already know. Either widen the band seasonally or, better, hold the threshold and accept that peak jobs get reviewed as an aggregate rather than one by one.

How to verify your threshold is set right

  • Count the flags against your actual review throughput for a month. If flagged jobs consistently exceed reviews completed, the threshold is too loose no matter how principled it sounds. A backlog is a threshold problem, not a discipline problem.
  • Check the hit quality. Of the jobs you reviewed, what share produced a real, recorded action? If it is under about half, you are flagging noise and the threshold should tighten. If it is near all of them, you may be missing findings just below the line and could afford to loosen.
  • Audit what the gate dropped, once a quarter. Total the hours in the jobs that tripped one test but not the other. If that pool is growing as a share of your total give-away, the gate is now filtering out something real.
  • Watch for the flagged list going empty. A threshold that stops catching anything usually means logging has drifted toward the estimate rather than estimating having become perfect. Cross-check against the shop-level spread: genuine improvement narrows the distribution and keeps it scattered both ways, while collapsed logging shows an implausible cluster right at the estimate.

References

  • U.S. Small Business Administration: job costing and cost control guidance for small contractors.
  • Standard construction practice on variance analysis and materiality thresholds in project cost control.
  • See related: The Post-Job Cost Review SOP.
  • See related: How to Compare Estimated Against Actual on Every Job.
  • See related: Why Estimates Miss and Which Misses Matter.