SLA Compliance Only Counts Jobs That Carry a Deadline

Why this matters

This is the most flattering operations number a shop keeps, and it is flattering by construction rather than by accident. A shop that records a deadline on one job in ten reports its compliance on that tenth and knows nothing at all about the other nine. Worse, the figure improves the fewer commitments you write down, so a shop can watch it climb quarter after quarter while its promise-keeping gets steadily worse, and every individual number in the calculation stays correct the whole way.

The denominator that has to be opted into

Of completed jobs that have a response or completion deadline recorded, the share finished by it. The assumption is that the jobs carrying a recorded deadline are the shop's promise set: that if you promised something, it is in there.

That is an opt-in denominator, meaning a rate computed only over the rows somebody chose to fill in. Every metric built that way shares one property, and it is worth stating plainly because it governs several numbers a shop watches: the rate improves as coverage falls, without anything improving in the work. This article is where that claim gets derived, because this is the number where it does the most damage.

An opt-in denominator is one species of a wider problem this card owns for the set: any metric whose denominator excludes a population, where the exclusion is never random, and where the repair is the same shape every time, a companion count over exactly the excluded population, printed beside the headline rather than folded into it. The opt-in case is sharpest because the figure improves the less you record, but the derivation below holds wherever rows leave a denominator for a reason connected to what is being measured, so siblings sharing the shape spend a sentence and point here.

Why the excluded rows are not a random sample

The general form first. If the rows missing from a denominator were a random slice, the reported figure would be a noisy estimate of the truth and nothing worse. They are almost never random, because whatever decides that a row gets recorded is usually the same thing that decides how it performs. That connection is the defect, and it gives the error a direction rather than only a size.

Here it runs through two ordinary mechanisms, and both push the same way.

A deadline gets recorded when somebody sits down with the job. A call taken at the counter in the middle of something else, a favour squeezed in at the end of the week, an emergency at four in the afternoon: the jobs most likely to be late are the jobs least likely to have had anyone enter a date against them. The conditions that cause a missing field are the conditions that cause a miss.

A deadline that is obviously going to be missed does not get entered. This is not dishonesty and nobody has to decide to do it. A dispatcher who already knows the part is three days out simply does not put a same-week date on the job, because entering one would be entering something they know is false.

Both mechanisms remove late jobs from the denominator at a higher rate than they remove on-time ones. That is the whole inversion, and once you can name the two mechanisms you can predict the direction without any arithmetic at all.

Two questions do that work on any figure of this shape. What decides whether a row gets recorded, and does that same thing decide how it turns out? Where the second answer is yes, the reported figure is a best case rather than an estimate.

The same quarter, at two levels of coverage

One shop, one quarter, 520 completed jobs, and the identical field performance read twice.

As it actually ran. 94 jobs carried a recorded deadline, which is 18.1 percent coverage of the 520 completed jobs. Of those 94, 88 met the deadline. Reported compliance: 88 / 94 = 93.6 percent.

After the shop starts recording a deadline on every job that carries a real promise. Same quarter, same work, same crew, reconstructed from the call notes and the agreements. 347 jobs carried a promise, which is 66.7 percent coverage of the 520. Of those 347, 274 met it. Compliance: 274 / 347 = 79.0 percent.

As it ran With real coverage
Completed jobs 520 520
Jobs carrying a deadline 94 347
Coverage of completed jobs 18.1% 66.7%
Met the deadline 88 274
Reported compliance 93.6% 79.0%

The compliance figure fell 14.6 points while the shop's performance did not change by one job and its record-keeping got better. That is the number to have in mind the next time somebody's compliance figure improves.

The newly-covered rows tell you why. 347 minus 94 is 253 jobs that previously carried no deadline, and 274 minus 88 is 186 of them that met one. That is 186 / 253 = 73.5 percent compliance on the previously invisible jobs, against 93.6 percent on the 94 that were always visible, a gap of 20.1 points between the two populations. Both of those percentages are shares of their own deadline-carrying population in the same quarter, so they are directly comparable. Neither is comparable to the 520 completed jobs, because even at the better coverage 173 of those, 33.3 percent of the quarter, still carry no promise at all.

That 20.1-point gap is the selection effect, measured. It is the cost of the two mechanisms above, and it was invisible while coverage sat at 18.1 percent.

Compliance is an upper bound, not an estimate

This is the practical form of the rule, and it is cheap to remember.

Where the unrecorded jobs perform at least as badly as the recorded ones, which the two mechanisms above make the default expectation, the reported compliance figure is an upper bound on the shop's true compliance over all promised work. It is not a point estimate with some noise around it. It is the best case, and the true figure is somewhere below it by an amount that grows as coverage falls.

There is one case where the bound flips, and it is worth checking rather than assuming. A shop that records deadlines only on its difficult contracted accounts, and leaves easy residential work bare, has a denominator enriched in hard jobs, so the unrecorded rows may genuinely run better and the reported figure would then understate the truth. The check is one question answered from your own records: what kind of job carries a deadline? If the answer is "the ones we have a contract with", you are in the second case. If the answer is "the ones somebody had time to enter", you are in the first, which is far more common in a small shop.

Run that question before you interpret the level, because it decides the direction of your own error.

The pair you have to quote

Compliance and coverage, in the same sentence, always. "93.6 percent on 18.1 percent coverage" and "79.0 percent on 66.7 percent coverage" are two completely different reports about one quarter, and only somebody holding both pairs can tell that the second one is better news than the first.

A compliance figure quoted alone is unreadable rather than merely weak, since a high one means either that the shop keeps its promises or that it rarely writes them down, and those two states are opposites. That is this figure's version of a rule a sibling card owns in general form.

Coverage stability is the comparability test

There is no defensible cross-shop benchmark for coverage here, and there cannot be one: it is set by how much of your book is contracted, emergency, or routine, so another shop's figure carries no information about yours. Benchmark it against your own trailing periods, and watch the stability rather than the level.

That is deliberately a different rule from the one that governs arrival coverage, and the difference is worth holding onto because the two figures look alike. Every scheduled job should carry a recorded arrival, so arrival coverage has a right answer of 100 percent and a level floor beneath which the figure is unreadable. Most jobs should not carry a deadline, so deadline coverage has no right answer at all, and the only thing you can hold it to is consistency with itself.

The rule that follows is concrete. A coverage move of more than about 5 points between periods disqualifies the compliance comparison until somebody explains the coverage move. Under that threshold, a compliance change is probably about the work. Over it, you cannot tell, because a 5-point coverage shift can move compliance by more than a real operational change usually does.

The two figures above are 48.6 points of coverage apart, nearly ten times the 5-point move that would disqualify a period-to-period comparison, which is exactly why they are not two readings of one shop getting worse. Note the rule is doing different work here: it was written for a coverage move between two periods, and this is one period measured two ways, so the gap is not a disqualifier but a demonstration of what the disqualifier is protecting you from. One quarter of work, measured twice, and the second measurement is the honest one.

Which jobs deserve a deadline, and which are made worse by one

The answer is not all of them, and a shop that reacts to this article by stamping a date on every job has made the number worse in a new direction. A deadline that was never a promise is a compliance failure waiting to be recorded against work nobody was late on.

Record a deadline where:

  • A time was said out loud to a customer. A callback response, a warranty visit, an emergency response, a return date given at the end of a first visit.
  • A written agreement names a response or completion time. A commercial service agreement, a property-management contract, a manufacturer's warranty administration requirement.
  • A third party's clock is running and a miss has a consequence outside your shop. A permit inspection window, an insurance adjuster's deadline, a tenant habitability clock, a scheduled shutdown the customer has planned around.
  • Your shop has decided to promise a standard whether or not the customer asked. A same-day response on a no-heat or no-water call is the common one, and it belongs in the denominator precisely because it is a promise you made to yourselves in public.

Do not record a deadline where:

  • The work is routine maintenance scheduled at the customer's convenience. There is no promise to miss, and a synthetic date turns your own scheduling preference into a recorded failure.
  • The job is waiting on a customer decision or a special-order part. The clock is not yours. If you must hold a date for planning, hold it as a separate field from the promise, and write the rule for what starts and stops that clock once rather than deciding it per job.
  • The customer asked you to come "whenever you are next out this way". Honouring that is a routing decision, not a commitment.
  • The date is an internal target nobody outside the shop has heard. That is a goal, and mixing goals into a compliance denominator turns the figure into a report on your ambition rather than on your promises. Track internal targets if you want them, separately, under a different name.

The operational default, then: record a deadline on every job where a time was said to a customer or a contract names one, and on nothing else. Expect coverage to settle wherever your own mix puts it, then treat a move in that settled figure as the first thing to explain in any period, before anybody reads the compliance number sitting next to it.

One last distinction, because shops conflate them and they answer different questions. On-time arrival is about whether you turned up when you said you would. Compliance is about whether the work was finished by when you said it would be. A shop can run high on one and poorly on the other, and the fixes have nothing in common.

References

  • See related: On-Time Arrival Is a Grace Window and a Check-In Habit - the same opt-in problem, where the missing rows are created by a check-in habit
  • See related: Estimate Conversion Rate and the Cohort Problem - the neighbouring defect, where outcomes are counted in the same window that produced the opportunities
  • See related: The Close Rate Improved Because They Stopped Quoting - owns the general rule that a ratio cannot be read without the counts underneath it
  • See related: The Satisfaction Score and Who Actually Answers a Survey - selection in a response set
  • See related: Technician Utilization and What the Denominator Assumes