The Callback Number Looked Fine and the Customers Did Not
Why this matters
A shop's rework figure held flat for three quarters while complaints tripled and survey scores slid. Both numbers described the same work, so one of them was wrong, and for two quarters the owner assumed it was the complaints - a few loud customers, a rough season. It was not. The rework figure was right about everything it could see and structurally blind to more than half of what was happening, because of one setting nobody had ever looked at: the length of the window it searches.
This is what that looked like from the inside, including the two quarters of wrong turns, because the wrong turns are the part that repeats.
The signal
Three quarters, three numbers, read side by side at the quarterly review.
| Quarter one | Quarter two | Quarter three | |
|---|---|---|---|
| Completed jobs | not stated | not stated | 396 |
| Repeat visit rate (7-day window) | 4.9 percent of that quarter's completed jobs | 5.1 percent | 4.8 percent |
| Written complaints | 5 | 9 | 16 |
| Written complaints per 100 completed jobs | not computable | not computable | 4.0 |
| Survey returns scoring in the top band | 78 percent of returned surveys | 71 percent | 61 percent |
The rework figure is flat and slightly improving. The other two are going the wrong way and accelerating. The survey line was treated as corroboration rather than as a measurement, because the people who answer a survey are not the people you served - see related: The Satisfaction Score and Who Actually Answers a Survey - but complaints are counted, not sampled, and 5 to 16 in two quarters is not a sampling artefact.
One caution on that row, because it is the only raw count sitting beside two rates. This write-up carries the completed-job base for quarter three alone, so 5, 9 and 16 are counts and only quarter three can be stated as a rate, 4.0 complaints per 100 completed jobs. Read the tripling as a count: it would take a collapse in volume to explain a move like that away and nothing here suggests one, but the first two quarters cannot be normalised from what is recorded.
Three candidates, killed on evidence
The crew. The first assumption, and the most expensive one to act on wrongly. The 16 complaint jobs in quarter three span 6 of the 7 technicians, including both of the longest-serving. The two newest technicians ran 87 of the quarter's 396 completed jobs, 22.0 percent of the quarter's completed work, and account for 3 of the 16 complaints, 18.8 percent of the quarter's complaints. Their share of the complaints is slightly below their share of the work. There is no person in this.
The parts. The second assumption, and the easiest to check. Manufacturer claims filed ran 7, 6 and 8 across the three quarters, flat, and parts returns were flat with them. The 16 complaints name 9 different components and no single component appears more than three times. A failing part or a bad batch does not scatter like that; it concentrates, and the warranty line moves with it.
Jobs not actually being finished. The third assumption: work left half done under time pressure. Every one of the 16 complaint jobs was marked complete with a captured customer signature. Checklist completion held at 88, 86 and 87 percent of checklists started in each of the three quarters, so the paperwork habit had not moved either. Whatever this is, it is happening after the technician leaves a job that was genuinely finished.
At this point two quarters had gone by and the shop had eliminated the three things it knew how to eliminate. The remaining move was to stop testing hypotheses about the work and start reading the complaints themselves.
Reading the sixteen complaints
Somebody sat down with all 16 quarter-three complaints and sorted them by what the customer actually said.
- 11 of the 16 are a version of "it is doing it again."
- 3 of the 16 are about scheduling or being kept informed.
- 2 of the 16 are about the bill.
Eleven of sixteen complaints are recurrence complaints. That is the same thing the rework figure claims to measure, and the rework figure says rework is flat. Two numbers about the same phenomenon, pointing opposite ways, which means the definitions differ somewhere.
The cut that worked
For each of the 16 complaints, count the days between the original job's completion and the date the complaint arrived.
- Inside 7 days: 3 complaints.
- Between day 9 and day 16: 11 complaints.
- Past 30 days: 2 complaints.
There it is. The repeat visit rate searches for a second completed job within roughly seven days of the first. Cross the two cuts and the eleven recurrence complaints are exactly the eleven that landed between day 9 and day 16. The three inside seven days are the scheduling complaints and the two past thirty days are the billing ones. So not one of the sixteen complaints that was actually about the work recurring arrived inside the window the measure searches, and the three that did could never have been rework in the first place. The eleven were not being counted as anything at all - not counted as rework, not counted as fine, simply outside the frame.
Nothing was broken and nothing was miscomputed. The number was answering a narrower question than the one the shop thought it was asking, and it had been answering that narrower question correctly for years.
The recount at thirty days
They re-ran quarter three by hand: all 396 completed jobs, same customer, same equipment, second completed job within 30 days instead of 7.
- Raw pairs at 30 days: 68.
- Classified by hand: 15 are different work at the same property, 9 are planned second visits for parts on order, and 44 are genuine comebacks on the original complaint.
- 15 plus 9 plus 44 is 68, and the 44 includes the 19 the standing figure had already counted.
True rework: 44 of 396 completed jobs, 11.1 percent, against the published 4.8 percent of the same 396 completed jobs. The standing figure had seen 19 of the 44 real comebacks, about 43 percent of them. The reading took one person most of a working day, with the job notes open.
That 11.1 percent is not comparable to a published callback-rate benchmark. The figures shops quote to each other, and the targets on the KPI cards in the references, are counted the way the standing measure counts: automatically, on a short window, without anybody reading job notes. A hand-classified recount finds rework those figures never see, so it lands higher by construction. Compare a recount to your own previous recount, and compare the standing figure to the published band, and never cross the two.
Then the same split the complaints had suggested. Of the 11 late complaints, 9 were on calls where the fault would not reproduce while the technician was on site - the customer describes it, the equipment behaves, the tech leaves. Those calls were 63 of the quarter's 396 completed jobs, 15.9 percent of the quarter's completed work.
| At 7 days | At 30 days, classified | |
|---|---|---|
| Faults that would not reproduce on site (63 jobs) | 4 of 63, 6.3 percent of that category | 17 of 63, 27.0 percent of that category |
| Everything else (333 jobs) | 15 of 333, 4.5 percent of that category | 27 of 333, 8.1 percent of that category |
Both slices clear the floor of about 30 completed jobs in the window that any category rate needs before it can be ranked, so the comparison holds.
On the published number the two categories sit 1.8 points apart, each measured against its own completed-job base. Nobody argues about 1.8 points. Recounted, they sit 18.9 points apart on the same two bases: 27.0 percent of the category's own 63 jobs against 8.1 percent of the other 333, more than three times as common. The shop's only rework measure was structurally incapable of showing that, because a fault that needs a condition to return needs more than a week to return it.
The two things they changed
The measurement. Keep the seven-day figure as the standing number. It is fast, it is consistent with its own history, and a sudden move in it still means something. Add a hand recount at 30 days, once a quarter, on the last closed quarter only, with every additional pair classified into genuine comeback, different work, or planned second visit. Record both figures so the pair builds its own history, and treat a widening gap between them as the finding rather than either number alone.
The diagnosis. On the no-reproduce category, they changed what finished means. A technician who cannot make the fault appear now does one of two things before leaving. Either they leave something behind that will capture the fault when it returns - a reading the customer is asked to take and send when it happens, a logger, a recorded value to compare against - or they book a short second visit at the shop's own cost, scheduled into the window when the triggering condition is expected to be present. They also record the triggering condition the customer describes as a field on the job, which is the only reason the category could be counted at all in the analysis above.
Neither change is a training programme and neither costs a hire. One is most of one person's day, once a quarter, and the other is a definition of done.
The objection that came up immediately, and it will come up in your shop too: the rework number just went from 4.8 percent to 11.1 percent, and somebody has to explain that to a partner, a lender or a crew that was being told it was doing fine. The answer is that those are two figures with different windows, not an old figure and a corrected one, so both get published with their window written on the face of them - a seven-day figure and a thirty-day classified figure, always labelled.
What you must not do is restate a single quarter at the new window and compare it against a trailing history built at the old one. That comparison is guaranteed to show a collapse in quality that did not happen, because the baseline was never recounted. If you want a trend at 30 days, recount at least the four prior quarters the same way before anybody reads the line, or state plainly in the same sentence that the earlier figures are seven-day numbers and the new one is not.
The quarter that confirmed it
Quarter four, 404 completed jobs.
- Standing seven-day figure: 19 of 404, 4.7 percent of that quarter's completed jobs, against 4.8 percent of quarter three's 396. It did not move, and it was never going to. That is the point rather than a disappointment.
- The no-reproduce category: 64 completed jobs, and at 30 days 8 of 64 came back, 12.5 percent of that category's completed jobs, against 27.0 percent of that category's 63 completed jobs in quarter three. That is the number that moved.
- Written complaints: 7, against 16 in quarter three.
- Survey returns scoring in the top band: 74 percent of returned surveys, against 61 percent in quarter three.
The confirmation is worth reading carefully, because it contains the lesson twice. The measure the shop had been running all along was flat before the fix and flat after it. Had they made this change without building the 30-day recount first, they would have had no way to tell whether it worked, and the standing number would have quietly told them it had not.
What else in the shop has the same blind spot
Any measure with a fixed lookahead shares this shape, and most shops have several without thinking of them as measurements at all. A satisfaction survey sent two days after the visit. A review request sent at day three. A courtesy call at the end of the week. Each of them asks the customer for a verdict before a slow-returning fault has had time to return, and each of them will therefore run persistently kinder than the truth on exactly the work that is hardest to diagnose.
That is the portable rule out of all of this. When two numbers describing the same thing disagree, check the definitions before you check the work, and start with whichever one has the narrower frame. A number cannot report something it was never looking at, and a flat line from a narrow measure is not evidence of stability. It is evidence of the measure.
References
- See related: The Repeat Visit Rate and the Seven Day Window It Uses, for the composition of the standing figure and the quarterly hand cut
- See related: Callback Rate Management, and Mining Your Callbacks for What They Reveal About Your Systems
- See related: The Satisfaction Score and Who Actually Answers a Survey, for why the survey line was corroboration rather than evidence
- See related: The Satisfaction Score Held While the Repeat Rate Fell, for the same disagreement running the other way