The Satisfaction Score Held While the Repeat Rate Fell
The signal
A residential service shop, four technicians, running a satisfaction survey on every job it closed cleanly. Across a full year the score sat at 4.6 out of 5 in the first quarter, 4.6 in the second, 4.7 in the third and 4.6 in the fourth. It moved by a tenth of a point all year and never once fell below where it started.
Over the same four quarters, the share of completed jobs done for a customer the shop had served before went 47 percent, 44 percent, 41 percent, 38 percent. Nine points off the repeat share while the feedback number held flat.
The owner's first instinct was that the repeat figure was broken, because the survey was the number she trusted. That instinct is the reason this took nine months to find. The survey was arithmetically perfect the whole time.
The first candidate: the work got worse
If the work had degraded, the repeat rate would fall and the survey would eventually follow. The test was the rework count: visits inside 30 days of the original job, for the same complaint, at no charge.
| Quarter | Jobs completed | Rework visits | Rework rate, share of that quarter's completed jobs |
|---|---|---|---|
| Q1 | 178 | 6 | 3.4 percent |
| Q2 | 186 | 5 | 2.7 percent |
| Q3 | 181 | 7 | 3.9 percent |
| Q4 | 174 | 6 | 3.4 percent |
Four values inside a 1.2-point band, in no order, on a job count that moved less than 7 percent across the year. Nothing there. This is the right first candidate because it is the only one that would have made the survey wrong rather than misleading, and ruling it out is what forced the rest of the investigation.
The second and third: price, and a competitor
Price out of line. If the shop had drifted expensive, the estimates would have said so first, because an estimate is where a customer compares. Estimate conversion, as a share of estimates issued in the quarter, ran 62, 61, 63 and 62 percent. Estimates lost with the reason "went with another quote" ran between 9.8 and 12.1 percent of estimates issued, in no order. Both flat. And more to the point, neither of those numbers touches a returning customer, who mostly does not shop the second job at all.
A competitor moved in. This one had a sharp test available and it is the test that pointed at the answer. If somebody had opened up nearby and started taking the market, new work would fall alongside repeat work. First-time customers ran 41, 44, 39 and 43 a quarter: flat.
That split the problem cleanly in half. Acquisition was fine. Retention was not. Whatever was happening was happening to people who had already met the shop, and the instrument that was supposed to be watching those people was reporting 4.6.
The cut that worked
The question that finally landed was not about customers at all. It was: which jobs get a survey.
The rule, set up two years earlier and never revisited, was that the request went out automatically on any job closed with no open complaint and no rework visit against it. Sensible-sounding, and universally understood in the office as "we do not pester people while we are still sorting their problem out."
Count what that rule did over the year. 719 jobs completed. 604 survey requests sent. 115 jobs, 16.0 percent of completed jobs, were never asked at all and never came back into the pool once their complaint closed, because nothing in the rule brought them back.
Those 115 are not a random 16 percent. They are the jobs with a complaint, a rework visit, a disputed invoice or a cancelled and rebooked appointment. They are, by construction, the worst 16 percent of the year's work, and they are exactly the customers who then did not come back.
The survey was not sampling the shop's customers. It was sampling the shop's good days, on a loop, and reporting the average of them with two decimal places of confidence.
The recount
For one quarter they sent the survey on every completed job, with one exception kept deliberately: a job with an open complaint was asked after the complaint was closed rather than never. That is a sequencing rule. The old one had quietly been a selection rule.
- 181 jobs, 181 asks, 46 responses. Response rate 25.4 percent of asks sent, down from 28.3 percent under the old rule, which is what you would expect when the newly-included group is less inclined to answer.
- 46 responses clears the 30-response floor a mean needs before it is worth quoting, and 25.4 percent clears the response-rate floor. See related: The Satisfaction Score and Who Actually Answers a Survey.
- Mean across all 46: 4.2.
Then they split those same 46 responses by whether the old rule would have asked that customer:
| Group | Respondents | Sum of ratings | Mean |
|---|---|---|---|
| Would have been asked under the old rule | 39 | 179 | 4.6 |
| Would not have been asked | 7 | 15 | 2.1 |
| All respondents | 46 | 194 | 4.2 |
The 39 who would have been asked returned 4.6, which is the number the shop had been reading for two years. It was never wrong about them. The seven who would not have been asked returned 2.1.
And the count that mattered more than either mean: 6 of the 46 respondents rated a 1 or a 2, and 5 of those 6 sat inside that group of seven. The shop's entire bottom-end signal, the thing the survey existed to catch, lived almost wholly inside the population the survey was built to skip.
Why the old number was not wrong, and why the new one is not a fall
This is the part that gets reported badly, so it is worth being exact about.
The 4.6 was a correct mean of the ratings it received. Nobody fudged it, nobody deleted a bad response, and it passed every check the shop knew how to run: enough responses, a healthy response rate, a stable series. A number can be computed perfectly and still answer a question nobody asked, and the question this one was answering was "how do customers feel about the jobs that went smoothly."
Which means 4.2 is not a drop from 4.6. It is a different measurement over a different population that happens to print in the same units. Reporting "satisfaction fell 0.4 points this quarter" would have compared a corrected figure against an uncorrected one and invented a decline that nothing in the work supports.
The shop closed the old series with a dated note saying what the collection rule had been, and started a new one at 4.2. Nobody draws a line between them. That is the honest handling, it costs nothing, and the alternative is a chart that lies for as long as anyone keeps it.
What they changed
Two things, and only the second one does any work on its own.
One: survey every completed job, with a sequencing exception and no selection exception. Written down, so that "we do not pester people" cannot re-acquire its old meaning the first time somebody is having a bad week. A job held back has a date by which it gets asked.
Two: watch the bottom-end count, and add an effort question. The mean went to the bottom of the report. The number at the top became the count of 1 and 2 ratings in the quarter, alongside the share of those customers contacted by a person within two working days. Target for the second one is 100 percent, because it is a count of about six calls a quarter and there is no defensible reason to miss any of them. The effort question went on the same form, because a customer who rates the outcome well and the process badly is the specific shape this shop had been producing and never seeing. See related: The Effort Score and What It Predicts That Satisfaction Does Not.
The one division that would have caught it in week two
Nine months of investigation came down to a division nobody had ever run: asks sent over jobs completed. 604 over 719 is 84.0 percent coverage. Any coverage figure under 100 percent means a selection rule is running in your shop, whether or not anybody remembers writing one, and the 16 percent it is selecting out is never a random 16 percent.
That figure belongs on the same line as the score, every time it is reported. "4.6 on 43 responses, 84 percent coverage" is a statement a reader can weigh. "4.6" on its own is not, and the reason nobody computed it is mundane and worth naming: the ask count and the completed-job count lived in two different places, and no single person was responsible for both.
The same rule hides in three other asks a shop makes, and they are worth checking in the same hour, because every one of them is usually configured by whoever set it up and never looked at again:
- The review request. Same gate, same effect, and the cost is more than a review profile built only out of the jobs that went well: where the request goes only to customers you expect to be positive, that is review gating, which the FTC treats as deceptive under Section 5 and which its Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, in force since October 2024, reaches where the resulting display is presented as representative. Federal and penalty-bearing, not only a platform terms question, so widen this ask on the same day you widen the survey.
- The loyalty or recommend question, wherever it rides. If it goes out on the same trigger as the satisfaction ask, it inherits the same population and the same blind spot.
- The referral or membership offer. Usually gated on the tech's judgement on the day, which is the loosest selection rule of the four and the hardest to audit, because nothing is recorded when somebody decides not to mention it.
For each one, the check is the same single division: how many did we ask, over how many could we have asked. If you cannot produce the denominator, that is the finding.
How they confirmed it, and what the confirmation does not prove
They held the widened rule for three further quarters and watched the repeat share, as a share of completed jobs in each quarter: 38 percent, then 39, then 42, then 44. It held for one quarter, then rose in each of the three that followed, and by the last of them it was within three points of where the year of decline had started.
The widened satisfaction mean over those same quarters sat between 4.2 and 4.3, and they deliberately did not report the tenth of a point as a rise, for the same reason they had not reported 4.6 to 4.2 as a fall.
Now the honest part. This is one shop's before and after, not a controlled test, and the survey change by itself fixed nothing. Widening the sample did not make a single customer happier; it made unhappy customers visible. The work that moved the repeat share was the calling: somebody ringing every 1 and 2 rating within two working days, listening, and fixing what could still be fixed. The survey change is what made that work possible, and it is the reason the shop found six of those customers a quarter instead of one.
The generalisable part is smaller than the story and sharper. When an opinion number holds steady while a behaviour number moves, check the opinion number's collection rule before you check anything else, because a behaviour number sees every customer and an opinion number only sees the ones you asked. In this shop the rule had been in place for two years, had been written by somebody who no longer worked there, and had never appeared in a single report.
References
- See related: The Satisfaction Score and Who Actually Answers a Survey
- See related: The Effort Score and What It Predicts That Satisfaction Does Not
- See related: The Review Funnel: Sent, Clicked, and What Happens After
- FTC Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, for the review-request gate described above