The Satisfaction Score and Who Actually Answers a Survey
Why this matters
You have a satisfaction score in front of you. It is the average rating from the customers who replied to your survey this period, and it is almost certainly the single most quoted number in your shop after revenue. It is also computed over a group that was never a sample of the people you served: it is the people who chose to answer, and the choosing is the whole problem. This card owns that point for the rest of the feedback numbers, because every one of them (loyalty, effort, review requests) inherits it and none of them can be read without it.
What the mean is actually averaging
The composition is simple and that is why it gets skipped. The numerator is the sum of the ratings received in the window. The denominator is the count of responses received in the window. Neither one contains a customer who did not answer, and the survey never asked whether the answerers resemble the silent ones.
They do not. Two kinds of customer reliably answer a service survey: the delighted and the furious. Answering costs a minute and nothing happens as a result, so a person needs a reason. Being pleased enough to want to say so is a reason. Being angry enough to want it on record is a reason. Being reasonably content, which is where most of your customers actually live, is not a reason, so those people close the message.
That means the ratings you receive are not spread evenly around a middle. They pile up at the ends. An average is a description of a single-peaked distribution; run it over a two-peaked one and it lands in the valley between the peaks, describing a customer who does not exist.
Two months of replies, counted properly
Take a shop running a 1 to 5 scale where 5 is best, asking on every completed job.
- Month one. 96 jobs completed, 96 asks sent, 26 replies. Response rate 26 of 96, 27.1 percent of asks sent.
- Month two. 104 jobs completed, 104 asks sent, 31 replies. Response rate 31 of 104, 29.8 percent of asks sent.
Rolled across both months: 57 replies against 200 asks, 28.5 percent of asks sent. Here is the rolled distribution.
| Rating | Replies | Share of the 57 replies |
|---|---|---|
| 5 | 34 | 59.6 percent |
| 4 | 9 | 15.8 percent |
| 3 | 4 | 7.0 percent |
| 2 | 3 | 5.3 percent |
| 1 | 7 | 12.3 percent |
The mean is (5 x 34) + (4 x 9) + (3 x 4) + (2 x 3) + (1 x 7) = 170 + 36 + 12 + 6 + 7 = 231, over 57 replies, which is 4.05.
The mean describes almost nobody
The number on the page is 4.05. Nine replies out of 57, 15.8 percent of the replies, actually sat at 4. Forty-one of 57, 71.9 percent of the replies, sat at 5 or 1, which are the two furthest-apart answers on the form.
So a shop reading 4.05 concludes "customers are broadly happy, slightly short of excellent, keep going." What the replies actually say is that most people who bothered to answer thought the job was excellent, and a solid block thought it was the worst answer available. Those two groups need completely different responses from you, and averaging them produces a number that calls for neither.
The correct read of that table is two facts, stated separately and never blended: 34 of 57 replies were top-box, and 10 of 57 replies were a 1 or a 2. Watch those two counts as a pair. If the top-box count climbs while the bottom-end count also climbs, you are polarising, which usually means the work is inconsistent rather than getting better or worse.
When the distribution is not two-peaked
Two peaks is the default shape, not a law, and one condition genuinely changes it. A shop asking after routine scheduled visits on a maintenance agreement is asking a different population: those customers are in an ongoing relationship and answer out of habit rather than out of feeling, so the replies come back clustered in the middle-to-upper ratings. Against a single-peaked set a mean is a fair summary, and on a window carrying at least 30 replies a shift of two tenths of a point that holds into a second window is worth a look.
The test is your own table, not your trade. Print the reply count at every rating for a full quarter, then compare the two end ratings against everything in between. In the set above, ratings 5 and 1 carry 34 + 7 = 41 replies while ratings 4, 3 and 2 carry 9 + 4 + 3 = 16, so the ends outweigh the middle by more than two to one and the mean is not describing this shop. If the middle carries more than the two ends combined, average away.
How little it takes to move it
Month one alone: 26 replies summing to 107, a mean of 4.12. Suppose three customers who had not bothered to answer change their minds and each rates a 1. Now 29 replies summing to 110, a mean of 3.79. Three replies out of a month's 96 completed jobs, 3.1 percent of the jobs, moved the headline by 0.33 on a five point scale.
That is not a defect in those three customers. It is what a mean does on a small, ends-heavy sample, and it is why a month-to-month score line will look like it is telling you something when it is telling you who happened to open their messages. The same mechanism makes the score look great after a month with fewer angry customers and no change in the work at all.
The bottom-end count, and the base it belongs to
Ten replies of 57 rated a 1 or a 2. Two bases are available and they mean different things, so name which one you are using every single time:
- 10 of 57 replies, 17.5 percent of respondents. This is a share of people who answered. It is the right base for asking "among people who tell us anything, how many are angry."
- 10 of 200 completed jobs, 5.0 percent of jobs. This is a share of work done. It is the right base for asking "how many households have we upset," and it is a floor, not an estimate. The 143 jobs that produced no reply contain an unknown number of unhappy customers, and nothing in this survey will ever tell you how many.
Mixing those two bases is the most common error in reading this number. "Five percent of our customers are unhappy" is defensible as a minimum. "Seventeen percent of our customers are unhappy" is not a statement about your customers at all, it is a statement about your respondents.
The comment box is where the usable content sits. On a two-peaked set the rating only tells you which peak a person stood on, and a 1 with no comment is unactionable. A 1 with "third visit for the same fault" is a work order. So keep the comment field optional rather than required (requiring it costs you replies from exactly the mild-opinion group you are already short of) and read every comment attached to a 1 or a 2 inside the week. Five bottom-end replies a month in the shop above, and not every one will carry a comment now that the box is optional, so this is well under an hour of reading a month, and it is the highest-yield reading you will do.
When the score is quotable at all
Two floors, and both have to clear before the mean goes in front of anybody.
At least 30 responses in the window. Below that, the arithmetic above governs: a handful of replies swings the figure by more than any real change in the work would. And a response rate of at least 25 percent of asks sent. Below that you are so deep into the self-selecting ends that the mean is measuring enthusiasm, not service. Both numbers are a starting point to tune, not a law; if your replies run steadier than this, tighten them.
That floor is high on purpose and an email-only programme rarely clears it: the Net Promoter Score card in the references puts email response at 10 to 20 percent and text at 25 to 50, and published channel figures vary enough that another source will put email nearer 30, so take your own rate from your own sends rather than from anybody's table and check it against the floor directly. If you cannot clear 25 percent, the honest move is to stop quoting a mean at all and read the count of bottom-end responses instead, or to change channel.
Run the shop above through its own rule. Month one: 27.1 percent response rate clears the rate floor, 26 replies fails the count floor. So that month's 4.12 is not a number anyone quotes. The fix is to roll rather than to wait: two months rolled gives 57 replies at 28.5 percent of asks sent, both floors clear, and 4.05 is quotable with its distribution printed beside it. Roll to a quarter if that is what it takes. A rolled figure that clears the floor beats a monthly figure that does not, every time.
When a period fails the floor, report what you do have: the count of replies, the count of 1s and 2s, and what those customers said. That is more useful than the mean was going to be anyway.
Raising response rate honestly, and the version that destroys the number
Three things genuinely raise response rate without touching who answers:
- Timing. Ask within about 24 hours of the work finishing, while the tech is still a person the customer remembers rather than a company. Wait a week and you lose the merely satisfied first, which sharpens exactly the bias you are fighting.
- Length. One rating question and one optional comment box. Every additional question costs replies, and it costs them unevenly: the person with a mild opinion quits first.
- Channel. Ask on the channel the customer already used to talk to you. A customer who booked by text and never opens email is not a non-responder, they are a person you asked in a language they do not read.
And one thing that raises the number while destroying it: asking only the customers you expect to be happy. Suppressing the ask after a rework visit, or after a complaint, or leaving it to the tech's judgement on the day, lifts the mean immediately and permanently. It also removes the only population the survey existed to find. A score collected that way cannot be compared to a score collected any other way, it cannot be trended against its own history once the rule changes, and it will hold steady while customers quietly stop calling. See related: The Satisfaction Score Held While the Repeat Rate Fell.
The only legitimate use of this number is as a trend against your own history on a fixed collection rule. The moment you change who gets asked, the series restarts. Say so out loud when it happens, and keep the old series labelled rather than pretending the new one continues it.
Which of your numbers can see a customer who never answers
The survey mean is one input with a narrow gate on it. Before you treat a quiet score as good news, know which of your other numbers would have seen the same customer.
| Number | Sees a customer who never answers a survey? | What it therefore catches that the mean cannot |
|---|---|---|
| Survey mean | No | Nothing. It only contains respondents. |
| Bottom-end reply count | No | Nothing, but it is honest about being a floor. |
| Repeat purchase rate | Yes, every customer | The silent customer who simply never comes back |
| Rework and callback count | Yes, every job | Work that failed, whether or not anyone complained |
| Review click-through | No, only those asked and only if they act | Nothing extra. It is a narrower gate than the survey. |
| Request-to-job conversion | Yes, every customer who asked for work | A customer who asked, waited, and went elsewhere |
Read down the right-hand column. Every number that can see a silent unhappy customer is a behaviour number, not an opinion number. That is the practical takeaway: when the survey mean holds and a behaviour number falls, believe the behaviour number, and go looking at who your survey is not asking.
References
- See related: Net Promoter Score for Service Businesses, for the per-channel response-rate benchmarks
- See related: The Loyalty Score and Why a Middling Answer Counts Against You
- See related: The Effort Score and What It Predicts That Satisfaction Does Not
- See related: The Satisfaction Score Held While the Repeat Rate Fell
- See related: Service Request Conversion and How Long You Take to Answer