The Loyalty Score and Why a Middling Answer Counts Against You

Why this matters

The loyalty score is the one feedback number that behaves in a way nobody expects on first contact. A customer who tells you they would probably use you again, and means it warmly, contributes exactly nothing to it. A shop that everybody is mildly happy with scores close to zero. Owners who do not know that read a middling score as a warning and start fixing things that are not broken, or read a strong score as proof of a loyal base when it was produced by a small group of enthusiasts sitting next to a growing block of angry customers. The banding is the whole article. Get it and the number becomes useful. Miss it and the number misleads you in a consistent direction.

How the banding works

The conventional form is one question on a 0 to 10 scale: how likely the customer is to recommend you. The answers are cut into three bands, and the score is one share minus another share:

  • Top band, 9 and 10. Counted as a positive.
  • Middle band, 7 and 8. Counted in neither total.
  • Bottom band, 0 through 6. Counted as a negative.

The score is the top band's share of all respondents, minus the bottom band's share of all respondents, stated in points. Both shares run over the same denominator: every respondent, middle band included.

So the middle band does one thing only. It enlarges the denominator and shrinks both shares. That is the mechanism behind every surprise this number produces.

One warning before any arithmetic: those cuts belong to that 0 to 10 question and travel nowhere else. If your form runs 1 to 5 or 1 to 10, the 9-10 and 0-6 boundaries are not yours, and applying them anyway produces a number that looks like the published version and is not comparable to anything, including your own earlier figures. Pick a scale, write the bands down next to it, and never change either without restarting the series.

Two response sets, one headline

Two shops, each with 40 respondents in the quarter. Forty clears the 30-response floor a survey figure needs before it is quotable at all. See related: The Satisfaction Score and Who Actually Answers a Survey.

Top band (9-10) Middle band (7-8) Bottom band (0-6) Respondents
Shop A 20 12 8 40
Shop B 24 4 12 40

Shop A. Top share 20 of 40, 50.0 percent of respondents. Bottom share 8 of 40, 20.0 percent of respondents. Score 50.0 minus 20.0, which is +30.

Shop B. Top share 24 of 40, 60.0 percent of respondents. Bottom share 12 of 40, 30.0 percent of respondents. Score 60.0 minus 30.0, which is +30.

Identical headline. Shop B has 12 actively unhappy respondents against Shop A's 8, half again as many on the same respondent count, and it is running a business that produces enthusiasts and enemies rather than customers. Shop A is running a steadier one. If both owners quote +30 at each other, neither learns anything.

The count that actually differed is the bottom band, and the score threw it away by netting it against a top band that also differed. So never report this number without the three band counts printed beside it. One line, three counts, then the score. It costs nothing and it is the difference between a number and a summary of a number.

The bottom band is seven answers wide

The other thing the banding hides sits inside the bottom band, which spans 0 through 6. A customer who answered 6 was slightly let down and would probably call you again for something small. A customer who answered 0 is telling their neighbour. The score counts them as the same negative.

So split that band when you record it. Two shops, each with 12 bottom-band respondents out of 40, contributing an identical 30.0 percent of respondents to the negative side:

  • One splits 9 answers at 4-6 and 3 answers at 0-3. That is a shop with a persistent small irritation, usually scheduling, communication, or a job that ran long. It is fixable with process and nobody is out for blood.
  • The other splits 3 answers at 4-6 and 9 answers at 0-3. That is a shop with a real failure happening to a minority of customers, and those nine are actively costing it work it will never see quoted.

Same contribution to the score, same band count, opposite jobs for the owner on Monday. Recording the bottom band as two numbers rather than one costs you a column and answers the only question the band raises.

What a happy middle costs you

Take Shop A exactly as it stands and add twenty more respondents, every one of them a 7 or an 8. Nothing else changes: not one customer got angrier, not one enthusiast was lost.

Now there are 60 respondents. Top band still 20, which is 20 of 60, 33.3 percent of respondents. Bottom band still 8, which is 8 of 60, 13.3 percent of respondents. Score 33.3 minus 13.3, which is +20.

Twenty additional customers who were all reasonably pleased took ten points off the headline. That is not a distortion or an edge case, it is the number doing exactly what it is built to do: it does not measure satisfaction, it measures the balance of advocacy against hostility, and a contented customer is neither.

This is why a shop with a broad, stable, moderately-pleased customer base scores lower than a shop with a noisy following and a churn problem. Both facts are true and only one of them is a business you want.

Two shops with a zero

Two degenerate sets prove the banding faster than any realistic example can, because each one collapses the arithmetic to nothing.

  • Every respondent in the middle band. Forty respondents, all 7s and 8s. Top share 0.0 percent, bottom share 0.0 percent. Score 0.
  • Nobody in the middle band. Twenty top, twenty bottom, no middle. Top share 50.0 percent, bottom share 50.0 percent. Score 0.

The first shop has not upset a single customer. The second has upset half of everyone who answered. Both report a flat zero, and a reader who sees only the headline cannot tell which one just handed them a business.

That is the boundary of what this number can carry, and it is worth knowing before you put it on a wall.

What it is legitimately for

One thing: a direction, against your own history, on a fixed collection rule. Same question wording, same scale, same bands, same rule about who gets asked, same channel. Under those conditions a move of several points across consecutive quarters is real information, and the cause is usually visible in the band counts (the top band grew, or the bottom band grew, and those are different stories).

Anything that changes the collection rule restarts the series. Widening who gets asked, switching the question from text to email, shortening the form, moving the ask from one day after the work to a week after: each of those changes the population answering, and the new figure is a new series that happens to be printed in the same units. Label it as a restart rather than drawing a line through the break.

What it is not for

Not a cross-industry comparison. Published figures are collected on other people's populations, with other people's response rates, other people's question wording and frequently a third party doing the asking. A shop comparing its own quarterly figure against a published industry number is comparing two different measurements that share a name.

The benchmark tables in the references exist and are useful for one thing: telling you whether your figure is wildly outside what shops in your trade report, which is worth knowing once. What they cannot do is tell you whether a two-point move in your own figure is real, because the populations, response rates and question wording behind a published band are not yours. Use them as a sanity check on joining the scale, and your own trailing history for everything after that.

Not a target to manage toward directly. This is the important one, because the number has an obvious exploit: the fastest way to raise it is to change who gets asked, not to change anything a customer experiences. Suppress the ask on jobs that went badly and the bottom band shrinks within one cycle. The score improves, the business does not, and the series you built is now worthless for the one purpose it had.

Not a per-customer judgement. The bands are a property of the respondent set. Reading an individual 7 as "this customer is a problem" gets the direction of the whole instrument backwards; that customer is fine, and what the band tells you is that they will not go out of their way for you.

How big a move counts as a move

The score moves in chunks set by your respondent count, and at a small shop those chunks are large. One respondent shifting between adjacent bands changes a share by 1 divided by the respondent count. At 40 respondents that is 2.5 points per head. At 80 it is 1.25 points. At 30, which is the minimum respondent count worth quoting at all, it is 3.3 points per head.

So take two respondents' worth as your noise floor, which is 200 divided by your respondent count, expressed in points. At 40 respondents that is 5.0 points. A quarter-to-quarter move smaller than that is two people answering differently, and you will invent a reason for it if you let yourself.

These are two separate rules doing two separate jobs on the same unit, and both apply. The 30-response floor decides whether the score is quotable. The 200-divided-by-count floor decides whether a change in the score is a finding. A shop with 31 respondents clears the first and carries a 6.5-point noise floor on the second, which means it can publish a number and still cannot read most of its own movement. Rolling two quarters together fixes both at once.

Run it against the sizing below: Shop A has 40 respondents, so its noise floor is 5.0 points, and the ten-point move that follows clears it with room. The zero-point difference between Shop A and Shop B does not, which is another way of saying those two scores are indistinguishable even before you look at the bands.

Two ways to buy ten points

Here is where the netting becomes concrete and useful. In a 40-respondent set, one respondent is 1 of 40, so moving one person out of the bottom band into the middle band lifts the score by 2.5 points, and moving one person out of the middle band into the top band lifts it by 2.5 points as well. The two ends are worth the same per head, because the shares share a denominator. Moving one person from the bottom band straight to the top band moves it by 5.0 points, because it changes both shares at once.

So take Shop A at +30 and buy ten points. Two routes, both arithmetically exact:

Route Move Resulting bands Score
Convert passives 4 of the 12 middles become top band, 33.3 percent of that band 24 top, 8 middle, 8 bottom 60.0 minus 20.0, +40
Convert detractors 4 of the 8 bottoms become middle band, 50.0 percent of that band 20 top, 16 middle, 4 bottom 50.0 minus 10.0, +40

The score cannot tell those two routes apart. Your shop can, and they are not remotely the same work. Converting four passives means giving four already-satisfied customers a reason to be enthusiastic, which is usually a gesture: a follow-up call, a small courtesy, somebody remembering their name. Converting four detractors means finding four people who were let down and fixing what happened to them, which is usually a real operational defect with a cost attached.

A shop chasing the headline takes the first route every time, because it is cheaper and the four passives are easy to find. A shop reading the band counts takes the second, because four angry customers talk to their neighbours and the passives do not. Decide which route you are on before you look at the score, then use the score to check you moved the band you meant to move. If the score rose and the bottom band did not shrink, you bought points and left the problem in place.

References

  • See related: The Satisfaction Score and Who Actually Answers a Survey
  • See related: The Effort Score and What It Predicts That Satisfaction Does Not
  • See related: The Satisfaction Score Held While the Repeat Rate Fell
  • See related: Net Promoter Score for Service Businesses, for the published trade benchmark bands this card says to use once