The Customer Health Score a Small Shop Can Actually Keep

Why this matters

Customers almost never announce that they are leaving. They go quiet, and quiet is indistinguishable from content right up until the moment they are somebody else's customer. Meanwhile the owner's attention flows to whoever is making noise, which means the loudest accounts get worked and the disengaged ones get lost. A health score is the mechanism that redirects attention from noise to risk. Every large operation runs one. Most small shops assume they need software for it, so they run nothing, and lose customers they would have kept for the cost of one phone call made a year earlier.

Health is not value, and mixing them ruins both

Value asks: is this relationship worth spending on? It is retrospective, level-based, and built from billed work, realization, and referrals. That is a separate audit and is covered in the reference at the end of this card.

Health asks: is this relationship at risk of ending? It is forward-looking, trend-based, and built almost entirely from dates and responses.

The two answer different questions and they combine, they do not merge. A high-value healthy customer needs nothing from you. A low-value declining customer is not worth chasing. The customer the whole exercise exists to find is the high-value declining one, and you can only see that customer if you have both numbers separately.

The specific failure from merging them: any score that includes total revenue as an input will never flag a big-spending customer as at risk, because the revenue input keeps carrying them. The customers most expensive to lose become the ones the score cannot see.

Five design constraints for a score you will still be running in a year

  1. Every input computable from data you already have. Dates on invoices, dates on jobs, contacts logged, estimates written. Nothing that requires you to have been running a program you were not running.
  2. No more than five inputs. Six is where a monthly recalculation starts getting deferred.
  3. An integer output on a fixed range, so two people scoring the same customer get the same answer and so trends over time are comparable.
  4. An action attached to every band. A score that does not trigger anything is a spreadsheet hobby.
  5. At least two trend inputs, meaning inputs that measure change rather than level. Level inputs tell you what a customer is. Trend inputs tell you what they are becoming, and that is the entire purpose.

The five inputs

Each scores 0 to 3. Maximum 15.

Input 3 2 1 0
Recency against expected interval Inside the interval 1.0 to 1.25 times the interval 1.25 to 1.5 times Over 1.5 times
Interval trend (current gap against the median of their prior completed gaps) Same or shorter Up to 25% longer 25 to 50% longer Over 50% longer
Response to contact (last three non-operational contacts) Responded or acted on two or more One None, but visits continue None, and a due date was missed
Quote acceptance trend (last three estimates) Accepted two or more One None, no reason given None, and price was cited
Behaviour change (payment and friction, change not level) Stable and inside their own normal Stable but outside terms Slower than their own history A new dispute, discount demand, or complaint

Read the last input carefully, because it is the one most often scored wrong. It measures change against that customer's own history, not against your terms. A commercial account that has always paid at 45 days and still does scores 2, because nothing changed. The same account moving from 45 days to 70 scores 1, even though 70 days is still within many portfolios' normal range, because the movement is the signal.

Inputs one and two do all the heavy lifting and cost nothing to maintain, because dates advance on their own. A customer's score degrades automatically as they go quiet, with no one doing anything. That is what makes this an early-warning system rather than a survey.

The bands and what each one triggers

Score Band Trigger
13 to 15 Healthy Nothing. Normal cadence. Do not spend retention effort here
9 to 12 Watch The next contact is a live phone call from a person, not a message
5 to 8 At risk Owner or senior call within 30 days, with a specific anchor from the file
0 to 4 Slipping away Treat as lapsed. Hand to the lapsed-customer process rather than the cadence

The Watch band's rule is the most useful line in the table. It costs almost nothing, because you were going to contact them anyway, and it swaps a broadcast for a conversation exactly where a conversation still works.

Three customers, scored

All values illustrative. Annual-interval trade, so the expected interval is 12 months.

Customer 1, the steady one. Last visit 10 months ago, which is 0.83 of the interval, so recency scores 3. Prior completed gaps of 12, 13, and 12 months give a median of 12, and the most recent completed gap was also 12, so interval trend scores 3. Responded to two of the last three contacts, so 3. Three estimates in the file, two accepted, so 3. Pays in about 8 days against net 15, unchanged for years, so 3. Total 15 of 15, Healthy. No action.

Customer 2, the quiet one. Last visit 16 months ago. That is 1.33 times the interval, which falls in the 1.25 to 1.5 band, so recency scores 1. Prior completed gaps of 11, 12, and 12 months give a median of 12; the current open gap of 16 months is about 33% longer, which lands in the 25 to 50% band, so interval trend scores 1. No response to any of the last three contacts, and a due date passed, so 0. Two estimates in the last three years, neither accepted, no reason recorded, so 1. Last invoice paid in 9 days, consistent with their own history, so 3. Total 6 of 15, At risk. Owner call within 30 days.

Customer 3, the loud one. Last visit 5 months ago, so 3. Gaps stable against their own median, so 3. Responded to all three of the last contacts, so 3. Last three estimates all declined with price cited on two of them, so 0. Last invoice disputed and a discount demanded, so 0. Total 9 of 15, Watch.

Now read what the three scores actually say, because this is the whole payload of the card.

Customer 2 has a perfect payment record and an entirely uneventful file. Nothing in it looks like a problem. The three inputs that flag them, recency, interval trend, and contact response, scored 1, 1 and 0, a combined 2 out of a possible 9. The two inputs that look at how they behave when they do transact scored 1 and 3, a combined 4 out of a possible 6. A shop reading only that second half, which is what most shops read, sees a customer who pays fast and never complains and concludes everything is fine. The customer is most of the way out the door.

Customer 3 occupies more of the owner's attention in a month than Customer 2 has in three years, and scores 9 against Customer 2's 6. That ordering is correct and it is counterintuitive. Customer 3 is engaged, arguing, and still transacting, which is a live relationship with a pricing problem you can address. Customer 2 is disengaged, which is the shape of a customer who has already made a decision they have not told you about.

Ranked by how much attention they consume, the order is 3, then 1, then 2. Ranked by risk of loss, it is 2, then 3, then 1. Those two orderings are close to inverted, and that inversion is why the score exists.

What deliberately stays out of the score

Total revenue or job count. That is the value axis. Including it here means your largest customers can never score at risk, which removes the only cases worth the score's existence.

Satisfaction survey results. Response bias runs exactly the wrong way: the customers drifting away are also the ones who do not answer surveys, so a survey-weighted score is blind to the population it is supposed to detect. Survey results are useful for improving service and useless for predicting churn in a small sample.

Tenure. A long relationship is a value signal, not a health signal. If anything, long-tenure customers churn more quietly, because both sides have stopped paying attention, and a tenure input would mask exactly that.

Anything requiring new data collection. An input that needs a field nobody fills puts a hole in the score within two months, and a score with a hole gets abandoned rather than corrected.

Gut feel, in any weighting. Gut feel is what produced the attention ordering in the worked example, where the loudest customer got the most concern and the quietest one was invisible. The score's job is to disagree with gut feel; giving gut feel a vote defeats it.

Keeping it current without it becoming a project

Recompute all five inputs quarterly for the whole list. Recompute the two date-driven inputs monthly for anyone currently in Watch or At risk, since those are the ones whose scores are moving and the two inputs update themselves.

For a list of a few hundred customers this is about an hour a month if the record already holds last-visit dates and an interval per customer. If it takes much longer, the bottleneck is almost always that intervals are not recorded per customer and someone reconstructs them each time. Fix that once as a data task and the ongoing cost collapses.

Keep the previous score. A customer moving from 14 to 11 is more informative than either number on its own, and a shop that overwrites scores loses the only trend data the system produces.

What changes the design

Genuinely episodic trades make recency and interval trend nearly meaningless, since there is no honest interval to measure against. Drop those two inputs to a maximum of 1 point each, rescale the bands, and put the weight on contact response and quote acceptance, which still work. Do not keep a 15-point scale whose first six points are noise.

Commercial and managed accounts break the behaviour-change input, since payment timing is set by an accounts-payable policy rather than by sentiment. Replace it with two things that do predict loss in that setting: proximity to a contract renewal or rebid date, and turnover in the contact role. A new facilities manager is a bigger churn risk than any payment pattern.

Plan and agreement members have a contractual visit schedule, so recency is not a choice they made. Replace the recency input with visit completion against entitlement: how many of the visits they are owed have actually been delivered this term. An underserved member scores as at risk, which is correct, and it is the situation that produces cancellations at renewal.

A shop under about 60 customers does not need this. The owner can genuinely hold sixty relationships in their head. Somewhere between about 100 and 200 records, the head-held version starts failing silently, and it fails first on exactly the quiet customers the score is built to catch.

How to verify the score predicts anything

Back-test it, once, and the test is cheap. Reconstruct the score you would have computed 18 months ago for a sample of 40 customers, using only the data you had at that time. Then check what actually happened to each of them.

You are looking for one thing: what share of the customers who actually churned were in the At risk or Slipping away bands 18 months earlier. If most of them were, the score works and you should run it. If churn was spread evenly across all four bands, the score is not measuring anything, and the usual cause is a wrong expected interval, which corrupts two of the five inputs at once. Recheck the intervals before you conclude the method failed.

The second check is for the opposite error. Count how many Healthy-band customers churned. A few is normal, since some losses are situational. A lot means your inputs are too generous, most often because contact response is scored against contacts nobody logged, so everyone gets credit for a response that was never recorded.

References

  • U.S. Small Business Administration (SBA), customer retention and small business analytics guidance
  • Standard cohort and survival analysis practice as applied to customer churn measurement
  • See related: How to Tell Which Customers Are Worth Keeping; The Warning Signs a Customer Relationship Is Going Bad; The Lapsed Maintenance Customer SOP; Grading Your Customers A, B, C, and D