How to Track Whether a Customer Actually Came Back
Why this matters
A shop runs a reactivation push, 31 old customers book work, and everyone agrees it worked. The number nobody has is how many of those 31 would have called anyway. If the honest answer is 18, then the push produced 13 jobs, not 31, and every decision made on the strength of the bigger number - how much time to spend, whether to discount, whether to hire - is built on an inflation of more than double.
This is the measurement discipline that separates outreach you can improve from outreach you just repeat. It costs almost nothing to run: a definition, a window, a holdout group, and the discipline to write down the non-returns as carefully as the returns.
Step 1: Define the return event so two people count it identically
Write the definition down before anything goes out. The usable one is narrow: a completed and paid job at that customer record, initiated within the attribution window.
Then define, explicitly, the three things that are not a return but get counted as one:
- A booked appointment. Bookings cancel and reschedule out. A booking is a milestone worth tracking on its own line, never the success line.
- An estimate delivered. An estimate is an opportunity, not a return. If your trade runs long from estimate to work, track estimates as a separate stage with their own conversion rate rather than folding them in.
- A reply. "Good to hear from you, I'll call when I need something" is a positive signal and zero revenue. Count it in the reply column.
Track the whole pipeline as five columns: records worked, people reached, replies, jobs booked, jobs completed and paid. The drop between booked and completed is where a lot of reported success quietly evaporates, and you will never see it if success is a single number.
Step 2: Set the attribution window before the campaign, not after
The window is how long after a touch a job still counts as caused by that touch. Pick it before you send anything, because picking it afterwards means picking whichever length makes the result look best, and everyone does that unconsciously.
A 90-day window is a reasonable default for most repeat service work and is what the example below uses. Tune it to your trade: if your typical time from decision to scheduled work runs longer than a month, widen it; if your work is same-week, a 30-day or 45-day window will be cleaner and less contaminated.
The tradeoff is real in both directions. A short window undercounts genuine returns from people who needed to wait for a paycheck or a weekend. A long window increasingly counts jobs that would have happened regardless, because eventually every customer who was going to call calls. The holdout group in step 5 is what protects you from the long-window version of that error, which is why the two decisions belong together.
Step 3: Record the touch on the customer record, not only in a campaign list
Every touch gets written to the customer's own record with a date, a channel, and who sent it. This sounds like bookkeeping and it is the load-bearing step, for three reasons.
The office needs it at the moment the customer calls back, so whoever answers knows why. A customer responding to a note they got two weeks ago, met with "can I get your address?", has just learned the outreach was not personal.
Second, it is the only way to avoid double-touching somebody who is already in another sequence, and a customer who receives two unrelated messages from you in one week reads it as automated regardless of how well each one was written.
Third, campaign lists get deleted, exported, and lost. Six months later the only durable place the touch exists is the customer record, and without it you cannot reconstruct which customers were in which arm.
Step 4: Know your baseline return rate
The baseline is what share of lapsed records come back on their own in a given window with no contact from you. Most shops assume it is zero. It is not, and in the just-past-due band it can be substantial, because those customers were always going to call when the thing broke.
You can get it two ways. Retrospectively, look at a past period with no outreach and count how many lapsed records initiated paid work in a 90-day stretch. Or prospectively, from the holdout group in the next step, which is better because it is contemporaneous and matched.
Compute the baseline separately for each overdue band. A just-past-due record and a three-cycles-gone record have very different natural return rates, and one blended baseline will overstate the lift on one and understate it on the other.
Step 5: Hold a group out, and keep the holdout honest
Randomly set aside a portion of the eligible list and touch none of them. They are your counterfactual: whatever they do is what the touched group would have done without you.
Three rules make a holdout valid:
- Random selection. Not the ones with the worst phone numbers, not the ones nobody liked, not the leftovers after the good ones were picked. Any selection rule that correlates with likelihood of returning destroys the comparison and always in the direction that flatters the campaign.
- Same measurement window and same eligibility rules. The holdout is measured over the identical 90 days, using the identical definition of a return.
- No accidental contact. If a holdout record gets a due reminder from another process, it is no longer a holdout. Flag them so other sequences skip them for the duration.
On size: a holdout of about 20% to 25% of the eligible list is the usual compromise, big enough to be readable and small enough that you are not withholding contact from a large share of the list. Below roughly 30 records the holdout rate swings so much on one or two customers that it will not settle an argument, and you should treat the result as directional only.
Yes, you are choosing not to contact people who might have booked. That is the price of knowing whether the program works, and it is cheap compared to running an ineffective program for three years because you never had a comparison.
Step 6: Work one cycle end to end
A shop has 180 eligible lapsed records after cleaning. It randomly holds out 40 and works the remaining 140. The attribution window is 90 days from the first touch, set in advance.
Touched arm. Of the 140 records worked, 31 produced a completed and paid job inside the 90 days. That is a 22% return rate on records worked.
Holdout arm. Of the 40 held out, 5 initiated paid work in the same 90 days with no contact at all. That is a 12.5% return rate on records held out.
Lift. 22% minus 12.5% is a lift of about 9.6 percentage points, measured on records in each arm over the same 90 days.
Incremental jobs. Apply the lift to the touched arm: 9.6% of 140 records is about 13 jobs. So of the 31 completed jobs in the touched arm, roughly 13 or 14 are attributable to the outreach and the other 17 or 18 would most likely have arrived anyway. Stated as a share, a bit over 4 in 10 of the bookings the campaign claimed were actually caused by it.
The naive reading and why it fails. Without the holdout, the shop reports 31 jobs from 140 records and concludes the program produces a 22% return. The real number is closer to 9.6%, and the difference is not a rounding issue, it is more than half the claimed result. Every downstream decision - the time budget, whether the offer was needed, whether to scale it - was going to be made on a figure more than double the truth.
The uncertainty, stated honestly. The holdout has 40 records, so a single customer moves its rate by 2.5 percentage points. If 6 of the 40 had booked rather than 5, the holdout rate would read 15% and the lift would be about 7.2 points instead of 9.6. So the finding is "there is a real positive lift, somewhere in the high single digits to low double digits," not "the lift is 9.6." Report it that way, and let two or three cycles of the same measurement tighten it.
Step 7: The three ways lift gets computed wrong
No holdout at all. The entire touched-arm return rate is credited to the campaign. This is the default in almost every shop and it is the error the whole article exists to prevent.
A holdout that was not random. If the holdout was built from the records with the worst contact data or the longest lapse, its natural return rate is lower than the touched arm's would have been, and the lift is inflated by the difference. The tell is a holdout whose composition, checked afterwards by overdue band, does not match the touched arm's.
Mismatched windows. Touched arm counted over 90 days, holdout counted over everything that happened since. Or the touched arm's window starting at first touch while the holdout's starts at the campaign date. Same start rule, same length, both arms, every time.
Step 8: Record the non-returns as carefully as the returns
The 109 touched records that did not book are the more useful half of the data, and they only become useful if each one carries a disposition: reached and declined, reached and deferred with a stated future date, no answer after the attempt cap, bad number, do not contact.
Dispositions are what make the next cycle cheaper. A record dispositioned "reached, replaced their equipment last year with somebody else" never needs to be called again for that service. A record dispositioned "no answer, number rings" is a data-quality task, not an outreach failure. Without dispositions, next year's list is the same list, and you will pay to learn the same things again.
A do-not-contact request is permanent, applies across every future list build, and is not a disposition anyone gets to overrule later.
Step 9: Feed the result into the next cycle
Compute the lift separately by overdue band. The bands almost never respond alike, and a blended lift will tell you to keep doing the same thing everywhere. Where lift is near zero, stop touching that band and put the attempts into a band where it is positive.
Keep the holdout on every cycle, not just the first. Programs decay: the same message to the same list produces less each time, and a shop without a running counterfactual will not notice the decay until the program is producing nothing.
What changes the answer
- Very small lists. Under about 60 eligible records, a holdout leaves both arms too small to read. Skip it, touch everyone, and instead compare against a retrospective baseline from a period with no outreach, labeling the result as weaker evidence.
- Long decision cycles. Where the work is a major replacement rather than a repair, 90 days is too short and a 6 to 12 month window is more honest, with the corresponding expectation that a larger share of the returns are natural.
- The touch was a hazard notification or a dropped-message apology. These are obligations, not campaigns. Never hold anyone out of them, and measure them separately or not at all.
- You changed two things at once. A new message and a new offer in the same cycle means the lift is real but unattributable to either. Change one variable per cycle or accept that you have measured the package.
How to verify you got this right
- The window length and the holdout list both have a timestamp predating the first touch. If either was decided afterwards, the result is a story.
- The holdout's composition by overdue band roughly matches the touched arm's. If it does not, the randomization failed and the lift is not readable.
- Every reported rate names its base and its window in the same sentence. "22% of records worked, over 90 days" survives being repeated; "22% return rate" does not.
- The completed-and-paid column, not the booked column, is the one being reported as the result.
- Every touched record carries a disposition, including the ones that went nowhere. A cycle that dispositioned only its wins has recorded the least useful half of what it learned.
References
- See related: The Customer Reactivation SOP (the cycle this measures)
- See related: How to Run a Win-Back Campaign That Does Not Feel Desperate
- See related: How to Decide Which Lapsed Customers to Call First
- Trade-standard practice for control-group testing in direct customer outreach