The Fault That Changed Character Instead of Going Away
Why this matters
The most expensive callback in a shop is not the one where the repair failed. It is the one where the repair half worked, everybody read the improvement as progress, and the system was left in a worse condition than before anyone touched it while the paperwork said "resolved." Nobody lied. The event count really did fall. It fell because the fault moved, and the thing it moved to was worse than the thing it moved from.
This is one job, followed through, because the reasoning is what transfers. The numbers matter less than the habit of asking what the improvement was measured on.
The call that looked like progress
A system was tripping out and restarting itself. Mid-morning, moderate demand, on 6 of the last 7 days. Each event was brief, the system recovered on its own, and the customer's actual complaint was the noise of it cycling rather than any loss of use.
A tech found a restricted path, cleaned the part of it that could be reached through the service access, and closed the job. The differential measured across that path had been 3.2 times the reference reading for a clean assembly of that type. After the cleaning it read 2.1 times the reference. That is a 34% reduction, and it is real work, honestly done.
Three weeks later the customer called back, and what they said on the phone is the whole article: "It hardly ever does it now, but when it does, it stays off."
The number everyone quoted
The office logged the callback and pulled the history. Event days had gone from 6 in 7 to 1 in 7, which is 83% fewer event days. On that number alone, the repair looks like a success that just needs finishing.
The dispatcher nearly booked it as a minor follow-up on that basis. What stopped them was one field in the intake note: whether the system recovered by itself. Before the repair, it recovered on its own every time. After the repair, it recovered on its own never, and required someone physically present to reset it.
So run the arithmetic on time out of service instead of on event count. Before, roughly 0.2 hour per event across 6 events a week, which is 1.2 hours a week unavailable. After, roughly 3.5 hours per event because the system stays down until somebody gets home, across 1 event a week, which is 3.5 hours a week unavailable.
Weekly downtime went from 1.2 hours to 3.5 hours. That is 2.9 times worse, on the same repair that produced the 83% improvement everyone was quoting. Both numbers are correct. They describe different things.
A fault has four dimensions, and a partial repair moves them independently
This is the portable part. Any fault can be characterized on four axes, and "better" is only meaningful once you name which axis you measured.
- Frequency. How often, per unit of exposure. Per run-hour or per cycle, not per calendar week, because a system that ran half as much last week will show half the events without anything having changed.
- Severity and recovery mode. What it costs when it happens, and critically, whether the system recovers unattended, needs a manual reset, or stays down until a truck rolls. A shift from self-clearing to lockout is a large worsening that shows up nowhere in an event count.
- Trigger condition. What has to be true for it to happen. Moderate demand versus peak demand, cold versus warm, one system running versus two. A trigger that narrows to peak-only means the fault now waits for the moment the customer needs the system most.
- Location or presentation. Where it shows up and what it looks like. A fault that stops presenting at one component and starts presenting at another has usually not been reduced, it has been relocated, and the new location is often further from the cause.
Score the repair on all four. This one improved frequency, worsened recovery mode, narrowed the trigger to the worst condition, and left the location alone. One better, two worse, one unchanged.
Why the axes moved in opposite directions
The mechanism is simple once you see it, and it is the same in every trade.
The path was restricted by an amount that put it over the protective trip point at moderate demand. Cleaning the accessible half took the differential from 3.2 times the clean reference to 2.1 times. The protective device on this system acts at about 2.5 times the reference.
At 2.1 the system now sits below the trip point at moderate demand, so the mid-morning events stopped. But demand itself adds to the differential, and peak demand on this system adds roughly 0.6 times the reference. At peak, 2.1 plus 0.6 is 2.7 times, which is above the 2.5 trip point. So it still trips, only now at peak.
And peak-load trips on this equipment are a lockout requiring a manual reset, not the self-clearing cutout that happens at moderate load, because the control treats a trip under high demand as a more serious event. That single fact converted an 83% frequency improvement into a near-tripling of downtime.
None of this was visible from the event count. All of it was visible from the differential, which was still more than double the clean reference after the repair, on a path where the trip point is 2.5 times. A reading of 2.1 against a trip point of 2.5 has only 0.4 of margin, and demand alone eats more than that.
Partial repairs cluster at access boundaries, and access is not a diagnostic boundary
Nothing about this job was careless, and that is worth sitting with, because it explains why this pattern is common rather than rare.
The path had two halves. One was reachable through the service access in a few minutes. The other required pulling an assembly, which is a longer job, sometimes a second visit, and sometimes a part you do not carry. The tech did the half that the equipment was designed to let him do. Every trade has this shape: the accessible half of a duct run, the reachable half of a drain, the terminations you can see versus the splice inside the wall, the surface of a coil versus its depth.
So partial repairs happen where access ends, and the fault does not care where access ends. Those two boundaries have no reason to line up, and when they do not, you get exactly what happened here: a measurable improvement that stops short of the acting threshold.
Two habits close the gap. First, when you can only reach part of a fault path, measure the result across the whole path, not across the part you worked on. Cleaning half a path and confirming that half is now clean tells you about your own work. Measuring the differential end to end tells you about the system, and it is the only measurement that can tell you that you are not finished.
Second, say the boundary out loud and put it in writing before you leave. Something like: "I got the part I can reach through this access, and the reading came down. It is still above where this system acts, which means the rest of it is in the section I would have to pull the assembly to reach. Here is what that takes." Written that way, the customer owns a real decision. Written as "cleaned and tested, operating normally," the customer owns a repair they think is finished, and your shop owns the callback that follows.
The commercial version of the same point: a job closed at the access boundary bills once and returns. A job closed at the diagnostic boundary bills the real scope and does not.
The reading that should have closed the first visit
The first tech did competent work and stopped one measurement short. The missing step was not more cleaning. It was comparing the post-repair reading to the trip point rather than to the pre-repair reading.
Measured against where it started, 2.1 is a big win. Measured against where the system acts, 2.1 is a fault waiting for a hot afternoon. The margin, not the improvement, is what determines whether you are done.
That gives a rule worth carrying: on any repair that reduces a measured quantity toward a threshold, record the post-repair value against the threshold, not against the pre-repair value, and require real margin before you call it complete. On a path like this one, coming back within about 1.5 times the clean reference leaves room for demand swings. Anything above roughly 2 times is a return visit with a date on it.
Finishing it, and how they knew this time
The second visit reached the inaccessible half of the path, which required removing an assembly the first visit had worked around. The differential came back at 1.2 times the clean reference. At peak demand, adding the same 0.6, that is 1.8 times, which sits 0.7 below the 2.5 trip point.
Confirmation was not "no events since." It was three things, in order:
- The margin was measured, at peak, not inferred from a moderate-load reading plus optimism. They created the peak condition and read it there.
- The recovery mode was tested deliberately. They confirmed the system's response at the loads the customer actually runs, so nobody was relying on the assumption that a lockout would not recur.
- The event count was tracked per run-hour for the following period, not per calendar week, so a cool stretch with low runtime could not produce a false clean result.
That third one is the trap on the back end of every one of these. A fault that fires once per 30 run-hours will produce zero events across two quiet weeks, and everyone involved will believe the repair worked.
The habit this leaves you with
When a customer says "it is better," treat that as a question, not an answer, and ask it back on all four axes: how often, how bad, when, and where. Then compare each answer to the same axis before the work, per unit of exposure rather than per week. If one axis improved and any other axis worsened, you did not reduce the fault. You moved it, and the burden is on you to show the move was toward safety and availability rather than away from them.
References
- Manufacturer documentation for protective trip points, reset behavior, and the reference values for a clean or new assembly
- Trade-standard practice for measuring a post-repair value against the acting threshold rather than against the pre-repair value
- See related: Reading the Order Symptoms Appeared In
- See related: How to Separate Runtime-Driven Faults From Ambient-Driven Ones