How to Tell Partial Progress From a Moved Symptom
Why this matters
"It is a lot better" is the most dangerous sentence a customer says, because it closes jobs. It closes them whether the fault shrank or whether it relocated into a form that fires less often and hurts more when it does. You cannot tell those apart from a conversation, and you cannot tell them apart from an event count either. You can only tell them apart by comparing a structured description of the fault before and after, on the same axes, over comparable exposure.
So build the description as an artifact. This procedure is that artifact, constructed one field at a time. Each field exists because a specific way of being fooled exists, and skipping a field reopens exactly that hole.
The sheet: eight fields, and why each one is there
Fill it once when you first meet the fault, and again after any work. Two filled sheets are what makes the comparison possible; one filled sheet is a note.
Field 1: Exposure denominator
The amount of running the system did during the observation window. Run-hours, cycles, starts, gallons, whatever the equipment counts natively. Calendar days only as a last resort.
Why it is first: every other field is meaningless without it. A system that ran half as much this month shows half the events with nothing having changed, and that is the single most common false positive in field service. Get the counter reading at the start and end of the window and record the difference.
Field 2: Event count within that window
Raw count, plus how it was obtained. Do not convert it to a rate yet. Keep the count and the denominator separate on the sheet so anyone can recompute.
Why: a rate written down as a single number cannot be audited later. A count over a denominator can.
Field 3: Conditions present at the event
What else was running, what the load was, ambient conditions, time of day, who was home, what had just finished. As specific as you can get.
Why: this is where a narrowed trigger shows up. A fault that used to fire at moderate load and now fires only at peak has not improved, it has learned to wait for the worst moment, and the field is the only place that becomes visible.
Field 4: Onset shape
Sudden with no warning, or gradual with a build-up. If gradual, what the warning was and how long it ran.
Why: a fault that loses its warning period has worsened even if the count fell, because the customer no longer gets the chance to act. This field also tells the next tech whether a live capture is realistic.
Field 5: Presentation and location
What the observer actually perceives, and where. A noise at one end, a reading out of range at a specific point, a leak at a specific joint, a device that stops responding.
Why: a symptom that jumps to a different component or a different part of the system is the clearest possible sign of relocation rather than reduction, and it is very easy to lose when the second person to write the sheet describes it in their own words instead of the original words.
Field 6: Recovery mode
Pick one: recovers unattended, needs a manual reset by an occupant, needs a truck. Nothing else goes in this field.
Why: this is the field that catches the moved symptom that everyone else misses. A shift from self-clearing to lockout is a serious worsening and it never shows up in a count. Constraining the field to three values stops it being filled with prose that hides the change.
Field 7: Out-of-service time per event
How long the system is unavailable each time, in hours. Estimate honestly and say it is an estimate.
Why: multiplied by the count and divided by the denominator, this is the number the customer actually experiences. Field 2 measures how often you are annoyed. Field 7 measures what it costs.
Field 8: Observer and method
Who saw it and by what means: the control's stored history, an instrument you left in place, a customer report, a photograph, a recording.
Why: two windows measured by different methods are not comparable, and this is the most common way an honest comparison goes wrong. Customer reports undercount because people stop reporting a thing they have gotten used to.
The comparison rule
Normalize first. Divide the event count by the exposure denominator, both windows, same unit. Then apply the rule as written:
Compare per unit of exposure, field by field. Call it progress only when at least one field improved AND no field worsened. If any single field worsened, it is a moved symptom regardless of how far the count fell. If nothing improved and nothing worsened, the fault is unchanged and the work you did was unrelated to it.
The unit of analysis is one fault on one system, not the system's general health, and the Boolean is AND, not OR. One improved field beside one worsened field is not a partial success, it is a relocation, and it gets handled as an open fault.
One more gate before you compare at all: the post-repair window must cover at least 3 times the pre-repair mean interval between events before a low or zero count means anything. Below that, a clean window is exactly what you would expect even if nothing had been fixed, so reporting it as success is a guess dressed up as a result.
The filled-in pair
A fault that shuts the system down. Both windows read off the control's own stored history, so field 8 matches on both sides.
| Field | Before the work | After the work |
|---|---|---|
| 1. Exposure denominator | 240 run-hours | 96 run-hours |
| 2. Event count and source | 8, control history | 2, control history |
| 3. Conditions at event | Moderate load, various times | Peak load only, late afternoon |
| 4. Onset shape | Gradual, a noticeable build over several minutes | Sudden, no warning |
| 5. Presentation and location | Same component each time | Same component each time |
| 6. Recovery mode | Recovers unattended | Needs a manual reset by an occupant |
| 7. Out-of-service per event | About 0.3 hour | About 2.0 hours |
| 8. Observer and method | Control history | Control history |
Now normalize and read it.
Mean interval. Before: 240 run-hours divided by 8 events is one event per 30 run-hours. After: 96 divided by 2 is one per 48 run-hours. The interval is 60% longer, which is a genuine improvement on field 2.
The gate. The pre-repair mean interval is 30 run-hours, so the post-repair window needs at least 3 times that, or 90 run-hours, before its count carries weight. The window ran 96 run-hours, so it qualifies. Had it run 60 run-hours, the 2 events would still be reportable but the improvement would not be, because a shorter window is not evidence.
Downtime, which is what the customer feels. Before: 8 events at 0.3 hour is 2.4 hours across 240 run-hours, which is 1.0 hour of downtime per 100 run-hours. After: 2 events at 2.0 hours is 4.0 hours across 96 run-hours, which is 4.2 hours per 100 run-hours. The customer's downtime is about 4.2 times worse on the same repair that made the interval 60% longer.
Field by field. Field 2 improved. Field 3 worsened, narrowing to peak load. Field 4 worsened, losing the warning period. Field 5 unchanged. Field 6 worsened, from unattended recovery to a manual reset. Field 7 worsened, from about 0.3 hour to about 2.0 hours.
One field improved. Four worsened. Under the rule, this is a moved symptom and the fault is open, and that verdict does not depend on anybody's judgment about how much a lockout matters. It falls straight out of the sheet.
What a shop without the sheet does here. It quotes the count, 8 down to 2, calls it a 75% improvement, and books a short follow-up. That is a defensible-sounding number computed on raw counts over unequal windows, and it is wrong twice: the windows differ by a factor of 2.5 so the counts were never comparable, and the axis it improved on is not the axis the customer is living with.
What the sheet cannot settle
Four situations produce a filled pair that looks clean and is not. Recognize them and say so on the sheet rather than letting the comparison stand.
- The two windows were measured by different methods. A pre-repair count from a control's stored history against a post-repair count from customer reports will nearly always show improvement, because people stop reporting. If field 8 does not match on both sides, the comparison on field 2 is void and you fall back to fields 3 through 7, which are less sensitive to reporting.
- The duty changed, not just the amount. Run-hours normalize how long a system ran, not how hard. A mild season and a severe season can produce identical denominators under completely different stress. When the duty shifted, note it, and weight field 3 more heavily than field 2.
- The repair changed the recording. A replaced control clears history and may log different events under different names. Counters reset. When the instrument itself was part of the repair, the pre-repair count is not comparable to anything and you need a fresh baseline window before you can claim any result.
- The customer's behavior changed. Once an occupant learns where the reset is, they stop calling and start resetting, and the fault vanishes from every record you have. Ask directly at the follow-up: "Have you been resetting it yourself?" That one question recovers a whole field that otherwise silently reads as improvement.
In all four, the correct entry is not a guess. Write "not comparable" in the field and say what you would need to make it comparable. A sheet with an honest gap in it is worth more than a complete one that quietly compares two different things.
References
- Manufacturer documentation for what the control's stored history records, how deep it goes, and what a replacement or reset clears
- Trade-standard practice for normalizing fault rates against a machine-side exposure counter rather than the calendar
- See related: The Fault That Changed Character Instead of Going Away
- See related: How to Log an Intermittent Fault Over Days