When Normal Behavior Signals a Real Problem Anyway
Why this matters
The most expensive miss in this whole category is not calling a fault normal. It is correctly identifying a behavior as normal and stopping there. A protective device that operates is working perfectly. A system that cycles is doing what it is built to do. A drain that carries water away is functioning. Every one of those can be simultaneously true and a report that something upstream is wrong.
The event is normal. The rate, the duration, the season, or the trend across time is the finding. A tech who only knows how to classify a single observation will close every one of these visits clean, and the failure that follows will look like it came from nowhere.
A normal event and a normal pattern are two different claims
When you tell a customer a behavior is normal, be clear in your own head about which claim you are making.
The single-instance claim is that the behavior you observed is a designed function of the system, produced by a mechanism you can name. That claim can be verified in one visit with the discriminating tests: timing within the cycle, comparison against an identical peer, trend across a run, persistence after shutdown.
The pattern claim is that this behavior, occurring at this frequency, for this duration, under these conditions, is what the system should be doing. That claim cannot be verified in one visit at all. It requires a baseline, and if you do not have one you should say the honest thing, which is that the behavior is a normal function but you cannot yet judge whether the rate is.
Most techs make the second claim while only having evidence for the first. That is the gap this card is about.
Five ways a normal behavior becomes a finding
Rate. The behavior is correct but happening far more often than it should. Every protective and relief function lives here. A device that opens, trips, discharges, or diverts is doing exactly its job, and the job description is "act when the system reaches a limit." Every activation is therefore a report that the system reached its limit. Once a year, that is a system with margin. Six times a year, that is a system living at its ceiling with no margin left, and the next thing to happen is a failure of whatever is being protected.
Duration. The behavior is correct but taking longer to complete than it used to, or longer than documented. A designed clearing, purging, or regeneration cycle that has stretched is telling you the thing it clears is accumulating faster or clearing more slowly. Compare against the published duration, not against your impression.
Season and context. The behavior is correct for a condition the system should not currently be in. A defrost or clearing cycle in the wrong season, a protective delay in mild weather, a full-output run at a load that should require a fraction of capacity. The behavior is textbook. The context is the fault.
Trend across visits. Nothing is wrong on any single visit, but the same number has moved in one direction across three of them. This is the most valuable finding a maintenance program produces and the one shops throw away by not recording numbers on routine visits.
Normal for a state the system should not be in. The behavior is exactly what the equipment does when a particular condition exists, and the real question is why that condition exists. Equipment behaving correctly in a bad state will read as fully normal to every test you run on the equipment.
The compensating-system trap
This is the dangerous one, and it deserves its own treatment because it defeats every discriminating test.
Many systems contain something that quietly absorbs a developing problem. A protective device that keeps resetting. A drain that keeps carrying away water that should not exist. A control that keeps modulating to hold a setpoint against a growing load. A redundant element carrying more than its share because its partner is degrading. In every case the compensating element is doing its job correctly, and because it is, the symptom never reaches the customer and never reaches your instrument as an out-of-range reading.
What you can observe is that the compensator is working harder. That is the finding, and it will always look like normal operation unless you specifically ask how hard it is working compared to before.
The rule worth carrying: when you find something operating correctly at the edge of its range, ask what it is compensating for. A system with no margin left is not a healthy system, it is a system whose failure has already been scheduled and just has not arrived.
The worked example
A commercial customer with a routine maintenance agreement. Nothing has ever failed on this system, no complaint has been made, and every previous visit closed clean.
Reading the file rather than the equipment produces three findings, none of which is a fault.
Rate. The protective discharge on this system had been noted in the maintenance record once per year for three consecutive years. In the most recent six months, it has been noted three times. Three events in six months is a rate of six per year, six times the prior rate. Each individual discharge was the device operating correctly. The rate is the finding.
Duration. The system's documented periodic clearing cycle is supposed to run at an interval close to 60 minutes under continuous cold-weather operation. Timed across three events on this visit, it is now running at roughly 35 minutes, about 40 percent shorter than the documented interval. Each cycle completes properly and the system resumes normally every time. The interval is the finding.
Trend. The current draw recorded at heat soak on the last three annual visits reads 4.9, 5.2, and 5.5 amps against a nameplate of 5.0 amps. Every one of those was written up as within tolerance at the time, and every one of those write-ups was correct. Across the three visits the reading has moved 0.6 amps, which is 12 percent of nameplate, in one direction, with no change in the installed load.
No single observation on any single visit was a fault. Together they describe a system that is doing more work to deliver the same output, reaching its protective limit six times more often than it used to, and clearing an accumulation roughly 40 percent more frequently than designed. Every one of those is what you would expect if the same underlying restriction were growing.
The deliverable is not a repair order for the protective device, and that is the point. Replacing the device that is reporting the problem removes the report and leaves the problem. The deliverable is a diagnosis of what is loading the system, backed by three independent trend lines that each point the same way, and a straightforward conversation with a customer who has never experienced a symptom.
What you need to make this call: a baseline
None of the above is available to a tech who takes no readings on visits where nothing is wrong. That is the entire practical lesson.
Record the same small set of numbers on every routine visit, at a stated runtime, at the same measurement points, whether or not anything is wrong. Current or load figure at heat soak. The differential across the working element. Duty cycle observed across at least two complete cycles. Any protective activation since the last visit, from the control history or the customer.
Five numbers, a few minutes of work, and they are worthless on the first visit and increasingly valuable on every visit after. That asymmetry is why shops skip them, and it is why the shops that do not skip them catch failures a season early.
What changes the answer
System age. A trend of 12 percent in one direction over three years on a system in its first decade is a developing fault worth chasing. The same trend on a system past its design life may be the expected slope of wear, and the honest conversation is about planned replacement rather than a hunt.
A change in load. If the building, the process, or the occupancy changed between baselines, the trend may be describing the load rather than the equipment. Check what the system is being asked to do before concluding something about what it is doing.
Who recorded the baseline. Numbers taken at different runtimes, at different measurement points, or with different instruments are not a trend, they are three unrelated numbers. A trend built from inconsistent records is worse than no trend, because it produces confident wrong conclusions.
Whether the protective activation was a true activation. A device that trips because it is itself failing is a component fault, not a report about the system. Verify the device operates at its rated point before you interpret its activations as evidence about anything else.
When you have no history at all
Most residential calls have no baseline, and you still have to make a call. Three substitutes work.
Compare against the published design values rather than against past behavior. Interval, duration, and rated load are all documented for most equipment, and a documented interval is a baseline somebody else recorded for you.
Compare against an identical peer on site. Two of the same unit, two of the same zone, two of the same leg. Asymmetry between things that should be identical needs no history.
Ask the customer for the history you do not have. They cannot give you a reading, but they can often give you a rate. "How often did this used to happen?" is a genuine trend measurement from an untrained instrument, and it is frequently the only baseline available. Take it as approximate and say so in the record.
The failure mode
It looks like this. A tech is called for a protective device that keeps activating. The device is confirmed to be operating at its rated point, correctly, every time. The tech explains, accurately, that the device is doing its job and there is nothing wrong with it, and closes the visit as normal operation. Everything he said was true.
Six weeks later the thing the device was protecting fails, and it fails hard, because the condition that had been driving the device to its limit never stopped growing. The record shows a visit that found nothing wrong, and it is technically defensible and practically indefensible at the same time.
The tell that was available on the day: the customer had said it used to happen about once a year and was now happening every couple of months. That sentence was a rate measurement, it was offered for free, and it was the whole diagnosis.
References
- Manufacturer documentation for rated activation points, documented cycle intervals, and permissible operating ranges
- ASHRAE and equivalent trade-standard practice for recording maintenance baseline readings on routine service
- OSHA guidance on protective and relief device inspection intervals and on treating repeated activations as a reportable condition
- See related: Normal Versus Abnormal: A Field Reference for Common Observations; The Baseline Reading You Should Always Take; Documenting a Maintenance Baseline