The Remote Call That Should Have Been a Site Visit
Why this matters
Some faults cannot be diagnosed remotely no matter how good the questions are, and the tell is usually in the timing of the symptom rather than in its description. This is one of those calls, followed from the first phone contact through two wrong turns to the measurement that ended it. The lesson is not that remote diagnosis is bad. It is that a fault which only appears after a long continuous run is defined by a measurement taken while it is running hot, and there is no phone question that substitutes for that.
Safety first: what could not be asked for over the phone
Partway through the second call the obvious next step was to have the customer open the equipment enclosure and check whether anything looked loose or discolored. That instruction was not given, and it should never be. Behind that cover is energized wiring. A customer with a cover off and a flashlight is one slip from a contact they cannot see, and there is no reset for that.
On site, before any of the work described below, the circuit was isolated at its disconnect, locked, and verified dead with a meter on a known-live source first, then on the conductors, then on the known-live source again. Live-dead-live is the only test that catches a meter that died between readings. The one measurement in this case that had to be taken energized is called out below with the precaution that made it acceptable.
There is a second reason the phone instruction was wrong, and it matters diagnostically: a customer wiggling terminals destroys the evidence. A loose or oxidized joint that gets disturbed will often work fine for a week and then fail again, and now nobody can prove what it was.
The call: the symptom as reported
Residential customer, an existing system that had worked normally for years. Her description, close to verbatim: "It runs fine most of the morning and then it just quits. If I leave it alone about forty-five minutes it comes back on by itself and runs again."
Three facts were extractable from that and they were the whole foundation of the case:
- It runs first. This is not a no-start. Whatever fails is available at startup and becomes unavailable later.
- It stops after a long stretch, not a short one. She put it at "most of the morning," which on the follow-up question converted to roughly two hours of running.
- It recovers on its own after a consistent interval. Roughly forty-five minutes, unattended, no reset performed.
Self-recovery after a fixed cooling interval is the signature of something thermal. Nothing else resets itself on a clock like that. Either a protective device is opening on temperature and closing again when it cools, or a component is changing its behavior with temperature and coming back when it drops.
The first hypothesis, and why it was reasonable
The leading candidate was a restriction causing the system to work harder and eventually trip a protective device on temperature. That is far and away the most common cause of a run-then-quit-then-recover pattern, it is cheap to test, and it is often something the customer can address.
The supporting evidence was thin but real. She had not changed the filter element in longer than she could remember, and the equipment was in a space with obvious dust. A restricted filter raises the work the system does, raises internal temperatures, and produces exactly this shape of fault in many trades: restricted airflow, restricted water flow, a clogged strainer ahead of a pump.
The remote instruction given was to replace the filter element with the correct size and rating and to report back after a day of running.
What the remote instruction actually proved
She replaced it. Two days later she called back: it still quits, and now it seems to happen a little sooner, "closer to an hour and a half."
That result is worth reading carefully, because it is where the case nearly went off the rails. The obvious interpretation is that the change made things worse, which points at the new filter being wrong. The correct interpretation is that the change proved nothing about the filter and one important thing about the fault.
- It did not exonerate airflow entirely on its own, but combined with the next fact it did.
- The shift from about 120 minutes to about 90 minutes tracked the weather. The second period ran on warmer days. A fault whose time-to-trip moves with ambient in the same direction as the load is heat-related, which was already suspected.
- The trip interval stayed consistent within each condition. Not random. Not tied to what she was doing. That consistency is the strongest single fact in the whole case, and it rules out an enormous amount.
A fault that fires at a repeatable elapsed runtime is being driven by something that accumulates with runtime. Heat accumulates with runtime. Random mechanical failures do not.
The dead ends and what ruled each one out
Restriction on the flow path. Ruled out by the filter replacement producing no improvement, and independently by her report that the system was still delivering normally right up to the moment it quit. A restriction severe enough to trip a protective device on temperature almost always degrades output noticeably before it trips. Nothing degraded. It ran normally and then stopped, cleanly.
Ambient conditions. Considered, because the second failure window was warmer. Ruled out as the primary cause because the original failures occurred on cool days as well. Ambient was moving the trip time, which is real, but a fault that fires on a cool day is not caused by the weather. Weather was a modifier, not a cause. Keeping those two roles separate is what stopped this from becoming a "your system is undersized" conversation.
Undersizing or overload. Ruled out by the consistency of the trip interval. An overloaded system trips sooner when demand is high and later when it is low, and her demand varied considerably across those days. The interval did not vary with demand. It varied with runtime.
A weak protective device. A protective device that has aged and now opens below its rating produces this pattern too, and it was a live candidate right up to the site visit. What eventually ruled it out was a measurement, not reasoning: the condition upstream of it was already out of tolerance before the device opened, so the device was doing its job correctly on bad information.
A control component failing when heat-soaked. This was the leading hypothesis on arrival, and it is a real failure family. Components change behavior at temperature and recover when they cool. It was ruled out by measuring what was arriving at that component rather than what it was doing.
On site: the measurement that ended it
The visit was scheduled as a long block on purpose, because by then everyone knew the fault needed roughly 90 to 120 minutes of continuous running to appear. That scheduling decision is the single most useful thing that came out of the remote calls.
Sequence on site:
- Cold baseline first, isolated and verified dead. Every accessible connection in the control enclosure was checked for tightness by feel and inspected for discoloration. Two terminals showed a slight darkening. Nothing conclusive. Photographs taken before anything was touched, and nothing was retorqued yet, because retorquing at this stage would have erased the fault.
- A cold energized reading. With the cover secured and probing done only at designated test points, the voltage drop across the suspect section of the circuit was in the noise, call it under a tenth of a volt. Illustrative, but the point is that it was indistinguishable from a good joint.
- Run it and wait. The system was started and left to run. Temperatures inside the enclosure were spot-checked with a non-contact thermometer at intervals without opening anything that had to stay closed.
- The hot reading. At roughly 90 minutes, with the enclosure well heat-soaked, the drop across that same section had climbed to well over a volt, illustrative again, and the non-contact reading at one terminal was running clearly hotter than the identical terminals beside it, on the order of tens of degrees Fahrenheit above its neighbours.
- The fault reproduced. Shortly after, the system stopped, exactly as she had described.
That is the whole diagnosis. A high-resistance connection at one terminal. Cold, its resistance is small enough to be invisible. Under sustained current it heats, and the resistance of a degraded joint rises as it heats, which makes it dissipate more, which heats it further. That is a positive feedback loop with a clock on it, which is why the trip interval was so consistent. Eventually the voltage available downstream falls below what the load needs to stay engaged, the system drops out, current stops, the joint cools, resistance falls, and it comes back on its own about forty-five minutes later.
Why remote could never have reached this
Every fact that identified the cause was a measurement taken under a condition that only exists after an hour and a half of running, inside an enclosure a customer must not open. There is no phone question that produces a voltage drop across a specific joint at heat soak. The remote calls could establish the shape of the fault, which they did well, and could not establish the cause, which they were never going to.
Count the cost honestly. Two remote calls, roughly fifteen and twelve minutes, so about 27 minutes of office time. One customer-purchased consumable that was not the problem, though replacing it was overdue anyway. Nine days of a customer living with an intermittent system. Against that, the remote calls did buy one genuinely valuable thing: the knowledge that this visit needed roughly double a standard diagnostic block. Arriving with a one-hour slot on a fault that takes ninety minutes to appear produces a visit that ends in "could not reproduce," which is the worst outcome available.
What would have changed the conclusion
- If the trip interval had been random rather than consistent, heat accumulation drops down the list and a mechanical or intermittent-contact cause moves up. Consistency was the load-bearing fact.
- If output had degraded before each stop, the restriction hypothesis survives and the filter result would have been read very differently.
- If it had never recovered on its own, this is a failed component, not a thermal one, and it becomes a straightforward on-site diagnosis with no run-time requirement.
- If the hot voltage drop had been clean, the heat-soaked control component becomes the answer, and the fix is a component replacement rather than a joint repair. That measurement is what separated two very different repairs.
- If this were a commercial site with maintenance staff, a trained on-site person could have taken the hot reading under direction, and the whole case resolves on day one without a truck. The limit here was the person available, not the phone.
References
- See related: When to Stop Diagnosing Remotely and Roll a Truck
- See related: Reading Rust and Corrosion Patterns
- OSHA 29 CFR 1910.147 on the control of hazardous energy, and 29 CFR 1910.333(b)(2) for verification of de-energization on electrical work
- NFPA 70E guidance on energized work justification and safe work practices
- Trade-standard practice for torque and inspection of electrical terminations