The No-Fault Call That Was a Design Behavior

Why this matters

The second visit on a complaint the first tech could not find is where shops lose money and customers lose patience. By then the customer has decided somebody is missing something, and the pressure to leave with a part replaced is real. This case is a walkthrough of one of those seconds visits, including the three theories that were wrong and exactly what killed each one, ending at a cause that was never a fault at all. The value is not the answer. It is the sequence of eliminations, because that sequence is what you can carry to a different system in a different trade.

Lead with the hazard, not the theory

The complaint involved a system shutting itself down and restarting, which is the exact profile that gets a tech killed reaching into equipment that is off but not dead. Before any of the diagnosis below, the disconnect was opened, the circuit proved dead at the point of contact with live-dead-live (meter proved on a known live source, circuit proved dead, meter proved live again), and a lock applied. Anything that self-starts on a timer or a control signal gets locked out, not switched off, because "it was off when I put my hand in there" is not a defense against a system designed to restart itself.

Stored energy got the same treatment. Capacitive components were discharged and verified at zero before anything was touched, and the pressurized side was isolated and relieved before any joint was opened.

The call as it arrived

Second visit, same system, same complaint. First tech had closed it as no fault found after about 40 minutes on site.

Customer's words: "Every so often it just stops. It makes a loud noise, there's steam or something, and then a few minutes later it starts back up on its own. It's been doing it since it got cold out and it's getting worse."

The first tech's record said: "Ran unit, operates normally, all readings within range, no fault found." No numbers, no runtime, no target. That record is the reason visit two started from zero and is worth noting as its own lesson.

What the customer told me before I touched anything

Three questions, before opening a single panel.

How long does it last? "A few minutes. Maybe five." Consistent, not variable. That mattered more than it seemed at the time.

Does it happen at any particular point? "Not that I can tell. It seems random." Customers are poor timers but good at ruling things out, and "not tied to when I do anything" removed the whole family of user-triggered causes.

What changed around when you first noticed? "It got cold. And we changed the filter ourselves in the fall, we got a cheaper one from the hardware store." That answer put two things on the board at once: a season, and a customer-sourced consumable.

Dead end one: the intermittent connection

The first theory was the obvious one for a system that stops and restarts. A loose or high-resistance termination heats under load, opens, cools, and remakes contact. It fits the symptom shape almost perfectly, and it is the single most common cause of a self-clearing shutdown across every trade that runs electrical controls.

What killed it: the interval. With the system locked out, every accessible termination in the control and load path was checked for discoloration, for a halo of corrosion around a lug, and for looseness against the stated torque. All clean, all tight. That is suggestive but not conclusive, because a marginal termination can look perfect cold.

The conclusive evidence came from the runtime observation later. A thermal-mechanical intermittent does not keep time. It opens when it happens to reach its threshold, and the interval between events wanders with load, ambient, and how long the system has been running. What was actually observed was three events spaced 44, 46, and 45 minutes apart, a spread of 2 minutes across a roughly 45-minute interval, under 5 percent variation. Nothing driven by a loose connection is that punctual. Regularity is the signature of a control, not a defect.

Dead end two: the customer-supplied consumable

The cheap replacement filter was the most attractive theory on the board, because the customer had volunteered it, the timing lined up, and off-spec consumables genuinely do cause faults that look like equipment failure. An overly restrictive element raises the pressure drop across the system, pushes the working element outside its intended operating envelope, and can trip a protective device that then resets when things cool.

It was measured, not assumed. Pressure drop across the installed element read close to twice the drop across a correct-spec element from truck stock, measured at the same operating point. So the filter was genuinely wrong for the system and genuinely making it work harder. That is a real finding.

It was not the cause of the shutdowns. Two things ruled it out. First, swapping in the correct element and running the system for a further 50 minutes did not change the interval or the duration of the events at all. Second, the pressure drop was elevated but still well inside the range where the protective device would act, so there was no mechanism connecting it to a shutdown.

This is the discipline worth taking from this case: a real finding is not automatically the cause. The off-spec filter went on the invoice as a corrected item with the measurement behind it, and the diagnosis kept going. A tech who stopped there would have "fixed" the complaint, charged for it, and been back inside two weeks with a customer who now trusted him less than the first tech.

Dead end three: a protective limit tripping

Third theory: the system was reaching a limit, shutting down to protect itself, cooling off, and resetting. This deserved real attention because it is the theory that separates a nuisance from a hazard, and if true it would have meant the system was operating outside its envelope for a reason not yet found.

Two things killed it. The control history carried no lockout or fault record for the events, and a protective trip on nearly all modern equipment leaves a trace even when it self-resets. More decisively, a protective shutdown is a stop, not a sequence. What the customer described, and what was then observed directly, included a noise and a visible vapor plume followed by an orderly restart. A limit trip does not produce a choreographed sequence of events. It produces a stop.

The observation that broke it open

The regularity was the finding, and it took a full runtime to see it. Across just under two hours of continuous observation, three events, spaced 44, 46, and 45 minutes apart, each lasting close to 6 minutes, each following the identical sequence: a change in operating sound, a brief reversal of normal flow direction, a vapor release, then resumption.

Nothing broken keeps a schedule. Something that repeats within 5 percent of the same interval, runs for the same duration, and follows the same ordered sequence every time is being commanded by a control. At that point the question stopped being "what is failing" and became "what is this system designed to do on an interval."

Confirming it

The installed documentation described exactly this: a periodic cycle in which the system interrupts normal operation, reverses part of its flow path to clear an accumulation that builds during cold-weather operation, then returns to normal service. Interval and duration were both consistent with the published sequence, and the trigger conditions explained the timing of the complaint precisely. The behavior only occurs in cold weather, which is why the customer first saw it when it got cold and had never seen it in the years before.

Confirmation was not reading the documentation. It was forcing the cycle on demand, in front of the customer, and watching the identical sequence occur when commanded. A behavior you can produce at will is a behavior you have identified. A behavior that merely matches a description is a behavior you have guessed at.

What actually got delivered

Two line items, one of which is not a repair.

The off-spec element was replaced with a correct-spec one, with the measured pressure drop on both recorded on the invoice. That was a real correction with a real justification, and it stands on its own regardless of the shutdown question.

The shutdown complaint was closed as designed operation, with the interval, the duration, and the sequence documented, plus a written note of what the customer should expect: the cycle occurs only in cold weather, roughly every 45 minutes during continuous cold-weather operation, lasting around 6 minutes, and it will not occur at all in warm weather. The customer was shown it on demand once and could describe it back.

The most valuable output was the escalation trigger, written in plain language: call back if the interval drops well under 45 minutes, if the events start lasting substantially longer than 6 minutes, or if the system does not resume normal operation on its own afterward. Any of those means the cycle is no longer normal, and each is something the customer can actually observe. That is what stops the next call from being another no-fault visit, and what makes sure a genuine problem gets reported rather than tolerated.

What would have changed the conclusion

If the interval had wandered. Events spaced 12, 40, and 61 minutes apart would have pointed straight back to the intermittent-connection theory, and the correct next step would have been thermal imaging under load at heat soak rather than a cold inspection of terminations.

If the control history had carried a fault record. Any logged limit or lockout event coinciding with the shutdowns would have made theory three live again, and the diagnosis would have moved to finding why the system was reaching that limit.

If the correct-spec element had changed the behavior. Had the interval stretched or the events stopped after the filter swap, the consumable would have been the cause rather than a coincident finding, and the deliverable would have been an entirely different conversation about who supplies consumables.

If the system had not resumed on its own. A self-clearing event and a permanent shutdown are different faults with different urgency. The orderly restart was load-bearing evidence throughout.

If the customer had reported it in warm weather. The seasonal gate was central. A cycle that only runs in cold weather, occurring in summer, is a control or sensor problem, not a design behavior.

References

  • NFPA 70E live-dead-live verification and lockout practice for equipment capable of automatic restart
  • Manufacturer installation and service documentation for periodic defrost, purge, or regeneration sequences and their published intervals
  • Manufacturer specifications for consumable elements, including acceptable pressure drop at rated flow
  • See related: How to Verify a System Is Operating as Designed; How to Document a No Fault Found Visit Defensibly; Why Normal Noise, Cycling, and Condensation Alarm Customers