The Repeat Failure That Was Never the Part

Why this matters

Three techs replaced the same protective device on the same appliance and all three were competent. Every one of them fixed the symptom and none of them fixed the cause, because the cause only shows itself after the machine has been running for a while, and nobody stays for a while. This walkthrough is one real-shape case followed end to end: what was checked, what each check ruled out, why four sensible theories were wrong, and what finally named the cause. The point is the reasoning, not the component.

Safety first: a safety device that keeps tripping is a safety device working

Before any diagnosis on a repeat safety trip: do not jumper, bypass, or defeat the device to get the customer running. A protection cutout that trips repeatedly is reporting a real condition upstream of it. Bypassing it removes the last barrier between a fault and an event, and it moves liability squarely onto you.

On this class of work that means: shut the appliance down, isolate the energy source, verify dead with the live-dead-live method on the electrical side, and relieve and verify pressure and let surfaces cool before opening anything. If the customer asks you to leave it jumpered "just until the weekend," the answer is no, and the reason you give is that the device is not the problem, it is the messenger.

The call

The customer reports the appliance shuts itself off and needs a manual reset, but only sometimes. Pressed for detail, they say it happens in the evening, never in the morning, and that it always comes back after they reset it.

The file shows two prior visits. Two different techs, two replacements of the same over-temperature cutout, two clean departures with the unit running. The device life across the three units:

  • Original device: about 14 months in service before the first trip complaint
  • First replacement: about 5 months
  • Second replacement: about 2 months

Each one lasted between a third and 40 percent as long as the one before it. That accelerating trend is the first real datum, and it arrives before anyone touches a tool.

First read: the failure shape, not the component

A component that fails and gets replaced with a new one should reset the clock. Life that keeps shortening across three separate devices means the environment the device sits in is getting worse. Three bad parts in a row is possible. Three bad parts in a row that each fail faster than the last is a story that requires the parts to have been progressively worse, and there is no mechanism for that.

The second datum is in the customer's own description, and it is the one nobody wrote down: evening, never morning. That is not a temperature-of-day pattern, because morning cold would if anything load the appliance harder. It is a duty-cycle pattern. Evening is when the appliance runs its longest continuous stretches.

So the working hypothesis before opening a panel: something degrades during a long run, crosses the cutout's threshold, and recovers when the machine sits. That immediately explains why two techs left with a working unit. A cold-start functional check passes on a fault that needs 40 minutes of continuous operation to appear.

Dead end 1: a batch of bad devices

The obvious theory, and the one that got acted on twice.

Ruled out by three facts together. The three devices came from two different suppliers. Each was tested for continuity and reset function on the bench before install and all passed. And the accelerating life trend has no mechanism under a bad-batch theory. A bad batch produces random short lives, not a monotonic decline.

What a tech who stopped here would conclude: that the third device will hold and the problem is resolved. It will not, and the customer's confidence is already thin at three visits.

Dead end 2: low flow

Sensible theory. If the medium moving through the appliance slows down, the heat has nowhere to go and the cutout trips legitimately.

Checked by measuring the differential across the appliance at startup and comparing to design. It read at design at start of run. The circulating device drew normal current, the isolation valves were fully open, and the strainer ahead of it was clean.

Ruled out for the startup condition, but only for the startup condition. That distinction matters and is the reason the theory stayed on the list rather than being crossed off: a flow reading taken in the first two minutes tells you nothing about minute 38. Note it as unproven for the running case, not disproven.

Dead end 3: a high-resistance electrical connection

Also sensible. A loose or corroded connection in the control circuit heats under load, raises resistance, and can drop enough voltage to make a control misbehave. It also produces exactly the "gets worse the longer it runs" signature the complaint has.

Checked properly: de-energize, verify dead, inspect every terminal in the control path for discoloration or a halo of corrosion, then re-energize and measure voltage drop across each connection under load. All drops were negligible against the circuit voltage and the terminals were bright.

Ruled out. Worth doing anyway, because the signature fit so well that leaving it unchecked would have left a plausible alternative alive in the write-up.

Dead end 4: control calibration

If the sensing element or the control reading it had drifted, the appliance could be running hotter than the control believes and the mechanical cutout would be catching a genuine overtemperature the control never saw coming.

Checked by comparing the control's displayed reading against a calibrated instrument on the same point. Agreement was within instrument tolerance. Ruled out.

At this point four reasonable theories are gone and the appliance still trips. That is the moment where a tech either swaps the device a third time or changes method.

The measurement that broke it open

The method change: stop testing components and start running the appliance and watching numbers move.

Full continuous run, instrumented, with the temperature rise across the heat exchange surface logged every five minutes and the surface itself read with a non-contact thermometer at the same spots each time.

  • At start of run, the temperature rise across the exchanger measured at design.
  • By 20 minutes of continuous run, it had climbed to roughly 120 percent of design rise.
  • By 35 minutes it was around 140 percent of design rise, and the surface reading at the hottest spot was climbing faster than the outlet reading.
  • At about 40 minutes the cutout tripped.

That set of numbers only fits one family of cause. Output temperature lagging while surface temperature climbs means heat is not crossing the exchange surface efficiently. The heat has to go somewhere, so the metal on the fired side gets hotter and hotter until the protective device sees it. An insulating layer on the exchange surface produces exactly this, and it produces it progressively during a run because the surface needs time to soak.

It also retroactively explains the accelerating device life. Each device sat in a hotter environment than the last, because the insulating layer was thicker each time. A cutout that spends its life cycling near its own limit degrades faster than one that never approaches it.

Confirming it

Shut down, isolate, relieve pressure, let it cool, open the access. The exchange surface carried a hard, light-colored, scaly deposit, thickest at the hottest section, thin at the inlet end. That distribution is the signature: deposition tracks surface temperature.

The remaining question was where it came from, and that is a supply question, not an equipment question. Two things answered it:

  • The treatment equipment on the incoming supply was in bypass. The customer could not say when it had been put there.
  • The customer had been topping the system up from a hose bib rather than through the treated feed, because it was faster.

So: raw supply water into a system designed for treated, over roughly a year and a half of top-ups, laying a progressively thicker deposit on the hottest surface in the appliance. The consumable in this case was the water itself.

What the fix actually was

Not a fourth cutout. Descale the exchange surface with the appropriate method for that surface, restore the treatment equipment to service and verify it is actually treating rather than just plumbed in, close off or clearly mark the untreated fill path, and re-run the same instrumented test to confirm the temperature rise now holds near design through a full run rather than climbing.

The verification run is the part shops skip. Without it you have a theory and a cleaned surface. With it you have a before-and-after on the same measurement, which is what makes the write-up hold up if the customer questions the bill.

What would have changed the conclusion

  • If the temperature rise had stayed flat through the full run and the device still tripped, the cause would be electrical or in the device circuit, and the high-resistance-connection check would move back to the top of the list with a longer soak.
  • If the surface had been clean on teardown, the next suspect would be reduced flow appearing only when hot, from a component whose clearances change with temperature, and the flow check would need repeating at 35 minutes rather than at start.
  • If the deposit had been soft and dark rather than hard and light, that is biological or degraded-fluid in character rather than mineral, which sends you to standing time and storage conditions instead of to supply water.
  • If the appliance were near end of design life with the same findings, the honest recommendation shifts from repair to replacement economics, because you are restoring a surface that has already been thermally abused.

What it cost to get here

Three visits, two devices, and a customer who now needs convincing. Every one of those visits ended with a functioning appliance, which is why nobody escalated. The single change that would have caught it on visit one is running the machine long enough to reproduce the customer's actual condition before declaring it fixed. The customer told two techs it happened in the evening, and both of them tested it in the morning.

When a complaint carries a time-of-day or duration pattern, that pattern is a test instruction. Reproduce the condition or you have not tested anything.

References

  • OSHA lockout/tagout and verification-of-de-energization practice for service work
  • Manufacturer documentation for design temperature rise, water quality requirements, and approved descaling methods
  • Trade-standard practice on never bypassing or defeating a protective device to restore service
  • See related: How to Inspect a Consumable for Evidence It Caused the Fault; Why Off-Spec Supplies Fail Slowly Instead of Immediately; Reading Rust and Corrosion Patterns