The Unit That Works All Morning and Quits Every Afternoon
Why this matters
This is the complaint that eats service departments alive. The equipment works every morning, quits every afternoon, and is working again by the time a tech gets there. Three visits produce three clean inspections and one very angry customer, and by visit four somebody starts guessing with parts. What follows is one real-shaped case, start to finish, including the four things it turned out not to be and exactly what ruled each one out - because the ruling-out is where the diagnosis actually happened.
Safety first: what got shut down before anything got measured
The complaint was a repeated trip of a thermal protective device, and the equipment had been reset by the customer's staff, by their own count, more than a dozen times over two weeks. That fact drove the first action of the visit, before any meter came out.
A protective device that has tripped that many times is either doing its job against a real overheat or it has been cycled past the point where it can be trusted. Either way the equipment was shut down and locked out for the initial inspection rather than left running while I looked, and the customer's staff were told, on the spot and in writing, to stop resetting it. Repeated resets on an unknown overheat are how a nuisance trip becomes a burned assembly or a fire.
The inspection with the power off found no discoloration, no smell, no melted insulation, and no obvious debris in the primary airflow path. Good news, and also the start of the problem, because the easy findings were all absent.
Turning the complaint into a number
The customer's version was "it quits in the afternoon." That is not a diagnosable statement, so the first hour went into their own records and a log sheet. All values below are illustrative.
Failures over the prior four working days: 14:10, 13:55, and 14:25 on days the equipment started at 06:00. Recovery in every case was 35 to 50 minutes, unassisted, with the customer simply pressing reset once it would take.
That gave two numbers worth having: a time-to-fault of about 8 hours, tight to within half an hour, and a recovery interval of 35 to 50 minutes. The recovery interval alone said thermal, because a latched control fault clears on a reset regardless of how long you wait, and this one refused to clear until it had sat.
Dead end one: the afternoon itself
The obvious theory is that something about early afternoon does it - peak ambient, sun on the west wall, the building's own load peak. Killing this theory cost nothing.
The fourth day in the customer's log, the facility opened late and the equipment started at 09:00. It failed at 17:05. That is about 8 hours again, three hours later on the clock. Elapsed run time is the driver and the clock is not.
A cool overcast day the following week confirmed it from the other direction: ambient peaked around 71 degrees F instead of the 88 degrees F of the failure days, and the unit still quit at 14:05 on a 06:00 start. Ambient moved 17 degrees F and time-to-fault did not move at all.
Two data points, no parts, no tools. Everything that happens at 14:00 on the clock was now off the list.
Dead end two: an afternoon load change
Second theory: something downstream demands more late in the day, the equipment works harder, and it overheats from genuine overload. This is a real mechanism and worth eliminating properly.
A clamp-on current logger on the main load ran the full day. The trace was flat within about 4 percent from the end of the first hour right through to the trip, with no step, no ramp, and no afternoon climb. At the moment of the trip, current measured about 88 percent of nameplate.
That number mattered twice. It killed the load theory, and it killed the possibility that the protective device was correctly protecting against an overload. A device tripping while the load sits comfortably under nameplate is responding to temperature, not to current.
Dead end three: the supply
Third theory: an afternoon supply sag, common enough in a building or on a feeder with a peak, forcing higher current and more heat. A voltage logger on the supply ran alongside the current logger for the same full day.
Supply held flat within about 2 percent all day with no afternoon depression. Off the list, and worth the trouble, because a supply problem is somebody else's repair and finding it late means you have already replaced parts that were never faulty.
Dead end four: the protective device itself
Fourth theory, and the one most likely to get somebody in trouble: the protective device has weakened and is nuisance-tripping below its rating.
It is a tempting theory because it produces a cheap part and a fast exit. It was tested rather than assumed, at temperature rather than cold, and it opened within its stated tolerance. On that evidence the device was fine, and it was reporting a condition that genuinely existed. Hold that conclusion loosely: the test window ran from cold up to the temperature reachable on the bench, and the drift that eventually explained this fault only showed above that, which is why a device can pass a good test here and still be the answer later. That is a limit of the test, not a reason to skip it.
Note what would have happened had this been assumed rather than tested: a new device goes in, the real overheat continues unmonitored, and the next failure is not a trip, it is damage. This is the single most expensive shortcut in this whole family of faults.
Where the readings finally pointed
With four theories dead, the remaining question was simple: what is actually getting hot, and does it ever stop getting hotter?
Three logged channels through a full run - control compartment air, ambient at the equipment, and the current trace already in place - produced this, as delta above ambient:
- 30 minutes: 14 degrees F
- 4 hours: 29 degrees F
- 7.5 hours: 44 degrees F
- at trip: 48 degrees F, still rising
No flattening anywhere. A system that rejects the heat it makes reaches equilibrium and holds; this one never got there in 8 hours. On an 88 degree F day that put compartment air at about 136 degrees F at the trip.
The primary airflow path had been checked cold and was clear. The compartment's own ventilation path had not, because it is a screen most people never look at. It was packed solid with fine debris shed by a nearby process, to the point that no air movement was detectable at the outlet by hand.
That is the mechanism. The compartment made heat all day with nothing carrying it away, the air temperature climbed steadily for 8 hours, and the protective device opened when it reached its threshold. Overnight the compartment cooled and the whole thing reset.
Confirming it before committing
Two confirmations, both cheap once the condition had been reached.
The warm restart: after a 30 minute partial cool-down the unit was restarted, and it tripped again in 2.0 hours instead of 8.0. A 4x reduction in time-to-fault from a partial cool-down is what a thermal accumulation does, and it is not what a timed control routine or a loading filter does - a timer would have fired at the same run hour, and a loaded filter does not unload when you let it cool.
The screen itself: cleared and re-run. The compartment delta flattened at about 26 degrees F by hour 4 and held there. That is the equilibrium that was missing, and it is a 22 degree F improvement in peak delta.
The second cause hiding behind the first
The verification run failed anyway. At 10 hours 30 minutes, on a 92 degree F day, it tripped again, with compartment air at 26 degrees F delta, so about 118 degrees F.
Read that carefully. The first trip happened at about 136 degrees F. The second happened at about 118 degrees F, which is 18 degrees F lower. Clearing the screen bought 2.5 hours of run time and did not fix the fault, because the component was now opening at a temperature it had tolerated before.
That is a degraded thermal-sensing element, and it explains the whole history better than the screen alone did. The screen had been loading up for a long time; the component had been living hot for that whole period and its own threshold had drifted down. One caused the other, and by the time the customer called, both needed addressing.
The component was replaced. The verification run has to clear the LATEST time-to-fault, not the first one: after the screen was cleared the machine was failing at 10 hours 30 minutes, not at 8. An 11 hour run beats 10.5 by about 5 percent, which is the "long enough to feel good about" margin this article warns against, so the run was extended to 14 hours, clearing the 10.5 hour mark by a third. Compartment delta went flat from hour 4 onward and there was no trip.
What would have changed the conclusion
- If the late-start day had still failed at 14:00, the whole run-hours theory collapses and the correct suspects become ambient, solar gain, or a scheduled event elsewhere in the building.
- If current at the trip had been at or above nameplate, the protective device was doing exactly its job and the diagnosis moves to the load: drag, restriction, a failing bearing, a supply problem. Replacing the device would have been actively dangerous.
- If the warm restart had produced the same 8 hour time-to-fault, the mechanism is not heat. Look at a control routine that fires on a run-hour count, or a consumable that loads progressively and does not reset on cooling.
- If the compartment delta had flattened by hour 3 well below the trip point, the heat is not accumulating and the fault is something else that happens to correlate with run time.
- If clearing the screen had produced a clean 11 hour run first time, there would have been no second cause and the component would have been left in service. It is the trip at a lower temperature that proved the second fault existed, which is why the verification run has to be longer than the original time-to-fault rather than just long enough to feel good about.
References
- OSHA general industry guidance on lockout and on equipment operated with a tripped protective device
- NFPA 70B, recommended practice for electrical equipment maintenance
- Manufacturer documentation for thermal protective device ratings and reset behavior
- See related: How to Test a System That Only Fails After Hours of Running; Heat-Soak Faults Versus Cold-Start Faults; How to Log Temperature Drift Across a Full Duty Cycle