Heat-Soak Faults Versus Cold-Start Faults

Why this matters

Two customers describe the same complaint - "sometimes it just will not go" - and the correct response is opposite in almost every practical respect: what time you schedule the call, what you tell them to do before you arrive, how many test attempts you get in a day, and what you must not let them do in the hour before you pull up. Get that backwards and you will spend a full visit measuring a system that is in the one state where it works perfectly. This card is about the logistics of catching each one, not about the component families behind them, which a sibling card already covers.

Define each one narrowly, because the loose version is useless

A cold-start fault lives in the dead-cold condition. The system has sat long enough to reach the temperature of its surroundings, and in that state something is too tight, too thick, too contracted, or too demanding to get moving. Once it is running, the fault is gone until the next long sit.

A heat-soak fault lives in accumulated heat with nothing carrying it away. The critical word is soak, not hot. Heat soak is the condition where a system has been making heat for a long stretch and the heat has saturated the surrounding mass rather than being rejected - and its worst moment is very often after shutdown, not during the run, because when the equipment stops, the fans, pumps, and flow that were removing heat stop with it while the stored heat keeps redistributing.

That last point is the single most useful distinction on this page, and it is the one most techs miss. A heat-soak fault is not just "fails when hot." It frequently fails at a moment when the machine is doing nothing at all.

The peak-after-shutdown trap

Picture the sequence. The equipment runs for hours and a hot component has been sitting in moving air the whole time. The unit shuts off. Airflow stops immediately. The hot component now dumps its heat into stagnant air and into everything it touches, and the temperature of the parts around it - the control compartment, the terminal block, the sensor mounted on the wall of the housing - climbs for the next ten to thirty minutes even though nothing is running.

Then it falls again over the next hour or two.

So a component with a marginal thermal limit can be perfectly fine at the moment of shutdown and open ten minutes later. What the customer experiences is a machine that ran, stopped normally, and then refused to restart for half an hour. What a tech experiences, if they arrive at the wrong point on that curve, is a machine that restarts on the first try with no fault stored.

The practical rule: on a suspected heat-soak fault, the highest-value measurement window is often 5 to 25 minutes after shutdown, not during the run. If you shut a unit down and immediately start testing, you may be measuring before the peak.

The comparison that drives the visit

Cold-start fault Heat-soak fault
Failure state System at surrounding temperature after a long sit Accumulated heat, often peaking after shutdown
Worst time of day First start of the day, coldest part of the season Late in a long run, or shortly after it stops
What resets it Running for a while, or ambient warming Sitting long enough to cool, typically 30 to 120 minutes
Attempts you get per day Usually one, unless you can force a cool-down Several, using partial cool-downs and warm restarts
What the customer must do before you come Leave it off overnight, do not start it Run it normally until it fails, then leave it alone
The way you lose the visit Customer started it to "make sure it works for you" Customer let it sit and cool, or reset it repeatedly
Where to measure At the first attempt to start, before anything warms At the failure moment and through the post-shutdown window

Read the "what the customer must do" row twice. It is the opposite instruction in each column, and it is given over the phone by whoever books the call, not by the tech standing in the driveway.

What each one costs you in attempts, and why that changes the plan

A cold-start fault gives you roughly one honest attempt per visit, because once you have started the machine the cold condition is gone and you cannot get it back inside a service window. Anything you did not measure on that first attempt is lost until tomorrow. That means the entire visit has to be set up before the first start: meters connected, leads on, camera running, sequence decided. There is no second take.

A heat-soak fault is the reverse. Reaching the condition is expensive the first time, because it takes hours of run, but once you are there a partial cool-down and restart puts you back at the failure in a fraction of the original time. The cost structure is high setup, cheap repeats. That is why the correct plan is a long instrumented run to establish the condition, then several quick warm-restart cycles to test candidate causes.

Same complaint category, and the two plans have almost nothing in common.

Booking the call correctly

For a suspected cold-start fault, the booking script is: do not run it, do not test it, do not warm the space near it, and expect us at the earliest slot we have. Then take the first appointment of the day. If the customer needs the equipment overnight, you have a conflict and you should say so plainly rather than showing up to a warm machine.

For a suspected heat-soak fault, the booking script is: run it exactly as you normally do, and the moment it quits, call us and leave it alone. Do not reset it. Do not cycle the power. Do not open the panel to let it breathe. Each of those actions destroys the state you are being paid to measure, and a customer who has "helped" three times in a row is the reason a fault has survived three visits.

Both scripts are one sentence and both belong on the dispatch card, not in the tech's head.

A worked example, carried through

Two calls, same words in the complaint field: "won't start sometimes." All values below are illustrative.

Call A. Intake asks the three questions and learns it only fails on the first attempt of the morning, that a second or third attempt usually gets it going, and that it never misbehaves once it has run. That is cold-start. Booked for 07:15, first slot, with instructions not to touch it. The tech arrives, hooks up before touching a control, and captures the first start attempt: starting current about 1.8 times the value seen on a warm start attempt an hour later, and the unit failing to break free on attempt one, catching on attempt three. One attempt, one clean data set, diagnosis in under an hour of attendance.

Had the customer started it once at 06:30 to check, that visit produces nothing at all. The whole diagnosis rested on a state that a single button press destroys.

Call B. Intake learns it runs from about 07:00 and quits around 13:30 most days, and that it will restart after roughly 45 minutes. That is heat soak. Booked for 13:00 with instructions to run normally and not reset it.

The tech arrives at 13:00 with the unit still running, gets baseline readings, and the unit trips at 13:20. Here is where the post-shutdown behavior earns its keep: the tech does not immediately try a restart. Instead they log the control-compartment temperature every 2 minutes. It reads 118 degrees F at shutdown, peaks at 131 degrees F at 14 minutes after shutdown, and is back to 118 at about 35 minutes. The peak is 13 degrees F above the shutdown reading, entirely after the machine stopped.

That single curve reframes the fault. The tech had assumed a run-time overheat and was looking for lost airflow during operation. The post-shutdown rise says the compartment has no way to purge heat once the machine stops, which points at the compartment's ventilation path rather than at the running airflow. Two warm-restart cycles at roughly 55 minutes each confirm the same peak, so three confirmations fit inside one afternoon against a 6.5 hour cost for the first one.

Note the arithmetic that matters commercially: the first confirmation cost 6.5 run hours, each subsequent one about 55 minutes. That ratio, roughly 7 to 1, is why you never abandon a heat-soak condition once you have reached it. Techs who shut down, pack up, and return tomorrow pay the 6.5 hours again.

How to verify which one you have

Ask one question that separates them cleanly: after it fails, how long before it will work again, and does anything have to happen during that time? A cold-start fault does not have this shape at all - it either starts on a retry or it does not, and time spent waiting makes it worse rather than better. A heat-soak fault has a characteristic recovery interval, usually 30 to 120 minutes, and the customer can usually name it within about ten minutes because they have lived it.

If the answer is "it works again immediately if I just cycle it," you are probably in neither family, and should be looking at a control or a latching fault rather than a thermal one.

What changes the answer

  • A fault that appears both on cold starts and deep into long runs is usually not thermal in either direction. Look for a marginal component that is simply failing across its range, where temperature is a modifier rather than the cause.
  • In a hot climate or an unconditioned space, "cold start" may never mean genuinely cold, and the cold-start family thins out to almost nothing. Weight your suspicion toward heat soak.
  • Equipment that runs continuously and never shuts down cannot produce the classic post-shutdown soak peak, so the failure window collapses back into the run itself and the warm-restart tool is unavailable until you force a shutdown.
  • If the customer's recovery interval is very consistent (say within 5 minutes, every time), consider that you may be looking at a fixed timer in a control rather than a cooling curve. A cooling curve varies with ambient; a timer does not.

References

  • Manufacturer documentation for thermal limits, restart delays, and recovery timings
  • NFPA 70B, recommended practice for electrical equipment maintenance
  • Trade-standard practice for diagnosing temperature-dependent intermittent faults
  • See related: Cold-Start Faults vs Warm-Start Faults: The Real Difference; What Changes Inside a System During a Long Run Cycle; How to Test a System That Only Fails After Hours of Running