What a System Does While It Is Recovering From an Outage

Why this matters

The minutes after power returns are the only period in a system's life when normal behavior and failure behavior look the same from the outside. A control that is deliberately doing nothing and a control that cannot do anything both present as a powered unit with no output. Every wrong condemnation in this category comes from a tech who knew the equipment was misbehaving and did not know what correct behavior looks like during recovery.

This card is the other half of the diagnosis: not the procedure for deciding, but what is actually happening inside the box while you wait, and more usefully, the list of things recovery never does. The exclusions are what let you stop waiting.

Anything mid-recovery intends to start without warning

A recovering system is a system that has already decided to run and is counting down to it. Before a hand goes anywhere near a rotating load, a damper, an accumulator, a spring, or a pressurized line, isolate the energy source, apply your lock and tag, and bleed or block the stored energy, which is the control-of-hazardous-energy duty at 29 CFR 1910.147. If the work is on a panel, a branch circuit or energized conductors, de-energize and lock or tag before contact under 29 CFR 1910.333(b)(2), and prove dead with the live-dead-live sequence in NFPA 70E-2021, 120.5. "It has not started yet" is not the same as "it is off," and a staged restart is specifically designed to bring loads on while nobody is watching.

The four jobs a control is doing in the first minutes

Recovery is not one behavior. It is up to four different jobs, and knowing which one you are watching tells you how long it should take.

Hold-off. A fixed timer that prevents an output from energizing for a set interval after power returns, protecting a compressor or a pump from restarting against a pressure it has not equalized. Purely a clock. It ends at its published interval whether or not anything else is happening.

Staging. Bringing multiple loads back one at a time rather than all at once, so the site does not present its whole connected demand to the service in the same instant. On a multi-load system the total sequence is the number of stages times the interval between them, which is why staging is the recovery job most often mistaken for a dead output: the last stage can be several times further out than the first.

Position relearn. A control that does not know where a modulating device sits after losing power, so it drives it to a hard stop to re-establish a reference. Valves, dampers, actuators, doors and travel-limited mechanisms all do this. You will often hear it before you see anything: a device driving to an end stop and stalling briefly is the sound of relearn, not of a failure.

Baseline re-derivation. A control that adapts to the installation and lost its stored history, so it runs conservatively and re-measures. Adaptive fan settings, learned run times, flow baselines and drift compensation all sit here. This one is the slowest and the least visible, because the equipment is producing output the whole time, just not optimal output.

What recovery is never: the exclusion list

This is the part worth memorizing, because it is what ends a wait. No recovery job in any of the four families above produces any of the following. If you see one, you have a fault, and the clock is irrelevant.

  • A protective device operating. A breaker tripping, a fuse opening, a high-limit or high-pressure switch cutting out, an overload dropping a contactor. Protective devices respond to physical conditions. There is no recovery routine that trips a protector on purpose.
  • A hard lockout code. Informational codes during recovery are common and often say so in the literature. A lockout that requires a deliberate reset is a decision made about a failed attempt, and it means an attempt already happened and failed.
  • The same failure repeating on successive attempts. Recovery is monotonic: it moves toward running and does not restart itself. Three identical attempts, three identical failures, is a fault with a retry loop wrapped around it.
  • A physical measurement that is wrong on an attempt. Current several times or far below rated on a closed contactor, no pressure rise on a running pump, no temperature split developing across a full cycle. A timer can withhold an output. It cannot change what happens when the output finally comes on.
  • Anything visible, audible or smellable. Discoloration, a burnt smell, a shorted-out sound, water where water should not be, a device too hot to touch. These are never recovery states. Anything in this group ends the wait immediately and starts a containment decision instead.
  • Behavior that never converges. A relearn drives to a stop and finishes. A device hunting back and forth for many minutes with no settling is a feedback or reference problem, not a learning routine.

The reason the exclusion list is the more useful half: the positive list tells you what you might be watching, which is a probability. The exclusion list tells you what you are definitely not watching, which is a decision. A tech armed only with the positive list waits and hopes. A tech armed with the exclusions stops waiting the moment one of them appears, and on most calls one of them appears in the first two minutes.

How long each job runs, and what sets the number

Do not guess these. They come from the equipment's own service literature, and where the literature is silent, the practical floor is 15 minutes before you call a no-output condition a fault.

Job Typical driver of the duration Ends when
Hold-off A published fixed interval, commonly a few minutes The clock expires
Staging Stage count times the published inter-stage interval The last stage energizes
Position relearn Full travel of the device, plus a settling allowance The device reaches a reference stop and reverses
Baseline re-derivation One or more complete operating cycles Output quality stabilizes across a cycle

Two of these compound. A staged system whose last stage also carries a hold-off does not run them in parallel, which is why doubling the longest published delay is a safer window than adding the ones you happen to know about.

The one recovery behavior that genuinely mimics damage

Baseline re-derivation is the honest trap. The system runs. It produces output. The output is measurably worse than the customer remembers, and every static reading you take is inside tolerance. A tech under pressure reads "runs but underperforms with normal readings" and reaches for the expensive explanation.

The separator is direction over time. A re-deriving control improves across successive cycles, because that is the entire point of the routine. A degraded system is flat or worse. So the measurement is not a single reading, it is the same reading taken on two consecutive complete cycles. If cycle two is better than cycle one, you are watching a control learn. If cycle two matches cycle one, the control has already settled and the shortfall is physical.

Reading two units on one site after one outage

A small commercial site loses power for about half an hour. Two pieces of equipment, same building, same outage, same complaint from the manager: "neither one is doing anything."

Unit A is powered, has no output, shows an informational code rather than a lockout, and is quiet. Its literature describes a three-stage restart with a 4-minute interval between stages, so the full sequence runs 12 minutes from restoration. Nine minutes have passed. Two of the three stages have energized on schedule and the third is not due for another 3 minutes. Check it against the exclusion list: no protector has operated, no lockout code, no repeated attempt, nothing audible or visible. Nothing on the exclusion list is present, and the observed behavior matches the published sequence. This unit is doing its job. The correct action is to leave it alone and come back to it.

Unit B is powered, and it attempts. Three times in about 6 minutes it pulls in, runs briefly, and drops out on a protective device, returning to the same code each time. Two exclusions are present at once: a protector is operating, and the same failure is repeating across successive attempts. The clock does not matter. This one gets diagnosed now.

The pairing is the lesson. Same site, same outage, same customer sentence, and the two units resolve in opposite directions on evidence available before any tool comes out of the bag. A tech who treats the outage as the diagnosis condemns both or waits on both, and is wrong on one of them either way.

How to verify you read a recovery state correctly

Whichever way you called it, prove it before you write it up.

  • If you called it recovery, watch it finish. Stay for output, and for a re-derivation case stay for a second complete cycle so you can state the direction of change rather than assume it. A recovery call left early is a callback with your name on it.
  • If you called it a fault, name which exclusion you saw. Write the specific one on the ticket: which protector operated, which code latched, which measurement was wrong and against what rating. "It did not come back" is not a finding and will not survive a warranty review.
  • Re-check the elapsed time against continuous restoration. If supply dropped again at any point, or if anyone reset anything, the clock restarted there. A window measured from the wrong start time is the quiet way a correct method produces a wrong answer.

References

  • 29 CFR 1910.147, OSHA control of hazardous energy, for isolating and blocking stored energy before service on equipment that may start
  • 29 CFR 1910.333(b)(2), OSHA electrical safe work practices, for de-energizing and locking or tagging before work on electric circuit parts
  • NFPA 70E-2021, 120.5, the live-dead-live process for establishing an electrically safe work condition
  • Manufacturer service literature for published restart delays, staging intervals, relearn routines and code definitions
  • See related: How to Tell a Reset State From a Real Fault After a Power Interruption; The Protective Devices That Should Have Prevented Power Event Damage