Why Equipment Fails After a Power Event Rather Than During It

Why this matters

A site takes a hit, everything comes back up, and the customer signs off. Then things start failing on no obvious schedule for the next two months and every one of those calls is argued about: was it the storm, or is this equipment just old. The argument is winnable, because damage mechanisms have time constants and the delay between the event and the failure is evidence. A tech who can name which mechanism produces which latency stops guessing and starts telling the customer, and their insurer, something defensible. The reverse also matters: surviving the event is not evidence of being undamaged, and treating it as evidence is how a site gets a second wave of failures nobody planned for.

The gate before any of this work

Everything below happens in a live building with equipment that may be damaged in ways you cannot see. De-energize before opening anything: 29 CFR 1910.333(a)(1) permits energized work only where the employer can demonstrate that de-energizing introduces additional or increased hazards or is infeasible due to equipment design or operational limitations. Lock and tag under 29 CFR 1910.333(b)(2), which also requires stored electric energy that might endanger personnel to be released and capacitors discharged; note that 29 CFR 1910.147 excludes this work at (a)(1)(ii)(C) and hands it to Subpart S, with 29 CFR 1926.417 as the construction counterpart. Where a reading genuinely can only be taken running, work to the shock and arc flash boundaries and PPE from the risk assessments at NFPA 70E-2021, 130.5 and 130.7, in the edition your employer's program adopts, with the instrument proved live-dead-live at 120.5.

Latency is a fingerprint

Delay from event Mechanism that produces it What you see
During the event Direct puncture of a semiconductor junction or insulation, fuse or breaker operation Hard failure, often with a visible mark or a cleared fuse
Minutes to hours Thermal runaway in a degraded voltage-clamping component; a shorted turn heating under load Something gets hot after the site is already back up
Next start, not next run Winding or contact insulation weakened by the event, stressed hardest by the steep voltage front at start Ran fine all day, will not restart in the morning
Days to weeks Capacitance loss in metallized-film capacitors, ripple heating of a stressed electrolytic, carbon tracking spreading across a surface Starting trouble that worsens, then a hard failure on a hot afternoon
Months Slow insulation degradation from partial discharge at a damaged site A failure nobody connects to the event at all, correctly or otherwise

Read that column of delays as a filter, not a verdict. It tells you which mechanisms are still on the table for a failure sixteen days out and which are already excluded, and it is the only part of this that works before you touch the equipment.

The call

Storm on day zero, utility interruption and restoration, no immediate complaints. On day sixteen a rooftop unit will not start. The contactor pulls in, the compressor hums, and the thermal overload opens after a few seconds. The customer's question, which is the real work order, is whether the storm caused it.

The hum with an overload trip is a locked-rotor signature, so the first question is whether the machine is trying to start and failing, or not trying at all. It is trying. That puts the fault in the starting path: supply voltage at the terminals during the attempt, the start assist components, or the compressor itself mechanically.

Supply first, because it is upstream of everything else and it is where the storm would have acted. Measured with the unit running under the energized-work gate above, terminal voltage during the attempt sits within the tolerance band on the nameplate. That does not clear the supply completely, but it removes the simple case.

The number that named the mechanism

Power off, locked out, capacitor terminals discharged and proved at zero volts before anything touches them: the run capacitor's nameplate reads 45 microfarad and the meter reads 31. That is 69 percent of nameplate, or 31 percent low. The acceptance band is printed on the can, commonly plus or minus 6 percent for a run capacitor, and the printed band governs rather than any rule of thumb. Either way, 31 percent low is not a marginal call.

The useful question is not whether the capacitor is bad. It is why it is 31 percent low sixteen days after a storm on a machine that started fine all summer.

Metallized-film capacitors self-heal. A local fault in the dielectric clears by vaporizing the thin metallization around it, which isolates the fault and costs a small amount of capacitance. That is normal and it is the reason these capacitors fail gracefully instead of shorting. An overvoltage event drives a large number of those clearing events at once, so capacitance takes a step down at the event and then continues its ordinary slow drift from there.

That mechanism fits the timeline in a way that ordinary aging does not. Ordinary aging is a slow decline that would have produced hard starting across weeks of increasingly warm mornings. A step change followed by drift produces exactly what happened here: no complaint at all, then a machine that crosses the threshold where it can no longer break away against a warm-day head pressure, and does it on one specific day.

The second failure, and the one that was not related

On day twenty-three, a week after the first call, a control board on a different unit at the same site fails. Same site, same event, different mechanism and a different latency. That board's low-voltage supply had been running on a stressed electrolytic capacitor, and ripple heating carried it the rest of the way over the following week. Consistent.

On day thirty-one, a condenser fan motor at the same site seizes with dry bearings and rust in the end bell. Nothing about that mechanism has a path back to a voltage event. It goes in the report as unrelated, and saying so is what makes the other two credible. A tech who attributes every failure in the following quarter to the storm is not making a stronger case for the customer, they are making a weaker one, because the first obviously unrelated item on the list is what an adjuster uses to dismiss the rest.

The protective device that reported the wrong thing

The site has a surge protective device at the panel, and its indicator is green. The customer reads that as proof the equipment was protected.

The indicator on a listed surge protective device reports the state of its internal thermal disconnect, which is what keeps a failing clamping element from becoming a fire. It does not report how much clamping capability remains. A device can have absorbed most of its useful life and still show a healthy indicator, because the disconnect has not operated and, from the disconnect's point of view, nothing is wrong. There is no field measurement that recovers remaining capability, which is why devices intended for sites that take real events are specified with surge counters and why the manufacturer's instructions, enforceable through NEC 110.3(B) in the edition your authority having jurisdiction has adopted, are the authority on when the unit is replaced after a known event. Surge protective device requirements themselves live in NEC Article 242 in that same adopted edition.

So the honest line to the customer is that the protective device is of unknown remaining capability after a known significant event, and that unknown is a reason to replace it, not a reason to argue about whether it worked. It very likely did work, which is why the failures were latent rather than immediate.

What the delayed failures have in common

Every mechanism in the latency table above shares one structural feature: the event did not complete the failure, it moved a component's margin. The component then continued to operate inside a duty cycle it no longer had headroom for, and ordinary service finished the job. That is why the failures cluster where duty is hardest, which in a mixed site means starting current, ripple current and thermal cycling, and why they surface on the first genuinely hard day rather than on a schedule.

It is also why the correct response after a significant event is not a walk-around looking for damage. A walk-around finds the hard failures, which have already announced themselves. It finds none of the moved margins.

What to actually do in the window after an event

Take the readings that would show a moved margin while the equipment still runs, so you have a comparison point later. On motor-driven equipment that is worth capturing under the same conditions each time: run capacitor microfarad against nameplate, measured with the unit locked out and the capacitor discharged and proved at zero volts; running current against nameplate full-load amps, taken with the clamp on a single conductor without opening the enclosure further than the reading requires; and insulation resistance where your scope of work and the equipment's design allow it, which is a test performed on an isolated, discharged winding and never on a connected drive.

Write the date of the event on that record. A capacitance reading with no baseline is a pass or fail. The same reading taken twice, three weeks apart, is a rate, and a rate is what distinguishes a component that took a step and stabilized from one that is still on its way down.

How to verify you got this right

Test your attribution against the mechanism, not against the timeline alone. For each failure you are attributing to the event, you should be able to name the mechanism, name the time constant that mechanism has, and show that the observed delay fits it. A failure at sixteen days attributed to direct puncture fails that test immediately, because direct puncture is a same-instant mechanism.

Then check the population. If a site has twelve similar units and one failed, the event is a weak explanation unless something distinguishes that unit's exposure, such as its position on the feeder or an unprotected control path into it. If eight of twelve failed with mechanisms and latencies that all fit, the event is a strong explanation and the four survivors are the ones now carrying unknown margin.

Finally, re-read your own report for anything you have attributed on sequence alone. "It failed after the storm" is a fact about a calendar. "It failed by a mechanism whose time constant is weeks, and here is the reading that shows the step" is a fact about the equipment.

References

  • 29 CFR 1910.333(a)(1) energized-work gate and (b)(2) electrical lockout and stored-energy release; 29 CFR 1910.147(a)(1)(ii)(C) exclusion of electric utilization equipment; 29 CFR 1926.417 construction counterpart
  • NFPA 70E-2021, 130.5 and 130.7 for boundaries and PPE and 120.5 for verification, in the edition adopted by your employer's electrical safety program
  • NEC Article 242 (surge protective devices) and 110.3(B) (installation per listing and manufacturer instructions), in the edition adopted by your authority having jurisdiction
  • Component manufacturer documentation for capacitor tolerance bands, surge protective device replacement criteria and insulation test procedures
  • See related: Correlating a Fault with a Recent Power Event; What a Surge Protective Device Actually Protects