How to Test a Restart Sequence Safely
Why this matters
Almost everybody who has "tested restart" has done the same test: switch it off, wait, switch it on, watch it start. That test tells you the machine starts from rest, which you already knew, and it tells you nothing about the states the machine will actually be interrupted in at three in the morning.
The claim: restart behaviour is a function of two variables, where in the cycle the interruption lands and how long it lasts, and a test that varies neither has not tested anything. This is how to build the small matrix that covers both, and how to run it in an order where the first thing that goes wrong is the cheapest thing that could have.
Before you deliberately interrupt anything
You are about to create an unplanned stop on live equipment, on purpose. That is a legitimate test and it carries obligations.
Get agreement first, in writing where the equipment is not yours. An interruption test can damage a driven load, spoil a process batch, or trip something downstream, and you need the owner to have accepted that before the first switch operation, not after.
Three categories where you do not run this test at all: equipment serving a life-safety function or interlocked to one; a process where a mid-cycle stop itself creates a hazard, such as a material that sets in place, a load that must not stop under weight, or a reaction that needs continuous cooling; and any machine where the restart behaviour you are trying to characterize is itself the suspected hazard, which has to be established from documentation and a de-energized inspection before you let it happen live.
Operate the disconnect the way it is designed to be operated. Use the machine's own disconnecting means, with the enclosure door closed and latched, standing to the hinge side and clear of the arc path, in arc-rated protective equipment selected per NFPA 70E-2021, 130.7 and the equipment's own incident-energy labelling, with the shock protection 29 CFR 1910.335(a) requires. Do not simulate an outage by pulling a fuse or lifting one conductor: removing one phase of a polyphase supply is a different test that damages motors, and it is not what an outage does.
Nobody's hands are in the machine during any interruption test, and no guard is off. The machine may restart automatically at any moment during this procedure, which is the entire point of running it. If a reading has to be taken inside the enclosure between tests, that is a separate operation, de-energized and locked out per 29 CFR 1910.333(b)(2) with live-dead-live verification per NFPA 70E-2021, 120.5, and mechanical isolation and stored energy per 29 CFR 1910.147 where a spring, an accumulator or a raised mass is involved.
Declare an abort before you start. One sentence, agreed with whoever is present: what condition ends the test immediately and who says so.
Build the matrix before you touch the switch
The matrix has interruption points down one axis and outage durations across the other. Both come from the machine, not from habit.
Interruption points are the boundaries of the sequence's own steps. Take the sequence of operation, or a derived timeline if there is no published one, and list the distinct states the machine occupies for long enough to be interrupted in. Four to six is typical. States lasting a fraction of a second are not testable by hand and come off the list.
Outage durations are set by four thresholds the machine already has, and each duration you test should sit on the far side of a different one:
| Threshold | Where the number comes from | What crossing it changes |
|---|---|---|
| Control-power ride-through | Published for the controller or power supply | Below it the controller never knows anything happened |
| Coastdown to standstill | Measured, once, with a stopwatch | Below it the load is still turning at restoration |
| Minimum off time | Published by the manufacturer | Below it the machine is required to hold off |
| Process settling time | Estimated from the process, usually long | Below it the machine is not thermally or hydraulically at rest |
Skip the ride-through row and you will miss the most informative failure in this whole procedure, for reasons the worked example shows.
Rank every cell by its worst credible outcome, then run ascending
This is the ordering principle and it is the only one that matters here. For each cell you intend to run, write down the worst thing that could plausibly happen: nothing, a nuisance fault requiring a reset, a protective device operating, damage to a driven component, a process loss.
Then run them from least consequential to most, and stop at the first cell whose result you did not predict. Not the first failure, the first surprise. A surprise means your model of the machine is wrong, and every remaining cell was ranked using that same wrong model, so the ranking below it is no longer trustworthy.
Running the cheap cells first is not caution for its own sake. It is what gives you a correct model to rank the expensive cells with.
Prune the matrix before you run it
A full grid is usually more tests than the equipment should absorb. The prune rule: a cell earns a run when its predicted behaviour differs from a cell already run. Where two cells should behave identically, run one and note the other as covered by inference, explicitly, so the record shows what was tested and what was reasoned.
The matrix, filled in for one machine
A packaged machine with a driven load and a heat source. Figures taken from its documentation and one stopwatch measurement, not from another machine:
- Published control-power ride-through: 100 milliseconds
- Measured coastdown to standstill: 20 seconds
- Published minimum off time: 300 seconds
- Estimated process settling: on the order of 15 minutes
Interruption points: at rest, during pre-start proving, during the timed purge, during steady run.
Durations chosen to sit on the far side of each threshold, expressed in seconds throughout: 0.05, 5, 60, and 600.
Pruned to seven cells, ranked by worst credible outcome:
| Order | Point | Duration (s) | Predicted | Worst credible outcome |
|---|---|---|---|---|
| 1 | At rest | 600 | Normal start | Nothing |
| 2 | At rest | 0.05 | Nothing observable | Nothing |
| 3 | Steady run | 0.05 | Controller rides through | Nuisance fault |
| 4 | Pre-start proving | 5 | Restart from beginning | Nuisance fault |
| 5 | Timed purge | 5 | Purge restarts from zero | Nuisance fault |
| 6 | Steady run | 60 | Holds off, restarts at 300 s from stop | Protective device operates |
| 7 | Steady run | 5 | Holds off for the full minimum off time | Restart into a turning load |
Cell 7 sits last despite being the shortest outage, because it is the one where an incorrect minimum-off-time implementation puts a start onto a load still turning. That ranking is the whole reason for the exercise.
What cell 3 found
Cells 1 and 2 behaved as predicted. Cell 3 did not, and it is worth walking because it is the failure the conventional off-wait-on test can never produce.
The interruption was 0.05 seconds, half the controller's published 100-millisecond ride-through. Prediction: the controller does not register a loss, the machine keeps running, nothing happens.
Observed: the machine faulted, on a proving input, roughly two seconds after the interruption.
The explanation is a mismatch of two ride-through numbers that nobody had put side by side. The controller's power supply holds up for 100 milliseconds. The power contactor's coil does not: an electromagnetic contactor drops out well inside that window, and dropout is fast enough that a 50-millisecond gap is comfortably sufficient to release it. So the contactor opened, the driven load lost power, and the controller never knew, because from its point of view the supply never went away.
The controller therefore continued commanding a running machine while the machine coasted, and faulted when the proving input for that device dropped out, which is exactly correct behaviour on its part. The fault code named the proving input. It is not a proving fault, and a tech reading only the code would replace a switch that did nothing wrong.
What this changes. The machine is not tolerant of short supply disturbances the way its published ride-through implies, because the published figure describes the control power only. Anything that produces brief gaps on the supply, including certain transfer schemes and utility reclosing, will produce this fault. That reframes an intermittent nobody could catch into something with a known trigger, and the sibling article on power loss and restoration covers what a sag between coil pickup and dropout does to the same hardware.
Where the procedure stopped. At cell 3, by the rule. Cells 4 through 7 were ranked assuming the controller and the contactor saw the same supply. They do not, so the ranking was rebuilt before anything longer was attempted.
What to write down
Per cell, in the record that stays with the machine: interruption point, duration in seconds, predicted behaviour, observed behaviour, and whether the machine ended in a state you can account for. The last field catches the quiet failure, where the machine restarted and ran but a damper finished somewhere it should not be or a counter did not increment. A test that ends with the machine running is not automatically a test that passed.
Also record which cells you covered by inference rather than by running, and the reasoning. The next person needs to know the difference between "tested and fine" and "reasoned to be the same as a cell we did test", because if your reasoning was wrong, that is where it is hiding.
References
- NFPA 70E-2021, 130.7 for arc-rated and other protective equipment selection, and 120.5 for live-dead-live verification
- 29 CFR 1910.335(a) for electrical protective equipment, 1910.333(b)(2) for de-energizing and lockout of electrical circuits
- 29 CFR 1910.147 for mechanical isolation and stored energy where a spring-return, accumulator or raised mass is involved
- Manufacturer documentation for published control-power ride-through, minimum off time, and the written sequence of operation, all of which are equipment-specific
- See related: What Happens on Power Loss and Restoration; How to Establish Where in the Sequence It Stopped; Why Short Cycling Protection Looks Like a Fault