The Sequence of Operation as a Diagnostic Instrument
Why this matters
Most techs treat a sequence of operation as background reading and then diagnose from the symptom, which is the expensive way round. A symptom maps to a long list of possible causes. A state boundary maps to a short one. The sequence is the only document on a job that converts "it does not work" into "it completed state ten and did not complete state eleven," and that conversion is worth more than any meter reading, because it tells you which meter reading to take. The shops that use it well are not better at electrical theory. They are better at not testing things that have already proven themselves.
Before you use a running machine as a test bench
Watching a sequence advance means the equipment is energized and moving. Take every observation you can from the controller's status display, diagnostic output, indicator lights, and audible or visible state changes without opening an enclosure. Where you must measure inside an energized control enclosure, use test leads and probes rated for the circuit, keep your other hand out, wear the protective equipment the task requires, and do not defeat a door interlock to hold a circuit made.
Where a check can be made with the equipment off, de-energize the branch circuit at the disconnecting means and apply your own lock and tag, which is the requirement at 29 CFR 1910.333(b)(2) in general industry and 29 CFR 1926.417 in construction, then prove dead per NFPA 70E-2021, 120.5 by testing your meter on a known live source, testing the conductors, and re-testing the meter. Where the check requires opening a machine rather than a panel, de-energize and lock or tag under 29 CFR 1910.147, discharge capacitors, relieve pressure to zero on a gauge, and block spring tension first. Do not force a state on a combustion path, a pressurized path, a refrigerant-bearing path, or any device serving a relief or protective function.
The instrument: a boundary that partitions the causes
A sequence is an ordered chain of states, each gated on the one before. That ordering does something no symptom can do: it splits every possible cause into two sets.
Everything the machine needed in order to reach the last completed state has already demonstrated itself. Not been checked, not been assumed, demonstrated, by the machine, in front of you. And everything the machine has not yet exercised cannot be responsible for the failure of a step it was never asked to support.
The divergence point is the boundary between those two sets, and finding it is the whole technique. A cause on the proven side is not a cause you have made unlikely. It is a cause you have eliminated, unless the fault is intermittent, which is the one condition that puts a proven state back on the table.
Why this beats starting from the symptom
Consider a machine that produces no output. Starting from the symptom, the candidate list is every component in the chain, and the usual method is to check them in order of how easy they are to reach, which correlates with nothing.
Starting from the sequence, you are not asking what is broken. You are asking how far it got. That question has a small number of answers, each of which you can observe, and each answer eliminates a block of candidates at once rather than one at a time.
The second advantage is subtler and matters more on a repeat call. A divergence point is a fact you can write down and hand to the next person. "It fails at the flow proof" survives being written in a job record. "It is probably the control board" does not, because the next tech has no way to know what it was based on.
Proven to the controller is not proven to you
One caution before the method, because it is where confident diagnoses go wrong. A state advancing means the controller received the proof it required. It does not mean the physical thing being proved is happening correctly, only that the proving device closed.
So a proven state eliminates the causes the controller was actually checking for, not the ones you might wish it were checking. If the proof is a pressure switch, the state advancing tells you pressure crossed the switch's setpoint at the switch's sensing point. It does not tell you the pressure is correct, that it is stable, or that it is present anywhere else in the system. Keep the referent straight: the quantity the controller proved is the quantity at the sensing point, not the quantity at whatever convenient port you can reach with a gauge.
A case: a machine that stops partway through
A packaged machine with a published twelve-state sequence runs, makes noise, and produces no output. No fault code is displayed. The customer reports it has been like this since the weekend.
Bisecting. Rather than watching from the start, the tech watches for evidence of a state in the middle. State 6 completes. That eliminates states 1 through 6 and everything they depend on, in one observation. Next, state 9 completes, leaving 10 through 12. Next, state 11 does not complete. Next, state 10 does complete. The divergence is at state 11: the machine finishes 10 and never finishes 11.
Four observations to isolate one boundary out of twelve, and the reason is arithmetic rather than luck. Halving a set of twelve takes four steps, because three halvings only resolve eight candidates and four resolve sixteen. Walking the sequence from state 1 would have taken up to twelve observations on the same machine, three times as many in the worst case, and every one of them would have re-proved something the machine had already proved by itself.
Reading state 11. The published sequence gives state 11 as a proof step: an output is energized, and a proof device must report within a window or the sequence holds. Two families of cause remain, and only two, which is what the boundary bought.
- The physical quantity being proved is not occurring.
- It is occurring and the proof is not being reported, which means the proof device, its sensing connection, or its wiring.
The measurement that looked decisive and was not. The tech measures the quantity at the machine, at the most accessible port, and reads it at roughly 1.8 times the proof device's stated setpoint. That looks like family 2 confirmed: the process is happening, so the proof device must be at fault. A meter across the device's contacts while running would be the obvious next move.
The measurement that actually decided it. Before condemning the device, the tech measures again at the point the device senses, rather than at the convenient port. There the quantity reads roughly 0.6 times the setpoint. The device is not failing to report. It is correctly reporting that at its sensing point, the quantity is well below the threshold, and it is below it because a manual isolating device between the two measurement points had been left partly closed.
Note what happened to the earlier reading. It was not wrong. It was a measurement of a different quantity than the one the proof device is judging, and the whole apparent contradiction came from treating two readings at two locations as if they shared a referent.
Confirming it. The isolating device is opened fully and the sequence is re-run from a cold start. It advances through 11 and 12 and produces output. That confirmation matters as much as the finding: the sequence is not only the instrument that located the fault, it is the test that proves the fault is gone, because completing the last state is a stricter check than the symptom disappearing.
Cost of the whole diagnosis: four observations to bisect, two measurements, one confirming run. Nothing was removed, nothing was replaced, and the cause was a position, not a part.
Three ways the sequence itself will mislead you
Undocumented internal steps. A published sequence describes intent. The controller may implement steps the document never mentions, particularly around initialization, self-check, and recovery from a prior fault. A state you cannot see in the document is not a state that does not exist, so an unexplained delay is a question rather than a defect.
State numbering that does not match the fault codes. Sequences and diagnostic code lists are usually written by different people at different times. Do not assume fault code five corresponds to state five. Map them explicitly once, on paper, and keep the mapping with the equipment record, because you will need it again and rebuilding it takes longer than reading it.
Conditionally skipped states. Some states are skipped by design when a condition is absent: a second stage that never runs because demand was satisfied, a purge that is bypassed after a normal shutdown, a defrost that never initiates because its own initiate condition was never met. "It never did X" is a finding only after you have checked X's own condition. Reporting a skipped state as a fault is the most common way this method produces a wrong answer, and it is entirely avoidable.
Confirming a divergence point is real
Three checks worth running before you act on a boundary, each catching a different way of being fooled.
Re-run it from a cold start and confirm the boundary lands in the same place. A boundary that moves between runs is telling you the fault is intermittent or thermally dependent, which changes the method completely, because on an intermittent fault the states before the boundary are no longer eliminated.
Confirm you observed a state completing, not merely starting. An output energizing means the state was entered. The sequence advancing to the next state means it completed. Techs routinely record the first as if it were the second, which shifts the apparent boundary one state late and sends the diagnosis into the wrong half.
Confirm the sequence you are holding has the same state count as the machine. Count the stages, circuits, or pumps in front of you against the count implied by the document before you trust any state number in it. A twelve-state sequence read against a machine that only implements ten is a numbering error waiting to happen, and every conclusion downstream of it inherits the offset.
References
- 29 CFR 1910.333(b)(2) and 29 CFR 1926.417, requirements to de-energize and lock or tag before work on or near electrical conductors in general industry and in construction
- 29 CFR 1910.147, control of hazardous energy for servicing where unexpected energization or release of stored energy is possible
- NFPA 70E-2021, 120.5, the test-before-touch sequence for verifying an electrically safe work condition
- Manufacturer control documentation conventions for sequences of operation, proof devices, and diagnostic code lists
- See related: How to Read a Control Sequence of Operation