The Extended Run Test SOP
Purpose
To give the shop one repeatable procedure for diagnosing a fault that only appears after a system has been running for a long stretch, so that these calls stop producing repeat "no fault found" visits. The procedure controls run time as a deliberate test variable, captures the failure while it is happening, and produces a documented result the customer can see. It also defines who authorizes the run, who watches it, and when it gets shut down, because forcing equipment to operate to the point of failure is not something a tech should be improvising alone in a customer's building.
Scope
Applies to any service call where the reported fault has a time-to-fault measured in hours of operation, recovers on its own after a period of sitting, and has already produced at least one inspection that found nothing.
Does not apply to faults present on a cold start, faults that appear within the first hour of running, faults tied to a specific outdoor condition rather than to run time, or any situation where the equipment has a known active safety defect. Those route to their own procedures.
Never applies where the suspected failure mode could produce an uncontained release of fuel, water, steam, or pressurized fluid, or where any safety device in the chain is defeated, damaged, or has been bypassed. In those cases the defect is repaired first and the extended run test happens afterward, as verification.
Roles and responsibilities
| Role | Responsibility |
|---|---|
| Dispatcher / office | Runs the intake questions, records start time, fail time and recovery interval, books the return visit against the expected failure window, and puts the customer instruction on the job card |
| Lead technician | Clears the safety gate, obtains written authorization, sets the abort criteria, installs and verifies instrumentation, performs the failure-moment capture, and signs the result |
| Customer contact | Remains on site for the duration of the run, watches for the listed abort conditions, operates the shutdown if any occurs, and does not reset or open the equipment |
| Shop owner / service manager | Approves any unattended run, reviews any test that reaches an abort condition, and approves the post-repair verification length |
No extended run test starts without a named customer contact for the full duration. If nobody on their side can hold that role, the test is attended or it does not happen.
Procedure
1. Qualify the call before booking the test
Dispatch asks and records four things: the daily start time, the time of failure, how tight that failure window has been across the last several days, and how long before it runs again. If the customer cannot answer, send a log sheet and re-qualify after three to four working days. Do not book an extended run test on a remembered description.
Two answers disqualify the call from this procedure. A recovery that happens instantly on a reset points at a latching control fault, not a thermal one. A failure time that stays fixed on the clock across days with different start times points at ambient or a scheduled external event, not at run hours.
2. Clear the safety gate on arrival
Before any instrument is placed, the lead technician confirms and records:
- Every protective device in the chain is present, undefeated, and has not been repeatedly reset by the customer. If it has been reset more than a couple of times, the equipment stays down until the cause of those trips is understood.
- Where combustion is involved, the flue, combustion-air path, and spillage check are verified this visit, not accepted from a prior one.
- Any condensate or drainage path that will carry many hours of flow is clear, with a catch placed if there is any chance of overflow near energized parts.
- The isolation and pressure-relief sequence for every point that will be opened during the test is known and stated before the run begins.
If any item cannot be cleared, the test does not start. That is a stop, not a judgment call.
3. Set and post the abort criteria
The lead technician writes the abort thresholds on a card and tapes it to the equipment before the run. At minimum: a temperature, a pressure or current where applicable, and the sensory conditions - smoke, burning smell, unusual noise, visible water, fuel odor. The card names the exact switch or breaker to operate and states that the customer contact operates it immediately without calling first.
Anything that reaches an abort condition ends the test. The result is still a data point and gets documented.
4. Obtain written authorization
Before the run, the customer signs off on four plain statements: what will run and for how long, that the equipment will be operated outside its normal control pattern, that the test is designed to reach the failure they have been reporting, and the attended hours it will cost them. Verbal agreement is not sufficient here, because the test deliberately drives equipment to fail and that must not be a surprise afterward.
5. Record the baseline and install instrumentation
Write down every relevant reading, the ambient at the equipment, the time, and the condition of filters, strainers, and screens. Do not change any of them; a filter swapped mid-procedure invalidates the run.
Install a minimum of three logged channels: the suspect parameter, a reference channel (ambient at the equipment or supply condition), and a run-state channel proving the equipment actually ran. Sample at 30 to 60 seconds for a multi-hour run. Set the logger clock against a phone. Before leaving the equipment, pull several minutes of recorded data off every channel and confirm it shows real, plausible, moving values.
6. Run the test and cover the failure window
Attend the first 30 to 45 minutes, which is where the fast masses settle and where a setup error announces itself. Return no later than 45 minutes before the expected time-to-fault.
At the moment of failure, capture in this order, without opening anything first: temperature at the suspect point and ambient, the electrical or flow reading under load, the state of every protective device and indicator, then photographs of every display and gauge. The window before things cool is short, often under two minutes.
On a suspected heat-soak fault, continue logging through the post-shutdown period as well. The peak frequently occurs 5 to 25 minutes after the equipment stops, because cooling flow stops with it.
7. Confirm the mechanism with a warm restart
After a partial cool-down, restart and record the new time-to-fault. A time-to-fault that collapses to a fraction of the cold-start value confirms thermal accumulation. A time-to-fault unchanged from cold points instead at a run-hour counter in the control, or at a consumable loading progressively, neither of which resets on cooling.
Use the shortened cycle to test candidate causes the same day. This is the point of the whole procedure: the first confirmation is expensive and every one after it is cheap.
8. Document, remove, and hand back
Remove every jumper, clamp, probe, and tag, and account for each against the list you made when you installed them. Restore all controls to their normal settings and confirm the equipment operates normally under its own control before you leave.
The job record carries: the confirmed time-to-fault, whether it tracked run hours or clock hours, the shape of the curve, the reading that was out of range hot and in range cold, and the exported log data. That package is the diagnosis, and it is what justifies the recommendation to the customer.
9. Verify the repair with a longer run than the original fault
After the repair, run the system from cold past the original time-to-fault by at least 25 percent. If it used to quit at 8 hours, run it 10. A verification run that stops at the old time-to-fault proves almost nothing, and it is exactly where a second contributing cause hides - a system with two faults will often clear the first failure point and trip somewhat later, which only a longer run reveals.
Show the customer the before and after curves side by side. A flat curve where there used to be a rising one closes the conversation better than any explanation.
A worked example, carried through
A commercial installation quits about 8 hours into every working day and recovers in roughly 40 minutes. All values below are illustrative.
Dispatch qualifies it and finds one day with a 09:00 start that failed at 17:05, versus 06:00 starts failing around 14:00. Run hours, not clock hours, confirmed before a tech is dispatched.
The safety gate holds up the test by half an hour: the customer's staff had reset the protective device eleven times in two weeks, so the equipment stays down until the trip cause is established as thermal rather than an overload. Current measured at a subsequent trip comes in at about 88 percent of nameplate, which settles it.
Cost of the test under this procedure: 0.75 hours setup and start-up verification, 1.25 hours covering the failure window, 0.5 hours teardown and data pull. That is 2.5 attended hours against 8 run hours, so roughly 30 percent of what a fully attended run would have cost, and the tech ran other calls in the gap.
The warm restart returns the fault in 2.0 hours against 8.0 from cold, a 4x reduction, confirming thermal accumulation and making three further checks fit into one afternoon.
Verification run: 10 hours against the original 8, which is 25 percent longer, and the suspect temperature flattens at hour 4 where it previously climbed all the way to the trip. Documented, exported, and shown to the customer.
Against three prior no-fault-found visits at roughly 1 hour each, the procedure cost about 2.5 attended hours plus the verification, and produced a repair that held. The comparison, stated in hours rather than in apologies, is also the strongest thing you can put in front of a customer who is deciding whether to keep using your shop.
Records and retention
Keep the exported log data with the job record, not on the tech's device. The curve is the evidence for the repair recommendation, it is the baseline for the next visit on the same equipment, and on a piece of equipment your shop maintains it is the reference that turns next year's ambiguous complaint into a five-minute comparison.
References
- OSHA general industry guidance on lockout, machine operation during testing, and test-equipment use
- NFPA 70E and NFPA 70B guidance on energized work and electrical equipment maintenance
- Manufacturer documentation on continuous-duty ratings, thermal limits, and protective device reset behavior
- See related: How to Test a System That Only Fails After Hours of Running; How to Set Up an Extended Run Test Without Camping On Site; How to Log Temperature Drift Across a Full Duty Cycle