How to Test a System That Only Fails After Hours of Running

Why this matters

A fault that needs six or eight hours of continuous operation to show itself will never appear during a normal service call. You arrive, run the equipment for twenty minutes, everything reads clean, and you leave. The customer calls back that afternoon. Two or three rounds of that and you have burned more labor hours on "no fault found" visits than the repair itself would have cost, and the customer now believes you cannot fix their equipment. The fix is not more meter skill. It is treating accumulated run time as a variable you deliberately control, the same way you would control voltage or pressure in any other test.

Lead with the safety gate, not the test plan

Forcing a system to run far longer than it normally would is not a neutral act. Before you commit to an extended run, walk the hazard list and close each item.

  • Combustion equipment: confirm the flue and combustion-air path are clear, and confirm the safety chain (flame proving, high limit, pressure switch) is intact and not jumpered. A long forced run on a unit with a defeated safety is how a nuisance fault becomes a fatality.
  • Water plus electricity: if the extended run will produce condensate, drain flow, or any chance of overflow near energized parts, verify the drain path first and put a catch in place. An eight-hour run is eight hours for a slow drip to find a control board.
  • Stored energy and pressure: know where you will need to break into the system for readings, and know the isolate-and-relieve sequence for that point before you start, not at hour six when you are tired.
  • Thermal protection stays in the circuit. If your test depends on bypassing a limit to keep the unit running, you no longer have a test, you have an uncontrolled experiment. Find another way to reach the condition.

Write down the abort criteria before you start: the temperature, pressure, current, or sound that means you shut it down immediately regardless of what you have learned. A test with no abort threshold is a test that runs until something breaks.

Step 1: Convert the complaint into a time-to-fault number

"It quits in the afternoon" is not a test parameter. "It quits after about eight hours of running" is. Get there by asking the customer three specific questions, in this order:

  1. What time does it start each day, and does that time ever vary?
  2. What time does it fail, and how tight is that window across the last several days?
  3. How long before it will run again, and does it come back on its own or does someone reset it?

The third question matters as much as the second. A unit that recovers on its own after twenty to forty minutes is telling you something is cooling back below a threshold. A unit that stays dead until someone cycles power is telling you something latched. Those are different faults with the same complaint.

If the customer cannot answer with confidence, do not guess. Leave a simple log sheet with start time, fail time, and recovery time, and come back after three or four days. Three clean data points beat one confident recollection.

Step 2: Prove it tracks run hours, not clock hours

This is the step almost everyone skips, and it is the one that decides which family of causes you are in. Two very different things produce "fails in the afternoon":

  • Run-hours driven: heat, wear-in drift, accumulation, or a consumable loading up. The clock is irrelevant; what matters is how long the machine has been working.
  • Clock-hours driven: peak ambient temperature, peak building load, a utility or supply condition, a scheduled event elsewhere in the building, solar gain on one wall.

The discriminator is free. Find a day when the start time was different, or create one. If the equipment normally starts at 06:00 and fails around 14:00, ask the customer to start it at 09:00 one day. If it now fails around 17:00, the elapsed run time is the driver and you can ignore everything that happens at 14:00 on the clock. If it still fails around 14:00, run hours are not your variable and you should be chasing ambient, load, or supply instead.

One shifted start time eliminates half the possible causes before you open a panel. Nothing else in this procedure is that cheap.

Step 3: Freeze every other variable

An extended run test is only readable if the run itself is repeatable. Before you start the clock:

  • Set controls to a fixed condition rather than letting a thermostat or process demand cycle the unit. A unit that runs 40 percent of the time on one test day and 70 percent on the next has not given you two comparable results.
  • Record ambient at the equipment, not the weather forecast. Equipment in a closet, attic, or machine room lives in a different climate than the building.
  • Note filter, strainer, and screen condition at the start, and leave them alone. Changing a filter mid-test destroys the baseline you are building.
  • Confirm supply voltage or supply pressure at the start and note it. If it drifts during the run, you want to know that it drifted rather than discover it later as a surprise.

Step 4: Set the reading cadence against the curve, not the clock

Temperatures and the readings that follow them do not rise linearly. They rise fast at first, then flatten as the system approaches equilibrium, which means evenly spaced readings waste effort early and miss the interesting part late. Sample dense at the start to capture the shape, then dense again around the expected failure window.

A cadence that works for an eight-hour time-to-fault: readings at 0, 10, 20, and 30 minutes, then hourly, then every 10 minutes starting one hour before the expected failure. You are looking for one of three shapes:

  • Rises and flattens well below the trip point. The system reached equilibrium and is stable, so the failure is probably not a simple thermal accumulation. Look for something that changes with run time other than temperature.
  • Rises and is still climbing at the moment of failure. Something is not rejecting heat, and the trip is the consequence, not the cause. Find the lost cooling path.
  • Flattens, then starts climbing again late. Something changed partway through the run - a fan slowed, a filter loaded, a control staged differently, a fluid level dropped.

The third shape is the most informative and the easiest to miss with hourly-only sampling.

Step 5: Take the failure-moment readings in a fixed order

You have a short window, often under two minutes, before the system cools enough to start healing itself. Decide the order before you are standing there. A workable default: the reading that vanishes fastest first, the reading that persists last.

  1. Whatever changes fastest on cooling (surface or air temperature at the suspect component).
  2. The electrical or flow reading under load at the failure point.
  3. The state of every protective device and indicator.
  4. Photographs of every gauge, display, and indicator light in one pass.

Then, and only then, start taking things apart. A tech who opens the panel first has already dumped the heat that was the whole point of the exercise.

Step 6: Use the warm restart to buy test cycles

Once you have one confirmed time-to-fault from cold, you rarely need to repeat the full run. Let the system cool only partially, restart it, and the time-to-fault collapses, because you started partway up the curve. That shortened interval is both a confirmation of the thermal mechanism and a practical tool: it turns a one-attempt day into a three or four attempt day, which is what lets you test a candidate fix the same visit instead of scheduling another.

A worked example, carried through

An air-moving system in a light commercial space quits most afternoons and comes back roughly half an hour later. All values below are illustrative.

The customer log shows failures at 14:10, 13:55, and 14:25 on days the unit started at 06:00, so about 8 hours of run each time, tight to within half an hour. On the fourth day the space opened late and the unit started at 09:00. It failed at 17:05, again about 8 hours. Run hours are the driver, and every afternoon-ambient theory is dead before a panel comes off.

Baseline readings during the controlled run, taken as air temperature inside the control enclosure minus ambient at the equipment: 18 degrees F above ambient at 30 minutes, 34 above at 4 hours, 46 above at 7.5 hours, and still rising when it tripped. That is shape two: no equilibrium, so the enclosure is not rejecting the heat it makes.

At the moment of failure, current draw measured within nameplate, which rules out a genuine overload driving a protective device to do its job correctly. So the trip is not the protection working, it is a component losing margin at temperature.

Then the warm restart. After 25 minutes of cooling, the unit restarts and fails again in 2.0 hours instead of 8.0. That is a 4x reduction in time-to-fault from a partial cool-down, which confirms the mechanism is accumulated heat rather than a timer, a counter, or a scheduled external event. It also means the rest of the diagnosis costs 2 hours per attempt instead of 8, so three candidate checks fit in one afternoon.

The economics matter here. The full cold test consumed 8 run hours but only about 1.5 tech hours of actual attendance, because the middle of the run was instrumented rather than watched. Compare that to three prior no-fault-found visits at roughly 1 hour each, which produced nothing. The extended run cost about half again as much attended time as one wasted visit and produced a confirmed mechanism.

How to verify you got this right

You have run this correctly when you can state four things without hedging: the time-to-fault from cold, whether it tracks run hours or clock hours, the shape of the rise curve, and the reading that was out of range at the moment of failure but in range cold. If any of those four is still a guess, the test is not finished, and swapping a part now is a coin flip you will pay for on the callback.

After a repair, prove it with the same instrument you diagnosed with. Run the full cold cycle past the old time-to-fault by at least 25 percent - if it used to quit at 8 hours, run it 10 - and show the customer the curve flattening where it used to keep climbing. That is a verification they can understand, and it ends the argument about whether the problem is really fixed.

What changes the answer

  • If the time-to-fault varies widely between days (say 4 hours one day and 9 the next) rather than clustering, run hours alone are not the variable. Look for a second factor stacking on top, most often ambient or duty cycle percentage.
  • If a warm restart does not shorten the time-to-fault, the mechanism is probably not thermal. Suspect something that accumulates and does not reset on cooling: a loading filter or strainer, a level dropping, a counter or a control routine.
  • In a residential setting where you cannot force an 8 hour run without making the space unlivable, hand the run off to the customer with a log sheet and instrument it, rather than abandoning the method.
  • On equipment under a manufacturer warranty, check whether a forced continuous run outside normal control is an excluded operating condition before you start. Diagnosing your way into a voided warranty is a bad trade.

References

  • OSHA general industry guidance on machine operation and lockout during testing
  • NFPA 70B, recommended practice for electrical equipment maintenance
  • Manufacturer documentation for continuous-duty ratings and thermal limits
  • See related: How to Set Up an Extended Run Test Without Camping On Site; What Changes Inside a System During a Long Run Cycle; The Extended Run Test SOP