Why a Short Test Cycle Passes a Failing System

Why this matters

Most "no fault found" tickets are not a failure of skill. They are a failure of duration. The tech measured correctly, the readings were genuinely good, and the system genuinely was working at that moment, because the fault needs run time to appear and the test did not give it any. The customer then gets a system that quits again two days later and a note in their hand that says it tested fine, which is the worst possible combination: you were wrong and you documented being wrong. Understanding which faults are born out of run time, and how long a test has to run to expose them, is the difference between a diagnosis and a coin flip.

The fault has an incubation time and your test has a duration

Every intermittent fault has a time constant: the amount of continuous operation it takes before the failing condition develops. A loose lug needs enough current-hours to heat, expand, and open. An undersized condensate path needs enough run time for water to accumulate faster than it drains. A weak start component fails on the tenth start of the hour, not the first.

The only question that matters for verification is whether your test duration exceeds that incubation time with margin. It does not matter how good your instruments are. A perfect reading taken at minute 12 of a fault with a 90 minute incubation is a true statement about minute 12 and a false statement about the system.

Write the rule as a ratio and it stays honest: test duration divided by reported time to fault. Under 1.0 you have proven nothing about the complaint. At 1.0 you are betting on the customer's estimate being exact. At 1.5 you have real coverage, because a customer's "about an hour" is genuinely somewhere between 40 and 80 minutes.

The four things that only change after long run

Short tests miss faults that live in one of four accumulating quantities. Naming which one you are hunting tells you what to watch while the clock runs.

Heat rise in a component. Windings, connections, semiconductors, and bearings all climb toward a steady-state temperature over tens of minutes, not seconds. A motor winding or a control enclosure often needs 30 to 60 minutes of continuous operation to reach its final temperature. Resistance climbs with temperature, so a marginal connection that measures acceptable cold can drop enough voltage hot to starve the load or trip a protector.

Accumulation of something physical. Ice on a coil, scale on a heat surface, debris at a strainer, water in a pan, sediment stirred into suspension, dust bridging a gap. All of these are functions of run time and none exist at minute five.

The load catching up with the capacity. Many systems look fine at light load and only fail when the demand has been sustained long enough that the system is no longer keeping up. This is where the true undersizing, the partly blocked path, and the mild loss of output hide. At minute five nearly every system meets a small load.

Cumulative starts and stops. If the fault is driven by cycling rather than continuous run, your continuous test hides it. A weak start component, a contactor with pitted faces, or a control with a marginal timing circuit can survive one start and fail on the eighth.

The runtime clock is not the ambient clock

The most common misread is to blame the weather for a runtime fault, because both show up "on hot days." They are separable, and separating them changes the entire repair.

An ambient-driven fault tracks the outside condition and appears at whatever run time the condition is present, including immediately after a cold start on a hot afternoon. A runtime-driven fault tracks the hour meter and appears after the same amount of operation regardless of whether it is a mild morning or a hot afternoon, though heat shortens the incubation because the component starts closer to its limit.

The clean separator: ask whether the system has ever failed in the first ten minutes of a cold start. If yes, at least part of the cause is condition-driven. If the answer is a firm no across many occurrences, you are looking at run time, and no amount of testing on a mild day at minute ten will find it. There is a full method for running that separation in the related HowTo.

A worked example carried all the way through

A residential cooling complaint: the system quits producing cool air "after about an hour and a half" on warm afternoons, then works again after it has been off for a while. First visit ran a 15 minute test cycle. Every reading was in range. The ticket was written up as operating normally.

Do the ratio before anything else. Reported time to fault is 90 minutes. Test duration was 15 minutes. That is 15 divided by 90, about 17 percent of the incubation time. The first visit tested less than a fifth of the way to the event and then made a claim about the whole of it.

The callback landed three days later. This time the tech planned a run of 1.5 times the reported figure, so 135 minutes, and logged readings every 10 minutes instead of once. Illustrative values, but the shape is what matters: the temperature difference across the air side opened at about 18 degrees F at minute 10, which is normal, held near that through minute 40, and then began narrowing. By minute 100 it was near 9 degrees F, half the starting split, with no change in the thermostat setting or the outdoor temperature during that window. Current draw on the outdoor fan motor tracked the same curve, sitting comfortably under nameplate early and creeping to just over nameplate by minute 95.

Two readings, same direction, same timing. That is not drift, that is a developing restriction to heat rejection that only exists once the machine has soaked. The confirmed cause was a partly blocked airflow path on the outdoor side, clean enough to look acceptable at a glance and to pass a short test, restrictive enough that once everything was hot the unit could no longer shed heat as fast as it made it.

Now the economics, in hours rather than currency. The original visit billed 1.25 hours. The callback consumed roughly 1.4 hours of drive time plus 1.0 hour on site, so about 2.4 hours, which is close to twice the labor the first visit billed and none of it recoverable. Staying on the first visit to complete a 135 minute run would have added about 1.75 hours net of the 15 already spent, which is less than the 2.4 the callback took, before you count the customer's three days without cooling and the note in their file that said it tested fine.

That comparison is the whole argument for staying. Not diligence for its own sake, arithmetic.

Deciding how long is long enough

Set the run time from the customer's own report, not from habit.

What the customer reports Minimum test duration What you are watching for
Fails after a stated time 1.5x that time, continuous A trend in two independent readings, not one
Fails "after a while", no number 60 minutes minimum, or until two readings stop moving Steady state reached, then held
Fails on the second or third cycle Full cycles, not clock time. Run at least 4 Start behavior degrading start to start
Fails only under heavy demand Long enough to reach and hold full load Whether output can hold, not whether it starts
Fails at startup, never after A short test is legitimate here Start sequence, inrush, control handoff

The last row matters as much as the others. Not every fault needs a long test. If the complaint is genuinely a startup event and the system has never quit mid-run, a 15 minute test is proportionate and a 135 minute one is time you are not billing anyone for.

When you genuinely cannot stay

There are real reasons you cannot sit through a 2 hour run: the next call, an after-hours slot, a customer who has to leave. Do not solve that by shortening the test and keeping the same conclusion. Solve it by changing the conclusion or changing the instrument.

Change the conclusion. Write the ticket to say exactly what you established: system operated within range through 15 minutes of run, complaint describes a fault developing near 90 minutes, verification incomplete. That is honest, it is defensible, and it sets up the return visit as planned work rather than a callback.

Force the condition instead of waiting for it. Some incubating faults can be induced faster by raising the load, blocking a normal relief path in a controlled and safe way, or restoring the exact operating mode the customer uses. This shortens the wait without shortening the coverage. It only works when you can name the mechanism you are accelerating.

Leave an instrument. A logging thermometer, a current logger on the supply, a pressure recorder, or in many cases the equipment's own runtime and fault history will cover the hours you cannot. You come back to data instead of to a fresh guess.

Delegate the observation. Give the customer a specific, timed observation task with a written form: note the clock time the system started, the clock time the symptom appeared, and one concrete observable at each. Two clean events from the customer will bracket the incubation time better than a third short test by you.

What doing this wrong looks like in the field

The signature is a file with two or three visits, all showing good readings, all short, and a complaint that never changes. Nobody in that chain is incompetent. Each one arrived, tested for the length of a normal check, found nothing, and left a note that made the next tech trust the previous readings and test even less.

The second signature is a parts-swap spiral. When a short test finds nothing, the pressure to do something replaces the most suspected component. The system works for a week because the fault was never in that component and the incubation clock simply restarted after the power cycle. That temporary success then gets recorded as a fix, which corrupts the history for everyone who follows.

The catch for both is simple and worth building into your ticket template: any ticket that closes with "no fault found" must state the test duration and the reported time to fault side by side. If the ratio is under 1.0, the ticket is not closed, it is paused.

How to verify you got this right

Three checks before you call a long-run verification good.

Confirm you reached steady state, not just a long clock. Two independent readings that have stopped moving for 15 minutes is steady state. A clock that says 90 minutes while a temperature is still climbing is not.

Confirm you ran the system the way the customer runs it. Same mode, same setpoint, same doors open or closed, same accessories energized. A verification run in a mode the customer never selects proves that mode works.

Confirm the fault would have shown. If you ran 135 minutes and nothing moved, say so explicitly and then ask what was different about the failure days: a different load, a different setting, a different consumable in the machine, someone else operating it. A clean long run is a real result and it should redirect the investigation, not end it.

References

  • Manufacturer documentation for equipment startup, run-in, and steady-state verification times
  • Trade-standard practice for post-repair verification and operational testing
  • See related: How to Separate Runtime-Driven Faults From Ambient-Driven Ones
  • See related: The Duty-Cycle Questions to Ask Before You Drive Out
  • See related: The Fault That Only Shows in Heat, Cold, Humidity, or Storm