The Fault That Only Exists Under Concurrent Demand

Why this matters

A concurrency fault is the one that survives a competent visit. The tech tests the equipment properly, gets clean readings, writes it up honestly, and leaves, and the failure returns the same evening. After two or three of those the customer stops believing the shop and starts believing the internet, and the third visit begins from a worse position than the first.

The reason it survives is structural, not sloppiness. A single-actor test cannot detect a fault whose necessary condition is a second actor, and no amount of care inside a single-actor test fixes that. What breaks the cycle is knowing that concurrency faults come in a small number of distinguishable mechanisms, each with its own signature, and that the signature is readable before you own a single measurement.

Two rules that apply before any of this

Do not defeat, jumper or hold closed a protective device to keep two actors running together. The protector operating is the finding; removing it converts an investigation into an ignition, rupture or shock risk with nothing left to stop it.

Thermal mechanisms mean components that are hot enough to burn on contact and enclosures that hold that heat. Read temperature with a non-contact instrument or a thermal imager rather than a hand, and do not open a hot enclosure to probe until the circuit is de-energized, locked, tagged and proven dead using the live-dead-live sequence in NFPA 70E-2021, 120.5, with the de-energize-and-lockout duty for work on electric circuit parts at 29 CFR 1910.333(b)(2). Where the shared resource is mechanical or pressurized, isolate and bleed or block it under 29 CFR 1910.147 before service.

Why a single-actor test is blind by construction, not by accident

This is worth stating plainly because it changes how you read a clean report. A test energizes a system and observes it. If the fault's necessary condition is "actor B is also running," then a test that does not run actor B is not a weak test of the fault, it is not a test of the fault at all. The clean reading it produces is a correct measurement of a condition that is not the failure condition.

That has a practical consequence: a clean single-actor result never lowers your suspicion of a concurrency fault. It carries no information about it in either direction. Techs treat a clean test as evidence of health, which is right for most faults and wrong for this family, and that mistaken inference is what sends the second visit down the same road as the first.

Five mechanisms, sorted by what is shared in time

Sorting by what is shared in supply (a circuit, a pipe) does not separate these, because several mechanisms can ride the same physical resource. Sorting by what the two actors contend for in time does.

1. Instantaneous budget. The two demands overlap at their peaks and the shared resource cannot deliver both at once. Current, pressure, flow, combustion air. The characteristic moment is a start or a peak, and the event is brief.

2. Cumulative thermal. Neither actor alone raises a joint, conductor, winding or enclosure past its limit. Both together do, and it takes minutes because heat accumulates. The characteristic moment is late, after dwell, and the failure is a thermal protector or a softened or opened connection.

3. Control arbitration. Two controls want the same actuator, valve, damper or priority resource, and the logic that resolves the conflict either does not exist or resolves it badly. No resource is exhausted; the system is simply issuing contradictory orders. This one produces the strangest symptoms, because the equipment does things nobody asked for.

4. Shared return path. The two actors share the way back rather than the way in: a common neutral, a shared drain, a common vent, a shared condensate line. The tell is that the fault appears on an actor that is not even the heavy one, because it is being pushed on from behind by the other actor's return.

5. Timing beat. Both actors cycle on their own schedules, so they only overlap when their periods drift into alignment. Any of the four mechanisms above can hide inside this one. The distinguishing feature is that the fault appears on a schedule nobody recognizes as a schedule.

The signature table

Read your reported symptom across this before you choose a test, because the mechanism determines the test and the wrong test returns clean.

Mechanism When in the overlap it fails Does dwell matter Which actor shows it What clears it instantly
Instantaneous budget At a start or a peak, seconds No Usually the last one to start Removing either actor
Cumulative thermal Minutes in, later each cycle Yes, decisively Whichever shares the hot joint Cooling, not just load removal
Control arbitration Whenever both call, no timing pattern No The actuator, not either actor Removing one call
Shared return path Tracks the heavier actor's cycle Sometimes Often the lighter actor Removing the heavy actor
Timing beat Only on alignment, otherwise never Depends on what hides inside Whichever the inner mechanism picks Nothing you can do on demand

The most useful column is the third. Dwell dependence separates the two mechanisms that get confused most often, instantaneous budget and cumulative thermal, and it separates them with a stopwatch rather than an instrument.

The beat, and why "random every few days" is a schedule

Customers describe timing-beat faults as random, and techs write them down as random, and random is where a diagnosis goes to die. It is usually not random. Two independently cycling devices overlap only when their cycles align, and the interval between alignments is set by the difference between their periods, not by either period.

Take two illustrative cycles: one device running about every 6.00 hours and another about every 6.25 hours. The alignment interval is the product of the two periods divided by their difference, which is 6.00 times 6.25 divided by 0.25, or 150 hours. That is 6.25 days. So a fault that requires those two to coincide shows up roughly once a week, at a different clock time each occurrence, with nothing in the customer's day to explain it.

That is why the useful question to a customer is never "what time does it happen." It is "what was running," and if they cannot say, the answer is a log with timestamps rather than another visit. Two or three timestamped occurrences will show an interval, and an interval that is stable but does not match anybody's habits is a beat, which tells you to go looking for two cycling devices with close periods.

Narrowing five mechanisms to one, on a real-shaped call

The complaint: a protective device on one piece of equipment trips, but only sometimes, and the customer thinks it is worse in warm weather. Single-actor testing on two previous visits was clean. The customer eventually reports that it only happens when a second piece of equipment is also running, which puts us in this family.

Rule out instantaneous budget. The pair is brought up together, twice, with a capture-capable meter watching the shared node through both starts. Nothing trips at start overlap either time, and the peak readings are unremarkable. An instantaneous-budget fault fails at the peak, and this one survives every peak it is given.

Rule out control arbitration. Only one control commands the actuator in question, and removing the second actor's control call while leaving its load energized by other means changes nothing. Arbitration faults care about the call. This one cares about the load.

Rule out shared return path. A return-path fault typically shows up on the lighter actor while the heavier one runs. Here the second actor runs completely normally throughout, including during the trip, and its own readings never move.

Rule out timing beat. The fault reproduces on demand, every time both actors run long enough. A beat fault cannot be reproduced on demand, because you would have to wait for the alignment.

That leaves cumulative thermal, and dwell confirms it. The first controlled run, on a cool morning, trips after about 11 minutes of both actors running. The same test repeated the same afternoon, with the space noticeably warmer, trips after about 6 minutes. That is roughly 45% shorter time to trip against the morning run, and dwell-to-failure shortening with ambient is the signature of heat accumulating faster from a higher starting point. The search then goes to the joints and conductors carrying both loads, with a thermal imager under load first and a physical inspection after the circuit is de-energized, locked, tagged and proven dead.

What the misdiagnosis would have cost. Read as instantaneous budget, the same symptom leads to upsizing the supply or splitting the load, which is real work with a real invoice and does not touch a degraded joint. The joint keeps heating, the trip point moves around with the seasons, and the shop that did the upsizing owns the callback. The distinguishing observation was free: the fault never happened at a start, and it always happened after minutes.

Why these get misfiled in your own records

Concurrency faults enter the shop's history as "intermittent, could not reproduce," and that phrase is where they go to hide. Two consequences worth designing around.

The ticket does not record what else was running, because no field asks. Add one: on any could-not-reproduce write-up, record what else in the building was operating at the time of the customer's reported failure, in the customer's words. It costs one question and it is the single field that makes the third visit shorter than the second.

The pattern lives across tickets rather than inside one. A concurrency fault looks like bad luck on any one visit and looks like a mechanism across three. Whoever reads the history before dispatching the next truck is the person who can see it, and they can only see it if the first two techs wrote down the second actor.

How to verify you have identified the right mechanism

  • You can state which mechanism, and name the observation that eliminated each of the others. Not a hunch about which is likeliest. Concurrency mechanisms are eliminated on evidence, and if you cannot name the eliminating observation you have chosen rather than diagnosed.
  • Your test matched the mechanism's timing. A dwell test for thermal, a capture test for instantaneous budget, a call-versus-load separation for arbitration. Wrong timing returns clean and reads as exoneration.
  • You reproduced it in both directions at least twice. Fails with both, fine with one, fails with both again. Concurrency faults are unusually easy to demonstrate once you know the pair, and a demonstration is what convinces a customer who has already been told twice that nothing is wrong.

References

  • NFPA 70E-2021, 120.5, the live-dead-live process for establishing an electrically safe work condition before probing a hot enclosure
  • 29 CFR 1910.333(b)(2), OSHA electrical safe work practices, for de-energizing and locking or tagging before work on electric circuit parts
  • 29 CFR 1910.147, OSHA control of hazardous energy, for isolating a shared mechanical or pressurized resource before service
  • Manufacturer service literature for protective-device ratings, published cycle intervals and thermal limits on the specific equipment
  • See related: How to Test Two Systems Running at Once; The Interaction Between Two Loads That Creates a Fault; How to Map Which Loads Share a Circuit or Supply