Intermittent Fault Investigation and Monitoring
Purpose
An intermittent fault is the only call a shop can lose money on twice with the customer's full agreement. Nothing is wrong when you arrive, so a part gets replaced on a hunch, and a replaced part is indistinguishable from a repair until the fault comes back, at which point the shop has spent both the visit and its credibility.
This procedure replaces the hunch with an instrumented window. It sets the trigger hypothesis before anything is deployed, matches the instrument to that trigger, states in advance what ends the deployment, and closes with either a named finding or an honest no-event result carrying a written recheck trigger. What it forbids is the middle path, where a part is swapped and the ticket says resolved.
Scope
Covers a reported electrical fault in a dwelling or small commercial occupancy that cannot be reproduced at the visit: a circuit that drops occasionally, a breaker that trips weekly, a device that works until it does not, lighting that misbehaves in certain weather.
Does not cover a fault present at the visit, which goes to the procedure for its symptom. Does not cover the supply-side recorder deployment and utility correlation that a confirmed voltage complaint needs, which the Power Quality Complaint Investigation SOP owns, and it does not cover the thermal survey methodology or its acceptance thresholds, which belong to the Infrared Scan of a Panel SOP.
Roles and handoffs
| Role | Owns | Hands off |
|---|---|---|
| Office | Intake, and issuing the customer log sheet before the visit | The log sheet and a clock reference to the customer, so their times and the logger's agree |
| Lead technician | The trigger hypothesis, the fixed screens, the instrument choice, the deployment and the finding | The written removal criterion to the office, with the retrieval date it implies |
| Office | Booking the retrieval, and chasing a customer log that stops being filled in | The retrieval visit onto the board the day the criterion is met, not a month later |
Procedure
Turn the complaint into a trigger hypothesis with a written name. Ask what the fault follows: minutes of running time, a door slamming, rain, a specific appliance, a time of day. Acceptance: one named trigger with at least three dated instances behind it, or an explicit statement that no trigger is apparent. Wrong looks like a symptom list with no pattern attempted. Stop rule: fewer than three instances and no pattern means the customer keeps the log for two more weeks before anything is deployed, because an instrument aimed at nothing records nothing. Hazard: none at this step, it is an interview.
Try to reproduce the fault under its own named trigger before spending anything. Run the appliance, load the circuit for the stated time, wet the exterior fitting with a hose where rain is the trigger, cycle the equipment that precedes it. Acceptance: either the fault reproduced, which ends this procedure and sends it to the matching diagnostic SOP, or a written record of what was attempted and for how long. Wrong looks like ten minutes of trying and calling it not reproducible. Stop rule: reproduction ends monitoring entirely; you do not log a fault you can make happen on demand. Hazard: deliberately wetting an energized outdoor fitting can energize the water stream and the surface you are standing on, so the circuit is opened and proved dead before water is applied and the test is run as a leakage measurement afterward, never live.
Run the fixed screens before any instrument is left behind. Regardless of the trigger, take the service neutral trio at the disconnect, compare the suspect breaker connection's temperature against an identically loaded reference, and take an insulation resistance reading on the suspect circuit. Acceptance: line-to-neutral readings inside the ANSI C84.1 Range A band of 114 V to 126 V on a 120 V nominal base and summing to line-to-line, the suspect connection within a few degrees of the reference at the same current, and insulation resistance above the 1 megohm field minimum for 600 V class wiring. Wrong looks like skipping these because the fault is "intermittent," when a failing neutral and a loose stab both present as intermittent. Stop rule: any screen failing ends monitoring and hands to that screen's own SOP. Hazard: energized readings with the deadfront off, so wear the arc-rated clothing and face protection your program assigns per NFPA 70E-2021, 130.5 and 130.7, work one-handed, and take the thermal reading from outside the enclosure.
Match the instrument to the trigger, and use more than one. A load-dependent trigger gets a logging clamp on the circuit conductor; a supply trigger gets an RMS voltage logger at the panel; a thermal trigger gets irreversible peak-recording temperature labels on the suspect terminations; a mechanical trigger gets a plug-in event recorder at the point of complaint. Acceptance: at least two independent instruments, at least one of which records without needing anyone present. Wrong looks like a single logger, which cannot separate a real event from an instrument fault. Stop rule: no instrument suited to the named trigger means the deployment does not happen and the trigger is re-examined instead. Hazard: none at this step, it is selection done at the truck.
Deploy with a synchronized clock and get the enclosure closed properly. Set every instrument's clock against the same reference the customer's log uses, land current transformers and leads so nothing is pinched, and put the deadfront back on. Acceptance: every clock inside a minute of the reference, written down, and the deadfront secured with all screws in and no lead crossing a bus or trapped under a cover. Wrong looks like leaving a panel open with a logger inside it, which leaves an occupied building with an exposed live enclosure for a fortnight. Stop rule: instrumentation that will not fit inside a closed enclosure goes on a cord-connected adapter instead; the deadfront does not stay off. Hazard: your hands are in a live panel landing sensors, so use current transformers and leads rated for that system, route them clear of the bus, and confirm nothing is trapped before the last screw goes in.
Write the removal criterion down before you leave the site. State the two conditions that end the deployment: one recorded event, or a stated elapsed period covering at least one full cycle of the trigger. Acceptance: both conditions written on the work order and read back to the customer, with the retrieval date on the office board. Wrong looks like "we will come back in a couple of weeks." Stop rule: a deployment removed before either condition is met is a void run, not a finding, and it is re-deployed rather than reported. Hazard: none at this step, it is a decision written at the truck, but the customer instruction that goes with it matters: nobody stands at the panel during an event, and a breaker that opens is left open and photographed rather than reset repeatedly.
Retrieve, correlate against the customer log, and say plainly which result you have. Acceptance: either an event recorded on at least two instruments with a matching customer entry, which is a finding, or a full trigger cycle with no event, which is also a finding and gets stated as one. Wrong looks like reporting one instrument's event with nothing corroborating it, which is as likely to be a loose sensor as a fault. Stop rule: one instrument showing an event and the other showing nothing at the same timestamp means the sensor is checked before the building is. Hazard: removing a clamp from a live conductor with the deadfront off is the same exposure as installing it, so the same PPE applies and the enclosure is closed again before anyone leaves the room.
Close with a repair and its verification, or with a written recheck trigger. Where a finding names a defect, repair it, restore, and prove the protective function you disturbed by loading the circuit to its recorded level and re-reading the screen that caught it. Acceptance on the no-event path: a signed sheet telling the customer what to call about, what to photograph, and what not to do. Wrong looks like closing a no-event call by replacing something. Stop rule: a customer who wants a part changed anyway gets it in writing that this is a trial rather than a diagnosis. Hazard: re-energizing after a repair is where an unfound fault announces itself, so stand to the hinge side, close with the flat of the hand, and confirm the deadfront and all covers are back before the load goes on.
The record this produces
One monitoring record per complaint: the named trigger with its dated instances; the step 2 reproduction attempt and its duration; all three step 3 screen values; the instrument list with serial numbers and clock offsets; the deployment and retrieval dates; the removal criterion as written; the raw logs; the customer log; and the correlated event or the explicit no-event statement.
The office reads the removal criterion to know when to book the retrieval. The next tech reads the step 3 screens, the three things they would otherwise repeat. The customer reads the recheck trigger, and a signed sheet naming what to watch for is the difference between a second call that arrives with information and one that arrives as a complaint about the first visit.
Worked pass: 1996 two-story, whole first floor drops for a second, a few times a month, worse in wet weather
Step 1: nine dated instances over four months, six of them within a day of heavy rain. Named trigger: moisture, secondary suspicion of a service or panel connection.
Step 2: the exterior fittings and the meter base are wetted with the affected circuits opened and proved dead, then insulation resistance is re-read. Nothing reproduced in 40 minutes. Recorded as attempted, not as absent.
Step 3: L1 to neutral 120 V, L2 to neutral 121 V, L1 to L2 241 V, and 120 plus 121 is 241, so the trio is consistent and both legs sit inside 114 V to 126 V. The suspect breaker connection reads 39 C against a reference at 36 C, a 3 C difference. Insulation resistance on the first-floor circuits reads 260 megohm, well above the 1 megohm floor. All three screens pass, so monitoring is the right call.
Step 4: an RMS voltage logger at the panel, a logging clamp on the first-floor feeder conductor, and irreversible peak-recording temperature labels on the service neutral lug and the two suspect breaker terminations. Three instruments, two of which record unattended.
Step 5: both logger clocks set within 30 seconds of the reference the customer's phone uses, written on the work order. Leads routed clear of the bus and the deadfront back on with all eight screws.
Step 6 fails on the first run and takes its stop rule. The criterion written was one recorded event, or 21 elapsed days covering at least two rain events. The customer asks for the panel closed up for a gathering and the instruments come out at day 7, which is 14 days short and with no event. That is a void run, not reported as a finding, and the deployment goes back in the following week under the same criterion.
Step 7: the second run reaches day 17 when the customer reports an event at 21:14. The voltage logger shows a 0.9 second collapse on one leg at 21:14, the clamp shows first-floor current going to zero across the same second, and the customer's log entry reads 21:15. Two instruments and the log agree, so this is a finding rather than a sensor fault. The temperature label on the service neutral lug has blackened at its 71 C step while both breaker labels are unchanged.
Step 8: the finding names the service neutral termination, supported by the label and the single-leg collapse, and hands to the Open Neutral Diagnosis SOP for the confirming load-shift test and the utility call rather than being repaired off a logger trace.
References
- ANSI C84.1, in the edition your utility references, for the Range A band of 114 V to 126 V on a 120 V nominal base used in the step 3 screen
- NFPA 70B, in the edition your shop's maintenance program has adopted, for insulation resistance practice and the 1 megohm field minimum used in the step 3 screen
- 29 CFR 1910.333(b)(2) for work practices, with NFPA 70E-2021, 120.5 for live-dead-live and 130.5 and 130.7 for the arc-flash risk assessment and PPE used with the deadfront off
- See related: Power Quality Complaint Investigation SOP, Infrared Scan of a Panel SOP, Open Neutral Diagnosis on a Service SOP