How to Tell Whether Your Instrument Can See the Difference You Care About

Why this matters

Half the diagnostics in this trade are differences, not values. Did it change after the adjustment. Is this side worse than that side. Has it moved since last visit. And a difference is exactly the measurement that instruments are worst at, because the difference is small and both readings that produced it are large.

A sibling card owns the rule for the moment of reporting: a difference is a finding only when it clears about three times the combined uncertainty of the two readings behind it. This one is the work you do before that moment ever arrives. The deliverable is a single number per instrument, established once on a bench, written on a card that lives with the instrument: the smallest difference this instrument and this method can defend. Below that number, no amount of care in the field produces a finding, and knowing it before you drive out is what stops a tech from reporting a change that was never there.

Before you run a trial that breaks the connection a dozen times

The procedure below deliberately makes and remakes the coupling ten or more times. That is a hazard the instruction creates, and it multiplies whatever exposure a single reading carried.

Run the trial on a bench source, a known reference or a permanently installed test point. Do not run it by repeatedly opening an energized enclosure: 29 CFR 1910.333(a)(1) permits energized testing only where de-energizing is infeasible, and a trial that exists to characterize your own instrument never meets that test because it can be done on a de-energized bench. Instrument, leads and probe tips carry a measurement category and voltage rating at or above anything you later use them on under IEC 61010-1, binding through the listing mark, and 29 CFR 1910.334(c)(2) requires inspecting the instrument, leads, cables, probes and connectors for external defects before use. Do not run a re-coupling trial by repeatedly breaking a pressurized joint; isolate, relieve to a controlled point and confirm zero at an installed indicator first, which is stored energy under 29 CFR 1910.147, and let a hot line cool or drain before unmaking a fitting so it cannot flash at the person holding the wrench. On a hot surface trial, seat the probe with a clamp or stand rather than a bare hand, and let the surface cool below burn temperature if you have to touch it between readings.

The rule this procedure produces

State it before running anything, because every step below exists to fill in one term of it.

Smallest defensible difference, D, equals 6 times s, where s is the half-width of the observed spread from a re-coupled trial of at least 10 readings, taken at the point on the range where you actually work, for one specific combination of instrument, accessory, coupling method and range. One number per combination, not one per shop and not one per instrument.

Three things about that rule, each of which is a parameter someone will otherwise guess:

  • Where the 6 comes from. A difference is made of two readings, each carrying up to s of scatter, so the combined worst case is 2s. Applying the three-to-one reporting rule to that gives 3 times 2s, or 6s. The two terms are combined by adding rather than in quadrature because s here is the half-width of an observed range from a small trial, which is not a standard deviation and does not behave like one. Quadrature is correct for published standard uncertainties, and the sibling accuracy card owns that case.
  • Why at least 10 readings. Fewer than that and the observed range is mostly luck. Ten is a working floor, not a statistical claim.
  • Why per range point. Instrument error terms scale differently across a range, so a trial run at the top of a range does not describe the bottom of it.

One more relation gets used later, so it belongs here: averaging n readings reduces the random part of the scatter by roughly the square root of n. It does nothing at all to a consistent bias, such as a probe that always seats slightly off in the same direction.

Step 1: Name the question and the point on the range

Write the actual field question first: "has this value moved by more than X since last visit," with a real X. Then pick a trial point on the range near where that question gets asked. A card built at a convenient bench value describes a measurement you do not make.

Skip this and you get a number that looks authoritative and applies to nothing, which is worse than having no card, because someone will use it.

Step 2: Build a source that is steadier than the instrument

The trial measures the instrument's scatter, so the source has to be stabler than what you are trying to detect. A thermal mass at room equilibrium, a stable bench supply, a fixed physical artifact, a dead-weight or reference standard. If the source drifts during the trial, you will characterize the source and call it the instrument.

Test for this cheaply: take the first and last readings of the trial and note whether the series shows a consistent walk in one direction rather than scatter around a level. A walk means the source moved, and the trial is void.

Step 3: Trial A, the coupling untouched

Take at least 10 readings without touching anything: probe stays put, clamp stays closed, fitting stays made up. Let the instrument settle to its own stated response time between readings. Record every value.

This spread is the instrument and its electronics alone. It is the number a spec sheet gestures at and rarely publishes, and on its own it is optimistic to the point of being misleading, because no field reading is ever taken this way.

Step 4: Trial B, the coupling broken and remade each time

Take at least 10 more readings, breaking and remaking the coupling between every one. Unclamp and reclamp. Lift the probe and re-seat it. Disconnect and reconnect. Reposition exactly as you would on a real call, including the sloppiness of a real call, because that is what you are trying to measure.

Skipping this step is the single most common way a shop ends up with a card that is off by a large multiple. Trial A alone tells you about the box. Trial B tells you about the measurement.

Step 5: Compare the two spreads before you compute anything

If the re-coupled half-width is more than about twice the fixed-coupling half-width, technique governs and a better instrument will not move your card number much. Under that ratio, the electronics are a real share of the total and an instrument upgrade is worth pricing.

This is the step that decides where money goes, and it is available to you only because you ran both trials.

Step 6: Compute D, write the card, and gate on it

Compute D as 6 times the trial B half-width. Write on a tag that stays with the instrument: the instrument and accessory identity, the range point, the trial B half-width, D, and the date. Then in the field, any difference smaller than D gets reported as no detectable change, in those words, rather than as a number.

Worked example: an adjustment that could not be proven

A tech is asked whether a routine adjustment moved a value. The instrument is a handheld with a clip-on accessory, working near 61 units, which is where the real question lives.

Trial A, coupling untouched, 12 readings. Values span 61.2 to 61.6. Range 0.4, so the half-width is 0.2 units. The series scatters around a level rather than walking in one direction, so the source held.

Trial B, re-coupled between every reading, 12 readings. Values span 60.7 to 62.1. Range 1.4, so the half-width is 0.7 units.

Step 5 test. 0.7 divided by 0.2 is 3.5, well over the ratio of about 2. Technique governs here. Buying a steadier instrument would leave most of that 0.7 in place, because most of it is created at the coupling, not inside the box.

Step 6. D equals 6 times 0.7, which is 4.2 units. That goes on the tag.

The field call. Before the adjustment the tech read 61.4. After it, 64.0. The difference is 2.6 units, and it is in the direction he expected, which is exactly the condition under which people believe a number. Against D of 4.2 units, 2.6 does not clear the gate. The honest report is no detectable change by this method, and a tech who writes "improved by 2.6" has reported his expectation with a number attached to it.

Three moves make the difference detectable, and they are not equally good.

  • Fix the coupling. Land the probe in a well, a fixture or a clamp so it seats identically every time. This attacks the term step 5 identified as dominant. It is the durable fix and it improves every future measurement with that instrument.
  • Measure the difference directly. Where the quantity allows it, span the two points with one instrument and read the small value rather than subtracting two large ones. This removes both large readings from the arithmetic at once.
  • Average. Take 4 re-coupled readings on each side and use the means. Note that this multiplies field exposure by four on each side, so it is available only where the point can be re-coupled without opening an energized enclosure or breaking a pressurized or hot joint each time; where it cannot, this move is off the table and the first two stand. The random part of the scatter drops by roughly the square root of 4, so s falls from about 0.7 to about 0.35 and D falls to 6 times 0.35, which is 2.1 units. The observed 2.6 now clears 2.1 - but by 0.5 units, which is a thin margin, and the square-root relation applies only to the random component. If part of that 0.7 is a consistent bias, a probe that always seats slightly proud in the same direction, averaging leaves that part untouched and the real D is worse than 2.1. Averaging is the weakest of the three because it buys margin by assuming something about the error you have not established.

The failure mode this prevents is specific and common: the tech reports 2.6 as an improvement, the office logs it, and six months later that number is a data point on a trend line that is entirely constructed out of coupling scatter.

How to verify the card is honest

Have a second tech run trial B on the same instrument, same range point, same source, without seeing your numbers. Compare the half-widths.

If the two agree within roughly a factor of two, the card describes the instrument and method and you can hand it to anyone. If the second tech's spread is several times yours, the card describes you, and it will not travel: either the coupling procedure is underspecified and needs writing down step by step, or one of you is doing something the other is not. Either way that is a training finding, and it is more valuable than the card.

Re-run the trial after any drop, heat event or water exposure, after the accessory is replaced, and when the instrument moves from bench use into daily field use, because all three change the thing the card describes.

References

  • Manufacturer published response time, settling behaviour and accessory specifications for the instrument and accessory used in the trial
  • 29 CFR 1910.333(a)(1) energized-testing gate; 29 CFR 1910.334(c)(2) pre-use inspection of instruments, leads, cables, probes and connectors; 29 CFR 1910.147 for isolating and relieving stored energy before a joint is broken
  • IEC 61010-1 measurement categories, binding through the instrument's listing
  • See related: The Difference Between Accuracy and Resolution, which owns the three-to-one reporting rule and the uncertainty budget; Accuracy, Resolution and Repeatability Are Three Different Things