How to Check an Instrument Against a Reference in the Field
Why this matters
Two instruments get held against the same thing, they read close, and everyone relaxes. That check almost never establishes what the people running it believe it establishes, and the reason is arithmetic rather than technique: a comparison can only detect an error larger than the comparison's own uncertainty. Where the reference is no better than the instrument being checked, the check cannot see an error of the size you care about, and a clean result is indistinguishable from no result at all.
A sibling HowTo owns the field-verification routine itself: inspecting the instrument and everything attached to it, proving on a known source before and after, using a two-point reference where no live source exists, and why a self-test is not a verification. This card is the design step that comes before any of it. Every instruction below is derived backwards from one number, the error size you need to detect, and the deliverable is a check that is honest about what it could have found.
Before two sets of leads go into one enclosure
A comparison check needs both instruments seeing the same quantity at the same moment, which means two probes, two sets of leads and two people's hands where there was previously one. That is a hazard the procedure itself creates.
Run the comparison on a bench source, a de-energized known source, or through permanently installed test ports and thermowells. Where there is no installed port and the comparison would require two sets of probes inside a live enclosure, do not run it live: 29 CFR 1910.333(a)(1) permits energized work only where de-energizing is infeasible, and a self-imposed instrument check is never infeasible to defer to a bench. Both instruments, their leads and their probe tips carry a measurement category and voltage rating at or above the circuit under IEC 61010-1, which binds through the listing mark, and 29 CFR 1910.334(c)(2) requires both to be inspected for external defects before use. Where the check is at a pressurized point, use existing test ports rather than breaking a joint, and isolate, relieve to a controlled point and confirm zero at an installed indicator before unmaking anything, which is stored energy under 29 CFR 1910.147. On a hot line, let it cool or drain to a closed receiver before breaking a fitting, because a hot pressurized joint flashes at the person holding the wrench.
The rule the whole procedure serves
A field check can conclude an instrument is out of tolerance only when the observed disagreement exceeds the check's own combined uncertainty, and the check can only be designed to detect an error of size T when that combined uncertainty is at or below T divided by 4.
The combined uncertainty of the check is the reference's own uncertainty, plus the spread your coupling introduces on the instrument under test, plus the spread it introduces on the reference, added rather than combined in quadrature because these are observed half-widths rather than published standard uncertainties. The 4-to-1 ratio is the same convention a laboratory calibration is designed around, and the sibling card on what a calibration establishes owns where it comes from and what a thin ratio does to a pass.
Note the unit of analysis: this is per check design, for one instrument against one reference at one point on the range. It is not the same question as the sibling HowTo on seeing a difference in the field, which asks whether your rig can resolve a difference between two field values. Here both readings are of the same quantity and the difference under examination is instrument error.
Step 1: Write down T before you pick anything up
T is the error size you need to detect, and it comes from one of two places: the tolerance the instrument is meant to hold, or the tighter tolerance your particular work needs, whichever is smaller. Write it with its unit.
Skip this and the check gets designed around whatever reference is convenient, which is how a shop ends up running a monthly ritual that could not detect a doubling of the instrument's allowed error.
Step 2: Compute the budget the check is allowed
Divide T by 4. That is the entire uncertainty budget for the comparison, covering the reference and both couplings.
This is the step that eliminates most field checks on paper, before anyone drives anywhere, and eliminating a useless check is a result rather than a failure.
Step 3: Pick a reference that fits inside the budget
The reference has to be materially better than the thing it is judging, and its published accuracy has to fit inside the budget with room for the coupling terms. Confirm its calibration is current and, where the check's outcome will be reported to anyone outside the shop, that its traceability is documented; the sibling card on traceability owns which of your instruments actually need that and which do not.
Two instruments of the same class are not a reference for each other. They can tell you the two have diverged. They cannot tell you which one moved, and a shop that treats the newer one as correct is making a guess with a number attached.
Step 4: Kill the coupling terms, because they are yours to control
Fix both probes so they cannot move. Use the same port, the same seating, the same orientation each time. Let both instruments settle to their published response times before reading, and read them at the same instant rather than one after the other, because a source that drifts between the two readings shows up as instrument error.
This is the only step in the procedure where more care produces a better number, which is exactly why it gets skipped in favour of arguing about the result.
Step 5: Compare at the point you actually work at
Instrument error terms vary across a range, so a check run at a convenient mid-scale value says little about the bottom of the range where fixed error terms dominate. Run the comparison near the value where your decisions get made, and note that value on the record.
Step 6: Read the result as one of three outcomes
- Disagreement smaller than the check's combined uncertainty. The check found nothing. Write it in those words. It does not mean the instrument is in tolerance; it means any error it carries is smaller than this check can see, which is a different and weaker statement.
- Disagreement larger than the combined uncertainty, with a reference materially better than the instrument. Something moved, and it is attributable to the instrument under test. Tag it out of service.
- Disagreement larger than the combined uncertainty, with a reference of similar quality. Something moved and you cannot say what. The next move is a third instrument or a real reference, not a judgment call about which one looks newer.
Step 7: Record what the check could have detected
The record carries three things: the reference identity, the smallest error this check could have detected, and the observed difference. A verification record that lists only the observed difference is unreadable a year later, because nobody can tell whether a clean result meant the instrument was good or the check was blind.
Worked example: the ritual and the redesign
A shop checks a handheld instrument monthly. The instrument's tolerance, and the error the shop needs to detect, is T equals 1.0 units.
The ritual as it was being run. The reference was a second handheld of the same class, published accuracy plus or minus 1.0 units. Handheld coupling on each side contributed about 0.3 units of spread. Combined uncertainty of the check is 1.0 plus 0.3 plus 0.3, which is 1.6 units.
Step 2 says the budget was T divided by 4, which is 0.25 units. The check is running at 1.6 against a budget of 0.25, so it is about six and a half times too coarse. In practice it could only flag a disagreement above 1.6 units, and 1.6 is already 60 percent past the tolerance the shop was trying to protect. Every clean result that ritual ever produced meant nothing, and it had been producing them for years, which is the worst version of this because it built confidence rather than leaving a gap someone might have noticed.
The redesign. The shop bought one bench reference with a published accuracy of plus or minus 0.1 units, ten times better than the tolerance being judged, and built a fixture so each probe seats identically, dropping each coupling spread to about 0.05 units. Combined uncertainty becomes 0.1 plus 0.05 plus 0.05, which is 0.2 units, inside the 0.25 budget. The check is now capable of the job it was always being asked to do.
The first run of the redesigned check. At the working point, the reference reads a known value and the instrument under test disagrees by 1.4 units. That is 7 times the check's combined uncertainty of 0.2, so the disagreement is unambiguous, and because the reference is ten times better than the instrument's own tolerance the error is attributable to the instrument. It is also outside the 1.0 unit tolerance, so the instrument is out of tolerance and comes out of service that day.
What the check does not tell you. It does not tell you the correction to apply, and a tech who starts subtracting 1.4 from field readings is inventing a calibration. It does not tell you when the error appeared, so the past work taken with that instrument has to be bounded, and the sibling card on calibration scheduling owns how to size and triage that review set. And it does not replace calibration: it is the thing that catches a failure occurring in the weeks since the last one, which no interval can.
The failure mode this prevents is precisely what the shop had before: a check that runs on schedule, gets recorded, satisfies an auditor, and is arithmetically incapable of detecting the failure it exists to detect. It is more dangerous than no check, because a shop with no check knows it is exposed.
How to verify the check itself is honest
Take your existing verification records and compute, for one of them, the smallest error that check could have detected. Compare it to the tolerance the instrument is meant to hold. If the first number is not comfortably under a quarter of the second, the record is documentation of an activity rather than evidence about an instrument.
Then run a deliberate test of the check: introduce a known offset roughly the size of T, by using a reference point you know differs from the nominal by that amount, and confirm the check flags it. A check that cannot find an error you planted will not find one that arrives on its own.
References
- Manufacturer published accuracy statements and response times for both the instrument under test and the reference
- 29 CFR 1910.333(a)(1) energized-work gate; 29 CFR 1910.334(c)(2) pre-use inspection of instruments, leads, cables, probes and connectors; 29 CFR 1910.147 for isolating and relieving stored energy before a joint is broken
- IEC 61010-1 measurement categories, binding through the listing mark on each instrument
- See related: How to Verify a Test Instrument Before You Trust It, which owns the field-verification routine; What Calibration Actually Establishes, which owns the 4-to-1 ratio; When Two Instruments Disagree, Which One to Trust