Why Two Instruments Disagree and Which One to Believe

Why this matters

Two instruments, one measurement, two numbers. The reflex is to decide which one is broken, and the trade's usual tiebreaker is calibration status: the one with the fresher sticker wins. A sibling card lays out that trust hierarchy and it is sound advice once you get there.

The problem is that most field disagreements never should get there, because they are not disagreements. Two instruments reading 0.49 and 0.50 have not contradicted each other if each is only good to a couple of hundredths. Arbitrating them is a waste of a morning and, worse, it produces a false conclusion: somebody decides an instrument is drifting when its published accuracy already covered the whole gap.

And when the difference genuinely is too large, the cause is usually not either instrument. It is the connection, the loading, or a parameter somebody typed in. Calibration is not the first question. It is roughly the fourth, and in field work it is rarely the answer.

First, if the disagreement is about a hazard

You do not arbitrate a safety reading. You act on the more dangerous one and resolve it afterward.

Two testers disagreeing about whether a conductor is live means the conductor is live. That is not a judgment call, it is the reason the proving sequence exists: prove the tester on a known live source, test the conductor, prove the tester again, per NFPA 70E-2021, 120.5 in the edition your employer's electrical safety program adopts. A tester that fails the second proof invalidates the reading you just took, and the work stays under 29 CFR 1910.333(b)(2) until absence of voltage is established properly.

Two gas detectors disagreeing means gas is present: everyone leaves, no switches or lights are operated, no phone is used inside, and the call goes out from outside. Two carbon monoxide instruments disagreeing means you act on the higher number and clear the space.

Everything below applies to diagnostic readings, not to hazard readings.

Question one: is the difference bigger than the two bands added together

Every instrument's accuracy is published as a tolerance, and most are written in one of two shapes: a percentage of the reading plus a fixed number of display counts, or a percentage of the reading plus a percentage of full scale. Both forms mean the same thing in practice, which is that the band is not a single percentage and it gets proportionally wider as you read further down the range.

Work out each instrument's band at the value you actually read, not at full scale, and pull the numbers from the instrument's own documentation rather than from a general rule.

Then combine them, and pick the combination deliberately:

  • Worst-case sum. Add the two bands. This is the bound neither instrument can be outside of, and it is the right choice when you are about to make a decision you do not want to revisit.
  • Root-sum-square. Square each band, add, take the root. This gives a realistic band rather than a bound, and it assumes the two errors are independent. That assumption fails if both instruments were calibrated against the same reference, if they share a sensor technology with a common environmental sensitivity, or if both are sitting in the same hot mechanical room. Where independence is doubtful, use the sum.

If the difference between the two readings is smaller than the combined band, there is no disagreement to resolve. Stop. You have two consistent readings and an answer that is an interval instead of a point.

The answer is an interval, and an interval is usually enough

When two readings are consistent, the honest result is the overlap of their two intervals: the range of values both instruments say the quantity could be. Report that, then check it against the decision you are trying to make. If the whole overlap sits on one side of the threshold, the question is answered and it never mattered which instrument was closer to true.

This is the same move as bracketing a flow measurement with two oppositely biased methods, and it is worth internalising for the same reason: field questions are almost always threshold questions, and a threshold question can be answered by a range.

Question two: what did each instrument do to the thing it measured

If the difference does exceed the combined band, do not go to calibration yet. Ask what each instrument changed by being connected. Every one of these has a direction:

  • A voltmeter without enough input impedance on a high-impedance circuit draws current from it and reads low. The higher-impedance instrument is closer to true.
  • A series ammeter develops a burden voltage that reduces the current it is measuring, so it reads low wherever the circuit's own impedance is not large against the meter's shunt. That is most low-voltage electronics and control loops; it is a function of source impedance rather than of supply voltage, so a stiff 24 V supply may show nothing and a high-impedance sensor loop a great deal.
  • A contact thermometer with mass pulls the surface toward its own temperature, so it reads low on something hot and high on something cold, and the smaller the measured object the worse it is.
  • A pressure gauge on a long connecting line adds volume and damps, so on a rising pressure it reads low and lags.
  • An insertion flow element adds restriction and reduces the flow it is measuring, so it reads low.
  • A probe in a small duct blocks part of the cross-section and raises the local velocity past itself, so it reads high.

The general shape: the less an instrument disturbs what it touches, the closer it is to true, and where two disagree, the more intrusive one is the more likely to be biased.

Question three: how many conversions stand between the sensor and the display

Rank the two readings by how much arithmetic happened after the physical sensing. Four depths, shallowest first:

  1. Sensed and shown. A thermocouple showing temperature, a clamp showing current.
  2. Converted through a physical relationship. Velocity pressure to velocity, pressure drop to flow.
  3. Converted through a parameter somebody typed in. Pipe wall thickness in a clamp-on flow meter, emissivity in an infrared thermometer, free area in a grille reading, the fuel selection in a combustion analyser.
  4. Computed from several of the above. Combustion efficiency, air-free carbon monoxide, real power from voltage, current and phase.

Prefer the shallower reading, and when a deep one disagrees with a shallow one, suspect the entered parameters before you suspect either sensor. Depth three is where the field's quiet errors live, because a wrong entered parameter produces a perfectly stable, perfectly repeatable, perfectly wrong number that no amount of re-reading will shake.

Where calibration finally comes in

By now you have eliminated a false disagreement, a loading effect and an entered parameter. What is left is a genuine instrument problem, and that is when calibration status, specification quality and range selection decide it - the hierarchy the sibling card sets out. Reaching for it earlier is not rigour, it is an expensive way to answer a question you had not yet shown existed.

Two manometers on one duct, twice

Both instruments in these cases are the same pair throughout. Instrument A's documentation gives its accuracy as 1 percent of reading plus 0.003 in w.c. Instrument B's gives 2 percent of reading plus 0.005 in w.c. Both figures are illustrative and both should be replaced with your own instruments' published specs.

First case. Both connected through a tee to the same static tap on a supply plenum. A reads 0.49 in w.c., B reads 0.50.

A's band at 0.49 is 0.0049 plus 0.003, or 0.008, giving 0.482 to 0.498. B's band at 0.50 is 0.010 plus 0.005, or 0.015, giving 0.485 to 0.515. The worst-case combined band is 0.023, and the root-sum-square combination is 0.017.

The difference is 0.01, under both combinations. There is no disagreement. The overlap of the two intervals runs from 0.485 to 0.498, and the question on the table was whether the system exceeds the equipment's rated maximum external static of 0.80 in w.c. The entire overlap sits well under it. Answered, in about two minutes, with neither instrument arbitrated.

Second case. Same pair, same day, a reading on the return side. A reads 0.61, B reads 0.72.

A's band is 0.0061 plus 0.003, or 0.009. B's is 0.0144 plus 0.005, or 0.019. Worst-case combined, 0.028 on the rounded bands above. The difference is 0.11, nearly four times that. This one is real.

Loading is not the suspect: neither manometer meaningfully disturbs a duct. Conversion depth does not separate them either, since both are depth one, sensing pressure and showing pressure.

That leaves what is between the instrument and the tap. B was connected through a long tubing run coiled on the floor with a low spot in it, and the shop had been reading a wet return plenum all morning.

Here the units hand you the answer. A reading in inches of water column is a height of water. A 0.11 in w.c. discrepancy is a column of water 0.11 inches tall sitting somewhere it should not be, which is a few drops in the bottom of a tube. The sign depends on which leg holds it: liquid in the leg connected to the higher pressure opposes it and reads low, liquid in the reference leg reads high.

Drain both legs, re-run with the tubing routed without a low point, and B reads 0.615 against A's 0.61. The difference is 0.005, inside both instruments' bands and inside any combination of them.

The wrong ending. The shop that goes straight to calibration status sends B out for service, gets it back with a certificate saying it was in tolerance, and learns nothing. On the next humid day the same tubing takes on the same water and produces the same 0.11 offset on a certified instrument. The failure was never in either box.

The swap test

One technique settles the instrument-versus-connection question faster than anything else, and it costs a minute: swap the two instruments between the two connection points and re-read.

If the difference follows the instrument, it is the instrument. If it stays with the port, it is the tap, the tubing or the location. That single move separates the two whole families of cause, and it works for manometers, gauges, thermometers, probes and meters alike.

How to verify you got this right

  • Write both bands down before you argue. Not the readings, the bands. Half the disagreements evaporate in that step and the other half get sharper.
  • State which combination you used and why. If you used root-sum-square, say what makes the two errors independent.
  • Do the swap before you condemn a box. An instrument sent out for calibration is out of your truck for a while, and the odds are it comes back in tolerance.
  • Check the entered parameters on any depth-three or depth-four reading, one field at a time, against a source rather than memory.
  • Record the interval, not a picked winner. A ticket that says the quantity sits between two values, with both instruments and both bands named, survives a challenge. One that says a number because one meter had a newer sticker does not.

References

  • NFPA 70E-2021, 120.5, as adopted through your employer's electrical safety program; 29 CFR 1910.333(b)(2)
  • Instrument manufacturer documentation for the accuracy specification, its form, and the range it applies over
  • See related: When Two Instruments Disagree, Which One to Trust; How to Reconcile Two Instruments That Disagree; The Difference Between Accuracy and Resolution; The Flow Measurement Methods and What Each One Assumes