Why a Drifting Sensor Is Worse Than a Dead One

Why this matters

A dead sensor gets you a phone call. A drifted one gets you nothing, for a season, while the loop runs faithfully against a target that is not the target anybody chose. The cost of drift is not paid at a moment, it is paid continuously in runtime, in cycles and in comfort, and it becomes visible only when the accumulated effect finally crosses somebody else's threshold - a utility bill, a tenant complaint, a limit that starts tripping. By then nobody connects it to a sensor, because the sensor is showing a perfectly reasonable number.

The gate that actually sorts these

The instinct is to sort faults by severity: dead is worse than degraded. That instinct is wrong here, and a single question sorts them correctly.

Does this fault produce a symptom that somebody who can act on it will report, and how soon?

Run that gate against two cases below. They start from the same place, a discharge-air sensor on a heating system controlling to a 105 degree discharge with 60 degree return air, and they resolve oppositely, which is the point.

Outcome one: the sensor drifts high by 6 degrees

Nothing on the screen looks wrong. The controller sees 105 and holds it, so the true discharge sits at 99.

Delivered capacity on that unit is airflow times air density times specific heat times the temperature rise across the coil. At constant airflow and roughly constant air properties over this range, capacity is proportional to rise, so the shortfall is straightforward: intended rise was 105 minus 60, or 45 degrees; actual rise is 99 minus 60, or 39. Six degrees of 45 is 13.3 percent of capacity gone. That proportionality is stated at constant airflow; on a unit whose modulation also changes fan speed it has to be recomputed with the new airflow, and the shortfall will not be the same number.

The space still gets warm, it just takes longer. To deliver the same total heat at 86.7 percent of the rate, the equipment runs 1 divided by 0.867, or about 1.154 times as long - roughly 15.4 percent more runtime. Against a baseline of 120 runtime hours in a month, that is about 138.5 hours, or 18.5 extra hours of running, every month, for as long as the drift stands.

Who reports it? Nobody. Fifteen percent more runtime does not feel like anything. The space reaches temperature. The equipment sounds normal. The bill moves by an amount that sits inside normal weather variation, and the only person who could have caught it is the tech who stood in front of the unit and read 105 on the screen and had no reason to doubt it.

Outcome two: the same sensor fails dead

Now consider the same sensor open-circuited, and run the same gate. The answer depends entirely on what the loop does when the input goes to its fault state, which is a design decision made years ago by somebody else.

If the loop shuts down on a failed input, the gate resolves immediately: no heat, a call within hours, one visit, one part, done. Total exposure is the downtime, and the downtime is bounded by how fast you can get there. This is the case people picture when they say a dead sensor is worse, and it is the cheapest fault in this article.

If the loop drives to full output on a failed input, the gate resolves the other way, and harder than drift did. The unit runs at full output whenever it is enabled. Against the same 120-hour baseline in a month with roughly 480 enabled hours, that is about four times the runtime, and the space is overheated the whole time. Occupants open windows and adjust nothing, because the thermostat says what it always said.

That is the real finding, and it inverts the title's own framing in a way worth carrying: drift is not dangerous because it is small, it is dangerous because it is silent. A dead sensor that is also silent costs more than drift does, and by a multiple rather than a percentage. Severity is not what sorts these faults. Announcement is.

What makes drift silent, and what would make it loud

Direction decides. A sensor drifting in the direction that makes the controller work harder creates a visible symptom: over-service, short cycling, a limit starting to trip, a bill that jumps. Those get found, because something in the building complains. A sensor drifting in the direction that makes the controller work less creates under-service, and people adapt to under-service. They put on a sweater, they run a space heater, they stop expecting the far office to be comfortable.

So the drift you find is the flattering half. The half you never find is the half that is quietly costing runtime and quietly leaving the system unable to hold design conditions on the worst day of the year, which is the day it will be discovered.

The cadence that actually catches it

Drift is not found by diagnosis, it is found by scheduled verification, and a shop that does not commit to an interval will not do it. Concrete defaults worth setting, and tune them to your equipment mix:

  • Verify every control sensor at commissioning and then annually, with a two-point check against a known condition where the element can be brought to one.
  • Verify any sensor feeding a protective function on the manufacturer's stated interval, not yours, because the trip point was selected with an assumed sensor accuracy inside it.
  • Verify opportunistically whenever a loop shows a persistent offset it cannot remove. A loop that sits 2 degrees off with its reset term fully wound is telling you something is wrong with the number, not with the loop.
  • Record the correction applied, every time, so the next tech can see whether the number moved by 0.5 degrees a year or by 6 degrees in one season. The rate of change is more diagnostic than the amount.

That last one is the one shops skip and the one that pays. A sensor drifting slowly and steadily is aging. A sensor that moved 6 degrees since last year did not age, something happened to it, and the something is usually mechanical or an environment it was never rated for.

Drift hurts a difference far more than it hurts a reading

Everything above treats drift on a single absolute reading. The moment the number that matters is a difference between two sensors, the same drift gets worse, and it gets worse in a way that is easy to miss because each sensor individually still looks fine.

Take two sensors on the same airstream, each within a stated plus or minus 1.0 degree. Individually, both pass any reasonable tolerance. But if one has drifted to its high limit and the other to its low limit, the error in the difference between them is 2.0 degrees, because the two errors add rather than cancel.

Against a design temperature rise of 45 degrees that is 2.0 of 45, or 4.4 percent, which most people would accept. Now run the same equipment at part load with a 15 degree rise: the same 2.0 degrees is 2.0 of 15, or 13.3 percent. Identical sensors, identical drift, three times the percentage error, purely because the quantity being measured got smaller. The ratio of those two percentages is exactly the ratio of the two rises, which is the general shape worth carrying: the percentage error of a difference scales inversely with the size of the difference.

This is why an approach temperature, a small delta across a heat exchanger, or any other measurement whose whole payload is a small gap between two larger numbers deserves tighter sensor tolerance than a plain absolute reading of the same quantity, and why the tolerance has to be stated against the smallest difference the measurement is used at, not the design one.

There is a way out, and it is not tighter sensors. If both sensors drift by the same amount in the same direction, the difference is untouched, because the common error subtracts out. That is the argument for a matched pair from the same production lot, and the stronger argument for using one instrument moved between the two points when you are taking a difference by hand. One instrument reading both points carries its own error into both numbers, where it cancels.

What flips the conclusion

Where the sensor feeds a protective device rather than a control loop, drift stops being an efficiency problem and becomes a safety one, and the ranking in this article does not hold. A high limit reading 6 degrees low trips early and looks like a nuisance. Reading 6 degrees high, it lets the process reach 6 degrees past the point somebody chose as unsafe, and it will still be showing a plausible number when it does. If a protective device is tripping, establish why it opened on evidence before touching anything: a device tripping because the process genuinely arrived at its trip point is working, and adjusting a sensor or widening a setpoint so it stops tripping reaches the same end state as jumpering it.

Where the sensor feeds a measurement somebody bills or reports on, drift is not a comfort issue at all, it is a records issue, and it propagates into documents you cannot retract.

Where the loop's failure position is well chosen and clearly annunciated, the dead-sensor case collapses to a short, cheap event and the ranking in this article holds cleanly. A sibling card covers how failure positions are chosen and what they do not protect against.

How to catch drift without a reference bath

Not every sensor can be pulled and put into a known condition, and you still need a read. Three checks, in the order that costs least.

Compare against a physically coupled quantity. If a discharge sensor reads 105 and the return reads 60, the rise is 45 degrees. Measure airflow and the delivered capacity has to reconcile with what the equipment is capable of. A number that requires the unit to deliver more than its rating is not a real number.

Compare two sensors that must agree under a known condition. With a system off and settled overnight, everything in one airstream should read the same within the accuracy of the instruments. Any sensor sitting apart from the pack under that condition has moved.

Watch the correction the loop is carrying. A reset term that has wound to a large steady value is the loop compensating for a wrong number. It is not proof of drift, but it is free, it is already logged, and it points you at which loop to check first.

References

  • Manufacturer documentation for sensor accuracy, drift specification and the verification interval required for any sensor serving a protective function
  • 29 CFR 1910.333(a)(1) - live parts to be de-energized before work, and the narrow conditions permitting energized troubleshooting, where a check requires reading a live terminal
  • See related: How to Verify a Sensor Against the Real Quantity; The Failure Position and Why It Was Chosen; The Sensor Failure Modes and What Each One Looks Like