How to Tell a Deadband Problem From a Sensor Problem

Why this matters

"It swings too much" and "it cycles too often" are the same complaint with two opposite causes, and the settings screen offers a tempting answer to both. Narrow the differential and the controller's own log gets tidier immediately, which feels like success. If the real cause was a measurement that cannot see what is happening, you have just doubled the starts on the equipment and improved the customer's actual comfort by almost nothing. The test that separates the two takes about forty minutes of watching and requires no parts.

Before you put an instrument on anything

  • A reference probe strapped to a hot pipe or pushed into a hot well is a burn route. Let the surface cool or use insulated gloves and eye protection, and pull the sensor out of the well rather than pulling the well out of the line. If the well itself has to come out, isolate the line, drain it to a safe point, and confirm zero on a gauge before a thread moves.
  • Do not disturb pipe or vessel insulation to reach a sensing point on older equipment. Thermal system insulation of unknown vintage is an inhalation hazard and should be treated as asbestos-containing until sampled; if it must be disturbed, that is respiratory protection selected under a program meeting 29 CFR 1910.134, not a dust mask.
  • Where a measurement must be taken inside an energized control enclosure, 29 CFR 1910.333(a)(1) permits energized work only where de-energizing would introduce additional or increased hazards or is infeasible due to equipment design or operational limitations, and live troubleshooting is a recognized case; use a meter and leads rated CAT III at or above the circuit voltage and work to the boundaries and protective equipment NFPA 70E-2021 assigns. Where the work can be done dead, open the disconnecting means, lock and tag under 29 CFR 1910.333(b)(2) in general industry or 29 CFR 1926.417 in construction, and prove dead per NFPA 70E-2021, 120.5.
  • On combustion equipment, a personal carbon monoxide monitor goes on your body before the appliance fires and stays there while you watch cycles.
  • Leave every limit and interlock in service for the whole test. If a protective device opens during it, that is the finding: establish why it opened before doing anything else, because a device that is doing its job does not become the fault by being inconvenient.

Step 1: get an independent trace at the same physical point

Put a logging reference instrument as close as physically possible to the controller's sensing element, in the same air, the same water, the same product. Not in the same room. Not six feet down the duct.

Skipping this leaves nothing to compare against, and every step after it becomes uninterpretable, so the whole visit produces an opinion rather than a finding. This is the single most expensive step to skip because it does not just weaken the diagnosis, it removes the possibility of one.

Step 2: capture at least three full cycles of both traces, timestamped

Both instruments running at the same time, both logging, with clocks you can line up afterwards.

Skipping this and taking a snapshot instead costs you the ability to see phase. A snapshot shows two numbers that differ, which you will read as a calibration offset. A trace shows whether the controller's reading is smaller, larger, or simply late, and late is a completely different repair from wrong.

Step 3: compare amplitude first, then phase

Three signatures, and each one sends you somewhere different.

  • The two traces agree in amplitude and in timing, and the swing is roughly the differential plus a little overshoot. The differential is doing exactly what it was set to do. This is a deadband finding, or a capacity finding, and the sensor is fine.
  • The controller swings more than the reference. The measurement is adding motion that is not in the process: electrical noise coupled into the sensing circuit, an intermittent termination, or a sensor sitting where two streams mix unevenly. Nothing about the differential caused this.
  • The controller swings less than the reference and its peaks arrive late. The measurement is smoothing motion that is real. The loop is blind, and the customer's complaint is true while the controller's own history looks clean.

Skipping this comparison and jumping to a settings change costs you the branch, and the two wrong branches fail in opposite directions, so a guess here has no better than a one-in-three chance of being right.

Step 4: prove the termination and the lead before you touch the element

Check the terminations at both ends, look for a chafed or shared-conduit run, confirm the sensor is mechanically bottomed in its well with heat transfer compound rather than sitting in an air gap.

Skipping this costs a part and a return trip. It does not cost the diagnosis, which is why it sits below the comparison steps rather than above them: a good sensor replaced with another good sensor produces the identical trace, and you will be back.

Step 5: read the settings actually in force, not the documented ones

Find the live differential, the live minimum on and off times, and any filtering or smoothing parameter on the input. Compare against the record. A software filter on the sensor input is the single most overlooked cause of the third signature above, and it lives in a different part of the menu from the differential.

Skipping this costs an adjustment made to a parameter that is not the active one, which is recoverable in a follow-up but wastes the change you were going to make.

Step 6: change exactly one thing and re-run the trace

One parameter, or one mechanical correction, then three more cycles logged the same way.

Skipping this and changing two things at once costs attribution. The system may improve and you will not know which change did it, so you cannot tell the next tech what to keep.

Worked example: the swing the controller could not see

A process vessel with an on/off heat input, differential set at 1.0 F, complaint that the product runs inconsistent even though the controller's own history shows a steady value.

Traces, three cycles each. The controller's display swung 1.1 F peak to peak with a period of 34 minutes. The independent probe, taped alongside the controller's element in the same well pocket, swung 3.2 F peak to peak over the same 34-minute period, and its peaks arrived about 5 minutes before the controller's.

The real swing is 3.2 divided by 1.1, or roughly 2.9 times what the controller sees, and the timing is late. That is the third signature: the measurement is smoothing and delaying a swing that is genuinely there. The controller is switching correctly on the information it has, and the information is stale by five minutes, so the process runs past the band in both directions before the output responds.

The tempting wrong move, run anyway to price it. Narrowing the differential from 1.0 F to 0.5 F was tried first, because that is what the previous ticket recommended. The result: the cycle shortened from 34 minutes to 25 minutes, which is 1.76 cycles per hour up to 2.4, a 36% increase in starts. The reference swing came down from 3.2 F to 2.8 F, which is a 12.5% improvement. Thirty-six percent more starts bought twelve and a half percent less swing, because the swing was dominated by the lag, not by the band. That is the whole argument against branching on a snapshot.

The actual repair. The element was found sitting loose in the well with no heat transfer compound and about a finger's width of air between the tip and the well bottom. Air in that gap is a poor conductor, and the element was effectively reading the well wall through an insulating layer, which is exactly what produces a small, late signal. With the element bottomed and the well filled with compound, and the differential returned to 1.0 F, the controller swung 1.2 F and the reference swung 1.5 F over a 33-minute period, with the peaks within a minute of each other.

Reading the series honestly. The period held near 34 minutes through the baseline and the sensor repair, 34 then 33, and the only thing that moved it was the differential change, to 25. That is the tell that the cycle length was a property of the process and the differential, and that the sensor fault was affecting what the loop could see rather than how fast the vessel moved.

What would have changed the conclusion. If the reference had swung 1.1 F alongside the controller's 1.1 F, there would have been no sensor finding at all and the honest answer would have been that a 1.0 F differential on this vessel produces a 1.1 F swing, which is what a differential does. If the controller had swung 4 F while the reference swung 1 F, the fault would have been in the opposite direction and the search would have moved to terminations and to whatever runs in the same conduit.

The failure mode of getting it wrong. Leaving the narrowed differential in place after fixing the sensor. Both changes improve the display, so on a busy day it is easy to keep the settings change as insurance. The equipment then runs at roughly 36% more starts forever, on a vessel that no longer needs it, and no document anywhere connects the eventual failure to a control adjustment made during a swing complaint.

Confirming the call before you leave

Re-run the traces at a different load. A sensor fault produces the same amplitude ratio at every load, because it is a property of the sensing path. A differential-driven swing changes with load, because on-time and off-time move with rise and fall rates. One extra half hour at a second condition separates them conclusively.

Move the reference and confirm the disagreement follows the sensor. If the spread changes when you move the reference to a slightly different spot, the finding was placement rather than the element, and the repair is a mounting location.

Write down the differential, the filter setting, and both trace amplitudes on the ticket. The next swing complaint on this equipment needs to know what the numbers were when it was right, and no controller history stores the reference trace.

References

  • 29 CFR 1910.333(a)(1) and (b)(2), OSHA general industry work practices for energized and de-energized electrical work
  • 29 CFR 1926.417, OSHA construction lockout and tagging of circuits
  • 29 CFR 1910.134, OSHA respiratory protection program requirements
  • NFPA 70E-2021, 120.5, verifying an electrically safe work condition
  • See related: What Deadband and Hysteresis Are For; How Sensors Fail