Why a Control Loop Hunts

Why this matters

A hunting loop gets treated as a controller that is trying too hard, so somebody softens it, the swing gets smaller, and the job closes. That is a real fix roughly half the time. The other half, the loop was hunting because of a delay that nobody looked for, and detuning it bought stability by making the system permanently slower at everything else, which shows up months later as a different complaint that no one traces back. The distinction is readable in about twenty minutes, and it comes down to one number that most techs never write down: how long the round trip takes.

Before you stand and watch a loop cycle

  • Keep hands, tools and pry bars out of a hunting actuator's travel path. It is moving under power and a spring-return unit stores energy. To touch a linkage, isolate the actuator, release or restrain the spring, and lock or tag under 29 CFR 1910.147.
  • Do not widen, jumper or replace a limit that is opening during the hunt. An oscillation that reaches a protective device is telling you the swing is now dangerous, not that the device is a nuisance. Establish why it opened first; replacing a correctly-operating limit reaches the same end state as jumpering it, one step slower.
  • A hot or pressurized line is a burn and scald route. Let it cool or use insulated gloves and eye protection, and take temperatures at existing wells rather than opening a pressurized line. If a line must be opened, isolate, drain to a safe point, and confirm zero on a gauge before a thread moves.
  • Where a measurement must be made inside an energized control enclosure, 29 CFR 1910.333(a)(1) permits energized work only where de-energizing introduces additional or increased hazards or is infeasible due to equipment design or operational limitations; use a meter and leads rated CAT III at or above the circuit voltage and work to the boundaries and protective equipment NFPA 70E-2021 assigns. Where the panel can be dead, open the disconnecting means, lock and tag under 29 CFR 1910.333(b)(2) in general industry or 29 CFR 1926.417 in construction, and prove dead per NFPA 70E-2021, 120.5.

The gate: round-trip time against correction rate

A loop hunts when the controller keeps correcting for a condition that has already changed. The controller acts, the effect takes time to arrive at the sensor, and during that time the controller sees no improvement, so it acts harder. When the effect finally lands, it lands too large, the error reverses, and the same thing happens in the other direction.

That is the entire mechanism, and the single gate that decides whether a given loop will hunt is: how long is the round trip from action to measurement, compared to how fast the controller corrects? A fast controller on a slow round trip hunts. The same controller on a fast round trip does not. Neither one is defective.

The consequence that makes this diagnostically useful: the period of the oscillation is set by the process, and the amplitude is set by the gain. Soften the controller and the swing gets smaller while the period barely moves. That is the fingerprint of a genuine tuning oscillation, and it is what tells you the delay is real and is where the money is.

Where the round trip actually comes from

Four contributors, and they are not interchangeable because the repairs differ.

Transport delay. Time for the fluid, air or product to physically travel from where the action happens to where the sensor sits. It is distance divided by velocity, it is pure delay with no partial response at all, and it is the hardest kind for any controller to handle. Nothing happens at the sensor for the entire duration and then everything happens.

Capacity lag. Time for a mass to change state: a coil, a vessel, a slab, a room. Unlike transport delay this one starts responding immediately, just slowly, which makes it far more forgiving.

Sensor time constant. Time for the sensing element itself to reach the value around it. A bare element in moving air is fast; the same element in a heavy well, in a thermowell with an air gap, or behind a software filter can be minutes.

Actuator stroke time. Time for the final device to reach the commanded position. A valve or damper with a long stroke adds delay on every correction, and a positioner with backlash adds an unpredictable amount.

Add them and you have the round trip. Most loops are dominated by one of the four, and finding which one is dominant is the whole job.

Two systems, one gate

Two nearly identical hydronic setups in the same portfolio, same controller model, same throttling range, same reset settings, commissioned by the same contractor. One holds beautifully and one hunts.

System A. The mixing valve is in the mechanical room and the control sensor is in the supply main 15 ft downstream. Flow velocity in that main, taken from measured flow and the pipe size rather than assumed, works out at about 3 ft per second. Transport delay is 15 divided by 3, or about 5 seconds. Add a 60-second valve stroke and a modest amount of mixing lag and the round trip is on the order of a minute. It settles.

System B. Same valve, same controller, same settings, but during a renovation the control sensor was moved to the return main, about 200 ft of pipe away. At the same 3 ft per second, transport delay is 200 divided by 3, or roughly 67 seconds, a bit over a minute of pure delay before the valve's own stroke time and the loads' thermal capacity are counted at all. It hunts, at a period of about 3 minutes and a swing of 6.0 F peak to peak.

Is a 3-minute period consistent with that delay? For a loop dominated by pure transport delay, an oscillation at the stability limit runs at a period on the order of twice the delay, so a bit over two minutes from transport alone, lengthened by the valve stroke and the mixing volume. Three minutes fits. Note the condition on that rule of thumb: it is derived for loops where pure delay dominates. A loop dominated instead by one large thermal capacity oscillates on a period set by that capacity, and the doubling rule does not apply to it.

Run the gate's own test on System B. Widen the throttling range from 10 F to 20 F, which halves the gain, and watch. The swing fell from 6.0 F to 2.5 F, a 58% reduction, while the period moved from 3.0 minutes to 3.2 minutes, about 7%. Amplitude collapsed, period held. That is a linear tuning oscillation on a real delay, confirmed, and it means the delay is worth removing rather than tolerating.

What removing the delay bought. The sensor was returned to the supply main 15 ft from the valve, and the throttling range was put back to the original 10 F. The loop did not oscillate, and it responded to load changes faster than the detuned version had. That is the argument in one sentence: detuning trades speed for stability, while removing the delay gives back both.

What would flip this. If return temperature genuinely is the variable that must be controlled - and on some systems it is, for a reason written into the design - then the delay is inherent and cannot be moved. The correct answer there is not a heroically detuned single loop but a cascade: a fast inner loop controlling supply temperature at the valve, taking its target from a slow outer loop that trims on return. The inner loop sees a short round trip and can be tuned tight; the outer loop is allowed to be slow because it should be.

The failure mode. Leaving System B detuned and moving on. It stops oscillating, the ticket closes, and every load change from then on is absorbed slowly. The follow-on complaint arrives as poor recovery after setback or as zones that never catch up in the morning, and nothing on either ticket connects the two, so the next tech tightens the loop back up and the hunting returns.

Oscillations that are not tuning oscillations

If softening the controller changes the period substantially, stop tuning. You are looking at something else, and four candidates cover almost all of it.

A limit cycle around a deadband on an on/off output. The period here is set by the differential and the rise and fall rates, not by any gain, so gain changes do nothing at all. This looks like hunting and is not; the cycle-rate arithmetic that governs it is covered in the deadband article.

Two controllers acting on the same process or the same actuator. Two devices with different targets, or a local controller fighting a supervisory command, produce a slow beat that no single-loop tuning resolves. The tell is that the actuator moves when the local controller's error says it should not.

Mechanical slop or a hunting positioner. Backlash means the actuator does not land in the same place for the same signal depending on direction, so the loop chases a moving target of its own making. These oscillations are often much faster than a process-driven hunt and can usually be seen at the actuator by eye.

Staging chatter. A stage cutting in and out around its own switching point, which is a staging differential problem living inside a modulating loop and gets missed because the modulating output looks reasonable between events.

Confirming which delay dominates

Time the step, do not estimate it. Put the loop in hand, make one deliberate output change large enough to see, and record two numbers: how long before the measurement moves at all, and how long until it finishes moving. The first is dead time, dominated by transport and sensor delay. The second is the capacity lag. A loop whose measurement does not twitch for a minute after a large output change has a delay problem no tuning parameter will fix.

Do that test with every protective device in service and stop it at the first sign a device is approaching its trip point. A step test is a deliberate disturbance, and it is not worth a nuisance trip that then needs its own diagnosis.

Compare the dead time you measured against the period you observed. If the period is many times the dead time and the process has no large mass to explain it, the oscillation is probably not this loop at all, and the two-controllers case deserves a look before anything else gets adjusted.

References

  • 29 CFR 1910.147, OSHA control of hazardous energy, for actuator spring tension and mechanical isolation
  • 29 CFR 1910.333(a)(1) and (b)(2), OSHA general industry work practices for energized and de-energized electrical work
  • 29 CFR 1926.417, OSHA construction lockout and tagging of circuits
  • NFPA 70E-2021, 120.5, verifying an electrically safe work condition
  • See related: How to Settle a Hunting Loop Without Guessing; What Deadband and Hysteresis Are For