Why a Failed-Open Trap Costs More Than a Failed-Closed One
Why this matters
A steam trap is a valve whose only job is telling condensate apart from steam. There are exactly two ways that discrimination can break, and they produce opposite experiences. Failed closed stops the heat and someone calls you within the shift. Failed open changes nothing anybody can feel and can run for years. The energy channel is dominated almost entirely by the second one, and not because it leaks faster. It leaks for longer, and how much longer is a number you set when you pick a survey interval.
The two directions have different cost structures, not different sizes of the same cost
Put the two side by side and the difference is not degree, it is kind.
| Failed open | Failed closed | |
|---|---|---|
| Steam lost | Continuous, at the seat's flow capacity | None |
| Effect on the process | None. Equipment still heats | Heat output falls, then stops |
| Who detects it | A survey, or nobody | The operator, usually within a shift |
| Typical exposure before repair | Months to years | Hours |
| Damage to equipment | Raises return pressure, stalls neighbours | Waterlogging, freeze split, hammer |
Neither column is the cheap one. The failed-closed column carries per-event damage that can total a coil, and that is a separate article's subject. What this one is about is the energy channel, where the two columns are not close, and the reason is in the exposure row.
Loss rate belongs to the trap, total loss belongs to your schedule
A failed-open trap's instantaneous loss rate is a property of the trap and the service: the open flow area at the seat, the absolute upstream pressure, and whether the flow is choked. You do not control any of those once the trap is installed.
What you do control is how long that rate runs. Total loss is rate multiplied by time, and time here is the interval between the moment the trap failed and the moment someone found it. Because a failed-open trap produces no process symptom, that interval is not set by the equipment or by luck. It is set by how often you look.
Under two conditions, that gives a clean and useful result. First, that failures arrive at a roughly uniform rate through the interval rather than clustering. Second, that detection happens only at survey. Under those two, a trap that fails at a random point inside a survey interval has been failed, on average, for half that interval when the surveyor arrives. Halve the interval and you halve the average undetected exposure, and therefore halve the standing loss the population carries at any moment. Both conditions have exceptions, and they are worth naming, so they get their own section below.
One failed-open trap can manufacture the appearance of more
This is the part that makes a blowing trap worse than its own flow rate. Live steam entering a condensate return raises the pressure in that return locally, and the trap next to it now has to discharge against a higher back pressure.
For a disc trap the consequence is direct: the disc is held down by a pressure difference across the seat, so as back pressure climbs toward inlet pressure, closing force falls and the trap goes to continuous blow. For mechanical traps the consequence is a loss of capacity, because the differential driving flow through the orifice has shrunk. Either way the direction is one-way: rising return pressure never helps a trap seat, it only ever pushes toward passing.
So a survey that finds six failed-open traps on one return header should not immediately generate six trap orders. Read the header pressure at a gauge port first. If it is elevated, some of those six are healthy traps reporting a system condition, and replacing them puts identical traps back into the same condition.
Worked example: 120 traps and one scheduling decision
A plant has 120 traps, all running continuously on year-round service. Survey history over the last three cycles shows about 5 percent of the population failing per year, and the failures splitting roughly two to one toward open. Both of those figures come from the plant's own survey sheets, which is where they should come from; another plant's numbers do not transfer.
That gives 120 times 0.05, or 6 failures per year. Two to one toward open puts 4 of those failing open and 2 failing closed.
Current interval, 24 months. Each failed-open trap sits undetected an average of half the interval, which is 12 months. Four per year at an average of 12 months each is 4 trap-years of continuous leakage carried at any given time. In hours, 4 times 8,760 is 35,040 trap-hours of leakage per year.
The two failed-closed traps in that same year are found by the process, not the survey, so their exposure is the response time. Say 4 hours each, giving 8 trap-hours.
The ratio between the two exposure figures is 35,040 divided by 8, which is about 4,380 to 1. That is the whole argument in one number, and notice what it is not: it is not a claim that the open failure leaks harder. It is a claim about how long each one runs.
Halve the interval to 12 months. Average undetected exposure per failed-open trap drops to 6 months. Four per year at 6 months each is 2 trap-years, or 17,520 trap-hours. The standing leakage halves, exactly as the uniform-arrival condition predicts.
Now price the decision in the only currency that matters here, technician hours. Say a full survey of 120 traps takes 16 technician-hours including access, tagging and write-up. At a 24-month interval that is 8 survey-hours per year. At 12 months it is 16, so the change adds 8 technician-hours per year.
Those 8 added hours remove 17,520 trap-hours of continuous leakage per year, which is about 2,190 trap-hours of leakage eliminated per added technician-hour. That ratio is an exposure comparison, not a return: the hours on the left are labour you spend and the hours on the right are equipment-hours of leakage you avoid, and they are two different kinds of hour. State it as what it is, an exposure reduction per hour invested, and it stays honest.
Re-read the direction words against the printed figures before using any of this. Failures per year held constant at 6 across both cases. Exposure halved, from 4 trap-years to 2. Survey hours doubled, from 8 to 16. Nothing in the example claims the failure rate changed, because shortening the interval does not prevent a single failure. It only shortens how long each one runs.
What changes the arithmetic
The uniform-arrival condition and the survey-only-detection condition are both defeatable, and each one moves the answer a different way.
Failures cluster instead of arriving uniformly. A population installed all at once, or one that took a single water hammer event across a header, fails in a batch. The half-the-interval average is then wrong in both directions depending on where the cluster falls relative to your survey date, and the fix is not a shorter interval but a survey timed to follow the event that caused the cluster.
The plant runs seasonally. The integral is in operating hours, not calendar months. A coil trap on a service that runs 1,400 hours a year and fails in the off season has an exposure of nearly zero regardless of your survey calendar. Rank on operating hours rather than calendar age, which is its own subject and has its own article.
The failures are systemic rather than independent. Elevated return pressure, undersized return piping, or a change in operating pressure puts many traps into passing at once. Shortening the interval then finds the same population repeatedly and repairs nothing, because the fault is upstream of every trap on the list.
Detection is not survey-only. Where a service has permanent trap monitoring, or where a receiver vent is visible from an occupied area and someone genuinely watches it, the exposure clock is not the survey clock and the halving argument does not apply to those traps.
How to verify the interval you picked is the right one
- Check the failure rate you assumed against your own last three surveys, not against a figure you read. If your rate is 2 percent rather than 5, everything above scales down with it and a longer interval may be correct.
- Check the open-to-closed split on your own sheets, and read it as a survey artifact. Failed-closed traps are removed from the population by the process before the surveyor arrives, so a survey list skewed toward open is partly a selection effect and not a statement about which mode is more common. The article on why failed-closed traps get found first owns that mechanism.
- Confirm return header pressures at gauge ports before ordering traps on any header that produced multiple failures in one survey.
- Where a blowing trap discharges to atmosphere through an open vent or an open-ended blowdown, keep the discharge routed away from walkways and do not stand in front of it; a steam jet is invisible for the first stretch out of the seat, and sustained exposure at close range falls under your hearing conservation program per 29 CFR 1910.95.
- Re-survey a sample of ten traps a month after a full survey. If several have changed state, your interval is longer than that population supports and the arithmetic above already tells you what shortening it buys.
References
- 29 CFR 1910.95, occupational noise exposure, where sustained work near a blowing trap or open blowdown places a technician in continuous high noise
- Trap manufacturer published data for maximum back pressure as a percentage of inlet pressure, which sets when a rising return stalls a given trap body
- Your own plant survey records for failure rate and open-to-closed split, which are the only defensible source for both figures
- See related: Why a Failed-Closed Trap Is Found First and Fixed First; How to Survey a Trap Population and Rank What to Fix