The Reading That Was Right and the Conclusion That Was Wrong
Why this matters
The expensive mistakes in this trade are rarely bad readings. They are good readings answering a question nobody asked. The number is accurate, the instrument is fine, the technique is fine, and the sentence written next to it in the report is wrong, because the number described a condition that was true at that moment and was never true at the moment that mattered.
This case is reconstructed backwards. By the time the third technician arrived, the original condition had been altered twice and a part had been replaced, so nothing that was measured on day one could be measured again. What was left was the record, and the record turned out to contain enough to find the fault and enough to show exactly where the reasoning went off. The gaps in it are part of the finding.
Before anyone measures anything on a fired appliance
This case runs on a fuel-fired heating appliance, and the method it ends with requires standing next to it while it fires for half an hour. That is the hazard the diagnosis creates, so it gets handled before the diagnosis.
Wear a personal carbon monoxide monitor, on your person and switched on, from the moment you enter the space until you leave, under whatever program your employer has established for it. If it alarms, shut the appliance down at its own controls, get everyone out of the space, and ventilate from outside the space before anyone re-enters or troubleshoots anything. Do not operate the appliance with any part of the flue or vent disconnected or with a panel removed that the manufacturer's instructions require in place for the appliance to vent correctly. If you smell gas at any point, nobody touches a switch, a light or a phone inside; everyone leaves immediately and the call to the gas utility is made from outside.
Take temperature readings at existing test ports where they exist. Where a port has to be made, put it where the manufacturer's instructions permit one, never into the combustion or flue path, deburr the opening, and seal it before you leave.
What the record contained
Three visits over five weeks, all in the same file:
Visit 1. Complaint of insufficient heat. Recorded: temperature rise across the appliance measured at 12, runtime at measurement 4 minutes, entering air temperature recorded, instrument identified with a check date inside its interval. The manufacturer's documentation for this appliance states an acceptable rise band of 16 to 22, at steady state, and defines steady state as after a stabilization period stated in that same literature. Conclusion recorded: rise below band, capacity fault. A component was replaced.
Visit 2. Complaint continued. Different technician. Recorded: rise 18, runtime at measurement 22 minutes. Conclusion recorded: within band, no fault found. Customer advised the system was operating correctly.
Visit 3. Complaint continued and the customer was now out of patience.
Two visits, two readings, opposite conclusions, and no dispute about either number. Both were taken competently. This is the situation that makes shops distrust their own techs, and the distrust is misplaced, because neither reading was wrong.
Reconstructing visit one from four fields
The third technician did not start at the appliance. He started at the file, and four recorded fields did the work.
The value, 12. Against a band of 16 to 22 it is low by 4 at the nearest edge, which is 25 percent below the bottom of the band. That is a real gap, not a rounding argument, and it is well outside any plausible uncertainty for the instrument named on the ticket.
The runtime at measurement, 4 minutes. This is the field that broke the case open, and it is the one most reports do not have. The band it was compared against is published at steady state. Four minutes is not steady state on this appliance by its own literature.
The entering air temperature. Recorded on visit 1 and on visit 2, and close enough between them that entering conditions cannot account for the difference between the two readings.
The instrument and its check date. Same instrument type on both visits, both inside their check intervals, so the 6-unit difference between visit 1 and visit 2 is not two instruments disagreeing.
Strip those together and the two visits stop contradicting each other. One measured the appliance at 4 minutes and one measured it at 22 minutes. They are two points on the same warm-up, not two opinions about one condition.
The number was correct and the comparison was not
Visit 1's 12 is what that appliance actually did at 4 minutes of runtime. Nothing about it is a measurement error. What went wrong is that it was compared against a band established under a condition the measurement did not meet, and the conclusion drawn - a capacity fault - is a statement about the appliance at steady state, which is a state the technician never observed.
The failure has a name worth carrying: the reading was valid and it was not relevant. Those are two separate tests and a number has to pass both.
Validity asks whether the number is what the instrument claims: right function, right point, good coupling, instrument within its check, resolution finer than the difference being argued.
Relevance asks whether the condition the number describes is the condition the decision is about. This is the test that gets skipped, because a valid number feels finished. It is not finished until you can name the operating state it describes and confirm the specification you are comparing it against was established in that same state.
Visit 2 failed the same test in the opposite direction. Its 18 at 22 minutes was valid and it was also not relevant, because the question by then was not "can this appliance make heat," it was "does this house get heat," and a reading taken at 22 minutes of runtime tells you nothing about a machine that never runs that long.
What the third visit measured
Same appliance, same instrument type, existing test ports, personal CO monitor on and the flue and panels intact throughout. Rise recorded at four points, with the setpoint driven well above room temperature to force a continuous run rather than on a normal call for heat: at 4 minutes, 11. At 10 minutes, 15. At 20 minutes, 18. At 30 minutes, 18.
Read that series as printed. It rose across the first three readings and then held between the third and the fourth, which is what a warm-up looks like when it finishes. The appliance stabilizes at 18, inside the published band of 16 to 22. On capacity it is fine, and the component replaced on visit 1 was replaced on the strength of a number that never said what it was read to say.
The 11 at 4 minutes is also worth noting against visit 1's 12 at the same runtime. Two techs, five weeks apart, one component swapped in between, and the readings at matched runtime differ by 1. That is the strongest evidence in the file that the swap changed nothing.
Then the actual fault, which the same visit found once the forced run was released. During normal operation this appliance was observed to satisfy and shut off at around 6 minutes. It spends its entire working life on the rising part of that curve and never reaches the number its own documentation is written about. Why it stops at 6 minutes is its own investigation - control differential, sizing, and where the controlling sensor is reading from are all live candidates, and a sibling article covers what a sensor's location does to what it reports.
The field that made this possible
Runtime at measurement. One field, five characters on a form, and without it this file is two technicians who disagree and a customer who has to pick one.
The general form of it: every reading taken on a machine that changes with time needs the elapsed time recorded beside it, measured from a stated event. Not the clock time, the elapsed time, and the event it is measured from. Clock time is what the ticket already has, and clock time cannot tell a later reader whether the machine had been running four minutes or forty.
What the file did not have, and should have: nothing on either visit recorded how long the appliance normally runs before it stops. That is a different quantity from runtime at measurement, and it is the one that turned out to be the fault. A record can be good enough to exonerate two readings and still be missing the field that would have found the answer, and both of those things were true here.
What it cost, in technician hours
All figures below are technician hours on site, one currency throughout, and the replaced component is on top of them.
Visit 1 ran 1.5 hours. Visit 2 ran 1.0 hour and produced nothing. Visit 3 ran 2.0 hours, most of it waiting through the warm-up. Total 4.5 hours across three trips.
Had visit 1 held the appliance through stabilization before concluding anything, it would have added roughly 0.4 hour to that first call, for 1.9 hours in total. The wrong conclusion therefore cost about 2.4 times the hours the right one needed, plus a component, plus five weeks of a customer's winter and whatever that does to the referral.
The tradeoff is not close, and it generalizes: on any machine with a warm-up, the time spent waiting for the condition the specification was written under is smaller than the time spent servicing the consequences of not waiting.
References
- Manufacturer's installation and service documentation for the appliance, which states the acceptable rise band, the conditions it was established under, and where test ports may be made
- 29 CFR 1910.132, OSHA general industry, employer assessment and provision of personal protective equipment, under which a personal carbon monoxide monitor program is established
- Trade-standard practice for combustion appliance service, including personal CO monitoring and intact venting during operation
- See related: What a Sampling Interval Hides; What to Write Down So a Reading Survives