A Checked Box Is Not a Passed Check

Why this matters

Somewhere in most shops there is a number called an item pass rate, and in most shops it is not one. It is computed from whether an item was marked off, not from whether the item recorded a pass. An item completed with a failing result counts identically to one completed with a passing result, so the number cannot fall when the work gets worse. It falls only when the paperwork gets worse.

That makes it a proxy for compliance wearing a quality label, and the blind spot is absolute rather than partial: a deteriorating check is invisible to it by construction, not by bad luck. This card is about finding out which number you have and building the other one.

The test that tells you which number you have

It takes one job and about two minutes.

Open a checklist on a live or test job. Find an item that can genuinely fail - a reading against a range, a condition against a standard - and record a failing result on it. Save it. Then look at the item rate for that period.

If the number does not move, you have a completion rate. The failure you just recorded was counted as a success, because the only thing being counted is that the item stopped being blank.

If the number drops, you have a pass rate, and the rest of this card is a sharpening exercise rather than a rebuild.

Run this test even when you are fairly sure of the answer. The distinction usually sits in a setting nobody chose, and shops that have been reading the number in a meeting for two years are exactly the ones who have never checked which one they were reading.

Why almost every system defaults to the completion reading

A checkbox has two states: done, not done. A verdict needs at least three - passed, failed, not applicable - plus a fourth state for not yet reached. Checklists ship as checkboxes because a checkbox is what a checklist has always been, and because two states is what a person building a template reaches for when the template is mostly tasks.

The practical consequence is that in a checkbox-only template there is nowhere to put a failure. The tech finds the condition bad, fixes it, and records that in a note. Notes are prose. Nothing counts them, nothing trends them, and the item still shows as done because it was done. The number is not lying about its inputs; it was never given the input that would make it a quality measure.

So the fix is not a reporting fix. It is a template fix, and it has to happen before any number gets rebuilt.

Verdict items and task items

Split every item on your templates into two kinds. This is the whole build.

Verdict items produce a result that could come back bad: a measured value against a published range, a pressure, a clearance, a torque, a visual condition against a defined standard, a functional test with a defined outcome. These are the only items a pass rate can be computed over.

Task items are things you do: install it, clean it, replace the filter, photograph the plate, tidy the area, fit the cover. They have no failing state. Not doing one is an omission, not a failure.

Putting both kinds in one denominator is what dilutes a pass rate into uselessness. On a template that is two thirds tasks, a verdict item's failure is diluted by every checkbox around it, and the aggregate barely twitches when something real goes wrong.

Convert the verdict items to a result field and leave the task items as checkboxes. Do not convert everything: a sixteen-item template where every line demands a pass or fail is slower on the job, and the predictable response is a crew that closes the whole thing in one pass at the end of the day, which is the exact behaviour that makes completion data worthless. See related: Checklist Completion Rate and What Completion Means.

Building the pass rate that earns the name

Three lines, published together, and each one answers a different question.

  • Coverage. Verdict slots assessed, over verdict slots expected. A slot is one verdict item on one checklist instance. This is where "not yet done" and "skipped" go.
  • Pass rate. Verdict slots that recorded a pass, over verdict slots assessed. Not over slots expected. A slot nobody assessed is not a failure and it is not a pass, and folding it into either direction hides the thing coverage exists to show.
  • Per-item failure rate. For each verdict item, failures over that item's own assessments. This is the line that actually finds anything, and the aggregate exists mainly to tell you when to go read it.

Give the per-item line a floor: about 30 assessments of that item inside the window before its failure rate is worth ranking against other items. Below that, roll the item up to a year before reading it, because at 20 assessments two failures and four failures are a rate of 10 and 20 percent of that item's assessments and neither is a finding.

Not applicable, and found versus left

Two decisions have to be made once, written down, and held, or the pass rate drifts without anybody touching the template.

Where not-applicable goes. A verdict item that genuinely does not apply to this job is neither a pass nor a fail. It belongs outside the pass numerator and outside the assessed denominator, and it shows up in the coverage line instead. Shops that count not-applicable as a pass inflate the rate by exactly the share of slots that were skipped, which is largest on the templates that fit the job worst. Shops that count it as an assessment dilute the rate downward for the same reason. Neither is worse than the other; both are silently wrong in a direction that moves with how badly the template fits.

Then watch the not-applicable rate itself, per item. It is the available escape hatch, and a technician under time pressure will find it before they will leave a blank, because a blank looks unfinished and an NA looks decided. One person's NA share on one item running several times everyone else's is worth a look, subject to the same 30-assessment floor before it is read as a rate.

Found or left. Decide whether a verdict item records what the tech found or what they left behind after correcting it. Record what was found. An item that logs the corrected state can never fail, which returns you to the problem this card exists to solve, one step further in. The corrective action belongs in the note and in the parts or labour on the job; the verdict belongs on the item. Written down once, this also settles the argument that otherwise arrives the first time a tech's pass rate is discussed out loud.

Worked example: the quarter the headline did not move

A shop runs one 16-item template on its largest job type. Five items are verdict items; eleven are tasks.

The headline the shop reads.

  • Last quarter: 116 checklist instances, so 1,856 item slots. 1,765 marked. 95.1 percent of item slots marked.
  • This quarter: 120 instances, so 1,920 item slots. 1,832 marked. 95.4 percent of item slots marked.

Flat, up a whisker. Nobody had a question.

One verdict item, read on its own. The item is a measured value checked against the published range.

  • Last quarter: assessed on 114 of 116 instances. 3 failures, so 2.6 percent of its 114 assessments failed. Pass 111 of 114, 97.4 percent of its assessments.
  • This quarter: assessed on 117 of 120 instances. 29 failures, so 24.8 percent of its 117 assessments failed. Pass 88 of 117, 75.2 percent of its assessments.

Both quarters clear the 30-assessment floor by a wide margin, so the comparison is a real one. That item's failure rate went up about nine times quarter over quarter, and its own pass rate fell 22.2 points on its own assessment base.

Why the headline did not move. The item was marked on 114 of 116 instances last quarter and 117 of 120 this quarter. To a marked-off count those two quarters are the same quarter. The 29 failures were all recorded, all fixed on site, and all counted as successes by the only number anybody was reading.

The verdict-only aggregate, built properly.

  • Last quarter: 580 verdict slots expected (5 items on 116 instances), 571 assessed, 14 failures. Pass 557 of 571, 97.5 percent of assessed verdict slots. Coverage 571 of 580.
  • This quarter: 600 verdict slots expected (5 on 120), 588 assessed, 41 failures, of which 29 sit on the single item above and 12 are spread across the other four. Pass 547 of 588, 93.0 percent of assessed verdict slots. Coverage 588 of 600.

Applying the not-applicable rule to the unassessed slots: of the 12 this quarter, 9 were marked not applicable and 3 were left blank; of the 9 last quarter, 7 were not applicable and 2 blank. None of those 21 slots entered either pass rate, in either direction, which is why the two quarters' pass rates are comparable to each other at all.

The 29 failures on the deteriorating item were every one of them corrected on site before the tech left. Under the found-not-left rule they are still 29 failures, and that is what makes the item's line readable.

So the properly built number fell 4.5 points on assessed verdict slots over the same two quarters in which the marked-off headline rose 0.3 points on all item slots. Those are two different bases and two different constructs, which is the point: they are not a disagreement to be reconciled, they are one number that can see the problem and one that cannot.

What the shop does with it. The aggregate 4.5-point fall is the alarm, not the finding. The finding is the per-item line, which puts 29 of the quarter's 41 verdict failures on one item, and the next question is whether that item's failures cluster on a component, a supplier, an install date range or a technician. The aggregate would never have told anyone where to look.

Two degenerate cases that show what the rule is doing

A template with no verdict items at all. Every item is a task. Nothing on it can record a failure, so a pass rate computed over it is pinned at 100 percent of marked items permanently, in every period, whatever happens in the field. If your pass rate has never printed a figure different from your completion rate, check the template before you check anything else - this is almost always why, and no amount of reporting work will fix it.

A verdict item that has never once failed. Across a full year of use, zero failures recorded. Two explanations produce identical data, and you cannot tell them apart from the number.

The check is not statistical, it is a document search. Go and find a job where that condition genuinely was bad: a comeback, a warranty claim, a customer complaint that named it, a part returned. Then read what the item said on that job.

  • If the item says done and the condition was bad, nobody is running that check. The item is recording attendance.
  • If you search a full year and cannot find a single job where the condition was bad, the check is real and it is being run at the wrong frequency. Move it to an annual audit or a commissioning sheet and take it off every visit. A check that has never caught anything is costing time on every job and buying nothing, and leaving it there is what makes a crew treat the whole template as theatre.

Both endpoints collapse the same rule. A pass rate only means something over items that can fail, assessed by people who would record it if they did.

References

  • See related: Checklist Completion Rate and What Completion Means, for the checklist-level rate and its own blind spot
  • See related: Service Quality Checklist Discipline, and What a Competency Checklist Should Actually Contain
  • See related: The Checklist That Grows Every Time Something Goes Wrong, for deciding which items earn a place on the template
  • Manufacturer commissioning and service documentation, for the published range or condition a verdict item is judged against