Measuring Whether Your Training Is Working
Why this matters
Most shops that train cannot tell you whether it worked, so the first slow quarter they cut it, because an expense you cannot defend is an expense you eventually stop paying. The shops that keep training through a downturn are the ones holding a number. The catch is that the obvious numbers are all wrong: hours delivered, sessions run, attendance, and how good the session felt measure your effort, not the crew's behavior. Training is only real when something different happens on a truck, and the whole measurement problem is finding the cheapest honest evidence that it did.
Activity is not evidence
Write this down and post it where you plan the training calendar: you cannot measure training, you can only measure behavior downstream of it.
The four metrics shops reach for first, and why each one lies:
- Hours delivered. Measures input. A shop can deliver 200 hours of training and change nothing. Reporting hours to yourself is how a training program becomes theater.
- Attendance. Measures compliance. Attendance is worth tracking as a hygiene check, because 60 percent attendance means the slot is not protected, but full attendance proves nothing about learning.
- Session feedback. Measures how the session felt. It has one legitimate use, catching a session that was genuinely confusing or badly pitched, and one dangerous failure mode: the most enjoyable sessions are frequently the least transferable, because entertainment and difficulty pull in opposite directions.
- A written test. Measures recall in a room. Recall is necessary and nowhere near sufficient. Techs who can name every step of a procedure on paper routinely skip step three at the equipment because step three is inconvenient and nobody is watching.
None of these are useless. All of them are upstream. Track them cheaply, and do not mistake any of them for the answer.
Four questions, in order, and only the last two decide anything
The classic training-evaluation ladder has four rungs, and it maps cleanly onto a shop that has no training department:
- Did they engage? Cheap to observe, weakest evidence. One line: did people ask questions, or wait for it to end.
- Can they do it cold? A demonstration, unaided, some days after the session, not during it. This is the first rung that carries real information, and it is the one shops skip most often.
- Did it change what happens on the truck? Behavior in the field, measured without the trainer present. This is the rung where training either exists or does not.
- Did the business number move? Callbacks, first-visit completion, warranty rate, rework hours. Strongest evidence, slowest to appear, most contaminated by everything else going on.
The practical rule for a small shop: measure rung two and rung three deliberately, watch rung four on a lag, and stop pretending rung one means anything. Rung two is nearly free. Rung three costs one query against records you already keep. Rung four is where you argue the case to yourself in a bad quarter.
Leading and lagging, and why you need both
A lagging indicator tells you what happened. A leading indicator tells you what is about to happen. Training programs die when a shop watches only lagging indicators, because the lag is long enough that people lose faith before the number moves.
- Leading: documented reading before a part is replaced, checklist completion rate, second-opinion calls to a lead tech (which should rise first and then fall), diagnosis time on a target fault family, number of skills on the matrix with more than one qualified name against them.
- Lagging: callback rate in the target family, rework hours, warranty claims, first-visit completion, average time on that job type, tech retention.
Note the counterintuitive one. Calls to the lead tech usually go up right after good training, because a tech who now understands what they do not know starts asking. Treat an early rise as a positive signal and get worried only if it has not turned down by three to four months.
The metric set worth actually keeping
Keep few. A shop that tries to track twelve numbers tracks zero by month three.
| Metric | Type | Where it comes from | What it tells you |
|---|---|---|---|
| Callbacks in the target fault family, as a share of jobs in that family | Lagging | Job records | Whether the training hit the thing it aimed at |
| Documented reading before a part swap | Leading | Job notes | Whether the method is being used, not just known |
| Cold demonstration pass rate at 30 days | Leading | Sign-off log | Whether it stuck past the session |
| Skills with only one qualified name | Leading | Skills matrix | Your single-point-of-failure exposure |
| Rework hours per month | Lagging | Time records | The hours side of the business case |
| First-visit completion rate | Lagging | Job records | Whether diagnosis plus truck stock is improving together |
The first two are the workhorses. If you only ever track two, track those, and always express the callback figure as a share of jobs in that family rather than a raw count, or a slow month will read as an improvement.
Attribute it honestly
Before you claim a result, hold it up against the three things that most often produce a fake improvement:
Seasonality. Comparing eight weeks of peak against eight weeks of shoulder compares two different job mixes. Compare against the same window last year where you can, and where you cannot, say so out loud rather than quietly banking the number.
Staffing change. One new hire or one departure can move a five-tech shop's average more than any training will. Always segment by tech before you look at the average.
Everything else you changed. Shops rarely change one thing. If you trained on a fault family the same month you added a part to truck stock, the improvement is shared and you cannot say in what proportion. That is fine to admit. It is not fine to attribute all of it to training and then be surprised when the next training with no truck-stock change does nothing.
Worked example: one skill, one full cycle
A five-tech shop targets one behavior: take and document a reading before replacing a part in its worst fault family.
Baseline, eight weeks before training. 62 jobs in that family. 23 had a documented reading before a part went in, so 37 percent. Callbacks from those 62 jobs: 11, or 17.7 percent.
The intervention. Two twelve-minute toolbox talks, one bench-rig practice round of about 45 minutes per tech, and a change to the job-note template that puts a reading field where a tech cannot avoid seeing it. Total delivered time is roughly 5 crew-hours plus the template change.
Measurement, the eight weeks after. 58 jobs in the family. 41 had a documented reading, so 71 percent. Behavior moved 34 percentage points, which is the rung-three result and the one worth celebrating, because it appeared within weeks.
The lagging number. Callbacks from those 58 jobs: 5, or 8.6 percent, down from 17.7 percent. Roughly half. In absolute terms the family produced 6 fewer callbacks across 8 weeks. At about 1.5 unbillable hours per callback that is 9 hours recovered in 8 weeks, or about 1.125 hours a week, which annualizes to roughly 54 hours a year against about 5 crew-hours of delivery plus a template edit.
Now the honest audit. Segmenting by tech shows the documentation rate rose for four of the five, and the fifth barely moved. That single fact changes what you do next: this is no longer a training question, it is a coaching or a will question for one person, and running the whole crew through the session again would be a waste of everyone else's time. Also worth naming: the note-template change may be carrying a large share of the documentation improvement on its own, and a fair reading is that the talk taught the method while the template made skipping it awkward. Both were needed. Claiming the whole 54 hours for the talk would set you up to be wrong the next time you run a talk without a system change beside it.
When the number does not move
Test these three explanations in this order, because they get progressively more expensive to act on.
- They cannot do it. Run the cold demonstration. If they fail it unaided at 30 days, the training did not land and the fix is more practice, better practice, or a different teacher - not more pressure.
- They can do it but the job does not let them. Far more common than shops expect. The method takes eight extra minutes, the schedule allows none, and the tech is being measured on stops per day. You trained a behavior your dispatch model punishes. The fix is in scheduling, not in training, and no amount of retraining will beat the incentive.
- They can do it, the job allows it, and they are not doing it. Now it is a will question and it belongs to a manager, one person at a time, not to the training program.
Running these in the wrong order is the classic waste: shops retrain, repeatedly, when the real answer was number two the whole time.
Verifying your measurement itself is sound
- Check the baseline was measured before anyone knew you were watching. A baseline collected the week you announce the training is already contaminated upward, and it will make a real improvement look like nothing.
- Check the same person is not scoring their own students. A cold demonstration signed off by the trainer runs a real risk of measuring the relationship rather than the skill. Have a second qualified tech sign off wherever you have one.
- Check for a definition drift. "Documented reading" has to mean the same thing in week 16 as it did in week 1. The most common quiet failure is that the standard loosened while the number improved.
- Re-measure at six months, not just at eight weeks. The rung-three number that holds at six months is the only one worth putting in front of yourself when you are deciding whether to protect the training slot through a slow quarter.
References
- Kirkpatrick four-level training evaluation model, widely used training-standard practice
- U.S. Department of Labor, Employment and Training Administration guidance on competency-based evaluation
- OSHA 29 CFR 1910.147(c)(7)(ii), demonstrated proficiency as the standard for training verification
- See related: Build a Skills Matrix: Who Can Do What, The Toolbox Talk SOP, The Hidden Cost of Under-Training a Tech, How to Decide When Someone Is Ready to Work Alone