The Annual Skills Review SOP
Purpose
To verify, once a year, what every person in the shop can actually do unassisted, correct the skills matrix against evidence rather than memory, and produce two outputs: a short development plan per person and a training plan for the shop that fits inside the capacity the shop actually has. The review exists because an unmaintained skills matrix does not stay neutral, it inflates. People get credited for things they did once, three years ago, and the shop makes coverage decisions off a document that is quietly wrong.
Scope
Covers every person whose work is on the task inventory, field and office. Covers technical competencies, qualification-gated tasks, and the back-office tasks that gate field work.
Explicitly out of scope: compensation, performance, attitude, attendance, and any discussion of whether someone should keep their job. Those are legitimate conversations and they belong in a different meeting, at a different time of year. The separation is the design point of this SOP, not an oversight. A person who suspects their pay is riding on the answer will not tell you honestly what they cannot do, and an inflated matrix is exactly what this review exists to prevent.
Roles and responsibilities
| Role | Responsibility |
|---|---|
| Program owner (owner or lead) | Schedules the review, refreshes the task list, owns the shop-level rollup and the year's plan |
| Reviewer | Runs the individual conversations; must be someone who knows the work, not only the people |
| Employee | Self-rates first, in writing, before seeing anyone else's rating |
| Verifier | Observes contested and high-consequence competencies on live work; must not be the person's mentor |
| Scheduler | Sequences the resulting pairings against real capacity and season |
Procedure
1. Time it away from performance and pay
Hold the skills review at least several weeks clear of any performance or compensation conversation, and preferably in a different quarter. Say plainly at the start of each session that nothing here feeds pay this cycle. Then honor it, because the crew will check.
Pick a slow stretch. The review takes real hours and it is the first thing to get cancelled if it lands in peak season.
2. Refresh the task list before rating anyone
Pull the last twelve months of dispatched work and update the recurring task inventory. Add what the shop started doing, retire what it stopped. Rating people against last year's task list is how a matrix drifts out of alignment with the business without anyone noticing.
Add the categories dispatch will not show you: back-office tasks that gate field work, qualification-gated tasks, and judgment calls like quoting a replacement or telling a customer no.
3. Apply the recency rule before the meeting
Use a four-level scale and no more. More levels sound precise and produce arguments about the difference between a 3 and a 4.
| Level | Meaning |
|---|---|
| 0 | Not trained. Does not perform this. |
| 1 | Can assist, or can perform with someone available. |
| 2 | Performs unassisted, to standard. |
| 3 | Performs unassisted and can teach and sign off others. |
Then apply the rule most matrices lack: a level 2 or 3 with no occurrences in the last twelve months automatically drops one level unless it is re-verified this cycle. Skills that are not exercised decay, and the decay is invisible precisely because the person still believes they can do it and so do you.
Run this against dispatch history before any conversation, so the meeting starts from a corrected draft rather than from last year's optimism.
4. Gather three evidence sources
Never rate from one source. Each of the three is wrong in a different direction, which is why you want all of them.
- Dispatch history. Who was actually sent to what, how often. This is the strongest signal for level 2 and the hardest to argue with. A person credited at level 2 who has never been dispatched to that task alone is not a level 2 in practice, whatever the matrix says.
- Callback and rework, by cause. Not raw callback count, which mostly tracks job volume. Cause, so you can see whether a person's rework clusters on one competency.
- Direct observation. Required for anything contested, anything qualification-gated, and anything where the consequence of being wrong is a safety or compliance event.
5. Self-rate first, then reconcile
The employee rates themselves in writing before seeing the reviewer's draft. Doing it in the other order produces agreement, not information, because most people will not argue down a rating their boss just handed them.
Then compare. The gaps are the whole value of the exercise, and they point in both directions:
- They rated themselves lower than the evidence. Common in careful people and in anyone who has been in the trade long enough to know how much they do not know. Fix the rating and tell them why, because underrating is why good people do not get dispatched to work they can handle.
- They rated themselves higher than the evidence. Do not argue about it. Route it to step 6 and let the observation settle it.
6. Verify on live work, not in conversation
Every contested cell and every high-consequence cell gets observed on a real job by the verifier, who is not the person's mentor. Watching one real occurrence resolves in twenty minutes what a discussion will not resolve at all.
Any lapsed qualification found during the review is a stop-work item, not a development item. If someone has been performing work that requires a current certification and the certification has expired, they come off that work immediately, today, and the fix is the renewal process rather than a line on a development plan. This is the one finding in the whole review that does not wait for the plan.
7. Set two to three competencies per person, each with a sign-off bar
Not a list of ten. Two or three, chosen where the person's interest and the shop's exposure overlap. A plan with ten items on it is a plan with zero items on it, because nothing is prioritized and everything gets deferred.
Each competency is written the way the pairing SOP requires: as a task the person will perform, with a stated bar. "Commission a new install unassisted, three consecutive times, judged by the verifier" is a competency. "Get better at commissioning" is not.
At least one of the shop's competencies for the year should move somebody from a 2 to a 3, because your level-3 count is what caps how much training you can run at all. Step 8 shows why.
8. Roll up and check against capacity
Aggregate the corrected matrix and read three things:
- Where coverage is one. Feed this straight into the single point of failure audit. The two documents are meant to be read together.
- How many level 3s exist, and on what. This is your entire mentor pool. Tasks with zero level 3s cannot be taught internally at all this year, and need an outside course, a qualified hire, or a standing sub relationship.
- Whether the year's plan fits. Total the competency moves and divide by how many pairings you can run at once, which is capped by your level-3 count and the mentor load cap of one active pairing each.
9. Sequence against season, not ambition
Hand the sequence to the scheduler with the seasonal calendar in front of you. Start pairings so their second month lands outside peak, because that is where training programs actually die. If a wave has to straddle peak season, write the pause and the restart date into the plan rather than letting it drift.
10. Record and set the next date
Store the corrected matrix with the date and the evidence basis for each changed cell. Store each person's two or three competencies with their bars. Store the shop rollup and the sequenced plan. Set next year's review date now, in the slow season.
Retain qualification-related records for as long as the qualification requires, and longer than you think, because a sign-off record can matter in an incident review.
Worked example: a six-tech shop
Six people against a refreshed inventory of 24 task types is 144 cells in the matrix.
Self-rating comes back with 61 cells at level 2 or above. Evidence checking knocks 14 of those down: 9 from the recency rule, where the person genuinely could do it once but has not touched it in over a year, and 5 where dispatch history shows the task has never gone to them unaccompanied. That leaves 47 verified cells at level 2 or above, meaning the self-rated matrix was overstated by about 23%.
That 23% is not dishonesty. It is what any matrix does when nobody re-tests it, and it is the number that would have driven this year's coverage decisions if the review had not run.
Level 3s across the whole shop: 8 cells, held by 4 people. Across 24 task types, that means 16 tasks have nobody who can teach and sign off. Those 16 are not candidates for internal pairing this year regardless of how much the shop wants them covered, and naming that early prevents a plan built on capacity that does not exist.
The plan: 6 people times 2 competencies each is 12 competency moves for the year. With 4 people holding level 3 somewhere and a cap of one active pairing per mentor, the shop can run 4 pairings at a time. Twelve moves at four concurrent is 3 waves. At roughly 10 weeks per pairing including sign-off, that is 30 weeks of the 52, leaving 22 weeks of slack for the peak-season pause and for pairings that need an extension.
That fits, but only just, and only because the number was checked. A shop that set 3 competencies each would be planning 18 moves, which at 4 concurrent is 5 waves and 50 weeks, with zero slack and a peak season in the middle of it. That plan fails in the second month of wave two and everyone concludes training does not work here.
Cost of running the review itself: about 45 minutes per person in conversation, so 4.5 hours, plus roughly 3 hours of evidence preparation and 1.5 hours for the rollup and sequencing. Call it 9 hours, once a year, mostly in the slow season. A single high-exposure gap left uncovered typically costs more than that in rework and rescheduling in one season.
What would change this plan. If the shop had 2 level-3 holders instead of 4, concurrent capacity halves to 2 pairings, 12 moves becomes 6 waves, and the honest response is to cut to one competency per person and spend the freed capacity growing level 3s instead. If a competency is qualification-gated, it leaves the pairing queue entirely and goes on an external timeline. If a key mentor is leaving inside the year, their pairings move to the front of the sequence regardless of exposure ranking, because their availability is the constraint that expires.
Review of this SOP
Reread the SOP itself each year before running it. The two numbers that tell you whether it is working: what fraction of last year's competencies actually closed with a sign-off, and whether the overstatement rate found in step 5 is falling year over year. A closure rate under about half means the plan was sized past capacity, which is a step 8 problem, not a people problem.
References
- U.S. Department of Labor, registered apprenticeship competency tracking and on-the-job learning records
- OSHA general industry standards on qualified persons and currency of required training
- See related: The Single Point of Failure Audit for Shop Skills
- See related: The Mentor Pairing SOP
- See related: Why Most Small-Shop Training Fails in the Second Month