How to Measure Whether You Are Actually Less Busy

Why this matters

You handed off two things, tightened the intake, and put a rule on interruptions. Six weeks later somebody asks whether it worked and the honest answer is a shrug. Feeling is a terrible instrument here: a single bad Thursday will convince you nothing changed, and a quiet fortnight will convince you everything did.

The stake is not curiosity. If a change did not work and you believe it did, you stop there and carry the same load for another year. If it did work and you cannot see it, you conclude that delegation does not pay and you stop handing things over, which is the more expensive mistake of the two.

Step 1: Pick one definition of "less busy" before you look at any numbers

Three definitions are in circulation and they can move in opposite directions at the same time. Pick one as primary, in writing, before measuring.

  • Fewer hours. You work less. Simple, and the least common outcome of a delegation.
  • Fewer of your hours spoken for by other people's work. Same total hours, different composition.
  • More usable hours. Same total, same composition, but arriving in longer stretches so real work can land in them.

Choosing after the fact is how every change gets declared a success. If your primary definition was fewer hours and the hours did not fall, the change did not deliver what you wanted, regardless of what else improved - and that is worth knowing, because the next fix should target load rather than structure.

If you skip this step, you will pick the measure that moved and call it the goal.

Step 2: Get a baseline, or admit you are reconstructing one

Compare against a measured baseline if you have one. If you do not, reconstruct it from the calendar and be explicit about the bias: a reconstructed baseline understates interruptions and understates fragmentation, because neither gets booked. That bias runs in one direction, which makes it usable - a reconstructed baseline will make your improvement look smaller than it was, never larger.

What is not acceptable is a remembered baseline. "It used to be constant" is not a number and cannot be compared to one.

Step 3: Pair the windows, and pair them on workload

The single most common way this measurement lies is a comparison between a busy old window and a quiet new one.

Use four-week windows, and pair them on position in your season, not on convenience. If the baseline was your spring ramp, the comparison window should be a spring ramp too, even if that means waiting a year. Where waiting is not practical, take the second-best pairing and record the workload difference so you can normalise for it in Step 6.

Four weeks because one week is dominated by whatever happened in it, and eight weeks blurs the change you are trying to detect.

Step 4: Measure four numbers, not one

One number can be gamed by circumstance. Four cannot, and the pattern across them is the actual finding.

Measure Unit of analysis What it tells you
Total working hours Hours per week, averaged over the 4-week window Load
Committed-hours share Hours already spoken for by someone else's work, as a share of total hours, same window Composition
Median longest unbroken block Longest unbooked gap per workday, medianed across the window's workdays Usability
Check-back queue depth Items awaiting a decision from you, counted on the same weekday of each window Whether delegation actually delegated

The fourth is the one people leave out and it is the one that catches a fake win. Work can leave your hands and stay on your plate as a stream of approvals, which reduces your task time and increases your interruption load. Total hours drop, the shop feels worse, and nobody can explain why.

Step 5: Look for substitution before you celebrate

Freed hours refill. That is the default, not the exception, and it happens without a decision being made.

The rule: if total working hours are flat and committed-hours share is flat, the freed hours were substituted, not saved. Something moved into the space. That is not automatically bad - it may be better work - but it is a different outcome from being less busy, and it should be named rather than absorbed.

To see what filled the gap, take the two or three largest categories in the new window and compare them to the same categories in the baseline. The category that grew is where your savings went.

Step 6: Normalise by workload before you read anything

If the shop did more work in the second window, flat hours are a gain, not a null result. If it did less, flat hours are a loss disguised as stability.

Divide your total hours in each window by a workload count you actually track - jobs completed is usually the cleanest. Owner-hours per job is the normalised measure and it is far harder to fool than raw hours.

Step 7: Set the minimum change you will call a result

Weekly variation in owner hours is large. Reading a 4% move as a win is how a shop ends up with a folklore of fixes that never did anything.

The rule: on paired four-week windows, treat a change under about 10% in any measure as noise, and require the direction to be consistent across the four weeks within the window rather than driven by one outlier week. Both conditions, not either. A 15% improvement caused entirely by one week where you were away is not an improvement, it is a holiday.

Step 8: Run a perception check with two people

Two questions to the office lead and the field lead, asked separately: has it got easier or harder to get a decision out of me in the last month, and has anything you used to bring me stopped reaching me at all?

The second question is the important one. A change that made you measurably less busy by making you unreachable is not a win, and it will not show up in any of the four numbers. It will show up as a problem two months later that nobody escalated.

A worked measurement

Five months earlier, an owner had handed permit and approval chasing to his office lead and replaced ad-hoc quoting with a price book covering the twenty most-quoted items. His stated primary definition, written down at the time: fewer hours.

Baseline, a measured four-week window in spring: 57 working hours per week, committed-hours share 84%, median longest unbroken block 35 minutes, check-back queue 11 items.

Comparison window, four weeks five months later, paired as closely as his season allowed: 54 working hours per week, committed-hours share 79%, median longest unbroken block 48 minutes, check-back queue 7 items.

Now apply Step 7 to each.

  • Total hours: 54 against 57, a fall of 3 hours per week, about 5% of the 57-hour baseline. Under the 10% line, so noise.
  • Committed-hours share: 79% against 84%, 5 percentage points, which against a base of 84 is a relative change of about 6%. Also under the line, so noise.
  • Median longest unbroken block: 48 minutes against 35 minutes, a gain of 13 minutes, about 37% of the 35-minute baseline, and it held in three of the four weeks with one weaker week. Real.
  • Check-back queue: 7 items against 11, a fall of 4 items, about 36% of the 11-item baseline. Real.

Step 6, normalise. The baseline window covered 62 jobs completed. The comparison window covered 71, about 15% more. Owner-hours per job: 57 hours a week over 4 weeks is 228 hours for 62 jobs, about 3.7 owner-hours per job. The comparison window is 54 over 4 weeks, 216 hours for 71 jobs, about 3.0 owner-hours per job. That is a fall of about 17% in owner-hours per job, comfortably over the 10% line. Real.

Step 5, substitution. Total hours flat and committed share flat, which is exactly the condition the substitution rule names. The freed capacity went somewhere, and Step 6 says where: the shop ran about 15% more jobs on effectively the same owner hours.

The reading. Against his stated primary definition, fewer hours, the change did not deliver. He is working 5% fewer hours, which is inside normal weekly variation and cannot be claimed. What actually happened is that his capacity improved substantially - 17% fewer owner-hours per job, a 37% longer usable block, a 36% shorter check-back queue - and every bit of that capacity was consumed by growth.

That is a genuinely good result and it is not the result he asked for. Which matters, because the next move depends entirely on which one he was after. If he wants the hours, capacity has to be converted deliberately: cap the intake, or the next capacity gain gets absorbed the same way, silently, by a decision nobody made out loud. If he is content to have grown 15% without adding owner hours, then the change worked and the definition he wrote down five months ago was the wrong one, which he should also say out loud rather than quietly revise.

Step 8, perception. The office lead reported decisions coming back faster. The field lead reported that two things he used to raise directly now went to the office lead and one of them had sat for a week. That is the reachability failure the second question exists to catch, and it was fixed with one line added to the escalation rule rather than by undoing anything.

When the measurement itself cannot be trusted

You changed how you record while you were measuring. Starting to log interruptions mid-window inflates the second half against the first. Any change to the instrument invalidates the comparison; restart the window.

The comparison window contains a week you were away. Remove it and use three weeks against three matched weeks, or the away week will carry the entire result.

You are the one counting, and you know what you are hoping for. This is not a reason to skip it, but it is a reason to write the four numbers down as you go rather than at the end, and to have someone else read the check-back queue count. Queue depth is the easiest of the four to shade, because deciding what counts as "awaiting a decision" is a judgment.

Nothing moved at all, on any of the four. Before concluding the change failed, check that the change actually happened. In roughly half of these cases the handoff quietly returned within a few weeks and nobody said so - the work is back on your plate and you are measuring a fix that was never in place.

References

  • U.S. Small Business Administration (SBA), small business performance measurement guidance
  • Trade-standard practice, owner workload measurement in small field-service shops
  • See related: How to Run a Time Study on Yourself, The Quarterly Owner Time Review SOP, The Check-Back Cadence for Delegated Work, Reading the Warning Signs in Your Own Calendar, How to Hand Off a Recurring Task Permanently