How to Tell Whether a Tool Actually Saved Time
Why this matters
Every tool a shop buys was justified by a time saving that nobody ever checked. The tech who wanted it says it is faster. The owner who bought it wants it to be faster. Six months later the tool is either a permanent fixture or gathering dust, and in neither case does anyone know whether it saved a single hour. So the next purchase gets decided the same way, on enthusiasm, and the shop never learns which kinds of tools pay off for the kind of work it actually does.
Measuring this is not hard. It is just easy to do badly, and doing it badly is worse than not doing it, because a wrong number gets quoted for years.
Step 1: Pick one job type and freeze the definition
Measure the tool against one job type, not against the shop's average day. "Average labor hours per job" will never move enough to see, because the tool touches a fraction of your work and everything else swamps it.
Write down what counts as that job type in one sentence, before you look at any data. If the definition drifts mid-measurement, the comparison is dead and you will not notice. The most common drift is scope: jobs of the type get quietly larger or smaller over the year, and the tool takes the credit or the blame.
Step 2: Get a baseline from history, not from memory
Pull labor hours on at least 8 completed jobs of that type from before the tool arrived. Fewer than 8 and one ugly job dominates.
Use the median, not the mean. Field labor hours are skewed: most jobs cluster and a few run long for reasons that have nothing to do with tools (access, a surprise, a customer conversation). A mean is a report on your worst jobs. A median is a report on your typical one.
If you do not have clean labor hours per job in your records, that is the real finding, and it is worth more than this measurement. Fix the time capture first. See related: How to Capture Actual Costs Without Slowing the Crew.
Step 3: Throw out the learning curve deliberately
The first few uses of any new tool are slower than the old method. That is not a verdict, it is a setup cost. A shop that measures uses 1 through 4 and concludes the tool is a dud has killed a good purchase.
The rule: exclude the first 3 uses per tech from the after-period, then require at least 8 jobs in the remaining sample before you read a result. Per tech matters here. A tool used competently by one lead and for the first time by three others is producing two different populations, and pooling them gives you a number that describes nobody.
The mirror error also exists. The tech who campaigned for the tool works harder on the first jobs with it, so the earliest post-curve numbers can flatter it. Watching the trend across the after-period, not just its median, is how you catch that: a genuine saving is flat, a honeymoon decays.
Step 4: Control for the three things that will fake a result
Crew mix. If the after-jobs were run by your strongest tech and the before-jobs by a mix, you have measured the tech, not the tool. Either match the crew across both periods or split the comparison by tech.
Scope drift. Compare something about the job size across both periods - fixture count, linear feet, unit count, whatever your trade uses. If the after-jobs are systematically smaller, the hours dropped for a reason that has nothing to do with the tool.
Seasonality and schedule pressure. Jobs run in a flat-out week take longer than jobs run in a slow one, because of interruptions and travel compression. If your before-period is a busy season and your after-period is not, the tool gets free credit.
You do not need statistics for any of this. You need to look at the three columns and ask whether the two groups are comparable. If they are obviously not, extend the window rather than reporting a number you know is dirty.
Step 5: Convert the saving honestly
This is where most tool evaluations quietly lie to themselves.
A saved field hour is not a billable hour and it is not money. It becomes money only if one of two things happens: the freed capacity gets sold to another job, or the job type is priced flat-rate and the hour comes straight out of your cost. If the schedule does not refill, the tool bought your techs an earlier finish, which is a real benefit for retention and overtime but not the benefit you costed.
So state the conversion explicitly: "the tool saves X hours a month of field time; on the flat-rate portion of this job type that lands directly as margin, and on the time-and-materials portion it lands only if we sell the freed hours." Two different currencies. Do not add them and call the total a return.
Worked example: a powered method replacing a manual one
A shop buys a powered tool that replaces a manual method on one recurring job type. The owner wants to know at 12 months whether it earned out.
Baseline. 11 completed jobs of the type in the 5 months before purchase. Median labor: 3.4 hours. Range runs from 2.7 to 6.1, which is why the median is the right statistic - that 6.1 job had an access problem and nothing to do with method.
After-period, raw. 17 jobs in the 9 months since. Two techs have used the tool: the lead, who has run it 12 times, and a second tech, who has run it 5 times. Under the step 3 rule the first 3 uses per tech come out, so 3 of the lead's and 3 of the second tech's, which is 6 jobs. That leaves 11. Eleven clears the 8-job minimum, so the sample stands.
After-period, clean. Median labor across the 11 remaining jobs: 2.6 hours.
The saving. 3.4 minus 2.6 is 0.8 hour per job. Against a 3.4-hour baseline that is a 23.5% reduction in labor hours on this job type. Not on all jobs, on this job type.
The excluded jobs, checked. The 6 learning-curve jobs had a median of 3.8 hours, which is 0.4 hour worse than the 3.4-hour manual baseline. That is the setup cost the shop paid, about 2.4 hours total across 6 jobs, and it is the number that would have condemned the tool if anyone had read it in month two.
The controls. Job size proxy across both groups is close: median unit count 6 before, 6 after. Crew mix is mixed in both. The before-period spans a busier season than the after-period, which the owner notes as a reason to treat 0.8 hour as the optimistic edge of the estimate rather than a precise figure.
The volume. The job type runs about 2 times a month. So 2 jobs at 0.8 hour is 1.6 hours a month of field time saved.
The earn-back. The shop's hurdle was that the tool return its cost, expressed as the equivalent number of shop labor hours, inside 12 months. That converted to about 14 hours. At 1.6 hours a month, 14 hours is reached at about 8.75 months, so just under 9 months, inside the hurdle. Add back the 2.4 hours of learning-curve cost and the true break-even is 16.4 hours of saving, reached at about 10.25 months. Still inside 12, so the purchase stands on either reading, and the honest version to record is a little over 10 months.
The conversion. This job type is quoted flat-rate, so the 1.6 hours a month comes out of cost directly rather than needing to be resold. If it had been time-and-materials, the correct statement would have been "1.6 hours a month of freed capacity, worth nothing unless we book it," and the owner's next question would have been whether the schedule was full enough to absorb it.
What would have made this a false positive. If the 11 after-jobs had a median unit count of 4 against the before-period's 6, the jobs got smaller and the 0.8 hour is mostly scope, not tool. If all 11 had been run by the tech who wanted the tool, the comparison measures him. And if the owner had counted the raw 17-job after-period instead of the clean 11, the learning-curve jobs would have dragged the after-median toward the baseline and understated a real saving, which is the opposite error and just as wrong.
What to do when the answer is no
A tool that did not save time is not automatically a mistake, and selling it should not be the reflex. Work through three questions before you act:
Did everyone actually use it? Check the attach rate: on how many jobs of the type did the tool go out at all? A tool used on 4 of 12 jobs has not been tested, it has been ignored, and the finding is about training or findability.
Did it save something other than time? Some tools buy accuracy, callback reduction, reach into a space, or a job you could otherwise not quote. Measure the thing it actually changed. A tool that cut your callback rate on a job type while leaving labor hours flat is a good buy measured badly.
Was the old method already good? Sometimes the honest answer is that the manual method was fine and the shop bought a solution to a non-problem. Record that plainly against the tool, with the numbers. That record is worth more than the tool, because it is what stops the same purchase in two years.
How to verify your own measurement
Before you circulate a result, do three checks on your own work. Re-read the percentage you stated and confirm the base is named in the same sentence and in the same unit - a saving stated per job cannot be compared against a monthly total without saying so. Re-do the subtraction by hand. And look at the two job lists side by side and ask whether you would have believed the comparison if someone else had handed it to you with the tool's advocate having chosen the sample.
Then put the result somewhere the next purchase decision will find it, filed against the tool type rather than the specific tool. The point of doing this once is to be better at deciding the next one.
References
- Trade-standard practice for job-level labor hour capture and post-implementation review
- Manufacturer documentation for expected cycle times, used as a claim to test rather than a result to assume
- See related: How to Decide What a Specialty Tool Has to Earn, Adopting a New Tool Without Disrupting the Week, The Tool Utilization Question, How to Compare Estimated Against Actual on Every Job