How to Evaluate a Tool You Have Never Used
Why this matters
A tool demo is a sales environment. The rep runs it, on a prepared sample, in good light, with a fresh battery, and it works beautifully. Then it arrives at your shop, gets handed to whoever is in the bay, and joins the pile of things that were going to change everything. The shop did not buy a bad tool. It bought a tool it had never actually tested against the way it works.
Deciding whether the volume justifies owning a tool is a separate question with its own arithmetic. This one comes first: does the thing do what it claims, in your hands, on your jobs, and can a normal tech get value out of it without the person who championed it standing there.
Step 1: Write the standard down before you touch the tool
One paragraph, before any demo, answering three things:
- What it replaces. The current method, named specifically. "We locate it by opening panels" is a method. "We do it the old way" is not.
- The measurable it has to beat. Time per task, first-visit resolution rate, callback rate, or number of people needed. Pick one primary measure. If you cannot name one, you are shopping, not evaluating.
- The baseline number for that measurable, taken from records, not memory. Pull 15 to 20 recent instances of the task and get the actual average.
Writing this after the demo is worthless, because the demo will have reset your sense of what normal is. This is the step everyone skips and the reason most tool trials end in a shrug.
Step 2: Get one in your hands without buying it
In descending order of quality:
- Rent it for a defined block. Best option, because rental terms force an end date and a decision.
- Supplier loaner or demo unit for a fixed period. Common and usually free. Insist on a return date in writing so the trial does not decay into ownership by default.
- Borrow from another shop. Fine for a look, weak for a trial, because you will be careful with it in ways that hide how it holds up.
- A single-day rep visit. Not a trial. Treat it as product research only.
Never let a demo unit stay past its return date "so we can keep testing." A trial with no end has no decision at the end of it.
Step 3: Run it on real jobs, in the hands of the wrong person
Three rules that decide whether the trial tells you anything:
Real jobs, not a shop bench. Bench trials measure the tool. Field trials measure the tool plus the access, the weather, the customer standing behind you and the battery that started the day at half charge.
The skeptic runs it, not the champion. The person who wanted the tool will make it work. That tells you the ceiling. You need the floor: hand it to a competent tech who is indifferent, give them the same instruction a new hire would get, and see what they get out of it.
Log every use, including the ones where it did not help. A one-line entry per use: job, task, time with the tool, whether it changed the call, and what got in the way. Trials fail on record-keeping far more often than on the tool.
Step 4: If it produces a number, verify it against a known source first
Any instrument that outputs a reading you will act on must be verified against something whose value you already know, before its first reading in the trial counts. Compare against a known-good instrument on the same point, a reference source, or a condition you have independently confirmed. An unverified instrument in a trial does not produce a fair test, it produces a random number generator with a display.
Do not let a new diagnostic instrument enter a safety verification path during a trial. A thermal camera, a non-contact indicator, a listening device or any other new instrument does not substitute for the established live-dead-live check performed with a properly rated meter that was itself verified on a known live source before and after the test. Test instruments and their leads get a visual inspection for external defects and damage before use, and any instrument showing damage that could expose someone to injury is removed from service on the spot (29 CFR 1910.334(c)(2)). A trial tool is exactly the instrument nobody has inspected yet.
Step 5: Score the four things that actually kill adoption
Time saved is the number everyone measures and the weakest predictor of whether the tool survives the year. Score all four.
| Dimension | What you measure | The kill signal |
|---|---|---|
| Does it do the thing | Trial reads that matched the confirmed cause, as a fraction of trial uses | Any pattern of confident wrong answers; a tool that is sometimes wrong and never uncertain is worse than no tool |
| Time delta | Average task time with the tool against the recorded baseline | Saving that disappears when the champion is not running it |
| Learnability | What the skeptic accomplishes after one hour of instruction | Needing the manual on every use after two weeks |
| Upkeep | Charging, consumables, calibration, storage, what it needs weekly | Anything requiring a habit nobody currently has |
The first row is the one to weight heaviest for anything diagnostic. Score it against the confirmed cause after the repair, not against whether the reading looked plausible at the time. A tool that agrees with your existing hunch is not adding information.
Worked example: a compact thermal imaging camera on trial
A six-tech shop trials a compact thermal camera for locating problems behind finished surfaces. Rented for two weeks.
Baseline (Step 1): they pull 18 recent calls of this task type. Average locate time, from arrival to confirmed location, is 0.6 hours. Primary measure: locate time. Secondary: whether the located point matched what the repair actually found.
Trial: 20 uses over two weeks, run by two techs, one of whom argued against buying it.
Results:
- Average locate time across the 20 trial uses: 0.35 hours. That is 0.25 hours saved per use against the 0.6-hour baseline, a reduction of about 42% of baseline locate time.
- Accuracy: on 17 of the 20 uses the located point matched what the repair found. On 3 of the 20, about 15% of trial uses, the camera pointed at a warm area that turned out to be incidental, and the tech opened the wrong spot first.
- Skeptic's result after one hour of instruction: usable readings on his second job. He stayed skeptical about the 3 misses and was right to.
- Upkeep: charges on a proprietary cable, needs a clean lens, and gives poor results in direct sun on an exterior wall.
Volume: the shop runs about 14 of these calls a month. At 0.25 hours saved per call, that is 3.5 hours a month, or 42 hours a year, of locate time removed.
The interpretation that matters. The 42-hour figure is technician labor hours freed, not billable hours won and not margin. Some of it becomes capacity, some of it becomes a shorter day, and none of it is money until the shop actually sells the freed time. The honest statement to the crew is "this removes about 3.5 hours a month of hunting," and then the separate question of what you do with that time.
The 15% miss rate is the finding, not a footnote. Three wrong openings out of 20 trial uses means the tool needs a rule attached: the camera narrows the search, and a second confirming indication is required before anyone cuts. With that rule written into the procedure, the shop bought it. Without the rule, they would have bought a tool that turns a careful tech into a fast one who is wrong 3 times in 20.
What they did with the sun problem: noted it as a known limitation on the tool's card in the crib, rather than discovering it independently four times.
The three ways a trial lies to you
Not restating the steps, these are specific to the trial itself and they are why honest shops still buy the wrong tool.
Selection of the easy cases. People reach for a new tool on the jobs where it will obviously help and revert on the hard ones. Check your log for whether the tool got used on the messy jobs. If every trial entry is a clean case, your time saving is measured on the wrong population.
Novelty speed. New tools get run with full attention. The comparison baseline was measured on a routine task the tech had done a hundred times, half asleep. Where the margin is thin, re-measure a few baseline instances during the trial week rather than trusting a number from three months ago.
The one great save. A tool that solved one nightmare job earns permanent goodwill regardless of what it did on the other 19 uses. Log the great save separately and evaluate it as a rare-event insurance question, which is a legitimate reason to own something, rather than letting it contaminate the average.
When no trial is available at all
Some tools cannot be trialled. Nobody rents them, there is one supplier, no loaner exists, and the only way to find out is to buy. Do not treat that as permission to skip evaluation. Substitute three things, in this order.
Reference-check two shops that already own it, and ask about failure rather than features. The questions that produce useful answers are narrow: what does it need that you did not expect, what has broken on it, how long was it out when it did, and who in your crew refuses to use it. "Do you like it" produces nothing, because everyone likes the tool they bought.
Ask the supplier what the failure and return pattern looks like. A supplier who cannot tell you what commonly fails on a tool they have sold for years either does not sell many, or is not the supplier you want on the phone the week it stops working. Their answer is data either way.
Stage the purchase. Buy one, not one per crew. Run it under the same logging discipline you would have used in a trial, for the same number of uses, with the skeptic operating it. Then make the fleet decision on your own data. A shop that buys six of something on a demo has removed its own ability to change its mind.
Set the review date at purchase, at roughly 20 uses or one quarter, whichever comes first. Without a date the evaluation never closes, and the tool is judged forever by whoever last used it.
The decision memo
Close the trial the day the rental ends, with five lines on one page: the baseline, the trial result, the accuracy fraction, the rule the tool needs attached, and buy or decline. Then hand it to whoever runs purchasing.
If the answer is buy, the volume arithmetic runs next, and it is a separate decision that a good trial result does not settle. If the answer is decline, keep the memo. The same tool will be pitched to you again in eighteen months, and knowing exactly why you passed is worth more than the trial cost.
References
- Occupational Safety and Health Administration, 29 CFR 1910.334(c)(2), test instruments and equipment, inspection before use
- Manufacturer documentation for stated accuracy, operating conditions and limitations of measuring instruments
- See related: How to Decide What a Specialty Tool Has to Earn; Adopting a New Tool Without Disrupting the Week; Verify the Tool Before You Trust the Reading