How to Avoid Optimism Bias in Your Own Estimates

Why this matters

Almost every estimator in the trades runs light, and almost none of them believe they do. The reason is not carelessness. When you walk a job and picture doing it, you simulate the work going the way it goes when it goes right, because that is the only version you can actually picture. The version where the fitting is seized, the customer is on a call, and the second trip happens is not a step you can visualize; it is a distribution of things that did not happen this particular time.

The consequence is a systematic one-directional error, which is the most correctable kind there is. A random miss cannot be fixed. A miss that leans the same way on eight jobs out of ten can be measured and multiplied out in an afternoon. What stops most shops is that they try to fix it with resolve instead of with a number.

This is about correcting your own numbers. Pricing confidence and how to hold a price are a separate skill; see the references.

Step 1: Measure your bias before you try to fix it

Pull your last twelve to twenty closed jobs where actuals exist. For each, record estimated hours, scope-corrected estimated hours (add the hours from approved change orders), and actual hours. Compute the ratio of actual to scope-corrected estimated hours.

Two things this must be. Scope-corrected, because uncorrected ratios include work the customer added and will overstate your bias, causing you to over-correct and lose bids. And your jobs, not the shop's, because bias is personal: two estimators working from the same template routinely run in opposite directions, and a shop-level factor applied to both makes one worse.

Twelve is the working floor. Below about eight jobs the ratio distribution is not stable enough to set a factor from, though it will show a direction.

Step 2: Read the direction from the count, not the average

Before you touch the ratios, count how many landed above 1.0 and how many below. A calibrated estimator should be roughly half and half, because a correct estimate is the middle of a distribution and jobs should overshoot it about as often as they undershoot it.

That count is the cleanest bias signal you will get and it depends on no statistic. Ten of twelve above 1.0 is a decision. Seven of twelve is noise. The count answers "am I biased" and the ratios answer "by how much."

Then use the median ratio to size the correction, not the mean. Ratio distributions are right-skewed by construction: the downside is bounded near 1.0 while the upside is not, so one job that ran nearly double drags a mean noticeably. Correcting to the mean builds your two worst jobs into every future bid.

Step 3: Start from the outside view, then check with the inside view

The inside view is how everyone estimates: walk the steps, add up the durations. It is where optimism lives, because you are adding up the steps you thought of.

The outside view is different: find the closest reference class you have, meaning your own closed jobs most similar to this one, take their median actual hours, and start there. Then adjust for the specific differences on this job.

Do the outside view first and write the number down before the step-by-step. If your bottom-up number lands materially below the median of comparable jobs you actually ran, the burden of proof is on the bottom-up number, and the question is not "did I add correctly" but "what is on this job that was not on those." Usually the answer is nothing, and the gap is the steps you did not think of.

The ordering matters because anchoring runs one way. Build the bottom-up number first and the historical median arrives as an annoyance you argue with. Put the historical median on the page first and the bottom-up number arrives as a claim that has to justify itself.

Step 4: Write the best case and the expected case as separate numbers

On any job worth more than a routine call, write two hour figures: the number if nothing goes wrong, and the number you would actually bet on. Then quote from the second one.

Write both rather than just the second, because without the best-case number on paper the best case is what you are quoting, you just have not admitted it. An estimator whose two numbers are always identical is not producing an expected case at all.

The gap between them is diagnostic. On a familiar job type at a familiar site it should be small, a few percent. A large gap on a job type you run constantly means the uncertainty is not in the work but in your information about that site, and the fix is scoping rather than padding.

Step 5: Do not quote on the day you scoped

Set an overnight minimum for anything above a routine service call, longer for a multi-day job.

The mechanism is specific: right after a walkthrough the job feels understood, because your memory of it is vivid and complete-feeling. The parts you did not notice have not had a chance to surface, and vividness is what your confidence is actually tracking. A day later, half the questions you should have asked will occur to you.

If a customer presses for a number on the spot, give a range with the spread you actually have and say a firm number follows tomorrow. A range you can defend beats a firm number you will regret, and the customer who cannot wait a day for the firm number is telling you something useful about the job.

Step 6: Apply your measured factor, do not add a buffer

"Add a little" is not a correction, because it is applied by feel and feel is the thing that is broken. The buffer lands largest on the jobs that worry you, which are the ones you already scoped carefully, and smallest on the routine ones, which is where your bias actually lives.

Multiply the base hours by your measured median ratio, uniformly, on every job of the types the measurement covered. Especially the ones you feel good about: confidence is not evidence of accuracy, and a factor applied selectively is a buffer with better paperwork.

Keep the factor visible as its own line, not folded into base hours. As a separate line you can re-derive it next quarter and watch it move toward 1.0. Folded in, you will lose track of what it was and eventually correct on top of a correction.

Step 7: Send a second reader to the jobs you are most confident about

A second set of eyes helps most where you are least worried, which is the opposite of how most shops use review. The jobs that get a second look are the big scary ones, which already had your full attention. The routine job you priced in four minutes from memory is where the unexamined assumption sits.

Give the reader one job: name the assumption that is not written down. Not to re-price it, which turns into an argument about hours. That question has an answer more often than an experienced estimator expects.

A worked correction

An estimator pulls 14 closed jobs from the last two quarters and computes scope-corrected ratios of actual to estimated hours. Sorted:

0.92, 0.97, 1.02, 1.05, 1.08, 1.12, 1.15, 1.18, 1.22, 1.25, 1.31, 1.38, 1.52, 1.74

The count first. Two of 14 came in under 1.0, so 12 of 14, about 86 percent, ran over. A calibrated estimator would sit near 7 of 14. That asymmetry alone establishes bias before any statistic is computed.

Median versus mean. With 14 values the median is the average of the 7th and 8th, which is 1.15 and 1.18, giving 1.165, call it 1.17. The mean of the same 14 is about 1.21. The 0.04 difference comes almost entirely from the two jobs at 1.52 and 1.74, both of which were single-cause blowouts with named reasons on the job record. Correcting to 1.21 would price those two events into all 14 future jobs; correcting to 1.17 does not.

The spread, read separately. The middle half of the ratios runs from about 1.05 to 1.31, a band 0.26 wide. That width is precision, and a bias factor cannot change it. Knowing that in advance is what keeps the estimator from concluding the correction failed when jobs keep landing off the number.

Applying it. A job this estimator would have called 20 hours becomes 20 times 1.17, which is 23.4, rounded to 23.5 on the estimate and carried as 20 base hours plus a 3.5 hour calibration line.

What the estimator expected, and got. After the correction the same middle-half band lands at roughly 0.90 to 1.12 in ratio terms, straddling 1.0 instead of sitting entirely above it. Jobs still miss by similar absolute amounts, and now they miss in both directions, which is the definition of the correction working.

The check, one quarter later. Fourteen more closed jobs, ratios computed against the corrected estimates. Six came in under 1.0 and eight over, with a median of 1.03. Six of 14 is inside the range of ordinary variation around an even split, and a 1.03 median on 14 jobs is not a bias worth chasing further. The estimator leaves the factor at 1.17 and re-checks in two quarters.

What would have signaled over-correction. Eleven or twelve of 14 landing under 1.0, or a median down near 0.90. That result means jobs are now routinely finishing early, which sounds harmless and is not: it means bids are going out high and the ones that are being lost are being lost invisibly, with no record anywhere.

The pricing decision is separate. Correcting the model raised this estimator's hours by 17 percent on every job. Whether the price rises by the same amount is a different call, made deliberately and visibly. If the shop chooses to hold the price on a competitive job type, it does that by accepting a thinner margin on a number it knows is honest, never by walking the hours back down.

What changes the answer

A new estimator. Do not set a factor from their first handful of jobs. Watch the count, coach the process, wait for eight to twelve closes. A factor set on four jobs is a coin flip you then defend for a year.

Bias that varies by job type. Fourteen jobs support one factor. With sixty across types you may find one runs at 1.05 and another at 1.35, and a blended factor makes the first uncompetitive while leaving the second exposed. Split only when each type has its own sample of twelve or more.

A factor above roughly 1.3. That is an information problem, not a calibration problem. Multiplying by 1.35 is a confession that you cannot see the job at bid time. The correct responses are better scoping, a conditions allowance, or a different contract shape, not a bigger multiplier.

The direction reversing. An estimator whose ratios sit consistently below 1.0 is not a hero, they are losing bids at the same rate a light estimator loses margin, and the losses are invisible because nobody writes down the jobs they did not get. Track win rate alongside ratio and read them as a pair. Note also that a crew, method, or tooling change alters actual hours without your estimating changing at all, so check for one before attributing any shift in your ratios to yourself.

How to verify the correction worked

  • Read the count, not the median, as the primary test. A near-even split above and below 1.0 is the target. A median near 1.0 achieved by half the jobs landing far above and half far below is a precision problem wearing a calibration success as a disguise.
  • Confirm the spread did not narrow. It should not, and if it did, something else changed in the same period. Attributing a narrowing spread to a bias factor is a claim the factor cannot support, and it will lead you to expect precision the method does not deliver.
  • Watch win rate as the counter-signal. A meaningful drop right after correction is expected and should stabilize. A continuing slide means either the correction is too large or your prices were carrying an unacknowledged discount that the honest number just exposed. Those need different responses, so read them separately.
  • Re-derive rather than adjust. Each quarter, recompute the factor from scratch on the newest window of jobs rather than nudging the existing one. Nudging reintroduces the feel-based judgment the factor exists to replace.

References

  • U.S. Small Business Administration (SBA), estimating and pricing guidance for small business
  • Trade-standard practice for estimate calibration and reference-class comparison in project cost estimating
  • See related: Estimating Confidence From Tracking Actuals, Padding the Estimate: The Temptation, How to Build an Estimate From a Cost Model, How to Price From History Instead of Instinct