In 1994, Buehler, Griffin and Ross asked students to predict when they would finish their honours theses. The average prediction was 33.9 days. The average actual was 55.5 days. Then they asked for a worst-case estimate, assuming everything went as badly as possible. The average worst case was 48.6 days. Fewer than half finished even by that pessimistic date[1]. The students were not guessing about strangers. They were predicting their own behaviour, on a task they understood well, with every incentive to be accurate.
Key Takeaway
The planning fallacy, identified by Kahneman and Tversky, is the systematic tendency to underestimate the time, cost and risk of a project while overestimating its benefits, even when the estimator has direct experience of similar projects running over. Critically, the bias survives explicit instruction to be pessimistic, and more granular bottom-up planning does not correct it. The evidence-backed countermeasure is reference class forecasting: ignoring the detail of the specific project and instead anchoring the estimate on the actual distribution of outcomes from a class of genuinely similar completed projects. For small businesses without formal databases, a workable version can be built from your own historical job records within a single afternoon.
The Finding
Kahneman and Tversky introduced the term "planning fallacy" in 1979 to describe forecasts that are unrealistically close to a best-case scenario, produced by people who could generate an accurate distribution of outcomes if asked about anyone else's project[2]. The definition matters: it is not general optimism, and it is not incompetence. It is a specific failure mode in which the forecaster reasons forward from the plan rather than backward from history.
The effect has three consistent characteristics. It is directional, errors cluster on the optimistic side rather than distributing symmetrically. It is persistent, experience with prior overruns does not reliably correct it, which is why an owner on their fifth over-budget build can still produce an aggressive estimate for the sixth. And it is resistant to introspective correction: as the thesis study showed, asking for a deliberately pessimistic estimate produces a number that is still, on average, too optimistic.
Why More Detail Doesn't Help
The instinctive response to a history of overruns is to plan more thoroughly: break the work into finer tasks, estimate each one, sum them. This feels rigorous and frequently makes things worse.
Bottom-up estimation captures the work you can foresee. Overruns are driven overwhelmingly by the work you cannot: the permit that takes eleven weeks instead of four, the supplier who fails mid-project, the scope change a client raises in month three, the key employee who leaves. These are individually unpredictable and collectively near-certain. Decomposing the foreseeable work in greater detail adds precision to the part of the estimate that was never the problem, while leaving the unforeseeable part at implicitly zero.
There is a second, subtler mechanism. Decomposition tends to encourage best-case estimation at each node, because each individual task genuinely can be completed quickly if nothing goes wrong. Summing forty best cases produces a total that assumes nothing goes wrong forty times consecutively. The arithmetic is correct and the result is fantasy.
The Evidence Base
The most systematic evidence comes from Bent Flyvbjerg's work on infrastructure. His 2002 study of public works projects found average cost overruns in real terms of approximately 45% for rail, 34% for fixed links such as bridges and tunnels, and 20% for roads, with overruns both prevalent and, importantly, predictable[3]. Later work extended this across 25 project types, reporting mean overruns ranging from roughly 1% to 238% depending on category[4].
The predictability is the actionable finding. If overruns were random, nothing could be done beyond adding generic contingency. Because they cluster by project type with characteristic distributions, the historical distribution for a project class is itself a forecasting instrument, and a better one than the detailed plan for any individual project within that class.
This insight produced reference class forecasting, which has since been endorsed by the American Planning Association and adopted in various forms by governments including the United Kingdom, the Netherlands, Denmark, Switzerland and others[5]. It is not a fringe technique; it is established public-sector practice in several jurisdictions, and almost entirely absent from private small-business planning.
Inside View vs. Outside View
Kahneman and Lovallo formalized the distinction that underlies the method[6]. The inside view builds a forecast from the specifics of the case at hand: this team, this scope, this schedule, these known risks. It is the default mode of essentially all business planning, it feels responsible, and it is systematically optimistic.
The outside view ignores the specifics almost entirely. It asks: what class of project is this, what actually happened across a set of genuinely comparable completed projects, and what does that distribution imply for this one? It treats the current project as an instance of a category rather than as a unique undertaking.
The outside view feels wrong to practitioners, and the reason it feels wrong is precisely why it works: it discards exactly the detailed knowledge that produces overconfidence. A contractor who knows this crew is excellent and this client is organized will forecast optimistically because those facts are true and salient. The outside view says: teams are usually good and clients are usually organized on the projects that still run 30% over.
Uniqueness Bias: "Our Project Is Different"
The standard objection to reference class forecasting is that the historical class does not apply because this project is unlike the others. Flyvbjerg and colleagues named this uniqueness bias, the tendency to disregard divergent information on grounds of distinctiveness regardless of actual similarity[4].
Every project is unique in details and typical in structure. The permit delay is unique in its cause and utterly typical in its existence. The specific supplier who failed is unprecedented; supplier failure is not. Uniqueness bias mistakes the unpredictability of which disruption occurs for evidence that disruptions in general cannot be forecast, when the historical distribution forecasts exactly that: not which thing will go wrong, but roughly how much going-wrong to expect in aggregate.
The practical test is whether the claimed uniqueness is structural or incidental. A genuinely novel construction method, a regulatory regime nobody has operated under, a technology with no deployment history, these are structural, and the reference class should be widened or treated cautiously. A familiar project with a new client, a different site, or an unusual timeline is incidental, and belongs squarely in the class.
The Uncomfortable Alternative Explanation
Intellectual honesty requires noting that optimism bias is not the only explanation for systematic overruns, and possibly not the dominant one. Flyvbjerg's own framing identifies strategic misrepresentation, the deliberate distortion of estimates to secure approval, as a distinct cause operating alongside the psychological one[4]. Where an accurate estimate would get a project rejected, the incentive to produce an inaccurate one is straightforward, and no amount of debiasing addresses a forecast that was never sincere.
A separate line of critique, advanced most forcefully by Love, Ika and colleagues, argues that cost and schedule overruns in large projects are better explained by genuine complexity, pervasive uncertainty and evolving scope than by behavioural bias at all, and that the assumption of bias itself risks being biased[7]. This debate is live in the academic literature and not settled.
For a business owner, the practical implication is the same regardless of which explanation dominates: if your estimates have historically run over, anchoring the next one on your own historical distribution will improve it, whether the original error was psychological, political, or genuinely epistemic. The method is agnostic about the cause. That is part of its appeal.
How To Actually Run Reference Class Forecasting
The method has three steps and is conceptually simple. The difficulty is entirely in the discipline of doing it.
Step one: identify the reference class. Assemble a set of completed projects genuinely comparable to the one being forecast. Comparable means similar in type, scale and complexity, not similar in every detail. The class should be large enough to have a distribution, and narrow enough to be relevant. For most small businesses, the right class is their own last ten to twenty jobs of the same type.
Step two: establish the distribution of outcomes. For each project in the class, record the original estimate and the actual result, then compute the ratio. This produces a distribution of estimate-to-actual ratios, not a single average. The distribution is the point: the mean tells you the typical overrun, the spread tells you your uncertainty, and the tail tells you the downside you should be capitalized against.
Step three: position the current project within that distribution. Produce your inside-view estimate as usual, then adjust it using the historical ratio. If your jobs of this type have historically come in at 1.28 times estimate on average, an estimate of $200,000 implies roughly $256,000 as a central expectation, with the spread of historical ratios defining a credible range around it.
The output should be a range with an explicit central expectation, not a single number. A forecast that cannot express uncertainty cannot communicate risk, and communicating risk is most of what a forecast is for.
The Small-Business Version
Formal reference class forecasting assumes a database that most Canadian small businesses do not have. They do, however, have something functionally equivalent and almost universally ignored: their own completed job records.
The practical build takes an afternoon. Pull the last twenty completed jobs of a given type. For each, record the original quoted or budgeted amount, the original expected duration, the final actual cost, and the final actual duration. Compute cost ratio and schedule ratio for each. Sort them, and look at the median, the range, and the worst three.
Two things typically emerge. First, the median ratio is almost never 1.0, and owners are usually surprised by how far from 1.0 it sits, because individual overruns are remembered as exceptions while the pattern is not aggregated. Second, the variance is usually wider than expected, which is the more important finding for pricing and for capitalization.
From there the application is direct: quote using the inside view, plan cash and capacity using the historical median, and stress-test the business against the worst decile. This is not sophisticated. It is simply using data you already own, which is why the main obstacle is not analytical capability but the absence of a habit of looking backward.
A Worked Case: The Shop Fit-Out
A specialty retailer plans a second-location fit-out. The inside-view estimate, built bottom-up from contractor quotes, is $180,000 over fourteen weeks. The owner has completed three prior fit-outs, at $95,000 against an $80,000 estimate, $140,000 against $110,000, and $210,000 against $175,000.
The historical ratios are 1.19, 1.27 and 1.20, a median of about 1.20 and a tight cluster. Applied to the $180,000 inside-view estimate, the outside view implies roughly $216,000 as a central expectation, with a plausible range of about $214,000 to $229,000 given the historical spread.
Note that the class here is only three projects, which is far too small for statistical confidence and would be dismissed as inadequate in the academic literature. It is nonetheless enormously more informative than the $180,000 figure the owner was about to finance against, because it is anchored on what actually happened rather than on what was planned. Three data points pointing the same direction beat zero data points, and the tightness of the cluster is itself informative.
The decision consequence is concrete: financing arranged at $180,000 leaves the project roughly $36,000 short on a central-case basis. Financing arranged at $230,000 costs slightly more in standby fees and does not run out of money in week eleven. The figures are illustrative and any real case turns on the specific project, but the structural error, capitalizing to the inside-view number, is extremely common and entirely avoidable.
The Forgotten Half: Benefit Overestimation
Almost all discussion of the planning fallacy concentrates on cost and schedule, because those are measured, visible and audited. The definition, however, includes a second and largely unmonitored component: overestimation of benefits. Flyvbjerg's work on transport projects addressed this directly through demand forecasting, finding systematic inaccuracy in projected ridership alongside the cost overruns[3].
In a small-business context, benefit overestimation is the more dangerous error, and it is almost never tracked. A business that budgets $180,000 for a second location and spends $216,000 has a 20% cost problem, which is uncomfortable but survivable. A business that projected the location would generate $400,000 in incremental revenue and achieves $240,000 has a 40% revenue problem, which is frequently fatal, and which compounds with the cost overrun rather than offsetting it.
The asymmetry in attention has a simple explanation. Cost overruns announce themselves through cash: invoices arrive, the account drains, someone notices. Revenue shortfalls are gradual, attributable to a hundred external factors, and can be explained away for several quarters before the pattern becomes undeniable. By the time a revenue projection is clearly wrong, the capital is spent and the escalation dynamics discussed elsewhere in this publication have taken hold.
The remedy is symmetrical treatment: apply reference class forecasting to the revenue side as well. If your last three expansions each achieved roughly 60% of projected first-year revenue, that ratio is at least as important as your cost ratio, and considerably fewer businesses have ever calculated it.
The Premortem
A complementary technique, developed by Gary Klein, is the premortem[8]. Before committing, the team is told to imagine that the project has been completed and has failed badly, then asked individually to write down the reasons why.
The mechanism is a shift in framing. Asking "what could go wrong?" invites defensive reassurance, particularly in front of a decision-maker who has visibly committed. Asserting that failure has already occurred and requesting an explanation licenses candour: participants are no longer criticizing a plan, they are explaining a stipulated outcome. Klein reports that the technique surfaces concerns that participants held but had not voiced.
The premortem complements reference class forecasting rather than substituting for it. Reference class forecasting tells you how much to expect things to go wrong, drawing on history. The premortem helps identify which things, drawing on the team's tacit knowledge. Together they address both the magnitude and the mechanism.
What This Means For How You Price
The forecasting discussion has a direct commercial consequence that is worth drawing out, because it changes the stakes from internal planning hygiene to margin.
A business that quotes from the inside view and delivers at the historical ratio is, by construction, systematically underpricing. If your jobs run 1.20 times estimate on average and you price at estimate plus a 15% margin, your realized margin on the average job is negative before overhead. This is not a hypothetical: it is the arithmetic consequence of quoting one number and incurring another, repeated across every job, and it explains a great deal of the puzzling phenomenon of busy trades and service businesses with poor profitability.
What makes it insidious is that it hides in the aggregate. Individual jobs that run over are experienced as bad luck, difficult clients, or unusual circumstances, each with its own story. Only the ratio distribution reveals that the overrun is the norm and the on-budget job the exception. An owner who has never computed the distribution can work extremely hard for years without ever seeing the structural pricing error, because every individual data point has a plausible individual explanation.
Correcting it does not necessarily mean raising quoted prices, which may not be commercially possible. It may mean tightening scope definitions, changing contract structures to allow variation recovery, declining categories of work where the ratio is worst, or accepting lower margins knowingly rather than discovering them retroactively. All of those are legitimate responses. Continuing to quote against a number the business has never once achieved is not.
Common Objections
"This just builds in failure." The most frequent objection, and it misunderstands the method. Reference class forecasting does not add padding to a plan; it corrects a known measurement bias in the estimate. A scale that reads 10% light is not made accurate by asking it to try harder. Whether the corrected number is acceptable is a separate, and now properly informed, decision.
"Our team will just spend to the budget." A legitimate concern, and Flyvbjerg has addressed it directly: contingency should not be freely accessible, but held and released against genuine need through a defined mechanism[9]. The forecast informs capitalization; it need not be published as the operational target.
"We don't have enough history." Usually false. Most businesses have more completed-job data than they think, sitting in accounting records rather than in any analytical form. Where the history genuinely does not exist, industry benchmarks or peer data are imperfect substitutes that still beat pure inside-view estimation.
"We'd never win a bid at those numbers." This is the most important objection, and it is not really about forecasting. If accurate costing makes your bids uncompetitive, that is information about your cost structure or your market, and it will express itself either as lost bids now or as unprofitable won work later. The forecast is not creating the problem; it is revealing it earlier, when it is still cheap to respond to.
The Limits Of This Analysis
Reference class forecasting is not a solved problem, and overselling it would repeat the overconfidence it exists to correct. Its accuracy depends entirely on the quality and relevance of the reference class, and class selection involves judgment that can itself be biased. It performs poorly for genuinely novel undertakings with no meaningful comparators. Empirical evaluations have found mixed results in some applications, with traditional methods occasionally outperforming it[5].
There is also an unresolved debate, noted above, about how much of the observed overrun pattern is behavioural at all. A method built on the premise of optimism bias may still work while being wrong about why, which is intellectually unsatisfying even where it is practically useful.
What can be said with confidence is narrower and still valuable: estimates anchored on the historical distribution of comparable completed work are, on average, closer to eventual outcomes than estimates built purely from the plan for the specific case. That is a modest claim, well supported, and sufficient to justify the afternoon it takes to build the underlying data.
Frequently Asked Questions
What is the planning fallacy in one sentence?
Why doesn't more detailed planning fix the problem?
How many past projects do I need before this is useful?
Isn't this just adding contingency to every budget?
What is a premortem and how is it different?
Is optimism bias definitely the cause of cost overruns?
References
- Buehler, R., Griffin, D., & Ross, M. (1994). Exploring the "Planning Fallacy": Why People Underestimate Their Task Completion Times. Journal of Personality and Social Psychology, 67(3), 366-381.
- Kahneman, D., & Tversky, A. (1979). Intuitive Prediction: Biases and Corrective Procedures. TIMS Studies in Management Science, 12, 313-327.
- Flyvbjerg, B., Skamris Holm, M. K., & Buhl, S. L. (2002). Underestimating Costs in Public Works Projects: Error or Lie? Journal of the American Planning Association, 68(3), 279-295.
- Flyvbjerg, B., & Gardner, D. (2023). How Big Things Get Done. Currency. See also Flyvbjerg, B. (2014). What You Should Know About Megaprojects and Why: An Overview. Project Management Journal, 45(2), 6-19.
- Flyvbjerg, B. (2008). Curbing Optimism Bias and Strategic Misrepresentation in Planning: Reference Class Forecasting in Practice. European Planning Studies, 16(1), 3-21. For a recent critical review of the method's promises and problems, see Production Planning & Control (2025), tandfonline.com/doi/full/10.1080/09537287.2025.2578708
- Kahneman, D., & Lovallo, D. (1993). Timid Choices and Bold Forecasts: A Cognitive Perspective on Risk Taking. Management Science, 39(1), 17-31.
- Love, P. E. D., Ika, L. A., & Ahiaga-Dagbui, D. D. (2019). On de-bunking "fake news" in the post-truth era: Why does the Planning Fallacy explanation for project overruns fall short? Transportation Research Part A, 126, 397-408.
- Klein, G. (2007). Performing a Project Premortem. Harvard Business Review, 85(9), 18-19.
- Flyvbjerg, B., Hon, C., & Fok, W. H. (2016). Reference Class Forecasting for Hong Kong's Major Roadworks Projects. Proceedings of the Institution of Civil Engineers, 169(6), 17-24.
This article discusses forecasting methods documented in academic research and is provided for informational purposes only. Worked examples use illustrative figures. Project financing, budgeting and capitalization decisions are fact-specific and should be confirmed with qualified advisors familiar with your circumstances.