Your business makes forecasts constantly. Whether that client will renew, whether the contract closes this quarter, whether the hire works out. Almost none of them are written down with a probability attached, which means almost none of them can ever be scored.

Key Takeaway

Tetlock recruited 284 people who made their living commenting on political and economic trends and collected their probabilistic forecasts over roughly two decades[1]. Aggregate accuracy was marginally better than chance on straightforward questions and often inferior to simple statistical baselines on complex ones[2]. Tetlock's own summary of the cognitive-style result is worth quoting exactly: "foxes were modestly better forecasters; hedgehogs were modestly worse; and confident television hedgehogs were the worst of all"[3]. Popular retellings describe foxes as having crushed hedgehogs[4].

Our Grades For These Claims

Applying the scheme from the first article in this series.

That expert forecasting accuracy is poor is Grade A. It rests on a two-decade longitudinal study with tens of thousands of probabilistic forecasts scored against verifiable outcomes, which is about as good as evidence about judgment gets.

That confidence is inversely related to accuracy is Grade B, reported consistently across our sources but characterised rather than quantified in anything we obtained.

The fox and hedgehog distinction is Grade B, with the important qualification that Tetlock's own description of the effect size is modest.

The popular "foxes crushed hedgehogs" version is Grade D, and the gap between it and the author's own wording is the most instructive thing in this article.

A Note On Method

Everything here is verified to August 2026.

We did not obtain Tetlock's book. We rely on an academic review hosted by a university faculty page[1], on encyclopedic and summary sources[2], and on secondary accounts, each identified.

Our sources conflict badly on the number of forecasts, giving figures that differ by roughly a factor of three, and on the date range. We set this out in its own section rather than picking one.

Several of our sources are commercial book-summary websites and blogs. We use them only where they quote, and we flag every such use.

We did not obtain the Good Judgment Project papers and report that programme only in outline.

All arithmetic on calibration and sample size is ours, marked as such.

This article reviews research on judgment. It is not investment, forecasting or business advice.

The Study

The design, which is the reason the findings carry weight.

Tetlock recruited 284 people whose professions included commenting or offering advice on political and economic trends, and asked them to forecast the probability that various situations would or would not occur, picking areas, geographic and substantive, within and outside their areas of expertise[1].

The participants were drawn from universities, governments, think tanks, foundations, international institutions, and the media[5], and were selected based on their prominence and frequent public engagement in forecasting political and economic outcomes[2].

Beyond the forecasts themselves, he asked questions aimed at understanding how the forecasters came to formulate their forecasts, how they dealt with the failure of their forecasts, how they responded to contradictory information, and how they evaluated the probable accuracy of others' theories and predictions[1].

Three features make this unusually strong, and this assessment is ours.

The forecasts were probabilistic and about verifiable events, which makes them scoreable. Most professional prediction is neither.

They were recorded in advance, which removes the hindsight problem entirely.

And experts forecast outside their own specialisms as well as inside them, which allows the effect of expertise itself to be isolated.

A Conflict In The Headline Number

Something our sources disagree about, and it is not a small disagreement.

An academic review states that by 2003, he had accumulated 82,361 forecasts[1].

Several popular sources give roughly 28,000 predictions[4][6], and an encyclopedic source gives approximately 28,000[2].

A book citing the study gives 27,451 predictions on world politics between 1988 and 2003[5].

The date range is also given variously as 1984 to 2003[4], 1988 to 2003[5], and simply as two decades[1].

Two observations, ours.

82,361 against 28,000 is a factor of roughly three, not a rounding difference. A plausible reconciliation is that the larger figure counts individual probability judgments, since each question may require estimates across several possible outcomes, while the smaller counts distinct predictions. That is our conjecture and no source states it.

And this is the fifth article in this series to find a famous study's central figures reported inconsistently by careful secondary sources. Anyone quoting a number here should be quoting the book.

What Was Actually Measured

The scoring, which is more sophisticated than the popular version suggests.

A source describes the methodology as soliciting probabilistic forecasts on verifiable events, scoring them for calibration, meaning the alignment of predicted probabilities with actual outcomes, and discrimination, meaning distinguishing likely from unlikely events[2].

Two observations, ours, and this distinction matters for anyone assessing their own judgment.

Calibration and discrimination are different virtues. A forecaster who says 60 percent to everything, when 60 percent of things happen, is perfectly calibrated and useless, because they never distinguish one case from another.

A forecaster who is highly discriminating but poorly calibrated, saying 95 percent when they mean 70, is informative but systematically overstated, and their numbers can be corrected downward once you know the pattern.

The second is the more common professional failure, and the more fixable.

The Result

The headline, stated carefully.

An encyclopedic source reports that aggregate accuracy was marginally better than chance on straightforward questions but often inferior to simple statistical baselines on complex ones, with many forecasters exhibiting overconfidence and poor calibration[2].

Another source puts it as forecasting judgment was poor, only barely beating random chance and running below that of simple extrapolation algorithms[7].

The dart-throwing chimpanzee comparison, which is how most people encounter this study, appears in secondary accounts[4][6].

Three observations, ours.

The comparison that should worry a business is not the chimpanzee but the algorithm. Losing to chance is embarrassing; losing to simple extrapolation means a spreadsheet that assumes next year resembles this year would have done better than the experts.

Note the conditional structure: marginally better than chance on straightforward questions, worse than baselines on complex ones. Expertise held up where it was least needed.

And we did not obtain the effect sizes, the Brier scores, or the specific comparisons, and state none.

The Finding About Confidence

The result with the most direct application to a boardroom.

A source states that the experts who were most confident in their predictions tended to be the least accurate[4]. Another that the most televised, most confident voices were typically the worst calibrated[3].

Three consequences, ours.

The signal a listener naturally uses to weigh a forecast, how sure the forecaster sounds, was on this evidence pointing the wrong way.

And the second formulation adds something: public prominence tracked inaccuracy too. The people most in demand to forecast were not the better forecasters.

This connects directly to the fifth article in this series, where experts insisted a number had not affected them while being demonstrably affected. In both cases, the practitioner's own confidence carries no information about their accuracy, and in this one it may carry negative information.

Hedgehogs And Foxes

The distinction the book is best known for, borrowed from Isaiah Berlin.

Hedgehogs know one big thing. They are described as strong ideologically driven thinkers who favor parsimony[7], who view the world through a single ideological lens[2], and who have a theory, a framework, a lens through which they see everything[3].

Foxes know many things. They integrate diverse perspectives and update beliefs flexibly[2], are eclectic, drawing on multiple frameworks without committing fully to any[3], and are comfortable with uncertainty[8].

Two observations, ours.

The distinction is about cognitive style, not intelligence or knowledge. Both groups were accomplished professionals.

And a source notes that ideology had little correlation to good forecasting judgment[7]. It was not that one political position forecast better. It was how tightly the person held whatever position they had.

The Word Tetlock Used

The most important sentence in this article, and it is a quotation.

A source reproduces Tetlock's own summary of the result: "In our data, foxes were modestly better forecasters; hedgehogs were modestly worse; and confident television hedgehogs were the worst of all."[3]

Three things about that sentence, ours.

Modestly. Twice. The author of the study describing his own headline finding chose a hedged adverb, and used it for both directions.

The structure is three tiers, not two. Foxes modestly better, hedgehogs modestly worse, and a third group, the confident and televised, distinctly worse than either. The dramatic part of the finding attaches to that third group rather than to the hedgehog category as a whole.

And "in our data" is a further hedge, scoping the claim to the study rather than to the world.

An Overconfident Retelling

What happened to that sentence in transmission. This section is our own analysis.

Set the author's wording against how the finding travels.

Tetlock: foxes were modestly better[3].

A popular account: "Foxes crushed hedgehogs."[4]

Another: foxes outperformed hedgehogs in prediction, often substantially[8].

A fourth, more carefully: foxes did slightly better than chance, although not dramatically[3].

Three observations.

The gap between modestly better and crushed is the whole distance this series has been documenting between a finding and its reputation.

It is especially pointed here because the subject of the book is overconfident expert claims. A study about people overstating their certainty is being transmitted by overstating its certainty.

And it is a live demonstration of the reading rule from the ninth article in this series. When you encounter a striking finding, look for the author's own words, because the popular version has usually had the hedges removed. Here the hedge was a single adverb, used twice, and dropping it changes the claim entirely.

Foxes Have Their Own Failure

A finding that complicates the simple lesson, and it is rarely reported.

An academic review notes that Tetlock reports that he was able to push more foxes than hedgehogs into forecasts that violated a fundamental axiom of probability: that the sum of a forecaster's forecast probabilities not exceed one[1].

Two observations, ours.

That is a coherence failure, and it is not trivial. If you assign probabilities to a set of mutually exclusive outcomes that sum to more than one, you are not merely wrong about the world; your beliefs are internally inconsistent before any evidence arrives.

And foxes were more susceptible to it than hedgehogs. The very openness to multiple scenarios that made them better forecasters made them worse at keeping the scenarios in a coherent probabilistic relationship.

Which means the practical lesson is not simply be a fox. It is closer to: consider multiple scenarios, and then check that your numbers add up, which is a discipline the fox style does not supply on its own.

Domain Expertise Barely Mattered

An uncomfortable result for anyone who buys expert advice.

A source notes that the specific domain of expertise had so little bearing on forecasting ability that this required explanation, while adding that this does not imply that forecasting is foolhardy, because there were cognitive traits that did show consistent ability[7].

Three consequences, ours.

Recall the design: experts forecast inside and outside their own areas[1]. So this is a within-person comparison, which is a strong test.

If domain expertise contributes little to forecasting accuracy, then hiring a specialist to predict is buying something different from hiring one to explain. Specialists know a great deal about how a system works. That is a separate capability from estimating what it will do next.

And the constructive half matters: cognitive traits did show consistent ability. Something predicted accuracy. It just was not subject-matter knowledge.

How Experts Absorbed Being Wrong

The part of the study concerned with what happened after the forecasts failed.

A source describes Tetlock exploring how experts protect their self-image through hindsight bias, being "I knew it all along", creeping determinism, and belief system defenses, the rhetorical moves that let a confident but wrong prediction be re-described, and our source truncates there[8].

Another notes that hedgehogs retroactively adjusted narratives rather than probabilities[2], and that the study revealed systemic shortcomings like belief perseverance and insufficient updating in response to new evidence[2].

Two observations, ours.

The distinction between adjusting the narrative and adjusting the probability is the sharpest practical idea here. A forecaster who was wrong and explains why the situation was exceptional has updated their story. A forecaster who was wrong and lowers their confidence next time has updated their model. Only the second improves anything.

And this is why an unrecorded forecast is worse than no forecast. If nothing was written down, the narrative adjustment is the only option available, because there is no probability left to compare against.

What Calibration Means

The concept, explained plainly, because it is the usable output of this whole literature. This explanation and the illustration are ours.

A forecaster is calibrated if, across everything they called seventy percent likely, about seventy percent happened.

That is a property of a track record, not of a single forecast. No individual prediction can be calibrated or miscalibrated; only a set can.

A perfectly calibrated record looks like this: things called 90 percent happen about 90 percent of the time, things called 70 percent happen about 70 percent, things called 50 percent happen about half the time.

An overconfident record looks like this: things called 90 percent happen about 70 percent of the time, things called 70 percent happen about 55 percent, and things called 50 percent happen about half the time.

Two consequences.

Calibration is measurable without any special expertise. It needs a list of predictions with numbers attached and a later check.

And it is correctable. A person who discovers their nineties are really seventies can adjust, which is a rarer property than it sounds among cognitive failings.

Overconfidence Lives At The Extremes

A structural point about where miscalibration shows up. Ours.

Look again at the two records above. At 50 percent, they are identical. The overconfident forecaster is indistinguishable from the calibrated one in the middle of the range.

Three consequences.

Miscalibration is only detectable at confident forecasts. The further from fifty percent, the more information each outcome carries.

Which means a forecaster who never commits, who keeps everything between forty and sixty, can never be caught being overconfident. They have also never said anything.

And that produces a professional incentive worth naming: vagueness is safe and useless. The adviser who says a deal is "quite likely" has taken no risk and given you nothing to score.

The Tournament That Followed

The successor programme, reported in outline because we did not obtain its papers.

A source describes the Good Judgment Project as an IARPA-funded forecasting tournament that ran from 2011 to 2015 under the Aggregative Contingent Estimation programme, in which several university teams competed against each other, against statistical aggregates, and against the United States intelligence community on hundreds of geopolitical questions with verifiable outcomes[3].

The associated book is Tetlock and Gardner, Superforecasting: The Art and Science of Prediction (2015), and an associated paper is Mellers and colleagues (2015), The psychology of intelligence analysis: drivers of prediction accuracy in world politics, in the Journal of Experimental Psychology: Applied[3].

We obtained none of these and report no results from the tournament.

What the design tells you, and this is ours: the competitors included the intelligence community, on verifiable questions, scored. Whatever it found, it was set up so that somebody could lose, which is more than most claims about forecasting ability permit.

The Vague Word Problem

One reported practice, and it is directly transferable.

A source describes practices distinguishing the best forecasters, noting that ordinary forecasters speak in vague terms, being "likely", "possible", "could happen", and these words mean different things to different people[3].

Two observations, ours.

This is free to fix. Replacing "likely" with a number costs nothing and requires no training.

And it changes what a disagreement is about. Two people who both say a deal is "likely" may believe 55 percent and 90 percent and will never discover it. Two people who say 55 and 90 have found a disagreement worth an hour.

We flag that this practice reaches us through a commercial summary of a book we did not obtain, and that we report no evidence of how much it improves accuracy.

How Many Forecasts Before You Know

The practical constraint, computed by us using standard binomial arithmetic.

To test whether the things you call 90 percent likely actually happen about 90 percent of the time, within roughly ten percentage points at conventional confidence, you need about 35 forecasts at that confidence level.

To narrow that to five percentage points, about 139.

At the 70 percent level, about 81 forecasts for a ten-point margin and about 323 for five.

Two observations, ours.

A business making one significant forecast a month reaches 35 in about three years. This is a slow instrument, which is precisely why it has to start now rather than when you need it.

And it must be written down at the time, because a forecast reconstructed after the outcome is known is not a forecast. That is the hindsight problem the study design was built to eliminate, and it applies with full force to any log you keep yourself.

The Forecast Log

The intervention, and it is ours rather than anything the literature prescribes.

Four columns.

The claim, stated so that it will be unambiguously true or false by a date.

The probability, as a number.

The resolution date.

And what actually happened, filled in later.

Three properties, ours.

It costs under a minute per entry and requires no system.

It is the only way to find out whether your judgment is any good, because unrecorded judgment is scored by memory, and memory is exactly what the belief-system defences operate on.

And it produces something no consultant can sell you: evidence about your own calibration in your own domain, which the first article in this series argued is stronger evidence for your purposes than any published study.

What To Do

Attach a number to every forecast that matters. "Likely" is unscoreable and hides disagreements; a percentage is scoreable and surfaces them.

Write it down before the outcome. A forecast reconstructed afterwards is a memory, and memory is where the self-protective defences operate.

Stop reading confidence as competence. On this evidence the most confident experts were the least accurate, and public prominence tracked inaccuracy too.

Separate predicting from explaining. Domain expertise had little bearing on forecasting accuracy, so a specialist who explains a system well is not thereby a good guide to what it will do.

Check your probabilities add up. Foxes were more prone than hedgehogs to assigning probabilities across outcomes that exceeded one, which is incoherent before any evidence arrives.

Update the probability, not the narrative. Explaining why a wrong call was exceptional changes your story; lowering next time's confidence changes your model.

Expect the log to take years. On our arithmetic you need roughly 35 forecasts at a given confidence level before the check means much.

Notice that vagueness is safe. An adviser who never leaves the middle of the range can never be shown to be overconfident, and has also never told you anything.

And go to the source on any striking finding. Tetlock said foxes were modestly better. The version you will encounter says they crushed hedgehogs.

The Limits Of This Analysis

Several caveats matter. This article reviews research on judgment and is not investment, forecasting or business advice. Everything is verified to August 2026. We did not obtain Tetlock's book and rely on an academic review, encyclopedic sources and secondary accounts. We report no effect sizes, no Brier scores and no specific accuracy comparisons, because we obtained none. Our sources conflict badly on the number of forecasts, giving 82,361, approximately 28,000 and 27,451, and on the date range; our proposed reconciliation is our own conjecture that no source states. Several sources are commercial book-summary websites and blogs, identified at each use; the Tetlock quotation about foxes being modestly better reaches us through one such source and we did not verify it against the book. We did not obtain the Good Judgment Project papers, the Superforecasting book, or the Mellers and colleagues paper, and report no results from that programme. The vague-word practice reaches us through a commercial summary and we report no evidence of its effect on accuracy. All calibration illustrations and sample-size arithmetic are ours, use standard binomial calculation, and appear in no publication. The forecast log, the calibration-versus-discrimination distinction as applied here, the observation about vagueness being professionally safe, and the analysis of the overconfident retelling are our own reasoning. The study concerns political and economic forecasting by public commentators; application to ordinary business forecasting is our own extension.

Frequently Asked Questions

Did the study really show experts are no better than chance?
Close to it. An encyclopedic source reports aggregate accuracy as marginally better than chance on straightforward questions and often inferior to simple statistical baselines on complex ones. The comparison that should concern a business is the second: losing to simple extrapolation.
Should I hire a fox instead of a hedgehog?
Be careful how strongly you read that. Tetlock's own words were that foxes were "modestly better" and hedgehogs "modestly worse," with confident television hedgehogs worst of all. Popular retellings say foxes "crushed" hedgehogs, which is a considerably stronger claim than the author made.
Is being a fox simply better, then?
Not unconditionally. Tetlock reports being able to push more foxes than hedgehogs into assigning probabilities across outcomes that summed to more than one, which is an incoherence rather than an inaccuracy. The lesson is closer to: hold multiple scenarios, then check the numbers add up.
Does subject-matter expertise help?
On this evidence, surprisingly little. Experts forecast both inside and outside their specialisms, and domain of expertise had little bearing on accuracy. That suggests hiring a specialist to explain a system and hiring one to predict its behaviour are different purchases.
What can I actually do about my own forecasting?
Keep a log: the claim, a probability as a number, a resolution date, and what happened. It takes under a minute an entry and is the only way to find out whether your judgment is any good, because unrecorded judgment is scored by memory.
How long before the log tells me anything?
On our own arithmetic, roughly 35 forecasts at a given confidence level to check it within about ten percentage points. A business making one significant forecast a month reaches that in about three years, which is why it has to start before you need it.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article's central observation is that a study about overconfident expert claims is routinely transmitted by overstating what it found, and it sets the author's own wording against the popular version.

References

  1. Academic review of Tetlock, P. E. (2005), Expert Political Judgment: How Good Is It? How Can We Know?, hosted by a university faculty page, on the book reporting a two-decade study of expert predictions; on Tetlock having recruited 284 people whose professions included commenting or offering advice on political and economic trends; on his asking them to forecast the probability that various situations would or would not occur, picking areas geographic and substantive within and outside their areas of expertise; on his also asking how forecasters formulated forecasts, dealt with failure, responded to contradictory information and evaluated others' predictions; on his having accumulated 82,361 forecasts by 2003; on his evaluating predictions against outcomes and against alternate predictions derived from simple algorithms; and on his reporting that he was able to push more foxes than hedgehogs into forecasts violating the axiom that a forecaster's probabilities not sum above one. Note: a published academic review; we did not obtain the book. Its forecast count conflicts with several other sources. faculty.wharton.upenn.edu
  2. Encyclopedic entry on Expert Political Judgment, on Tetlock analysing approximately 28,000 predictions from 284 experts tracked over nearly two decades; on aggregate accuracy being marginally better than chance on straightforward questions but often inferior to simple statistical baselines on complex ones, with many forecasters exhibiting overconfidence and poor calibration; on the methodology involving probabilistic forecasts on verifiable events scored for calibration and discrimination; on revealing systemic shortcomings including belief perseverance and insufficient updating; on the Berlin-derived hedgehog and fox classification with hedgehogs performing worse; on hedgehogs retroactively adjusting narratives rather than probabilities; and on experts having been selected based on prominence and frequent public engagement in forecasting. Note: an encyclopedic reference source, not peer-reviewed. grokipedia.com
  3. Runaric almanac. Tetlock's Good Judgment Project: how forecasting becomes measurable, quoting Tetlock's summary of Expert Political Judgment that "in our data, foxes were modestly better forecasters; hedgehogs were modestly worse; and confident television hedgehogs were the worst of all"; on experts as a group performing barely better than crude statistical baselines and the most televised, most confident voices typically being the worst calibrated; on foxes doing slightly better than chance although not dramatically; on the Good Judgment Project being an IARPA-funded forecasting tournament running from 2011 to 2015 under the Aggregative Contingent Estimation programme, with university teams competing against each other, against statistical aggregates and against the United States intelligence community on hundreds of geopolitical questions with verifiable outcomes; and giving citations for Tetlock (2005), Tetlock and Gardner (2015) and Mellers and colleagues (2015). Note: a blog. The Tetlock quotation is central to this article and we did not verify it against the book. runaric.com
  4. Commercial book-summary site, on Tetlock having asked 284 people between 1984 and 2003 to make predictions; on his collecting 28,000 predictions and recording confidence levels; on the experts' predictions being on average barely more accurate than random chance, with a dart-throwing chimpanzee doing about as well; and on the experts who were most confident tending to be the least accurate. A further such source states "foxes crushed hedgehogs" and that "the people who sounded best on TV were the worst at actually predicting what would happen", and describes the practice distinction that ordinary forecasters speak in vague terms such as likely, possible and could happen, which mean different things to different people. Note: commercial book-summary websites, not peer-reviewed. The "crushed" characterisation is quoted in this article specifically to contrast it with the author's own wording. ideasthesia.org
  5. Reader annotation to Gaddis, On Grand Strategy, recording that Tetlock and his assistants collected 27,451 predictions on world politics between 1988 and 2003 from 284 experts in universities, governments, think tanks, foundations, international institutions and the media. Note: a reader's book annotation on a social reading platform. We record it solely because its figures conflict with our other sources and the conflict is material. goodreads.com
  6. Life Itself Collective, notes on Tetlock and Gardner's Superforecasting, on Tetlock beginning a research programme in the mid-1980s to learn what sets the best forecasters apart; on his recruiting experts whose livelihoods involved analyzing political and economic trends; on the experts making roughly twenty-eight thousand predictions; on the final results appearing in 2005; on the average expert having been roughly as accurate as a dart-throwing chimpanzee; and on averages obscuring, with two statistically distinguishable groups of experts in the results. Note: a blog summarising a book we did not obtain. lifeitself.org
  7. Reader summary of Expert Political Judgment on a book cataloguing platform, on forecasting judgment being poor, only barely beating random chance and running below simple extrapolation algorithms; on the radical sceptics not being fully validated because the research discovers consistent patterns in good judgment; on the specific domain of expertise having little bearing on forecasting ability, which does not imply forecasting is foolhardy because cognitive traits did show consistent ability; on ideology having little correlation to good forecasting judgment; and on the divide between foxes and hedgehogs, with hedgehogs knowing one big thing and being strong ideologically driven thinkers who favour parsimony. Note: a reader-written summary on a book platform, not peer-reviewed. goodreads.com
  8. Commercial book-summary site, Expert Political Judgment: Summary and Discussion Questions, on foxes knowing many small things, drawing on multiple frameworks, updating more readily and being comfortable with uncertainty; on foxes having outperformed hedgehogs in prediction, often substantially, especially over longer time horizons; on Tetlock being a professor of psychology and management at the University of Pennsylvania's Wharton School and co-director of the Good Judgment Project; and on his exploring how experts protect their self-image through hindsight bias, creeping determinism and belief system defenses. Note: a commercial book-summary website, not peer-reviewed. Its "often substantially" characterisation conflicts with the author's own "modestly" at reference 3, and we quote it to mark that gap. superbook.ai

This article reviews research on judgment and is not investment, forecasting or business advice. Tetlock's book was not obtained; no effect sizes or Brier scores are reported. Sources conflict on the number of forecasts by roughly a factor of three and on the date range, and the conflict is reported rather than resolved. Several sources are commercial book-summary websites, blogs or reader annotations, each identified; the central Tetlock quotation reaches this article through one such source and was not verified against the book. All calibration illustrations and sample-size arithmetic are the authors' own. Application to ordinary business forecasting is the authors' extension.