A file is reviewed after something has gone wrong. The reviewer reads the same documents the original decision-maker read, and the warning signs are obvious. This article is about why that obviousness is not evidence.

Key Takeaway

Fischhoff's abstract: outcome knowledge "was found to increase the postdicted likelihood of reported events and change the perceived relevance of event descriptive data, regardless of the likelihood of the outcome and the truth of the report. Judges were, however, largely unaware of the effect that outcome knowledge had on their perceptions. As a result, they overestimated what they would have known without outcome knowledge, as well as what others actually did know without outcome knowledge."[1]

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the original findings. We obtained the abstract verbatim, the paper has two meta-analyses behind it, and an editorial records more than 150 articles and book chapters addressing the phenomenon.

Ungraded for magnitude, because we could not obtain either meta-analysis and report no effect size anywhere in this article.

Grade B that the effect is eliminated when outcomes are attributed to chance, on a single source describing one experiment.

Grade B for the debiasing result, from a paper whose abstract we obtained in full.

Our position: the existence and direction are as solid as anything in this series, and we cannot tell you how large it is, which matters because the fortieth article was about exactly that failure.

A Note On Method

Everything here is verified to August 2026.

We obtained the 1975 paper's abstract verbatim, through a 2003 reprint in a healthcare quality journal rather than from the original[1]. We did not obtain the paper's methods or results.

We could not obtain either of the two meta-analyses and therefore state no effect size. This is the fifth time in this series that the most decision-relevant document was the one we could not get.

We obtained one full abstract for a 2022 paper on debiasing[2], and one truncated abstract for a study of professionals[3].

The chance-attribution finding reaches us through a law repository's description of an experiment[4], not a primary source.

All arithmetic is ours and uses a model we invented.

This article discusses research on judgment. It is not legal, audit, professional conduct or liability advice.

The Original Paper

The citation and a note on its framing.

Fischhoff published Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty in the Journal of Experimental Psychology: Human Perception and Performance, 1(3), 288–299, in 1975, DOI 10.1037/0096-1523.1.3.288[3][5].

Its opening frames the problem: "One major difference between historical and nonhistorical judgment is that the historical judge typically knows how things turned out."[1]

Two observations, ours.

The framing is about historical judgment, which is a wider category than it sounds. Any assessment of a past decision is historical judgment, including every file review, every post-mortem, and every claim against a professional.

Fischhoff's own name for the effect was creeping determinism[5], described as the tendency to perceive reported outcomes as having been relatively inevitable. The more familiar knew-it-all-along effect is attributed to Wood in 1978[6].

Three Experiments, Three Findings

The structure of the paper, which is tighter than most things in this series.

Experiment 1. Outcome knowledge "was found to increase the postdicted likelihood of reported events and change the perceived relevance of event descriptive data."[1]

Experiment 2. Judges "overestimated what they would have known without outcome knowledge."[1]

Experiment 3. They also overestimated "what others actually did know without outcome knowledge."[1]

Three observations, ours.

The sequence is cumulative rather than repetitive. The first establishes the effect, the second shows it applies to your own past self, the third shows it applies to other people.

The third is the one with teeth. Judging what someone else should have known is the same operation as judging what you would have known, and it is subject to the same distortion.

And note that Experiment 1 found outcome knowledge changed the perceived relevance of the evidence, not just the probability attached to it. The facts themselves are re-sorted. Details that were background become warning signs.

Regardless Of The Truth Of The Report

The clause we find most striking, and it is easy to read past.

The effect occurred "regardless of the likelihood of the outcome and the truth of the report."[1]

Three observations, ours.

The truth of the report. Being told an outcome occurred produced the effect whether or not it had. The mechanism does not require the outcome to be real; it requires only that you believe it.

Which means the effect is not learning in any useful sense. If a false report reorganises your reading of the evidence just as a true one does, what happened was not the incorporation of information.

And regardless of the likelihood of the outcome is the second half. Even improbable outcomes, once reported, came to look more probable in retrospect.

Largely Unaware

The sentence that makes the rest a problem rather than a curiosity.

"Judges were, however, largely unaware of the effect that outcome knowledge had on their perceptions."[1]

Three observations, ours.

Unawareness is what turns a bias into a trap. A known distortion can be corrected for. This one presents as clarity.

It is also why trying harder does not help. The experience of reviewing a file after a failure is not the experience of straining to be fair. It is the experience of things being obvious.

And it explains the characteristic sentence that follows every business disaster, which is that the signs were all there. They were. They were also there in every case that did not fail.

The Scale Of The Literature

How much work sits behind this.

An editorial records that since the 1975 paper, "more than 150 journal articles and book chapters, two meta-analyses (Christensen-Szalanski & Willham, 1991; Guilbault, Bryant, Brockway & Posavac, 2004), and one special issue" have addressed hindsight phenomena[6].

The editorial is Blank, H., Musch, J., and Pohl, R. (2007), Hindsight bias: on being wise after the event, Social Cognition, 25(1), 1–9[6].

Two observations, ours.

The literature is domain-spanning. One source lists documentation in auditory processing, labour disputes, medical diagnoses, consumer satisfaction, personnel management, sporting events, political strategy, legal proceedings, and nuclear accident analysis[7].

And two meta-analyses fourteen years apart is the signature of a finding that has been taken seriously rather than assumed.

The Meta-Analyses We Could Not Obtain

The gap, stated in the body because the fortieth article in this series was about the cost of not doing so.

We obtained neither meta-analysis and report no effect size for hindsight bias anywhere in this article.

Three observations, ours.

This means we can tell you the effect exists and is well studied, and we cannot tell you how large it is, which is precisely the deficiency we criticised in ourselves three articles ago.

It matters here more than usual, because the practical question is entirely about magnitude. A small distortion in retrospective judgment is a curiosity; a large one makes after-the-fact review unreliable as evidence.

And a reader who needs the number should go to Christensen-Szalanski and Willham (1991) and Guilbault and colleagues (2004), which are the two documents this article most wanted and did not have.

Ninety-Five Professionals

The study closest to what a professional firm actually faces.

A 2018 study reports: "Using an international sample of 95 mental health professionals the current study explored the impact of outcome knowledge on the decision-making process in forensic evaluation." Its design was "a 2 (suicide/self-harm v. homicide/other-harm) x 2 (outcome provided v. no outcome provided) between groups design" in which participants reviewed a hypothetical patient's chart and gave an opinion on risk[3].

The result: "Participants provided with outcome information were significantly more likely to indicate that they would have predicted the outcome than those who w"ere not, with our source truncating[3].

The paper also notes that before it, "there is no empirical research that explores hindsight bias in forensic evaluation"[3].

Three observations, ours.

These are credentialed professionals exercising domain judgment on a file, which is structurally the same act as a partner reviewing an engagement after a loss.

We have no effect size, only that the difference was significant, and our source truncates before any magnitude.

And the observation that no empirical research existed until 2018 is worth pausing on. A field where practitioners are routinely second-guessed after bad outcomes had not tested whether the second-guessing was reliable.

Why You Cannot Score Yourself

The bridge to the previous article. Our own model, invented, illustrative; no source states these figures and we assert no magnitude.

The previous article described forecasters stating an 80 percent interval about 9.4 points wide in a world whose volatility implied it should be roughly four times that. Such an interval should contain the outcome about 24 percent of the time.

Now suppose memory of the forecast drifts toward what happened. We modelled a remembered centre equal to the original plus some fraction of the distance to the outcome.

At a drift of zero, the recalled hit rate is the true one, about 24 percent.

At 0.2, it becomes 29.3 percent.

At 0.3, 33.3 percent.

At 0.5, 45.4 percent.

Two observations.

A memory that drifts even thirty percent of the way toward the outcome turns a badly calibrated forecaster into an apparently adequate one, without a single forecast having improved.

And the drift fraction is ours and entirely invented. We could not obtain a meta-analytic magnitude, so this demonstrates a direction rather than a size, and we would not want it read as an estimate.

At Full Drift

The limiting case, which is the clearest statement of the problem. Ours.

At complete drift you remember predicting exactly what happened. Your recalled hit rate is 100 percent, every time, permanently.

Three observations.

Nobody is at complete drift. The point of the limiting case is the direction of the error, which is always toward a flattering reconstruction.

Which produces the trap in this article's title. The instrument you would use to measure your forecasting error is the instrument the error acts on. Memory cannot audit memory.

And it is why the previous article's advice was to write the range down. A contemporaneous record is not a convenience. It is the only version of your forecast that is not subject to this.

The Condition That Eliminates It

A moderator, and it is the most practically hopeful finding here.

A description of one experiment reports: "When the outcome was attributed to unforeseeable 'chance' factors, such as an unexpected storm or an earthquake, the hindsight effect was virtually eliminated."[4]

This reaches us through a repository's description of an experiment, not a primary source, and we did not obtain the study.

Three observations, ours.

Virtually eliminated is a strong claim for a single moderator, and we grade it B on its sourcing.

If it holds, the mechanism is about narrative availability rather than probability. An outcome you can tell a causal story about becomes inevitable in retrospect; one you cannot does not.

And that is directly actionable in a review. Asking what part of this outcome was chance is not a rhetorical defence; on this finding it is the specific intervention that suppresses the effect.

A By-Product Of Learning

The most sympathetic account of why this happens.

The same source describes findings "as providing evidence that the hindsight effect is a by-product of adaptive learning from feedback"[4].

Two observations, ours.

If accurate, this is not a defect bolted onto cognition. Updating your model of the world when you learn an outcome is exactly what you should do, and losing access to the previous model is the cost of doing it efficiently.

Which sits awkwardly beside the finding that the effect occurs regardless of the truth of the report. A learning process that updates just as readily on false information is adaptive only when the information is good, and we cannot reconcile these two and are not going to pretend otherwise.

Saying Is Believing

The most recent work we found, and it is more encouraging than the rest.

A 2022 paper reports that the authors "reversed the classic memory design and showed that subjective probabilities also decreased when participants encountered foresight instructions after hindsight instructions, demonstrating that previously induced outcome knowledge did not prevent unbiased judgments."[2]

And a longer-run result: "The constructive impact of self-generated and communicated judgments ('saying is believing') was apparent after a 2-week consolidation period: Not outcome knowledge, but rather the last pragmatic response (either biased or unbiased) determined judgments at the third measurement."[2]

The authors conclude their findings "highlight the short-term malleability of hindsight influences in response to task pragmatics and has major implications for debiasing."[2]

Three observations, ours.

The first result matters because it says outcome knowledge does not lock you out. Knowing how it turned out did not prevent participants from producing unbiased judgments when asked the right way.

The second is stranger and more useful: after two weeks, what determined the judgment was the last thing the participant had said, not the outcome they knew.

And that suggests a practical rule for any review. The reconstruction you articulate becomes the one you keep, so a review that first produces a careful account of what was knowable at the time has done something durable, and one that opens with the outcome has also done something durable.

A Sixteenth Variant

The running sourcing note.

Two independent records date the 1975 paper to 2006, giving Journal of Experimental Psychology: Human Perception and Performance, 2006;1(3), while correctly reproducing volume 1, issue 3[5].

One observation, ours. A volume 1 issue 3 dated 2006 is internally inconsistent on its face, and it appears on a government-affiliated patient safety resource. That brings this series' count to sixteen, and this one is the most obviously wrong yet.

What This Does To A Review

The application to file review and post-mortems. Ours, untested.

Four consequences.

The obviousness of warning signs is not evidence. Experiment 1 found outcome knowledge changed the perceived relevance of the evidence, so the signs looking obvious is a predicted consequence of knowing the outcome.

A reviewer cannot correct by trying. Judges were largely unaware of the effect, so the experience of being fair is not evidence of being fair.

Only cases selected on outcome get reviewed, which compounds it. Files are pulled because something went wrong, so the base rate of these signs in files that went fine is never observed. That is the thirtieth article's selection problem sitting on top of this one.

And a review that asks what was knowable is a different exercise from one that asks what happened. Only the first is answerable in a way that improves anything.

And To An Adviser

The connection to the thirty-seventh article, which is ours.

That article reported that reputation as an adviser is slow to build and fast to lose, and that a client leaving after one bad outcome is applying a normal update rule.

Two observations.

This is a candidate mechanism for the harshness of that update. After a bad outcome the client does not merely see a bad result; on this literature the evidence itself re-sorts, and the advice comes to look as though it should obviously have been otherwise.

And it is why that article recommended recording the reasoning rather than only the recommendation. A contemporaneous note of what was uncertain at the time is the only artifact that survives this, because it was written before the outcome existed to reorganise it.

What To Do

Write the forecast down. Memory drifts toward the outcome, and the instrument you would use to measure the error is the one the error acts on.

Do not treat obvious warning signs as evidence of negligence. Outcome knowledge was found to change the perceived relevance of the evidence, so their obviousness is a predicted consequence of knowing.

Ask what part of the outcome was chance. On one experiment, attributing the outcome to unforeseeable factors virtually eliminated the effect, which makes this a specific intervention rather than a rhetorical move.

Sample files that went well. Reviews triggered by bad outcomes never observe how often the same signs appear in cases that were fine.

Ask what was knowable, not what happened. They are different questions and only the first can improve a process.

Record the reasoning, not just the decision. A note of what was uncertain at the time is the only artifact written before the outcome existed to reorganise it.

Notice that trying harder will not work. Judges were largely unaware of the effect, so the feeling of being fair carries no information about whether you are.

Do not assume you are stuck with it. A 2022 paper found outcome knowledge did not prevent unbiased judgment when participants were asked the right way, and that what people had last articulated determined their judgment two weeks later.

The Limits Of This Analysis

Several caveats matter. This article discusses research on judgment and is not legal, audit, professional conduct or liability advice. Everything is verified to August 2026. We obtained the 1975 abstract verbatim through a 2003 reprint rather than the original, and none of the paper's methods or results. We could not obtain either of the two meta-analyses and report no effect size for hindsight bias anywhere, which means this article establishes direction and not magnitude; a reader needing the size should go to Christensen-Szalanski and Willham (1991) and Guilbault and colleagues (2004). The study of 95 professionals reaches us as a truncated abstract with no magnitude, only that a difference was significant. The chance-attribution finding reaches us through a repository's description of an experiment and is graded B accordingly. All arithmetic is ours; the memory-drift model was invented for this article, no source states a drift fraction, and it demonstrates a direction rather than an estimate. We note without resolving the tension between the effect occurring regardless of the truth of the report and its characterisation as a by-product of adaptive learning. The applications to file review and to advisory relationships are our own reasoning, untested, and the connection to the thirty-seventh article in this series is ours rather than any source's.

Frequently Asked Questions

What did the original study find?
That outcome knowledge increased the postdicted likelihood of events and changed the perceived relevance of the evidence, regardless of the likelihood of the outcome and the truth of the report; that judges were largely unaware of this; and that they consequently overestimated both what they would have known and what others did know.
Why does "the truth of the report" matter?
Because the effect occurred whether or not the reported outcome had actually happened. Being told something occurred reorganised the reading of the evidence either way, which means what happened was not the incorporation of information in any useful sense.
How large is the effect?
We cannot tell you. Two meta-analyses exist and we obtained neither, so this article reports no effect size at all. That is a real limitation, and it matters here because the practical question is entirely about magnitude.
Why can't I just score my own forecasts from memory?
Because the instrument you would measure with is the one the error acts on. On our own invented model, a memory that drifts thirty percent of the way toward the outcome turns a 24 percent hit rate into an apparent 33, without any forecast improving.
Is there anything that reduces it?
Two things, on sources we did not obtain in full. Attributing the outcome to unforeseeable chance factors virtually eliminated the effect in one experiment. And a 2022 paper found that outcome knowledge did not prevent unbiased judgments when participants were asked under foresight instructions.
What does this mean for reviewing a file that went wrong?
That the obviousness of the warning signs is a predicted consequence of knowing the outcome rather than evidence about the original decision, that trying to be fair does not correct it, and that the only way to learn the base rate is to review files that went well too.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article reports a finding it could not size, says so in the body rather than the footnotes, and names the two documents a reader would need instead.

References

  1. Reprint of Fischhoff, B., Hindsight is not equal to foresight: the effect of outcome knowledge on judgment under uncertainty, in Quality and Safety in Health Care, 12(4), 304–312, 2003, DOI 10.1136/qhc.12.4.304, reproducing the abstract, on one major difference between historical and nonhistorical judgment being that the historical judge typically knows how things turned out; on receipt of outcome knowledge in Experiment 1 having been found to increase the postdicted likelihood of reported events and change the perceived relevance of event descriptive data, regardless of the likelihood of the outcome and the truth of the report; on judges having been largely unaware of the effect that outcome knowledge had on their perceptions; and on their consequently having overestimated what they would have known without outcome knowledge in Experiment 2, as well as what others actually did know without outcome knowledge in Experiment 3. Note: a 2003 reprint in a healthcare quality journal, not the 1975 original. We obtained the abstract verbatim and none of the paper's methods or results. pmc.ncbi.nlm.nih.gov
  2. Publisher record for a 2022 paper on pragmatic, constructive and reconstructive memory influences on the hindsight bias, reproducing the abstract, on hindsight judgments of outcome probabilities exceeding foresight judgments of the same probabilities without outcome knowledge; on a three by two within-participants design with three sequential judgments of outcome probabilities in two scenarios replicating both the within-participants hindsight bias of the classic memory design and the between-participants hindsight bias of a hypothetical design simultaneously; on the authors having reversed the classic memory design and shown that subjective probabilities also decreased when participants encountered foresight instructions after hindsight instructions, demonstrating that previously induced outcome knowledge did not prevent unbiased judgments; on the constructive impact of self-generated and communicated judgments, described as saying is believing, being apparent after a two-week consolidation period, with not outcome knowledge but rather the last pragmatic response determining judgments at the third measurement; and on the findings highlighting the short-term malleability of hindsight influences in response to task pragmatics with major implications for debiasing. Note: a publisher record; we obtained the abstract in full and not the paper. link.springer.com
  3. Publisher record for Beltrani, A., Reed, A. L., Zapf, P. A., and Otto, R. K. (2018), Is Hindsight Really 20/20?: The Impact of Outcome Information on the Decision-Making Process, on previous research having reviewed the issue of biasability in forensic evaluation and anecdotally related different forms of bias to forensic work while there being no empirical research exploring hindsight bias in forensic evaluation; on hindsight bias being the tendency for an individual to believe that a specific event, in hindsight, was more predictable than it was in foresight; on the study using an international sample of 95 mental health professionals to explore the impact of outcome knowledge on the decision-making process in forensic evaluation, using a two by two between groups design crossing suicide or self-harm against homicide or other-harm with outcome provided against no outcome provided, in which participants reviewed a hypothetical patient's hospital chart and indicated their opinion regarding the risk of harm; and on participants provided with outcome information having been significantly more likely to indicate that they would have predicted the outcome, our source truncating; together with its reference list giving Fischhoff, B. (1975), Journal of Experimental Psychology: Human Perception and Performance, 1(3), 288–299. Note: a publisher record; the abstract truncates before any magnitude and we report only that the difference was significant. journals.sagepub.com
  4. Law school repository record describing an experiment on hindsight, on people being usually unable to reproduce the judgments they would have made without outcome knowledge when they know how an event turned out; on their being unaware of their inability to recapture their pre-outcome state of mind; on that tendency to overestimate what they would have known being called hindsight; on an experiment exploring the moderating effects of the type of cause to which the outcome was attributed on the magnitude of the hindsight effect; on the hindsight effect having been virtually eliminated when the outcome was attributed to unforeseeable chance factors such as an unexpected storm or an earthquake; and on the findings being interpreted as supporting Fischhoff's creeping determinism hypothesis and as providing evidence that the hindsight effect is a by-product of adaptive learning from feedback. Note: a repository's description of an experiment, not a primary source. We did not obtain the study, and the chance-attribution finding is graded B on this basis. repository.law.umich.edu
  5. Patient safety resource record for Fischhoff, B., Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty, Journal of Experimental Psychology: Human Perception and Performance, 1(3), DOI 10.1037/0096-1523.1.3.288, on the article presenting three psychology experiments examining how knowledge of an outcome affects the perception of the past; on subjects having suffered from creeping determinism, the tendency to perceive reported outcomes as having been relatively inevitable in hindsight; on being informed an outcome had occurred having biased their hindsight view of events; and on an accompanying commentary extrapolating the findings to medical errors and arguing that we cannot accurately reconstruct the past. Note: a government-affiliated patient safety resource. This record dates the paper to 2006 while giving volume 1, issue 3, which is internally inconsistent; the paper is from 1975. psnet.ahrq.gov
  6. University research portal record for Blank, Hartmut, Musch, J., and Pohl, R. (2007), Hindsight bias: on being wise after the event, Social Cognition, 25(1), 1–9, reproducing the abstract, on Baruch Fischhoff's 1975 paper having opened up a whole new research field; on more than 150 journal articles and book chapters, two meta-analyses by Christensen-Szalanski and Willham (1991) and Guilbault, Bryant, Brockway and Posavac (2004), and one special issue of Memory in 2003 edited by Ulrich Hoffrage and RĂ¼diger Pohl having addressed hindsight phenomena; and on the editorial aiming to provide a roadmap to the research landscape and to place thirteen articles in historical and systematic perspective. Note: a research portal record reproducing an editorial's abstract. Our source for the scale of the literature and for the existence of the two meta-analyses, neither of which we obtained. researchportal.port.ac.uk
  7. Academic preprint literature section, on hindsight bias as described by Fischhoff (1975) being the tendency to adjust the estimates of various likelihoods of possible event outcomes in uncertain situations after the event has occurred and the outcome is known; on it also being known as creeping determinism, attributed to Fischhoff (1975), and the knew-it-all-along effect, attributed to Wood (1978); and on it having been documented in auditory processing, labor disputes, medical diagnoses, consumer satisfaction, personnel management, sporting events, political strategy, legal proceedings, and nuclear accident analysis, citing reviews including Hawkins and Hastie (1990) and Roese and Vohs (2012). Note: an academic preprint's literature section, not a primary source. Our source for the terminology attributions and the list of domains. arxiv.org

This article discusses research on judgment and is not legal, audit, professional conduct or liability advice. The 1975 paper was obtained only as an abstract through a 2003 reprint. Neither of the two meta-analyses was obtained, and no effect size for hindsight bias is reported anywhere in this article. All arithmetic is the authors' own; the memory-drift model was invented for this article and demonstrates a direction rather than a magnitude.