A decision is made under uncertainty and turns out badly. The reviewer has the file, knows exactly what was and was not knowable at the time, and is trying to be fair. This article is about the finding that none of that is sufficient.

Key Takeaway

From the abstract: "Subjects understood that they had all relevant information available to the decision maker. Subjects rated the thinking as better, rated the decision maker as more competent, or indicated greater willingness to yield the decision when the outcome was favorable than when it was unfavorable." And: "Although subjects who were asked felt that they should not consider outcomes in making these evaluations, they did so."[1][2]

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A. Five studies in a leading journal, abstract obtained verbatim from four independent sources including a national medical database and the lead author's own page, and two independent replications indexed as successful in a replication database.

That is the highest grade this series has awarded, and it is worth saying why: most findings covered here have one supporting paper and a dispute. This one has an original, an explicit replication record, and no critique we could locate.

Ungraded for magnitude. We did not obtain an effect size from the original or either replication, and the original's sample was twenty.

Our position: the existence of the effect is about as well established as anything in this series, and we still cannot tell you how big it is.

A Note On Method

Everything here is verified to August 2026.

We obtained the abstract verbatim from four independent sources, including a national medical literature database, the lead author's own institutional page, and a bibliographic service[1][2][3][4]. We did not obtain the paper.

One version of the abstract truncates mid-sentence on the counterfactual finding, and we report that finding incompletely as a result[1].

Replication details come from a replication paper's own account of its target[5][6] and from a replication database entry[7]. We did not obtain either replication's results.

One conceptual framing is taken from an encyclopedia entry[8] and flagged as such.

All arithmetic is ours, uses invented parameters, and states no real-world rates.

This article discusses research on judgment. It is not legal, audit, professional conduct or liability advice.

Not The Same As Hindsight Bias

The distinction, stated early because the two are routinely merged and the previous article covered the other one.

Hindsight bias is about what you believe was predictable. You misremember your own prior estimate, and you overestimate what was knowable.

Outcome bias is about how you rate a decision. It survives even when the predictability question is settled for you.

Three observations, ours.

The experimental design makes the difference concrete. Subjects were told the probabilities and told they had everything the decision-maker had. There was nothing to misremember.

Which means outcome bias is not a memory failure at all. It is an evaluation that incorporates information the evaluator agrees is irrelevant.

And practically, the two require different remedies. Hindsight bias is addressed by contemporaneous records. Outcome bias is not, because the record does not stop you weighting the result.

Five Studies

The paper.

Baron and Hershey published Outcome bias in decision evaluation in the Journal of Personality and Social Psychology, 54(4), 569–579, in April 1988, DOI 10.1037/0022-3514.54.4.569[1][3].

The design: "In 5 studies, undergraduate subjects were given descriptions and outcomes of decisions made by others under conditions of uncertainty. Decisions concerned either medical matters or monetary gambles."[1]

An encyclopedia entry describes one case as involving a surgeon deciding whether to perform a risky operation with a known probability of success, with subjects shown either a good or bad outcome and asked to rate the quality of the pre-operation decision[8]. This description is from an encyclopedia, not the paper.

Two observations, ours.

The two domains matter. Medical decisions and monetary gambles are very different in emotional weight, and the effect appeared in both.

And the gambles are the cleaner test, because a gamble's odds are stated. There is no ambiguity about what was knowable in advance.

All Relevant Information

The clause that carries the paper.

"Subjects understood that they had all relevant information available to the decision maker."[1]

Three observations, ours.

This is the control that removes the obvious alternative explanation. Without it, one could say the outcome was simply informative about what the decision-maker must have known or failed to find out.

With it, that route is closed. The subject and the decision-maker have the same information set, and only the outcome differs.

And it is why this study answers a different question from the professional-judgment study in the previous article, where clinicians reviewed a chart. Here the informational position is stipulated rather than inferred.

Three Measures, One Result

What was actually rated, and it escalates.

Subjects rated "the quality of thinking of the decisions, the competence of the decision maker, or their willingness to let the decision maker decide on their behalf."[1]

And on all three: "Subjects rated the thinking as better, rated the decision maker as more competent, or indicated greater willingness to yield the decision when the outcome was favorable than when it was unfavorable."[1]

Three observations, ours.

The first measure is about the process. Subjects rated the reasoning itself as better, though the reasoning was identical.

The second generalises from a decision to a person. One outcome moved the assessment of competence.

And the third is behavioural rather than attitudinal. Willingness to delegate is a decision with consequences, not a rating on a form, and it moved too.

They Knew They Should Not

The finding that removes the easy remedy.

"Although subjects who were asked felt that they should not consider outcomes in making these evaluations, they did so."[2]

A replication describes the same thing: "outcome bias occurred despite participants indicating that they believe outcomes should not impact their judgment."[6]

Three observations, ours.

This is different from the previous article's finding, and worse in one specific way. There, judges were unaware of the effect. Here they were aware of the principle and it did not help.

Which forecloses the intervention people reach for first. Telling evaluators that outcomes should not count does not work, because they already think that.

And it should change what a firm builds. A stated policy of judging decisions on process is not a control; it is a description of what everyone already believes they are doing.

The Road Not Taken Also Counted

A finding we can report only partially.

The abstract continues: "In monetary gambles, subjects rated the thinking as better when the outcome of the option not chosen turned out poorly than when it turned ou"t well, with our source truncating mid-word[1].

Two observations, ours.

Even reading it conservatively, the judgment moved on what would have happened under the option that was not taken, which is information the decision-maker could not have had and which does not exist until afterward.

A related line of work points the same way. A summary describes three experiments in which "decisions resulting in considerable amounts of profit, but missed alternative outcomes of greater profits, were rated lower in quality and produced more regret"[2]. We did not obtain that paper and report the description.

The Mechanism The Authors Propose

Their own explanation, which is modest.

The effect of outcome knowledge "may be explained partly in terms of its effect on the salience of arguments for each side of the choice."[2]

Three observations, ours.

Note the hedging: may be explained partly. The authors are not claiming a complete account, and we report it at that strength.

The mechanism is about which arguments come to mind. Knowing the operation failed makes the reasons against operating easier to retrieve, and those reasons were available beforehand too.

And it matches the previous article's finding that outcome knowledge changed the perceived relevance of the evidence. Two literatures, two designs, the same underlying story about re-sorting rather than re-remembering.

A Sample Of Twenty

The weakness, reported because a later paper reports it.

A replication team explains their choice of target: "the original study was based on a sample size of 20, which resulted in effect size estimates with relatively wide confidence intervals. A larger sample would hence allow us to obtain an estimate of effect size with higher precision."[6]

Two observations, ours.

Twenty subjects is small, and a finding this widely cited resting on it is exactly the pattern this series has criticised elsewhere. Experiment 1 used a within-participants design with 15 cases[6], which recovers some power, but the number is still twenty people.

And this is why the replication record matters more here than the original does. What makes us grade this A is not the 1988 paper. It is what happened to it afterward.

And It Replicated Twice

The record.

A replication database entry states that the paper "has been replicated by Yuk et al. (2021), described as successful" and "has been replicated by Aiyer et al. (2023), described as successful."[7]

The 2023 study is Aiyer, Kam, Ng, Young, Shi and Feldman, Outcomes Affect Evaluations of Decision Quality: Replication and Extensions of Baron and Hershey's (1988) Outcome Bias Experiment 1, in the International Review of Social Psychology[5], described by its authors as "an independent pre-registered and well-powered replication of a classic article on outcome bias."[6]

We did not obtain either replication's results and report the database's characterisation and the authors' own description of their design.

Three observations, ours.

Pre-registered and well-powered, by an independent team, is the standard this series has spent forty-four articles asking for.

The 2023 team also examined extensions, testing whether outcome bias is informed by perceptions about the responsibility of decision-makers or social norms about using outcomes[6]. We did not obtain those results.

And the existence of a public, searchable replication database that records this is itself worth noting. Several earlier articles in this series would have been easier to write if every finding had one.

The Unbiased Benchmark

What correct evaluation would look like, stated crisply by the replication team.

"In an unbiased situation, the judge processes only the information available to the decision-maker at the time of a decision. Considering that outcomes may not be related to the quality of the decision, outcome information should not affect the judgment of the decision in most cases."[6]

And the conceptual point, from an encyclopedia entry and flagged as such: "no decision-maker ever knows whether or not a calculated risk will turn out for the best. The actual outcome of the decision will often be determined by chance." Those influenced by outcome bias are "seemingly holding decision-makers responsible for events beyond their control."[8]

One observation, ours, and it is a caveat on the benchmark rather than an endorsement. The phrase "in most cases" is doing real work. Outcomes are not always uninformative about decision quality, and the next section is about how much they are actually worth.

How Much Is An Outcome Worth

Because saying outcomes should be ignored is too strong, and a number is more useful than a principle. Our own arithmetic, standard Bayes, invented parameters, no real-world rates asserted.

Suppose good decisions succeed with one probability and poor ones with another, half of decisions are good, and you observe only the outcome.

Where a good decision succeeds 70 percent of the time and a poor one 40, a success should move you to 63.6 percent confidence the decision was good, and a failure to 33.3. A swing of about 30 points.

Where the edge is more modest, 55 against 45, a success moves you to 55.0 and a failure to 45.0. A swing of 10 points.

Three observations.

The outcome is evidence. It is not nothing, and a benchmark that says ignore it entirely is wrong.

How much evidence depends entirely on how separated the two success rates are, which is precisely what a reviewer after the fact does not know and rarely estimates.

And in the modest case, a single outcome justifies moving from 50 to 55. The studies found people rating the thinking itself as better, the person as more competent, and saying they would be more willing to delegate, on evidence worth five points.

Ratios, Not Differences

A small technical note, included because our own table looks wrong at a glance. Ours.

Success rates of 60 against 50 and 55 against 45 both differ by ten points, but the second is more informative per observation.

One observation. What governs updating is the ratio of the two rates, not the gap: 55 divided by 45 is about 1.22, while 60 divided by 50 is 1.20. The same ten-point difference carries more information at lower absolute rates, which is unintuitive and is the sort of thing worth checking rather than eyeballing.

Eleven In A Row

What it would actually take to learn something about a person from outcomes. Ours, same assumptions.

Starting from an even prior, the number of consecutive successes needed to reach 90 percent confidence that someone is the better decision-maker.

At 70 against 40: four in a row.

At 55 against 45: eleven in a row.

Three observations.

At a modest edge it takes eleven consecutive successes, which most professional situations will never supply.

Which means a single outcome is not a verdict on a person. It is one weak observation, and the studies found it moving assessments of competence and willingness to delegate.

And the twenty-sixth article in this series made the same point from the other direction: sorting people on a noisy measure produces apparent differences that are mostly luck. Outcomes are a noisy measure.

What This Does To An Adviser

The application, ours, and it completes an argument this series has been building.

Three points.

The thirty-seventh article found that reputation as an adviser is slow to build and fast to lose. The forty-third offered hindsight bias as one mechanism. This is the second and it is more direct, because it does not require the client to misremember anything.

A client with the whole file, who understands exactly what was knowable, and who believes outcomes should not determine their assessment, will still rate the advice as worse when it turns out badly. That is the finding, in a setting stipulated to remove every other explanation.

And it explains why the contemporaneous record recommended in article thirty-seven is necessary but not sufficient. The record defeats the misremembering. It does not defeat this.

And To Judging Your Own Team

The other direction, which is the one a firm can actually control. Ours, untested.

Three observations.

Performance review, promotion, and who gets the next difficult file are all evaluations of decision quality made after outcomes are known.

The third measure in the study was willingness to let the decision maker decide on your behalf, which is close to a description of delegation. On the finding, that willingness moves with outcomes rather than with reasoning.

And the compounding is the risk. Someone handed a genuinely harder set of problems will produce worse outcomes on identical reasoning, and on this literature will be rated less competent and given less discretion, which is the twenty-sixth article's sorting problem operating on careers.

The Defence That Actually Works

What to do given that stating the principle does not work. Ours, and offered as reasoning rather than evidence.

Four measures.

Evaluate before the outcome is known. This is the only intervention that removes the cause rather than fighting it. A decision reviewed at the time of the decision cannot be contaminated by a result that does not exist.

Ask what the odds were, in numbers. The gambles in the study were the cleanest test because the probabilities were stated. Requiring an explicit estimate at the time creates the same clarity.

Review a sample of decisions that went well. Files are pulled because something failed, so the process quality of successful decisions is never examined, and the comparison that would calibrate the reviewer never happens.

And separate the two questions explicitly. Was this a good decision, and did it work, are different questions with different answers, and a review that produces one number has answered neither.

What To Do

Do not rely on telling people outcomes should not count. Subjects who believed exactly that showed the effect anyway.

Record the decision review at the time of the decision. It is the only measure that removes the cause rather than resisting it.

Write down the odds you assign. The clearest results came from gambles, where what was knowable in advance was explicit.

Treat one outcome as one weak observation. On our own arithmetic, at a modest skill edge a single result justifies moving from 50 to 55 percent confidence, and reaching 90 would take eleven consecutive successes.

Do not let an outcome move your view of a person. The studies found competence ratings and willingness to delegate moving on identical reasoning.

Check who gets the hard files. Harder problems produce worse outcomes on identical reasoning, and this effect converts that into a judgment about the person.

Review some decisions that worked. Without them the reviewer never sees how often good process accompanies bad results.

Ask the two questions separately. Whether it was a good decision and whether it worked have different answers.

The Limits Of This Analysis

Several caveats matter. This article discusses research on judgment and is not legal, audit, professional conduct or liability advice. Everything is verified to August 2026. We obtained the abstract verbatim from four independent sources but did not obtain the paper, and report none of its methods, statistics or effect sizes. One version of the abstract truncates mid-word on the counterfactual finding, which we therefore report incompletely. We did not obtain either replication and rely on a replication database's characterisation of both as successful plus the 2023 team's own description of their design; their extension results are unreported here because we do not have them. We report no effect size anywhere, and note that the original study's sample was twenty. One description of the surgeon scenario is taken from an encyclopedia entry rather than the paper, and the conceptual framing about chance is from the same source. The related work on missed alternative outcomes is reported from a bibliographic service's summary, not the paper. All arithmetic is ours, uses standard Bayes with invented success rates, and asserts no real-world figures; the parameters were chosen to illustrate a range, not to describe any profession. The applications to advisory relationships and to performance review are our own reasoning, untested, as is the connection drawn to earlier articles in this series.

Frequently Asked Questions

How is this different from hindsight bias?
Hindsight bias is about misremembering what you thought was likely. Outcome bias is about rating a decision. In these studies subjects were told the probabilities and told they had everything the decision-maker had, so there was nothing to misremember, and the effect appeared anyway.
Didn't the subjects know they shouldn't do this?
Yes, and that is the most important part. Subjects who were asked felt they should not consider outcomes, and did so anyway. Which means telling evaluators that outcomes should not count does not work, because they already believe it.
How strong is the evidence?
Unusually strong for this series. Five studies, an abstract we verified across four independent sources, and two independent replications indexed as successful, one of them pre-registered and well-powered. The original's sample was only twenty, which is why the replication record does the work here.
So should outcomes be ignored entirely?
No, and that benchmark is too strong. On our own arithmetic an outcome is real evidence: with a large skill edge a single result can move you thirty points. With a modest edge it moves you five. The error is not using outcomes, it is using them as though they settled the question.
How many outcomes would it take to judge someone?
On our own illustrative figures, at a modest skill edge you would need eleven consecutive successes to reach ninety percent confidence that someone is the better decision-maker. Most professional situations never supply that.
What actually helps?
Reviewing the decision before the outcome exists, writing down the odds you assigned at the time, sampling decisions that went well rather than only those that failed, and asking whether it was a good decision separately from whether it worked.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This is the first article in the series to award an A grade on the strength of a replication record rather than an original paper, and it says plainly that the original rested on twenty subjects.

References

  1. National medical literature database record for Baron, J., & Hershey, J. C., Outcome bias in decision evaluation, Journal of Personality and Social Psychology, 54(4), 569–579, April 1988, DOI 10.1037/0022-3514.54.4.569, reproducing the abstract, on five studies in which undergraduate subjects were given descriptions and outcomes of decisions made by others under conditions of uncertainty; on decisions concerning either medical matters or monetary gambles; on subjects rating the quality of thinking of the decisions, the competence of the decision maker, or their willingness to let the decision maker decide on their behalf; on subjects having understood that they had all relevant information available to the decision maker; on subjects rating the thinking as better, rating the decision maker as more competent, or indicating greater willingness to yield the decision when the outcome was favorable than when it was unfavorable; and on subjects in monetary gambles rating the thinking as better when the outcome of the option not chosen turned out poorly than when it turned out well, our source truncating mid-word. Note: a national medical literature database. We obtained the abstract and not the paper; the final sentence truncates and the counterfactual finding is reported incompletely as a result. pubmed.ncbi.nlm.nih.gov
  2. Bibliographic service record for Baron and Hershey (1988), reproducing the abstract and its closing passage, on subjects who were asked having felt that they should not consider outcomes in making these evaluations while doing so anyway, and on the effect of outcome knowledge on evaluation possibly being explained partly in terms of its effect on the salience of arguments for each side of the choice; together with a description of related work by Seta, Seta, Petrocelli and McCormick in which three experiments demonstrated that decisions resulting in considerable amounts of profit, but missed alternative outcomes of greater profits, were rated lower in quality and produced more regret. Note: a bibliographic service. Our source for the closing passage of the abstract and for the related work, which we did not obtain. semanticscholar.org
  3. Academic database record for Baron, Jonathan, and Hershey, John C., Journal of Personality and Social Psychology: Attitudes and Social Cognition, volume 54, issue 4, April 1988, 569–579, DOI 10.1037/0022-3514.54.4.569, reproducing the abstract identically and noting that full text access requires institutional login. Note: an academic database; independent corroboration of the abstract text and of the journal section, volume, issue and pages. Access restricted. proquest.com
  4. Lead author's own institutional page listing his published papers, reproducing the abstract of Baron, J., & Hershey, J. C. (1988), Journal of Personality and Social Psychology, 54, 569–579, and separately citing Fischhoff, B. (1975). Note: the lead author's own page; a fourth independent reproduction of the abstract text, and confirmation that the paper is presented alongside the hindsight bias literature. sas.upenn.edu
  5. Journal record for Aiyer, S., Kam, H. C., Ng, K. Y., Young, N. A., Shi, J., & Feldman, G., Outcomes Affect Evaluations of Decision Quality: Replication and Extensions of Baron and Hershey's (1988) Outcome Bias Experiment 1, International Review of Social Psychology, including its reference list identifying the 1988 original and related work on positive-outcome bias in peer review and in research abstracts. Note: a journal record for the 2023 replication. We did not obtain its results. rips-irsp.com
  6. Full-text repository copy of the 2023 replication paper, on an unbiased judge processing only the information available to the decision-maker at the time of a decision and on outcome information not affecting the judgment of the decision in most cases given that outcomes may not be related to decision quality; on the authors' first goal being an independent pre-registered and well-powered replication of a classic article on outcome bias and their second being extensions examining whether outcome bias is informed by perceptions about the responsibility of decision-makers or social norms about using outcomes; on Baron and Hershey having demonstrated outcome bias using five experiments, with Experiment 1 presenting participants with 15 cases of medical decisions in a within-participants design differing only by decision-maker and outcome; on participants having given higher ratings to medical decisions that resulted in positive outcomes despite the decisions being identical outside their outcome; on outcome bias having occurred despite participants indicating that they believe outcomes should not impact their judgment; and on the target having been chosen because the original study was based on a sample size of 20, resulting in effect size estimates with relatively wide confidence intervals. Note: the replication team's own account of the original and of their design. We did not obtain their results or extension findings. ncbi.nlm.nih.gov
  7. Replication database entry for DOI 10.1037/0022-3514.54.4.569, recording that the paper has been replicated by Yuk and colleagues (2021), described as successful, and by Aiyer and colleagues (2023), described as successful, and that it is indexed in a replication atlas maintained as part of an open research reform project. Note: a replication database entry. Our source for the replication record; we obtained neither replication's results and report the database's characterisation. forrt.org
  8. Encyclopedia entry on outcome bias, on no decision-maker ever knowing whether a calculated risk will turn out for the best and the actual outcome often being determined by chance, with some risks working out and others not; on individuals whose judgments are influenced by outcome bias seemingly holding decision-makers responsible for events beyond their control; and describing a Baron and Hershey scenario involving a surgeon deciding whether to perform a risky surgery with a known probability of success, in which subjects presented with bad outcomes rated the decision worse than those presented with good outcomes. Note: an encyclopedia entry, not a peer-reviewed source. Used only for the conceptual framing and the scenario description, both flagged in the body. en.wikipedia.org

This article discusses research on judgment and is not legal, audit, professional conduct or liability advice. The paper was not obtained; its abstract was verified across four independent sources and one version truncates mid-word. Neither replication's results were obtained. No effect size is reported anywhere in this article, and the original study's sample was twenty. All arithmetic is the authors' own, uses invented parameters, and asserts no real-world rates.