Every list of cognitive biases leads with this one. It is invoked in audit training, due diligence checklists, and hiring guidance. And the experiment everyone cites for it has a structural property that almost nobody mentions.
Key Takeaway
Klayman and Ha proposed "that many phenomena of human hypothesis testing can be understood in terms of a general positive test strategy. With this strategy, there is a tendency to test cases that are expected to have the property of interest rather than those expected to lack that property."[1] A scholarly source describes this as a "structural re-reading inside the cognitive-science literature itself", in which "the strategy is information-bearing under most realistic task environments."[2] We built the structure ourselves. In the famous task's configuration, a positive test resolves zero bits of uncertainty.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A that people use a positive test strategy. The proposal is stated in the paper's own abstract, which we obtained from an educational research database, and it is the framing subsequent literature uses.
Grade A that the classic task's configuration makes positive testing uninformative. This is not an empirical claim but a structural one, and we verified it by construction rather than taking anyone's word.
Grade C that positive testing is efficient in most real environments, which is the load-bearing interpretive claim and which our own check does not reproduce as it is usually worded. That has its own section.
Ungraded for the size of confirmation bias in the wild, because a 2024 paper states the field lacks a strong understanding of its boundaries.
Our position: the phenomenon is real, the famous demonstration is rigged, and the reappraisal is more often cited than checked.
A Note On Method
Everything here is verified to August 2026.
We did not obtain the 1987 paper. We have its abstract from an educational research database[1] and its argument as characterised by others.
The clearest statement of the reframing comes from an educational psychology website[3], which is not peer-reviewed. We use it only where a scholarly source corroborates the same point[2], and we flag it at every use.
We did not obtain the 1960 task paper, the 1998 review beyond a truncated first sentence, or any of the follow-up studies.
All arithmetic and set constructions are ours, use invented illustrative sets, and reproduce a structural property rather than anyone's data.
This article discusses research on reasoning. It is not audit, assurance or professional conduct advice.
What The Term Means
The standard definition, from the standard review.
Nickerson published Confirmation bias: A ubiquitous phenomenon in many guises in Review of General Psychology, 2(2), 175–220, in 1998[4].
It opens: "Confirmation bias, as the term is typically used in the psychological literature, connotes the seeking or interpreting of evidence in ways that are partial to existing beliefs, expectations, or a h"ypothesis in hand, with our source truncating mid-word[4].
Two observations, ours.
Note the careful hedge in the opening clause: as the term is typically used. A review that begins by describing how a term is used rather than what it denotes is signalling that the usage is loose.
And the definition covers two distinct activities, seeking evidence and interpreting it. Those have different mechanisms and different remedies, and the rest of this article is largely about the first.
The Famous Task
The demonstration everyone cites.
Reference lists identify Wason, P. C. (1960), On the failure to eliminate hypotheses in a conceptual task, Quarterly Journal of Experimental Psychology[5].
The task, as described by a source we flag as an educational website rather than a paper: participants are given the triple 2-4-6, told it follows a rule, and asked to discover the rule by proposing further triples and being told whether each fits. The true rule is "any ascending triple"[3].
We did not obtain the 1960 paper and report this description at one remove.
Two observations, ours.
The typical participant forms a narrower hypothesis, such as ascending by twos, and proposes triples that fit it. Every one is confirmed, because every triple ascending by twos is also an ascending triple.
And that is the entire result. People propose cases their hypothesis predicts, receive confirmation, and become confident in a wrong rule. Which sounds like a devastating indictment until you look at the geometry.
The Reappraisal
The paper that reframed it.
Klayman and Ha published Confirmation, disconfirmation, and information in hypothesis testing in Psychological Review, 94(2), 211–228, in 1987, DOI 10.1037/0033-295X.94.2.211[6].
Its proposal, from the paper's own abstract: "It is proposed that many phenomena of human hypothesis testing can be understood in terms of a general positive test strategy. With this strategy, there is a tendency to test cases that are expected to have the property of interest rather than those expected to lack that property."[1]
A scholarly source characterises the contribution: the authors "provided the structural re-reading inside the cognitive-science literature itself: participants are executing a positive test strategy, generating cases the hypothesis predicts and reading the feedback against the prediction, and the strategy is information-bearing under most realistic task environments."[2]
Two observations, ours.
The abstract's claim is descriptive rather than evaluative. It says what people do, not that it is good or bad.
And the same source adds a precise reformulation of what the phenomenon actually is: "The 'bias' the literature has named is the participant's reliance on the held hypothesis as the frame within which the next case is generated and interpreted."[2] That is a considerably narrower claim than wanting to be right.
A Goal And A Method
The distinction that does the work, and it is the reason this article exists.
An educational source states the argument: the classic task "conflates two different things. Seeking confirmation is a goal. Using a positive-test strategy, checking cases you expect to be positive, is just a method."[3]
This phrasing is from a psychology education website, not a peer-reviewed source, and we report it as its characterisation. The underlying distinction is corroborated by the scholarly source quoted above.
Three observations, ours.
A goal is motivational. Wanting your hypothesis to be true is a preference about the answer.
A method is procedural. Testing cases your hypothesis predicts is a way of gathering data, and it can be used by someone with no preference whatever about the outcome.
And the two produce identical behaviour in the classic task, which is precisely why the task cannot distinguish them. Observing the method does not establish the goal.
We Built The Structure
Because a structural claim can be checked rather than cited. Our own construction, invented illustrative sets in a universe of ten thousand items, standard information theory; no source states these figures.
A test is informative only to the extent you cannot predict its answer. We measured, for various relationships between a hypothesis and a true rule, how much uncertainty a single positive test resolves.
Hypothesis strictly inside a very broad true rule, which is the classic task's configuration: probability of confirmation 1.00, information gained 0.000 bits.
Hypothesis strictly inside a narrow true rule: also 1.00, also 0.000 bits.
Hypothesis partially overlapping the true rule: probability 0.50, information 1.000 bits, the maximum available from a yes-or-no question.
Hypothesis broader than the true rule: probability 0.25, information 0.811 bits.
Hypothesis roughly equal to the true rule, both rare: probability 0.83, information 0.650 bits.
Zero Bits
What that table means. Ours.
Three observations.
In the classic task's configuration, a positive test cannot fail. Every triple ascending by twos is an ascending triple, so the answer is yes before you ask, and a question whose answer you already know carries no information by construction.
That is not a claim about psychology. It is arithmetic, and it holds regardless of who is doing the testing or why. A perfectly rational agent using positive tests in that configuration learns nothing.
And in three of the five configurations we built, exactly the same strategy yields substantial information, in one case the maximum a single question can carry.
Where Our Check Disagrees
Reported because our construction does not reproduce a source's wording, and the previous articles in this series would be hypocritical if we hid it.
The educational source states: "In most real environments, a hypothesised category is narrower than the true category. Testing positive cases is then an efficient way to expose errors."[3]
Our own construction does not give that result. When we make the hypothesis strictly narrower than the true rule, positive testing yields zero bits, not efficiency. The strictly-narrower case is the failing case in our table, not the working one.
Two possible resolutions, ours, and we settle neither.
Real hypotheses are rarely strict subsets. A hypothesis that is roughly the right size but slightly misplaced overlaps the truth partially, and that is our third row, which gives the maximum possible information. The word doing the work may be approximately rather than strictly.
Or the source has compressed the argument imprecisely, which would be unsurprising for a summary on an educational website.
One observation. We did not obtain the 1987 paper and cannot settle which. What our construction does establish, independently of the source, is that the classic task sits in the configuration where the strategy is guaranteed useless, and that most other configurations are not like that.
What The Task Selects For
The conclusion we draw from our own arithmetic. Ours.
Three points.
The task uses a true rule, any ascending triple, that is close to maximally general. Almost any hypothesis a participant forms will sit strictly inside it.
Which means the experiment selects the one configuration in which positive testing is guaranteed to be uninformative, and then reports that people use positive testing.
And an educational source draws the same conclusion: "What the classic tasks expose may be less a flaw in human reasoning. It looks more like a mismatch between an adaptive heuristic and an artificially rigged task."[3] We flag that this is that source's phrasing, and note our construction supports the structural half of it.
This Is Not A Defence
The necessary correction to the correction, because a reader could take the above too far. Ours.
Four points.
Nothing here shows people are unbiased. It shows one famous task cannot distinguish a motivational bias from a procedural habit.
The interpreting half of Nickerson's definition is untouched by any of this. How people weigh evidence once they have it is a separate question and this article does not address it.
Reference lists point to substantial work on that half, including Lord, Ross and Lepper (1979), Biased assimilation and attitude polarization: The effects of prior theories on subsequently considered evidence, Journal of Personality and Social Psychology, 37(11), 2098–2109[3][6]. We did not obtain it and report the citation only.
And the reappraisal does not say positive testing is always fine. It says the strategy is information-bearing under most realistic conditions, which is a claim about a distribution of environments, not a guarantee about yours.
A Contrary Result
A finding that cuts against the rehabilitation, reported for balance.
A study using a variant of the classic task in which some feedback was deliberately erroneous reports: "Our results show that, in contrast to previous research, people are equally adept at identifying false negatives and false positives; further, successful subjects were less likely to use a positive test strategy than were unsuccessful subjects."[7]
We did not obtain the paper and report this passage.
Two observations, ours.
The second clause matters: in that experiment, people who used the positive test strategy did worse. The strategy is not universally benign, and the source explicitly cites the 1987 reappraisal while reporting this.
And it introduces a variable the classic task lacks entirely, which is unreliable feedback. Real professional environments have exactly that, since the data you get back is often wrong.
Content Changes Everything
A related finding in the neighbouring task, recorded because it is the strongest evidence that these tasks measure something other than general reasoning.
Reference lists identify Griggs, R. A., and Cox, J. R. (1983), The effect of problem content on strategies in Wason's selection task, Quarterly Journal of Experimental Psychology, 35, 519–533[7].
We did not obtain it and report the title only.
One observation, ours. A title asserting that problem content changes strategy is doing the same work as our construction, from the empirical side. If performance moves with the dressing rather than the logic, the task is measuring the dressing.
What A Recent Paper Admits
The state of the field, in a recent paper's own words.
A source states: "Despite decades of research showing its influence on human reasoning and its association with polarization, we do not have a strong understanding of the factors that influence it, its boundaries, or its developmental trajectory."[2]
Two observations, ours.
That is a candid statement about the best-known cognitive bias there is, from within the literature studying it.
And its boundaries is the phrase that matters commercially. Knowing a tendency exists without knowing when it operates is not enough to design a control around it, which is why this article offers process suggestions rather than a magnitude.
Myside Bias
A terminological note worth having.
The same source records that the tendency to favour information supporting existing beliefs "has been termed as a 'myside/confirmation bias'", citing a list including Baron and colleagues (1993), Klayman (1995), Klayman and Ha (1987), Mercier, Nickerson (1998), Stanovich, and Wason (1960)[2].
Two observations, ours.
The slash in myside/confirmation is doing real work. It signals that a literature has two names for something and has not fully settled whether they are one thing.
And this is the twenty-eighth article's problem again, and the forty-second article's. Three times now this series has found a single popular label covering distinguishable phenomena, and each time the practical advice built on the label transferred badly.
The Pattern This Series Keeps Finding
Placing this alongside the rest. Ours.
Four instances, all from earlier articles.
The twenty-sixth reproduced a famous chart from noise with no underlying effect. The thirty-first reproduced a famous decline from rational time management. The thirty-fourth found a famous debunking rested on a statistical error. And this one finds a famous demonstration built on the single configuration where the behaviour it demonstrates cannot possibly work.
One observation. The common thread is not that the phenomena are fake. It is that the canonical demonstration of each turned out to be a weaker piece of evidence than its fame implies, and in every case the weakness was structural and checkable rather than requiring new data.
What Survives For Practice
The residue, ours, and it is smaller than the popular version but more usable.
Three statements we think hold.
People frame the next enquiry around the hypothesis they hold. That is the scholarly source's precise reformulation, and it is a claim about how attention is directed rather than about wanting to be right.
Whether that costs you anything depends on the configuration. If your hypothesis is narrower than the truth in the strict sense, confirming evidence is guaranteed and worthless. If it merely overlaps, the same tests are highly informative.
And the diagnostic question is not "am I being open-minded". It is "could this test have come out the other way". That is answerable, it is structural, and it is what our construction actually measures.
What To Do
Ask whether the test could have failed. A check whose answer you can predict resolves zero bits, and that is arithmetic rather than psychology.
Notice when your hypothesis is a subset of a broader truth. That is the configuration where confirming evidence accumulates without meaning anything, and it is the classic task's configuration.
Do not treat looking for supporting cases as automatically wrong. On our own construction the same strategy carries substantial information in most configurations, in one case the maximum a yes-or-no question can carry.
Separate seeking from interpreting. The standard definition covers both, they have different mechanisms, and this article addresses only the first.
Assume feedback may be wrong. One study introducing erroneous feedback found positive testers performed worse, which is the condition most professional environments actually have.
Do not build a control on a magnitude nobody has. A recent paper states the field lacks a strong understanding of the effect's boundaries.
Watch for a single label covering several things. This is the third time in this series, and each time the transferred advice was wrong for some of what the label covered.
Check famous demonstrations structurally. Four times now the canonical evidence has been weaker than its reputation, and each time it was checkable without new data.
The Limits Of This Analysis
Several caveats matter. This article discusses research on reasoning and is not audit, assurance or professional conduct advice. Everything is verified to August 2026. We did not obtain the 1987 paper, and have only its abstract from an educational research database plus others' characterisations of its argument. We did not obtain the 1960 task paper, and the description of the task reaches us through an educational website. We did not obtain the 1998 review beyond a truncated opening sentence, nor any of the follow-up studies, and report several by title and citation only. The clearest statement of the reframing comes from a non-peer-reviewed educational website, flagged at every use and relied on only where a scholarly source corroborates the same point. Our own construction does not reproduce that source's claim that a hypothesis narrower than the truth makes positive testing efficient; we get zero information in that case, offer two possible resolutions, and settle neither. All arithmetic and set constructions are ours, use invented illustrative sets, and demonstrate a structural property rather than reproducing anyone's data. We report no effect size for confirmation bias anywhere, and note a recent source stating the field lacks a strong understanding of its boundaries. This article addresses the seeking half of the standard definition and not the interpreting half.
Frequently Asked Questions
Is confirmation bias real?
What is wrong with the classic task?
So is positive testing fine?
Did your check agree with your sources?
How large is the effect?
What is the practical test?
References
- Educational research database record for Klayman, J., and Ha, Y.-W., Confirmation, Disconfirmation, and Information in Hypothesis Testing, Psychological Review, 1987, reproducing the paper's abstract, on it being proposed that many phenomena of human hypothesis testing can be understood in terms of a general positive test strategy, and on there being, with this strategy, a tendency to test cases that are expected to have the property of interest rather than those expected to lack that property. Note: an educational research database reproducing the abstract. We obtained the abstract only and not the paper, its analysis or its results. eric.ed.gov
- Repository page for the 1987 paper carrying scholarly citing text, on Klayman and Ha having provided the structural re-reading inside the cognitive-science literature itself, with participants executing a positive test strategy, generating cases the hypothesis predicts and reading the feedback against the prediction, and the strategy being information-bearing under most realistic task environments; on the bias the literature has named being the participant's reliance on the held hypothesis as the frame within which the next case is generated and interpreted; on the tendency to favour information supporting existing beliefs having been termed a myside or confirmation bias, citing Baron and colleagues (1993), Klayman (1995), Klayman and Ha (1987), Mercier, Nickerson (1998), Stanovich and Wason (1960); and on the field, despite decades of research showing its influence on human reasoning and its association with polarization, not having a strong understanding of the factors that influence it, its boundaries, or its developmental trajectory. Note: a repository page reproducing text from papers citing the 1987 article. Our scholarly corroboration for the reframing; we obtained none of the citing papers in full. researchgate.net
- Educational psychology website article on confirmation bias, on Klayman and Ha having argued that the classic task conflates two different things, with seeking confirmation being a goal and using a positive-test strategy, checking cases you expect to be positive, being just a method; on a hypothesised category being narrower than the true category in most real environments and testing positive cases then being an efficient way to expose errors; on the classic task being unusual because its true rule, any ascending triple, is maximally general, so positive tests of a narrower hypothesis can never disconfirm a rule that broad; and on what the classic tasks expose looking less like a flaw in human reasoning than a mismatch between an adaptive heuristic and an artificially rigged task; together with its reference list identifying Klayman and Ha (1987), Psychological Review, 94(2), 211–228; Lord, Ross and Lepper (1979), Journal of Personality and Social Psychology, 37(11), 2098–2109; Mynatt, Doherty and Tweney (1977), Quarterly Journal of Experimental Psychology, 29(1), 85–95; and Nickerson (1998), Review of General Psychology, 2(2), 175–220. Note: an educational website, not peer-reviewed. Flagged at every use in the body. Its claim that a narrower hypothesis makes positive testing efficient is not reproduced by our own construction, which is reported in the body rather than resolved. simplypsychology.org
- Publisher record for Nickerson, R. S. (1998), Confirmation Bias: A Ubiquitous Phenomenon in Many Guises, Review of General Psychology, 2(2), 175–220, reproducing the opening of the article, on confirmation bias, as the term is typically used in the psychological literature, connoting the seeking or interpreting of evidence in ways that are partial to existing beliefs, expectations, or a hypothesis in hand, our source truncating mid-word; together with reference list entries including Klayman and Ha (1987) and Kirby (1994) on the four-card selection task. Note: a publisher record; we obtained a truncated opening sentence and the reference list, and not the review. journals.sagepub.com
- Repository copy of the 1998 review's reference list, identifying Wason, P. C. (1960), On the failure to eliminate hypotheses in a conceptual task, Quarterly Journal of Experimental Psychology, together with Klayman, J., and Ha, Y-W. (1987), Psychological Review, 94, 211–228, and Koriat, Lichtenstein and Fischhoff (1980), Reasons for confidence, Journal of Experimental Psychology: Human Learning and Memory, 6, 107–118. Note: a reference list; citations only. We did not obtain the 1960 paper and the description of its task in this article comes from reference 3. academia.edu
- Encyclopedia of decision-making reference list confirming Klayman, J., and Ha, Y. (1987), Confirmation, disconfirmation, and information in hypothesis testing, Psychological Review, 94, 211–228, and identifying Lord, C. G., Ross, L., and Lepper, M. R. (1979), Journal of Personality and Social Psychology, 37, 2098–2109; Rabin, M., and Schrag, J. L. (1999), First impressions matter: A model of confirmatory bias; and Lefebvre, Deroy and Bahrami (2024), The roots of polarization in the individual reward system, Proceedings of the Royal Society B, 291. Note: a reference work's citation list, used to confirm the 1987 citation independently. We obtained none of the works listed. link.springer.com
- Repository copy of a study using a variant of the 1960 rule discovery task in which some feedback was subject to system error, with hits reported as misses and vice versa, on the potential for error meaning that evidence evaluation must include decisions about when to trust the data; on the results showing that, in contrast to previous research, people are equally adept at identifying false negatives and false positives; and on successful subjects having been less likely to use a positive test strategy, citing Klayman and Ha (1987), than were unsuccessful subjects; together with its reference list identifying Griggs, R. A., and Cox, J. R. (1983), The effect of problem content on strategies in Wason's selection task, Quarterly Journal of Experimental Psychology, 35, 519–533, and Wason, P. C. (1960). Note: a repository copy; we obtained these passages and not the study's methods or magnitudes. academia.edu
This article discusses research on reasoning and is not audit, assurance or professional conduct advice. No paper discussed was obtained in full. The clearest statement of the reappraisal comes from a non-peer-reviewed educational website and is flagged at every use. All set constructions and arithmetic are the authors' own, use invented illustrative sets, and demonstrate a structural property rather than reproducing anyone's data; one source's claim is not reproduced by that construction, which is reported in the body and left unresolved. No effect size is reported anywhere.