A bat and a ball cost a dollar ten. The bat costs a dollar more than the ball. Most readers of this publication already know that the answer is not ten cents, and that fact is the subject of this article.
Key Takeaway
A 2016 study of 142 online volunteers found that "over half of the sample had previously seen at least one of the problems" and that these participants "produced substantially higher CRT scores than those without prior exposure (2.36 vs. 1.48), with the majority scoring at ceiling level", concluding the test "may have been widely invalidated"[1]. A 2018 paper across six studies with roughly 2,500 participants and 17 variables reports: "we did not find a single case in which the predictive power of the CRT was significantly undermined by repeated exposure"[2].
The Verdict, Stated First
Five claims, in descending order of confidence.
One. The contamination is real and large. Scores rise substantially with prior exposure. This is not disputed by anyone in the literature; the 2018 defence explicitly reports replicating the score increase while denying its consequences.
Two. The contamination appears not to destroy the test's predictive validity, and our own arithmetic supports that. We simulated contamination independently and found that inflating scores toward ceiling in three quarters of a sample degraded the observed correlation with an outcome by roughly 18 percent, not by anything approaching invalidation. That was not the result we expected when we built the simulation.
Three. The defenders' explanation for why is the most interesting idea in this article. They propose the test survives repeated exposure because the people it identifies as unreflective fail to reflect on the fact that they have seen the questions before. The contamination is itself diagnostic.
Four. A three-item test is a coarse instrument regardless of contamination. Three binary items produce exactly four possible scores. On our own model, even with strong item loadings, such a scale captures a limited share of an underlying continuous trait, and no amount of defending it against exposure changes that.
Five. A firm should not use this to assess anyone. Not because the research is bad, which it is not, but because the specific conditions under which it works as a research instrument are absent in a workplace, and we set those out below.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A that prior exposure raises scores. Reported by the critical paper, replicated by the defending paper, and recorded by at least three further sources.
Grade A that exposure does not undermine predictive validity in the studies conducted. Six studies, roughly 2,500 participants, 17 outcome variables, abstract obtained verbatim from two independent sources.
Grade B for the contamination prevalence figure of over half a sample, which rests on 142 online volunteers in a single study.
Grade B that self-reported and actual exposure behave differently, from a paper we did not obtain whose title asserts non-effects.
Grade C for the numeracy confound, where we have a commentary and several citations and none of the underlying experiments.
Our position: this is a genuinely well-behaved dispute in which both sides are right about different things, and it is the clearest example in this series of a criticism that is factually correct and practically wrong.
A Note On Method
Everything here is verified to August 2026.
We obtained the critical paper's abstract verbatim from two independent sources plus its opening paragraphs[1][3]. We did not obtain its results section.
We obtained the defending paper's abstract verbatim from two independent sources plus part of its introduction[2][4]. We did not obtain its studies.
We obtained the three test items verbatim from a journal's published supplementary materials, with correct and intuitive answers as printed there[5].
We did not obtain the 2005 original paper and report its citation, confirmed across at least six independent reference lists, without quoting it[12].
We did not obtain the papers on actual versus self-reported exposure, the numeracy critique's underlying experiments, or any successor test, and report all from citation records and citing descriptions.
All arithmetic and simulation is ours, uses invented parameters, and cannot adjudicate the published dispute.
This article discusses a psychometric dispute. It is not hiring, assessment, human resources or employment advice, and the sections on workplace use are our own reasoning.
The Test Itself
All three items, reproduced as printed in a journal's supplementary materials, with the answers as given there.
"(1) A bat and a ball cost $1.10 in total. The bat costs a dollar more than the ball. How much does the ball cost? ____ cents [Correct answer: 5 cents; intuitive answer: 10 cents]"[5]
"(2) If it takes 5 machines 5 minutes to make 5 widgets, how long would it take 100 machines to make 100 widgets? ____ minutes [Correct answer: 5 minutes; intuitive answer: 100 minutes]"[5]
"(3) In a lake, there is a patch of lily pads. Every day, the patch doubles in size. If it takes 48 days for the patch to cover the entire lake, how long would it take for the patch to cover half of the lake? ____ days [Correct answer: 47 days; intuitive answer: 24 days]"[5]
The same source reproduces a fourth item from a later expansion: "(4) If John can drink one barrel of water in 6 days, and Mary can drink one barrel of water in 12 days, how long would it take them to drink one barrel of water together? _____ days"[5]
Two observations, ours.
Note that the source prints an intuitive answer alongside the correct one. That is the design: each item has a specific wrong answer it is engineered to produce, and the score is not merely accuracy but resistance to a particular pull.
And note that we have just reproduced the test in full, which is itself the phenomenon this article is about. A publication writing about contamination contributes to it, and we return to that.
Why These Three Questions
The mechanism the items exploit, described by the critical paper's own framing.
The 2016 paper opens by locating the test within dual-process theory: "Dual process models of human cognition typically make a distinction between fast and autonomous Type 1 thinking and slower, consciously controlled Type 2 thinking. The advantage of Type 1 thinking is that it produces quick and approximate solutions to a given problem at little computational expense. However, while this system often provides good enough responses, it is susceptible to being misled. One of the main roles of Type 2 thinking is to reflect on and override such intuitive, but incorrect responses."[1]
Three observations, ours.
The test is built directly on the framework the previous article in this series examined at length, where a 2018 critique called the two-type typology a convenient and seductive myth and five researchers replied that the contested features were never definitional.
Which means an instrument's validity can be assessed independently of the theory that motivated it. Whether or not two types exist, the CRT either predicts things or it does not, and the papers in this article are about the second question.
And the design is genuinely clever regardless. Each item presents an answer that arrives before you have done any work, and getting it right requires noticing that an answer has arrived and declining to use it.
What It Claims To Measure
The construct, from the defending paper's own description.
It describes the test as "a widely used measure of the propensity to engage in analytic or deliberative reasoning in lieu of gut feelings or intuitions", and the items as "unique because they reliably cue intuitive but incorrect responses and, therefore, appear simple among those who do poorly."[2]
On its evidential standing, the same paper states: "Although the CRT obviously requires some degree of numeracy, it is thought to also capture the propensity to think analytically. That is, those who do well on the CRT are also less prone to rely on heuristics and biases even after measures of cognitive ability have been taken into account. Moreover, the CRT predicts a wide range of variables after controlling for numeracy."[4]
Three observations, ours.
The clause "appear simple among those who do poorly" is the design's cleverest feature and, as we will see, the defenders' whole explanation for why contamination fails to break it.
"Even after measures of cognitive ability have been taken into account" is the claim that makes the test interesting. If it merely measured intelligence it would be redundant.
And note the concession in the first clause: the test "obviously requires some degree of numeracy." That admission sits at the centre of a separate critique we come to below.
How It Became Ubiquitous
The reason contamination became possible at all.
The defending paper explains: "Perhaps because it is short and strongly predictive of a variety of outcome variables, the CRT has become a nearly ubiquitous psychological test."[4]
The critical paper describes the test as "a hugely influential problem solving task"[1].
Two observations, ours.
Short and predictive is a rare combination, and it is exactly the combination that guarantees overuse. A three-item instrument costs almost nothing to administer, which is why it ended up in hundreds of studies and, eventually, in newspapers.
And there is a structural irony here worth naming, because it recurs across this series. The property that made the instrument valuable is the property that degraded it. A test nobody used would remain uncontaminated and useless.
The Contamination Finding
The paper that raised the alarm.
Haigh, M. (2016), Has the Standard Cognitive Reflection Test Become a Victim of Its Own Success?, Advances in Cognitive Psychology, 12(3), 145–149[1][3].
Its abstract, in full: "The Cognitive Reflection Test (CRT) is a hugely influential problem solving task that measures individual differences in the propensity to reflect on and override intuitive (but incorrect) solutions. The validity of this three-item measure depends on participants being naïve to its materials and objectives. Evidence from 142 volunteers recruited online suggests this is often not the case. Over half of the sample had previously seen at least one of the problems, predominantly through research participation or the media. These participants produced substantially higher CRT scores than those without prior exposure (2.36 vs. 1.48), with the majority scoring at ceiling level. Participants that had previously seen a specific problem (e.g., the bat and ball problem) nearly always solved that problem correctly. These data suggest the CRT may have been widely invalidated. As a minimum, researchers must control for prior exposure to the three problems and begin to consider alternative, extended measures of cognitive reflection."[1]
Four observations, ours.
The second sentence contains the paper's real argument and it is a conditional: "The validity of this three-item measure depends on participants being naïve to its materials and objectives." Everything else follows from that premise, and the 2018 defence attacks precisely that premise.
"Predominantly through research participation or the media" names the two contamination routes, and the second is the one nobody can control.
The finding that exposed participants "nearly always solved that problem correctly" is item-level rather than aggregate, which makes it harder to explain away as a selection effect.
And the recommendation is measured rather than apocalyptic: control for prior exposure, and consider extended measures. The paper's title is more dramatic than its conclusion.
Two Point Three Six Against One Point Four Eight
The headline gap, and what it means on a four-point scale.
Exposed participants averaged 2.36 out of 3. Unexposed averaged 1.48[1].
Three observations, ours.
The gap is 0.88 of a possible 3, which is 29 percent of the entire range of the instrument. On a scale with only four possible values, that is close to a full category.
The unexposed mean of 1.48 sits almost exactly at the midpoint, which is what a well-designed test should produce and is a point in the instrument's favour.
And the exposed mean of 2.36 is close enough to the maximum that the distribution must be compressed against it, which is the ceiling problem the paper names and the next section quantifies.
What A Ceiling Does
The consequence that matters most and which the paper states but does not develop. Ours.
The paper reports the majority of exposed participants "scoring at ceiling level"[1].
Three observations.
A person scoring 3 out of 3 cannot be distinguished from any other person scoring 3 out of 3. The instrument has no resolution above its maximum.
So if 55 percent of a sample is at ceiling, the test discriminates within only 45 percent of that sample. At 70 percent, only 30 percent. The proportion at ceiling is the proportion about whom the test says nothing.
And this is a distinct problem from the score inflation. Inflation biases a mean; a ceiling destroys variance, and it is variance that correlations are made of. That distinction is why the next two sections matter.
Four Buckets
A limitation that applies even to a perfectly uncontaminated administration. Our own simulation, invented parameters, and it is a general property of short scales rather than a finding about this test.
Three binary items produce exactly four possible scores: 0, 1, 2 and 3.
We modelled a continuous underlying trait and three items each reading it with noise, at an invented loading of 0.75, across 200,000 simulated people.
The correlation between the latent trait and the three-item score came out at 0.781, meaning the score captured about 61 percent of the trait's variance. The score distribution was roughly even: 26.8 percent at zero, 23.4 at one, 23.2 at two, 26.7 at three.
Three observations.
Even under generous assumptions, roughly two fifths of the underlying variation is lost simply by squeezing a continuum into four buckets. That is not a criticism anyone in this dispute makes and it is not contamination.
The 0.75 loading is entirely ours and invented. No source we obtained states item loadings, and a different loading gives a different answer. This shows what short binary scales do in principle.
And note the flat distribution, with over a quarter at each extreme. A test where 26.7 percent score at maximum has a ceiling problem before anyone has ever seen the questions.
We Simulated The Contamination
Because the central question of this dispute is what contamination does to a correlation, and that is answerable by construction. Our own simulation, invented parameters throughout; this cannot adjudicate the published dispute and is not offered as doing so.
We generated a latent trait, measured it with three noisy binary items, generated an outcome the trait genuinely predicts, and then contaminated a share of the sample by inflating their scores toward ceiling independently of their true trait, which is the most damaging form contamination could take.
With no contamination, the observed correlation between score and outcome was 0.388.
With a quarter contaminated: 0.356.
With half: 0.334.
With three quarters: 0.318.
Three observations, and the first is the one we did not expect.
Contaminating three quarters of the sample with score inflation entirely unrelated to the trait reduced the observed correlation from 0.388 to 0.318. That is a real degradation and it is nothing like invalidation.
The reason is arithmetic rather than psychological. Contamination adds noise to the measure; noise attenuates a correlation; attenuation is gradual. A measure has to be almost entirely noise before its correlations vanish.
And this result independently reproduces the direction of the empirical defence published in 2018, which we had read before building the simulation but whose magnitude we had not anticipated. We report that we expected the contamination to matter more than our own arithmetic says it does.
And We Corrected Our Own Percentages
A note on our process, continuing a practice from the previous three articles.
Our simulation script printed percentage changes computed against a hardcoded baseline of 0.379, while the baseline it actually produced in the same run was 0.388. The printed comparisons were therefore slightly wrong.
Recomputed against the correct baseline, the degradations are 8 percent at a quarter contaminated, 14 percent at half, and 18 percent at three quarters. Those are the figures used above and in the summary at the top of this article.
One observation, ours. The error was in a comparison, not in a measurement, which is the most common kind of arithmetic mistake and the easiest to publish without noticing. It was caught by comparing the printed baseline against the printed percentages, which took under a minute.
Two Thousand Five Hundred Participants
The empirical answer to the contamination alarm.
Bialek, M., and Pennycook, G. (2018), The cognitive reflection test is robust to multiple exposures, Behavior Research Methods, 50, 1953–1959[2][6].
Its abstract sets up the question fairly: "By virtue of being composed of so-called 'trick problems' that, in theory, could be discovered as such, it is commonly held that the predictive validity of the CRT is undermined by prior experience with the task. Indeed, recent studies have shown that people who have had previous experience with the CRT score higher on the test. Naturally, however, it is not obvious that this actually undermines the predictive validity of the test."[2]
And its result: "Across six studies with ~ 2,500 participants and 17 variables of interest (e.g., religious belief, bullshit receptivity, smartphone usage, susceptibility to heuristics and biases, and numeracy), we did not find a single case in which the predictive power of the CRT was significantly undermined by repeated exposure. This occurred despite the fact that we replicated the previously reported increase in accuracy among individuals who reported previous experience with the CRT."[2]
Four observations, ours.
Not a single case out of seventeen variables is a strong claim, and the scale, roughly 2,500 participants against 142, is close to eighteen to one.
The paper replicates the score increase. Both sides agree on the fact and disagree on its consequence, which is the cleanest form a scientific dispute can take.
The framing sentence deserves attention: "it is not obvious that this actually undermines the predictive validity." That is the logical gap the critical paper leaves, and the defenders walk straight into it.
And the introduction states the gap even more directly: "Although there is strong agreement that multiple exposures to the CRT invalidate it as a test, remarkably, no studies (that we are aware of) have empirically tested this claim."[7] A field-wide consensus that nobody had checked.
The Explanation Is The Elegant Part
Why the test survives, in the defenders' own words, and it is the best idea in this article.
"We speculate that the CRT remains robust after multiple exposures because less reflective (more intuitive) individuals fail to realize that being presented with apparently easy problems more than once confers information about the task's actual difficulty."[6]
Four observations, ours.
Read it slowly, because the argument turns on itself. Seeing the questions a second time is itself a cue that they are harder than they look. Noticing that cue requires exactly the reflectiveness the test measures.
So the reflective person is helped by exposure and the unreflective person is not, which preserves the ordering even as it lifts the scores. The contamination inflates the mean without scrambling the ranking, and correlations depend on ranking.
This is why our own simulation understates the defenders' case rather than overstating it. We modelled contamination as inflation applied at random, independently of the trait. If the defenders are right, contamination is not random; it is correlated with the trait, which would degrade the correlation less than our figures suggest.
And the authors mark it clearly: "we speculate." This is a proposed mechanism offered after the empirical result, not a tested one, and we grade it accordingly.
Is That Argument Circular
The obvious objection, which we think fails but not trivially. Ours.
The worry: the explanation assumes the test measures reflectiveness in order to explain why the test still measures reflectiveness. If someone doubted the construct entirely, this would persuade them of nothing.
Three points on why we think it survives.
The argument is offered as an explanation of an already-established empirical result, not as evidence for it. The result is that predictive power held across seventeen variables; the speculation is about why. Explanations are permitted to use the theory.
It also makes a testable prediction: exposure should benefit high scorers more than low scorers, so the exposure effect should interact with baseline reflectiveness. A separate study is described as having found that exposure increased the test's predictive power in heuristics and biases tasks, albeit only among high need-for-cognition individuals[8], which is an interaction of that general shape though not the same variable.
And the honest caveat: we did not obtain that study either, and one interaction in one sample is not a confirmation.
The Creator's Own Paper
A development that complicates both sides, and it carries an unusual signature.
Reference lists identify Meyer, A., Zhou, E., and Frederick, S. (2018), The non-effects of repeated exposure to the cognitive reflection test, Judgment and Decision Making, 13, 246–259[9].
A later paper summarises its finding: "nor does actual (not self-reported) prior exposure increase performance on it (Meyer, Zhou, & Frederick, 2018)."[10]
We did not obtain the paper and report its title, citation, and one citing source's summary of its finding.
Three observations, ours.
The third author is Shane Frederick, whose 2005 paper introduced the test. The instrument's creator co-authored a paper reporting that repeated exposure has non-effects, which is an unusual and creditable thing for an author to publish about their own instrument.
The distinction it draws is subtle and important: actual exposure, as opposed to self-reported exposure. Those are different variables and they need not behave alike.
And if that finding holds, it partly dissolves the original alarm rather than answering it. The contamination effect may be substantially an artefact of who reports having seen the questions, which is the subject of the next section.
Self-Reported Versus Actual
The measurement subtlety this whole dispute rests on, and which almost no summary of it mentions. Ours.
Every contamination estimate quoted above comes from asking people whether they had seen the problems before. The 2016 study recruited volunteers and asked; the 2018 defence reports replicating the increase "among individuals who reported previous experience"[2].
Four observations.
Remembering that you have seen a puzzle before is itself a cognitive act, and there is no reason to assume it is uncorrelated with the trait being measured. A more reflective person may simply be better at recognising and reporting prior exposure.
If that is so, then the exposed group is not a random subset of the sample. It is a group selected partly on the very attribute the test measures, and the 2.36 against 1.48 gap would then be partly a selection effect rather than purely a learning effect.
Which is exactly what a finding of non-effects for actual exposure alongside real effects for self-reported exposure would imply. The two variables diverge because one of them is contaminated by the trait.
And we want to be careful about how far we push this. We did not obtain the paper making the actual-exposure finding, we have it through one citing sentence, and this reasoning is ours rather than any source's. It is a hypothesis consistent with the reported pattern, not a demonstrated account.
The Numeracy Confound
A separate line of criticism that has nothing to do with exposure.
A 2015 commentary describes the debate: it compares "(i) the hypothesis that Cognitive Reflection mirrors the human ability of suppressing automatic answers in favor of deliberate ones, with (ii) the hypothesis that numerical ability alone is able to predict superior decision making and to account for Cognitive Reflection Test results."[11]
The defending paper concedes the premise, as quoted earlier: the test "obviously requires some degree of numeracy", while maintaining that it "predicts a wide range of variables after controlling for numeracy."[4]
And a 2021 paper developing an alternative states the problem persists, citing "the issue of numeracy confounding" as a reason to build a verbal replacement[10].
Three observations, ours.
This critique is logically prior to the contamination one. If the test measures arithmetic ability under a different name, its exposure properties are a secondary concern.
The defenders' answer is a statistical control, which is a legitimate reply and not a complete one. Controlling for measured numeracy removes the variance numeracy explains, not the variance it shares with the construct.
And the fact that researchers built a verbal version specifically to remove the confound tells you the field did not regard the control as settling it. We did not obtain any of the numeracy experiments and grade this C.
The Successor Tests
What the field built in response, which is the constructive part.
The defending paper records: "As a consequence, much effort has gone into finding newer versions of the CRT."[7] Reference lists identify Primi, Morsanyi, Chiesi, Donati and Hamilton (2016), Thomson and Oppenheimer (2016), and Toplak, West and Stanovich (2014) as sources of such versions[7], and a 2021 paper in the Journal of Behavioral Decision Making presents a verbal cognitive reflection test designed to address numeracy confounding while exhibiting "low recognisability"[10].
We obtained none of these instruments and report their existence and stated purposes.
Three observations, ours.
Low recognisability is now an explicit design criterion. That is a field learning from a specific failure, and it is the correct response.
The 2014 expansion adds items rather than replacing them, which addresses the four-buckets problem our simulation identified as separate from contamination. More items, more resolution.
And there is an obvious tension nobody can escape: a successor test that becomes popular will become recognisable. The contamination clock restarts with each replacement, and the only structural fix is item pools large enough that exposure to any one item matters little.
What Actually Survives
Our reading, stated directly.
Five statements.
Prior exposure raises scores substantially. Both sides agree. The reported gap is 2.36 against 1.48 on a three-point scale.
It does not appear to break predictive validity. Six studies, roughly 2,500 participants, seventeen variables, not one case of significant undermining. Our own independent simulation gives a degradation of about 18 percent at three-quarters contamination, which is the same qualitative answer.
The mechanism proposed for that is speculative but coherent and testable. Noticing that a familiar puzzle is a warning requires the disposition being measured.
The instrument is coarse independent of all of this. Four possible scores, over a quarter at ceiling on our own model before anyone has seen anything.
And the numeracy question is unresolved and more fundamental. A confound in what the test measures matters more than a confound in who has seen it.
Should A Firm Use This
The practical question, and our answer is no. Ours, and the reasoning is ours rather than any source's.
Five reasons, and note that none of them is that the research is weak, because it is not.
The defence is about aggregate predictive validity, not individual assessment. A measure can correlate usefully across 2,500 people and tell you almost nothing about one person, and that distinction is the whole gap between a research instrument and an assessment tool.
Four possible scores cannot support an individual judgment. On our own model a three-item binary scale captured about 61 percent of an underlying trait's variance under generous assumptions, which is a serviceable research measure and an indefensible basis for a decision about a person.
Your candidates are the contaminated population. The 2016 study found exposure came predominantly through research participation or the media[1]. Anyone who has read a popular behavioural economics book has seen all three items.
The defenders' protective mechanism does not apply. Their explanation is that unreflective people fail to notice the significance of having seen the puzzle before. A candidate who knows they are being assessed has every incentive to prepare, which removes the condition the defence depends on.
And the numeracy confound is a live legal and practical problem in selection, quite apart from its scientific status, because a measure that partly indexes arithmetic ability is a measure of something you should be assessing openly if you intend to assess it at all.
And In Hiring Specifically
Extending that, because this is where such instruments actually get used. Ours, untested, and not employment advice.
Four points.
This series' sixth article covered a large revision of the hiring validity literature in which structured approaches to selection were central. Nothing in the present article displaces that, and a three-item puzzle set is not a structured selection procedure.
The thirty-second article's argument applies directly: measure the decisions, not the decision-makers. If you want to know whether someone reflects before answering, the evidence is in their work product, which you can examine.
And there is a specific failure mode worth naming. A candidate who answers "five cents" instantly has demonstrated recall, not reflection, and the instrument cannot tell those apart. The score is identical.
Which means, on the defenders' own logic, the test works precisely where the taker does not know it is a test. That condition is unobtainable in a selection process by definition.
What To Do
Do not use the CRT to assess anyone. The defence concerns aggregate predictive validity across thousands of people, not individual measurement, and your candidates are drawn from the contaminated population.
Do not conclude the research is unsound. Six studies, roughly 2,500 participants, seventeen variables, no case of undermined prediction. This is a criticism that is factually correct and practically overstated.
Separate the two critiques. Contamination is about who has seen the items. Numeracy confounding is about what the items measure. The second is more fundamental and less resolved.
Notice that both sides agree on the fact. Scores rise with exposure. The dispute is entirely about whether that matters, which is the cleanest form a disagreement can take.
Distrust a field-wide consensus nobody has tested. The defending paper records strong agreement that exposure invalidates the test, and that no study had empirically examined the claim.
Ask whether a measure has enough levels to support the decision. Four possible scores is a research instrument. It is not a basis for a judgment about a person.
Watch for self-report standing in for a variable. Every contamination estimate here rests on people remembering and reporting prior exposure, and remembering is not obviously independent of the trait being measured.
Expect successor instruments to have the same lifecycle. Low recognisability is now a design criterion, and a replacement that becomes popular becomes recognisable.
The Limits Of This Analysis
Several caveats matter. This article discusses a psychometric dispute and is not hiring, assessment, human resources or employment advice; the sections on workplace use are our own reasoning and are supported by no study we obtained. Everything is verified to August 2026. We obtained the critical paper's abstract verbatim from two independent sources and its opening paragraphs, and not its results section, so we report its headline figures without having seen the analysis producing them. We obtained the defending paper's abstract verbatim from two independent sources and part of its introduction, and none of its six studies; we report no effect size from it. We did not obtain the 2005 paper that introduced the test, and report its citation, confirmed across at least six independent reference lists, without quoting it. We did not obtain the paper reporting non-effects of actual exposure, which is a load-bearing item in this article, and have it through a single citing sentence in a later paper. We did not obtain any of the numeracy experiments, any successor instrument, or the study reporting an interaction with need for cognition. All arithmetic and simulation is ours and uses invented parameters throughout, including an item loading of 0.75 that no source states; it demonstrates general properties of short binary scales and of measurement noise, and cannot adjudicate the published dispute. Our simulation models contamination as inflation applied at random, which the defenders' proposed mechanism suggests is the wrong model and which would, if they are right, overstate the damage. One set of percentage figures in our own output was computed against a wrong baseline and is corrected in the body. The reasoning about self-reported versus actual exposure is our own hypothesis, consistent with the reported pattern and demonstrated by nothing we obtained. And this article reproduces the three test items in full, which contributes to the exposure it describes; we judged that a reader cannot assess the dispute without seeing what is disputed, and we note the cost.
Frequently Asked Questions
Has the test been ruined by familiarity?
Why would contamination not break it?
What did your own arithmetic find?
Is there a problem beyond contamination?
Should I use it to screen candidates?
Did the test's creator weigh in?
Doesn't this article make the problem worse?
References
- Haigh, M. (2016). Has the Standard Cognitive Reflection Test Become a Victim of Its Own Success? Advances in Cognitive Psychology, 12(3), 145–149, DOI 10.5709/acp-0193-5, published 30 September 2016, author at the Department of Psychology, Northumbria University. Full-text repository copy reproducing the abstract and introduction: on the Cognitive Reflection Test being a hugely influential problem solving task that measures individual differences in the propensity to reflect on and override intuitive but incorrect solutions; on the validity of this three-item measure depending on participants being naïve to its materials and objectives; on evidence from 142 volunteers recruited online suggesting this is often not the case; on over half of the sample having previously seen at least one of the problems, predominantly through research participation or the media; on those participants having produced substantially higher CRT scores than those without prior exposure, 2.36 against 1.48, with the majority scoring at ceiling level; on participants who had previously seen a specific problem nearly always solving that problem correctly; on these data suggesting the CRT may have been widely invalidated; on the recommendation that researchers must as a minimum control for prior exposure and consider alternative, extended measures; and on the introduction's framing of dual process models distinguishing fast and autonomous Type 1 thinking from slower, consciously controlled Type 2 thinking, with one of the main roles of Type 2 thinking being to reflect on and override intuitive but incorrect responses. Note: a full-text repository copy. We obtained the abstract and introduction and not the results section, so the headline figures are reported without our having seen the analysis producing them. ncbi.nlm.nih.gov
- Bialek, M., & Pennycook, G. (2018). The cognitive reflection test is robust to multiple exposures. Behavior Research Methods. Publisher record reproducing the abstract: on the CRT being a widely used measure of the propensity to engage in analytic or deliberative reasoning in lieu of gut feelings or intuitions; on CRT problems being unique because they reliably cue intuitive but incorrect responses and therefore appear simple among those who do poorly; on it being commonly held, by virtue of the items being trick problems that could in theory be discovered as such, that the predictive validity of the CRT is undermined by prior experience; on recent studies having shown that people with previous experience score higher; on it not being obvious that this actually undermines predictive validity; and on the authors, across six studies with approximately 2,500 participants and 17 variables of interest including religious belief, bullshit receptivity, smartphone usage, susceptibility to heuristics and biases, and numeracy, not having found a single case in which the predictive power of the CRT was significantly undermined by repeated exposure, despite having replicated the previously reported increase in accuracy among individuals who reported previous experience. Note: the publisher's record. We obtained the abstract in full and none of the six studies, and report no effect size from the paper. link.springer.com
- Journal record for the same 2016 paper, Advances in Cognitive Psychology, 12(3), 145–149, PMID 28115997, reproducing the abstract identically and recording the article's keywords as Cognitive Reflection Test, CRT, bat and ball problem, validity, and test security. Note: a second independent reproduction of the abstract, used to confirm the figures and the journal, volume, issue and page details. pmc.ncbi.nlm.nih.gov
- Publisher record for the 2018 defending paper reproducing part of its introduction: on the CRT obviously requiring some degree of numeracy while being thought also to capture the propensity to think analytically; on those who do well on the CRT being less prone to rely on heuristics and biases even after measures of cognitive ability have been taken into account, citing Toplak, West and Stanovich (2011, 2014); on the CRT predicting a wide range of variables after controlling for numeracy; on the CRT having become a nearly ubiquitous psychological test, perhaps because it is short and strongly predictive of a variety of outcome variables; and on it being widely held that trick problems such as the bat-and-ball problem will not be robust to multiple testing because participants will realise the problems only seem easy at first blush, citing Haigh (2016) and Stieger and Reips (2016). Note: the publisher's record, from which we obtained introductory passages. We did not obtain the paper's studies or analyses. link.springer.com
- Published supplementary materials for a journal article, reproducing the three Cognitive Reflection Test items as taken from Frederick (2005), with correct and intuitive answers as printed: the bat and ball problem with a correct answer of 5 cents and intuitive answer of 10 cents; the machines and widgets problem with a correct answer of 5 minutes and intuitive answer of 100 minutes; and the lily pad problem with a correct answer of 47 days and intuitive answer of 24 days; together with a fourth item taken from Toplak, West and Stanovich (2014) concerning two people drinking a barrel of water. Note: a journal's published supplementary file. Our source for the item wording; the attribution of items one to three to Frederick (2005) and the fourth to Toplak, West and Stanovich (2014) is as printed there. frontiersin.org
- National medical literature database record for the 2018 defending paper, PMID 28849403, reproducing the abstract and its closing passage: on the authors speculating that the CRT remains robust after multiple exposures because less reflective, more intuitive individuals fail to realise that being presented with apparently easy problems more than once confers information about the task's actual difficulty; and recording the article's keywords as CRT, Cognitive reflection test, Dual-process theory, Intuition and Reflection. Note: a bibliographic record; our source for the proposed mechanism, which the authors themselves mark as speculation. pubmed.ncbi.nlm.nih.gov
- Repository page for the 2018 defending paper reproducing further introductory text: on prior experience with the CRT being associated with higher scores, citing Haigh (2016) and Stieger and Reips (2016); on this casting doubt on whether the test can continue to be used as a valid tool for assessing analytic cognitive style; on much effort consequently having gone into finding newer versions of the CRT, citing Primi, Morsanyi, Chiesi, Donati and Hamilton (2016), Thomson and Oppenheimer (2016) and Toplak and colleagues (2014); and on there being strong agreement that multiple exposures to the CRT invalidate it as a test while, remarkably, no studies that the authors are aware of having empirically tested this claim. Note: a repository page carrying the paper's introduction. We obtained none of the successor instruments named. researchgate.net
- Repository page carrying the abstract of a follow-up study testing whether the exposure relationship is moderated by analytic thinking, on participants numbering 365 having completed the CRT, a Need for Cognition scale and a battery of heuristics and biases problems; on the CRT retaining its predictive power in that performance regardless of self-reported thinking dispositions and exposure; and on both factors nonetheless moderating the relationship, such that exposure increased the CRT's predictive power in heuristics and biases tasks, albeit only among high need-for-cognition individuals. Note: a repository page reproducing a study abstract. We did not obtain the study, and one interaction in one sample is not a confirmation of the proposed mechanism. researchgate.net
- Academic preprint reference list identifying Meyer, A., Zhou, E., & Frederick, S. (2018), The non-effects of repeated exposure to the cognitive reflection test, Judgment and Decision Making, 13, 246–259; and Bialek, M., & Pennycook, G. (2018), The cognitive reflection test is robust to multiple exposures, Behavior Research Methods, 50, 1953–1959. Note: a reference list; citations only, used to establish the journal, volume and pages of both papers. We did not obtain the 2018 non-effects paper. arxiv.org
- Publisher record for a 2021 article in the Journal of Behavioral Decision Making presenting a verbal cognitive reflection test, on an increasing proportion of participants tested with the CRT already being familiar with it, both in terms of prior exposure and knowledge of the items; on self-reported prior exposure substantially increasing performance, citing Bialek and Pennycook (2017), Haigh (2016) and Stieger and Reips (2016); on recent studies suggesting that self-reported prior exposure does not diminish the predictive validity of the test, citing Bialek and Pennycook (2017) and Šrol (2018b); on actual, as opposed to self-reported, prior exposure not increasing performance on it, citing Meyer, Zhou and Frederick (2018); on the effect of knowledge of the items on predictive validity remaining unclear despite the reassuring findings on prior exposure; and on there being a need for alternative measures that would address the issue of numeracy confounding while exhibiting excellent psychometric characteristics and low recognisability. Note: a publisher record. Our source for the actual-versus-self-reported distinction, which reaches this article through a single citing sentence; we did not obtain the paper making that finding. onlinelibrary.wiley.com
- Mastrogiorgio, A. (2015). Commentary: Cognitive reflection vs. calculation in decision making. Frontiers in Psychology, DOI 10.3389/fpsyg.2015.00936, published 3 July 2015. Full-text repository copy, on the commentary addressing a debate comparing the hypothesis that cognitive reflection mirrors the human ability of suppressing automatic answers in favour of deliberate ones against the hypothesis that numerical ability alone is able to predict superior decision making and to account for Cognitive Reflection Test results; and on the author's caution that evaluating reasoning abilities requires the ecological caveat that tasks are embodied in specific, real-world environments; together with its reference list confirming Frederick, S. (2005), Cognitive reflection and decision making, Journal of Economic Perspectives, 19, 25–42, DOI 10.1257/089533005775196732. Note: a commentary, not an experimental paper. We obtained none of the underlying numeracy experiments it discusses and grade that critique C accordingly. ncbi.nlm.nih.gov
- Bibliographic database records independently confirming Shane Frederick (2005), Cognitive Reflection and Decision Making, Journal of Economic Perspectives, American Economic Association, volume 19, issue 4, pages 25–42, Fall, as cited across multiple unrelated reference lists; together with Campitelli, G., and Labollita, M. (2010), Correlations of cognitive reflection with judgments and choices, Judgment and Decision Making, 5(3), 182–191, whose abstract reports measuring 157 participants and concluding that cognitive reflection is a thinking disposition including more characteristics than originally proposed, related to actively open-minded thinking; and Jimenez, N., Rodriguez-Lara, I., Tyran, J.-R., and Wengström, E. (2018), Thinking fast, thinking badly, Economics Letters, 162, 41–44. Note: bibliographic records and one abstract; citations used to confirm the 2005 paper's details, which we did not obtain and do not quote. ideas.repec.org
This article discusses a psychometric dispute and is not hiring, assessment, human resources or employment advice. Neither principal paper was obtained in full: the critical paper's abstract and introduction were obtained but not its results, and the defending paper's abstract and part of its introduction were obtained but none of its six studies. The 2005 paper introducing the test was not obtained and is not quoted. All arithmetic and simulation is the authors' own, uses invented parameters throughout, and cannot adjudicate the published dispute; one set of percentage figures in that output was computed against a wrong baseline and is corrected in the body. This article reproduces the three test items in full, which contributes to the exposure it describes.