Every firm that screens for anything faces the same arithmetic, and it is deeply counterintuitive. A test that is genuinely good will produce mostly false alarms whenever the thing being tested for is rare, and no amount of care by the person reading the results can change that.
Key Takeaway
Asked "If a test to detect a disease whose prevalence is 1/1000 has a false positive rate of 5%, what is the chance that a person found to have a positive result actually has the disease?", only a minority provided the correct answer, with most clinicians overestimating the positive predictive value[1]. The answer is about 2 percent. The study was replicated 36 years later, yielding similar results[1]. In a firm, at a 1 percent underlying error rate, a test with 90 percent sensitivity and a 5 percent false positive rate produces about 59 flags per 1,000 files of which 9 are real.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the arithmetic. The positive predictive value calculation is not a finding; it is a definition, and we work it below.
Grade A that trained professionals get this wrong, on a 1978 study in a leading medical journal that was replicated thirty-six years later with similar results.
Grade B for natural frequency formats improving performance, from a paper in Psychological Review, noting that the claim drew published replies.
Grade C for the strong version of base rate neglect, because a published reconsideration exists and one source states the original story was oversold.
Our position: this is the only article in this series where the central content is arithmetic rather than a finding, and it is therefore the one least likely to be overturned.
A Note On Method
Everything here is verified to August 2026.
We obtained the 1978 question verbatim and its result summary from a peer-reviewed journal article describing it[1], and passages from the 1995 frequency formats paper itself[2].
We did not obtain the 1978 paper, the 2014 replication, or any of the other studies discussed, and report citations and characterisations only.
We found two inconsistencies in how the 1978 paper is cited, and report both.
All arithmetic is ours except the 1 in 1,000 prevalence and 5 percent false positive rate, which come from the published question.
This article discusses probability and screening. It is not audit, assurance, tax, medical or professional standards advice, and nothing in it describes any regulator's or professional body's requirements.
The Question
The item, verbatim, from a peer-reviewed source.
"If a test to detect a disease whose prevalence is 1/1000 has a false positive rate of 5%, what is the chance that a person found to have a positive result actually has the disease, assuming that you know nothing about the person's symptoms or signs?"[1]
The paper is Casscells, Schoenberger and Graboys (1978), Interpretation by physicians of clinical laboratory results, New England Journal of Medicine, 299(18), 999–1001, DOI 10.1056/NEJM197811022991808[3].
Two sourcing notes, ours.
Sources disagree on the page range, most giving 999–1001 and two giving 999–1000[4][5].
And on the third author's name, which appears as Graboys in most sources and Grayboys in others, including in the 1995 paper's own citation of it[2][4]. This is the ninth time in this series a famous paper's details have been recorded inconsistently by careful sources.
The Answer
The arithmetic, in the format that makes it easy. Ours, using the published question's figures.
Take 10,000 people.
10 have the disease, at a prevalence of 1 in 1,000. Assume the test finds all of them.
9,990 do not have it. Five percent of those test positive anyway, which is about 500 people.
So there are about 510 positive results, of which 10 are genuine. That is 2.0 percent.
The reported result: only a minority provided the correct answer, with most clinicians overestimating the positive predictive value[1]. A commercial source reports that most physicians confused the test's 5% false-positive rate with the patient's odds, answering 95% instead of about 2%[4], and another gives the sample as 60 physicians and medical students with only 18% arriving at the correct answer[6]. We could not verify either of those two figures against the paper.
Two observations, ours.
The common wrong answer, 95 percent, is the complement of the false positive rate. It treats a property of the test as a property of the patient.
And notice that the correct arithmetic above required no formula. That fact turns out to matter enormously, and it is the subject of the section after next.
And It Replicated
The detail that separates this from most of the series.
A peer-reviewed source records: "The study was replicated by Manrai et al. 36 years later, yielding similar results, highlighting that medical statistics continue to challenge healthcare professionals, irrespective of grade, despite advancements in medical education."[1]
We did not obtain the replication and report only this description.
Three observations, ours.
Thirty-two articles into a series that has repeatedly found famous results failing to replicate, this one held, in a different cohort, decades later.
Irrespective of grade is the important phrase. It was not a student problem.
And despite advancements in medical education is the discouraging one. The obvious remedy, teaching people more statistics, was tried across an intervening generation and the result did not move.
The Original Demonstration
Where the phenomenon was first shown, and it was not medical.
Reference lists identify Kahneman, D., and Tversky, A. (1973), On the psychology of prediction, Psychological Review, 80(4), 237–251[5].
A behavioural science reference describes the design: "told a personality sketch came from a pool of thirty engineers and seventy lawyers, people judged the profession from how engineer-like it sounded and barely moved when the proportions were reversed."[5]
We did not obtain this paper.
Two observations, ours.
The design's strength is the reversal. Changing the population from 30/70 to 70/30 should move any correct judgment substantially, and reportedly barely moved it at all.
The literature also includes Bar-Hillel, M. (1980), The base-rate fallacy in probability judgments, Acta Psychologica, 44(3), 211–233; Tversky and Kahneman (1982), Evidential Impact of Base Rates; and Eddy, D. M. (1982), Probabilistic Reasoning in Clinical Medicine[4]. The 1995 frequency paper records that Eddy (1982) reported that 95 out of 100 physicians estimated the posterior probability incorrectly, with our source truncating[2].
The Fix That Works
The most useful finding in this article.
Gigerenzer and Hoffrage published How to improve Bayesian reasoning without instruction: Frequency formats in Psychological Review, 102(4), 684–704, in 1995[5].
The title states the claim: performance improves without instruction, purely by changing how the numbers are presented.
They followed it with Hoffrage and Gigerenzer (1998), Using natural frequencies to improve diagnostic inferences, Academic Medicine, 73(5), 538–540[7].
Two observations, ours.
This is a remedy that does not require anyone to become better at anything, which is rare. Most of what this series has covered offers awareness as a remedy, and the fifth and thirtieth articles both found awareness insufficient.
And it is free. Restating a probability as a count costs nothing and changes no substance.
Why Frequencies Help
The authors' own explanation, which is historical rather than psychological.
From the paper itself: "probabilities and percentages took millennia of literacy and numeracy to evolve; organisms did not acquire information in terms of probabilities and percentages until very recently." And: "Percentages became common notations only during the 19th century (mainly for interest and taxes), after the metric system was introduced during the French Revolution."[2]
Three observations, ours.
The argument is that the format is the problem, not the person. A single-event probability is a recent notational invention; a count of cases is not.
Note the parenthesis, which is pleasing in a publication like this one: percentages spread mainly for interest and taxes. The notation that defeats professionals became common because of the work accountants do.
And we would separate the evolutionary framing from the practical claim. You do not have to accept the account of why frequencies work to use the finding that they do.
The Fix Was Itself Contested
Reported because leaving it out would repeat the failure this series keeps documenting.
A reference list records Gigerenzer, G., and Hoffrage, U. (1999), Overcoming difficulties in Bayesian reasoning: A reply to Lewis and Keren (1999) and Mellers and McGraw (1999), Psychological Review, 106(2), 425–430[7].
We did not obtain the reply or either critique and report only that they exist.
Two observations, ours.
A reply to two separate critiques in the same journal indicates the frequency format claim was seriously contested at the time, which the popular version never mentions.
And we grade the claim B rather than A on exactly this basis. It is well-published and disputed, which is a normal and healthy state for a finding rather than a mark against it.
The Story Was Oversold
The balancing critique.
A behavioural science reference states plainly: "But the story that people simply ignore base rates was oversold."[5]
The literature includes Koehler, J. J. (1996), The Base Rate Fallacy Reconsidered: Descriptive, Normative, and Methodological Challenges[4], and Christensen-Szalanski, J. J. J., and Beach, L. R. (1982), Experience and the base-rate fallacy, Organizational Behavior and Human Performance, 29, 270–278[8].
We obtained neither and report titles only.
Two observations, ours.
The Koehler title's three words, descriptive, normative, and methodological, signal a comprehensive challenge: what people actually do, what they ought to do, and how it was measured.
And the second title, Experience and the base-rate fallacy, raises the question that matters most commercially: whether people who make a judgment repeatedly, with feedback, still make this error. We do not know what it found.
The Version That Matters In A Firm
The translation. All figures in this section and the next two are ours, use invented parameters, and come from no study.
Replace the disease with an error, and the test with a review procedure.
Suppose a procedure catches 90 percent of files that genuinely contain an error, which is good, and wrongly flags 5 percent of clean files, which is also reasonable.
The question a partner actually asks is not how good the test is. It is: when this thing flags a file, how likely is it that something is actually wrong?
And the answer depends almost entirely on something the test does not know: how common errors are in the first place.
What A Good Test Produces
The result, per 1,000 files. Ours.
At a 0.1 percent true error rate: about 51 flags, of which about 1 is real. Positive predictive value 2 percent.
At 1 percent: about 59 flags, of which 9 are real. 15 percent.
At 2 percent: about 67 flags, of which 18 are real. 27 percent.
At 5 percent: about 93 flags, of which 45 are real. 49 percent.
At 25 percent: about 263 flags, of which 225 are real. 86 percent.
Three observations, ours.
At a 1 percent error rate, nine out of ten flags are false, and the test is not bad. The base rate is low.
The same test looks excellent or useless depending purely on where it is pointed. Nothing about the instrument changed between the first row and the last.
And this explains a familiar experience: a control that everybody stops trusting, described as producing noise, when it is performing exactly to specification.
The Remedy That Does Not Work
The instinctive response, and why it fails. Ours.
When most flags turn out to be nothing, the usual reaction is that reviewers are being careless, or that the procedure needs to be applied more diligently.
Three points.
The false positives are produced by the test, not by the reviewer. They existed before anyone looked at them.
Which means more care per flag cannot reduce their number. It can only make each one more expensive to dismiss.
And there is a predictable second-order failure: when nine in ten alarms are nothing, people stop taking any of them seriously, which converts a precision problem into a detection problem. The tenth flag is the one that mattered.
The Lever That Does
What actually moves the number. Our arithmetic, at a 1 percent error rate, holding sensitivity at 90 percent.
At a 5 percent false positive rate: 59 flags, predictive value 15 percent.
At 2 percent: 29 flags, 31 percent.
At 1 percent: 19 flags, 48 percent.
At 0.5 percent: 14 flags, 65 percent.
Two observations.
Precision buys far more than sensitivity when the base rate is low. Cutting the false positive rate from 5 percent to 1 percent raised predictive value from 15 to 48 percent, while catching the same share of real errors.
And the second lever is to raise the base rate deliberately by pointing the test at a subpopulation where the thing is more common. That is what risk-based selection is, and the arithmetic above is the reason it works.
The Problem With All Of This
The caveat that keeps the article honest, and it is the strongest objection.
A source puts it well: "Laboratory problems hand you a clean, certain base rate. The real world rarely does. Actual base rates are often unstable, contested, or simply unknown, and a 'relevant' base rate for one decision can be the wrong reference class for" another, with our source truncating[4].
Three consequences, ours.
Every number in our tables is invented. A firm that does not know its true error rate cannot compute its predictive value, and most firms do not know it.
Which means the practical use of this arithmetic is often qualitative: knowing that a low base rate guarantees mostly false alarms, even without knowing the exact rate.
And it points at something worth doing regardless. Recording what proportion of your flags turn out to be real is a measurement most firms could take and few do, and it is the same kind of cheap internal measurement the thirty-second article recommended for consistency.
Choosing The Reference Class
The subtler half of the objection. Ours.
Even where a base rate exists, which base rate is a judgment.
Three versions of the same question. What proportion of all files contain this error? What proportion of files in this industry? What proportion prepared by this client?
Two observations.
These can differ by an order of magnitude, and the choice between them is not itself a statistical question. It is a judgment about relevance, made before any arithmetic begins.
And a narrower reference class is more relevant and less reliable, because it rests on fewer observations. That trade-off has no general solution, which is worth stating plainly rather than pretending the arithmetic settles it.
And The Inverse Error
The error that runs the other way, which this article should not encourage. Ours.
Everything above argues that a positive result means less than it appears when the thing is rare. There is a matching mistake.
Three points.
Where the base rate is high, the same arithmetic makes a positive result highly informative, and the last row of our table shows 86 percent.
And a negative result is not the mirror image. When a condition is rare, a negative result is highly reassuring precisely because almost nothing is there, which tells you very little you did not already know.
So the general lesson is not that tests mean little. It is that a test result cannot be interpreted without the prior, in either direction, and dropping the prior produces overconfidence in one direction and false comfort in the other.
What To Do
Convert probabilities to counts. Ten in ten thousand, five hundred false alarms, ten real. The 1995 paper's title says performance improves without instruction, purely from the format.
Never interpret a flag without the base rate. The same test produces a 2 percent or an 86 percent predictive value depending only on where it is pointed.
Expect mostly false alarms when the thing is rare. At a 1 percent error rate, our arithmetic gives nine false flags for every real one from a good test.
Do not respond by demanding more care. The false positives were produced by the test before anyone read them, and more care per flag only raises the cost of dismissing each.
Watch for alarm fatigue. When most alerts are nothing, people stop attending to all of them, which turns a precision problem into a detection failure.
Buy precision before sensitivity at low base rates. On our figures, cutting the false positive rate from 5 to 1 percent moved predictive value from 15 to 48 while catching the same share of real errors.
Or raise the base rate by aiming the test. That is what risk-based selection is for, and this arithmetic is why it works.
Record what proportion of your flags are real. Most firms have never measured it, and without it none of the above can be applied to your own numbers.
The Limits Of This Analysis
Several caveats matter. This article discusses probability and screening and is not audit, assurance, tax, medical or professional standards advice; nothing in it describes any regulator's or professional body's requirements. Everything is verified to August 2026. We did not obtain the 1978 paper, only its question and result summary as quoted in a peer-reviewed article; the reported answer of 95 percent and the figures of 60 participants and 18 percent correct come from commercial sources we could not verify. We did not obtain the 2014 replication, the 1973 paper, the 1995 frequency formats paper beyond two passages, the 1998 follow-up, the 1999 reply, either critique it replies to, the Koehler reconsideration, or the paper on experience, and report titles and characterisations only. We found two inconsistencies in how the 1978 paper is cited, on its page range and on the third author's name, and report both without resolving them. All arithmetic is ours, apart from the 1 in 1,000 prevalence and 5 percent false positive rate taken from the published question; every parameter in the firm tables is invented and none is an estimate of any real error rate. The sections on remedies, alarm fatigue, reference class selection and the inverse error are our own reasoning, not findings.
Frequently Asked Questions
What is base rate neglect?
Did that finding hold up?
What is the fix?
How does this apply to my firm?
So should I make reviewers more careful?
What is the catch?
References
- Peer-reviewed journal article on artificial intelligence and medical statistics, reproducing verbatim the question posed in the 1978 study, being whether a test to detect a disease whose prevalence is 1 in 1000 has a false positive rate of 5 percent, what is the chance that a person found to have a positive result actually has the disease, assuming nothing is known about the person's symptoms or signs; on the results showing that only a minority provided the correct answer, with most clinicians overestimating the positive predictive value; and on the study having been replicated by Manrai and colleagues 36 years later, yielding similar results, highlighting that medical statistics continue to challenge healthcare professionals irrespective of grade despite advancements in medical education. Note: a peer-reviewed article describing the 1978 study; our best source for its question and result. We did not obtain the 1978 paper or the replication. ncbi.nlm.nih.gov
- Gigerenzer, G., & Hoffrage, U. (1995). How to Improve Bayesian Reasoning Without Instruction: Frequency Formats. Psychological Review, 102(4), 684–704, hosted copy, on staff at Harvard Medical School and others having equally great difficulties with this and similar medical disease problems; on Eddy (1982) having reported that 95 out of 100 physicians estimated the posterior probability incorrectly, our source truncating; on probabilities and percentages having taken millennia of literacy and numeracy to evolve, with organisms not acquiring information in terms of probabilities and percentages until very recently; and on percentages having become common notations only during the 19th century, mainly for interest and taxes, after the metric system was introduced during the French Revolution. Note: a hosted copy of the paper; we obtained two passages only. This source cites the 1978 study's third author as Grayboys. researchgate.net
- Medical ethics journal article reference list, confirming Casscells, W., Schoenberger, A., & Graboys, T. B. (1978), Interpretation by physicians of clinical laboratory results, New England Journal of Medicine, 299(18), 999–1001; Gigerenzer, G., & Hoffrage, U. (1995), Psychological Review, 102(4), 684–704; Hoffrage, U., & Gigerenzer, G. (1998), Using natural frequencies to improve diagnostic inferences, Academic Medicine, 73(5), 538–540; and Gigerenzer, G., & Hoffrage, U. (1999), Overcoming difficulties in Bayesian reasoning: A reply to Lewis and Keren (1999) and Mellers and McGraw (1999), Psychological Review, 106(2), 425–430. Note: a peer-reviewed journal's reference list; citations only. journalofethics.ama-assn.org
- Commercial behavioural design website, on most physicians in the 1978 study having confused the test's 5 percent false-positive rate with the patient's odds, answering 95 percent instead of about 2 percent; on the disease being rare so that far more healthy people tested positive by accident than sick people tested positive truly; on the base rate fallacy being the human failure to apply Bayes' theorem, typically by dropping the base rate entirely and reasoning from the evidence alone; on laboratory problems handing you a clean, certain base rate while the real world rarely does, with actual base rates often unstable, contested or simply unknown, and a relevant base rate for one decision potentially being the wrong reference class for another, our source truncating; and its reference list identifying Casscells and colleagues (1978), Bar-Hillel, M. (1980), The Base-Rate Fallacy in Probability Judgments, Acta Psychologica, 44(3), 211–233, Tversky and Kahneman (1982), Evidential Impact of Base Rates, Eddy, D. M. (1982), Probabilistic Reasoning in Clinical Medicine, Gigerenzer and Hoffrage (1995), and Koehler, J. J. (1996), The Base Rate Fallacy Reconsidered: Descriptive, Normative, and Methodological Challenges. Note: a commercial website, not peer-reviewed. The 95 percent figure could not be verified against the paper. This source gives the page range as 999–1001. yukaichou.com
- Behavioural science reference dictionary, on Kahneman and Tversky's 1973 studies having made the phenomenon famous, in which people told a personality sketch came from a pool of thirty engineers and seventy lawyers judged the profession from how engineer-like it sounded and barely moved when the proportions were reversed; on Casscells and colleagues having found in 1978 that most clinicians at Harvard teaching hospitals badly overstated the chance of disease after a positive test, ignoring a low prevalence; on the story that people simply ignore base rates having been oversold; and confirming citations for Gigerenzer and Hoffrage (1995), Bar-Hillel (1980), Casscells and colleagues (1978) with DOI 10.1056/NEJM197811022991808, and Kahneman, D., & Tversky, A. (1973), On the psychology of prediction, Psychological Review, 80(4), 237–251. Note: a reference website, not peer-reviewed; used for the 1973 design description and for the balancing statement that the story was oversold. becisions.com
- Data science commentary article, stating that in 1978 researchers at Harvard Medical School posed a base-rate problem to 60 physicians and medical students, of whom only 18 percent arrived at the correct answer; and confirming the citations for Casscells and colleagues (1978), New England Journal of Medicine, 299(18), 999–1001, and Gigerenzer and Hoffrage (1995), Psychological Review, 102, 684–704. Note: a commentary article, not peer-reviewed. Neither the sample size nor the 18 percent figure could be verified against the paper. towardsdatascience.com
- Reference list in a medical ethics journal article, identifying Hoffrage, U., & Gigerenzer, G. (1998), Using natural frequencies to improve diagnostic inferences, Academic Medicine, 73(5), 538–540, and Gigerenzer, G., & Hoffrage, U. (1999), Overcoming difficulties in Bayesian reasoning: A reply to Lewis and Keren (1999) and Mellers and McGraw (1999), Psychological Review, 106(2), 425–430. Note: citations only. We obtained neither the 1998 follow-up, the 1999 reply, nor either of the two critiques it replies to; their existence is reported to show the frequency format claim was contested. journalofethics.ama-assn.org
- Reference list in an academic book chapter on improving the diagnostic inferences of medical experts, identifying Casscells, W., Schoenberger, A., & Grayboys, T. (1978), New England Journal of Medicine, 299, 999–1000; and Christensen-Szalanski, J. J. J., & Beach, L. R. (1982), Experience and the base-rate fallacy, Organizational Behavior and Human Performance, 29, 270–278. Note: citations only. This source gives the page range as 999–1000 and spells the third author Grayboys, both differing from other sources; we report the inconsistency without resolving it. link.springer.com
This article discusses probability and screening and is not audit, assurance, tax, medical or professional standards advice. No paper discussed was obtained in full. The reported common answer of 95 percent, and the figures of 60 participants and 18 percent correct, come from commercial sources and could not be verified. All arithmetic is the authors' own apart from the prevalence and false positive rate taken from the published question, and every parameter in the firm tables is invented rather than an estimate of any real error rate.