Forty-nine articles have applied a grading scheme[2] to behavioural research. This one applies arithmetic to the forty-nine articles, and the first thing to report is that the exercise produced a figure we had to correct before publishing.
Key Takeaway
Across 49 articles this series has awarded 170 graded claims: 31.8 percent A, 28.8 percent B, 27.6 percent C, 11.8 percent D. It has stated 488 times that it could not obtain a source, roughly ten times per article, and noted a source truncating mid-passage 135 times. The median effect size it has reported is 0.43. All figures extracted by script from our own published pages.
A Note On Method
Everything in this article is about our own corpus, not about the state of behavioural science.
The figures were produced by running regular expressions over the 49 published articles on this site and counting matches[1]. That method has obvious limitations, and we list them rather than presenting the output as authoritative.
It counts text patterns, not meaning. A grade mentioned in a sentence explaining the grading scheme counts the same as one awarded.
Effect sizes were extracted by matching a specific numeric format, so any reported in another form are missing from the count.
And the counts include every occurrence, so an article discussing the same effect size three times contributes three matches to the raw tally, which is why we report distinct values separately.
These are statistics about what we published. They describe this publication's output and nothing beyond it.
One Hundred And Seventy Graded Claims
The distribution.
Grade A: 54 claims, 31.8 percent.
Grade B: 49 claims, 28.8 percent.
Grade C: 47 claims, 27.6 percent.
Grade D: 20 claims, 11.8 percent.
Which gives 60.6 percent at A or B, and 39.4 percent at C or D.
Three observations, ours.
We expected the distribution to be worse. A series whose recurring subject is overstated findings might be assumed to hand out mostly low grades, and it has not.
But two in five graded claims sit at C or D, which for a body of research routinely presented to business audiences as settled is a substantial fraction.
And the grades are per claim, not per finding. A single article frequently grades the existence of an effect A and its magnitude C, which is where much of the C count comes from and is the single most common shape in the series.
The Number That Surprised Us
The count we did not anticipate.
Across the 49 articles, the phrase did not obtain and its variants appear 488 times. That is almost exactly ten per article.
Three observations, ours.
Ten admissions per article that a source was unavailable is more than we would have guessed, and we had written every one of them.
It means the typical article in this series is built from abstracts, working papers, reference lists and third-party descriptions rather than from papers. That is a real limitation on everything published here and it is worth stating in one place rather than only in individual footnotes.
And it is a fact about access, not about the research. Most of what we could not obtain sits behind subscription paywalls, which is a structural feature of the literature rather than a failing of the studies.
Every Effect Size We Have Reported
The extraction found 21 distinct values in the format we searched for, ranging from 0.02 to 1.39.
At the top: 1.39 and 1.35 from the brainstorming articles, 1.32 and 1.25 from the ease-of-retrieval article, 1.18 from choice overload, 1.09 and 0.85 and 0.80 from the article on burnout measurement.
In the middle: 0.54 mere exposure, 0.43 the disputed nudge headline, 0.41 feedback interventions, 0.31 framing.
At the bottom: 0.09, 0.08 growth mindset, 0.05, and 0.02 choice overload's meta-analytic mean.
One observation. That choice overload appears at both 1.18 and 0.02 is not an error in the extraction. It is the finding of the third article in this series, which is that a famous large effect in one study sits beside a meta-analytic mean of essentially zero.
The Median Is 0.43
The summary statistics of that set.
Median 0.43. Mean 0.61. Six of the 21 values fall below 0.20, eleven below 0.50, and twelve below 0.80.
Two observations, ours.
By conventional benchmarks, roughly half of everything this series has quantified is below a medium effect, and more than a quarter is below what is usually called small.
That figure would make a satisfying headline. It is also partly wrong, and the next section explains why.
A Correction To Our Own Chart
Because the fortieth article in this series was about the cost of not checking your own numbers, and this one nearly repeated the error.
Three of the largest values in that list are not effects anyone measured.
1.32 and 1.25 are from the forty-sixth article, where they are the smallest effects a badly powered study could have detected. They are thresholds, not observations, and including them inflates the distribution's top end with numbers that represent a study's weakness rather than a finding's strength.
1.39 is a penalty, being the amount by which interacting groups underperform the same people working separately. It is a real measured quantity but it is a deficit, and reading it as a large positive finding would be wrong.
Two observations.
We caught this by looking at what each number actually referred to rather than trusting the extraction, which took ten minutes and changed the article.
And it is exactly the failure this series has documented repeatedly in others: a figure lifted from a table, correct in itself, misleading once separated from what it measured.
What That Leaves
The corrected set.
Restricting to values that are reported effects rather than detection thresholds or penalties leaves 13 values.
Their median is 0.41, and 8 of the 13 fall below 0.50.
Three observations, ours.
The correction barely moved the median, from 0.43 to 0.41, which is worth reporting because it would have been convenient to imply the fix mattered more than it did.
What it changed was the top of the distribution, which is where the misleading impression sat.
And the honest summary is that the typical quantified finding in this series is around d = 0.4, which is real, useful, and considerably smaller than the way these findings are usually presented in business writing.
Ten Times An Article
What the 488 figure means in practice. Ours.
Three consequences.
Most of what this series reports is at one remove or more from the underlying work. We have consistently said so per article, and saying it once at scale is more honest than distributing it across forty-nine footnotes.
It changes what the grades mean. A Grade A here means the claim is well attested across independent sources, not that we read the study and checked its statistics, because in most cases we could not.
And it is why this series does its own arithmetic so often. Nineteen articles contain a simulation or calculation we built ourselves, largely because building the structure was possible where obtaining the paper was not.
The Truncation Problem
A smaller number with an outsized effect on what we can say.
The word truncat and its variants appear 135 times, about 2.8 per article. Each instance marks a place where a source cut off before finishing a sentence we needed.
Three observations, ours.
Several of these sit on the single most important sentence in an article. The forty-first reports a paper finding "further doubts" about informational cascades and cannot say what they were. The forty-fourth loses the counterfactual result mid-word.
The pattern is not random. Abstracts are truncated by preview systems at a fixed character count, which means the sentence lost is systematically the last one, and in an abstract the last sentence is often the conclusion.
And we have consistently chosen to report the existence of a finding without its content rather than infer the ending. That is the right call and it produces some frustrating paragraphs.
Pattern One: The Structural Reinterpretation
The most common recurring shape, appearing in 11 of 49 articles.
A famous result turns out to be reproducible from something other than the mechanism it is famous for.
The twenty-sixth reproduced the Dunning-Kruger chart from pure noise, with no metacognitive deficit in the simulation at all. The thirty-first reproduced the hungry judges decline from rational time management. The thirty-fourth found the hot hand debunking rested on a selection artefact, and its correction reversed the conclusion. The forty-fifth found the confirmation bias task uses the one configuration in which the strategy it demonstrates cannot possibly work.
Two observations, ours.
In every case the reinterpretation was checkable without new data. These are not disputes requiring a replication budget; they are arithmetic.
And in every case, the reinterpretation arrived years or decades after the finding entered general circulation. The hot hand correction took thirty-three years.
Pattern Two: One Name, Several Things
The second recurring shape.
The twenty-eighth found three distinct effects filed under framing, with independent literatures and different magnitudes. The forty-second found overconfidence is three separate measurements that can move in opposite directions on the same task. The forty-fifth found a goal and a method sharing the name confirmation bias. The forty-eighth found a literature using myside and confirmation interchangeably with a slash between them.
Two observations, ours.
The practical cost is consistent: advice built on one member of the family transfers badly to the others, and nobody notices because the label is the same.
And the diagnostic is simple. If a term covers behaviours that could move in opposite directions, it is covering more than one thing.
Pattern Three: The Broken Second Link
The shape with the most direct commercial consequence.
The thirty-eighth found a network meta-analysis of 492 studies in which interventions changed the target measure but the change did not carry through to behaviour. The forty-sixth found a replication in which the manipulation did make a task feel harder while the downstream judgment effect did not appear.
Three observations, ours.
Both are two-link chains where the first link held and the second did not, arrived at from completely unrelated literatures.
The commercial version is that demonstrating a programme moved something is not demonstrating it moved the thing you are buying, and most evaluations stop at the first measurement.
And it generalises: an intervention working through an intermediate step inherits the product of both stages, so a two-stage theory of change needs both stages strong.
Pattern Four: The Missing Meta-Analysis
The most frustrating shape, and the one most under our control.
On at least five occasions the document that would have settled a question was the one we could not get. The forty-third reports no effect size for hindsight bias anywhere, because both meta-analyses were unobtainable. The fortieth withdrew a claim rather than revise it downward, for the same reason.
Two observations, ours.
Our response has consistently been to report the gap and name the documents a reader would need, rather than substitute a number from a citing source.
Where we have used a citing source's figure, as with the effect size in the forty-eighth article, we have said so at every use. That is the best available compromise and it is not a good one.
Eighteen Bibliographic Variants
The running count, which began as a footnote and became a finding.
Across the series we have recorded eighteen instances of careful sources reproducing a citation inconsistently: page ranges differing, author initials changing, surnames misspelled, titles altered, and in one case a 1975 paper dated to 2006 on a government-affiliated resource while correctly giving volume 1.
Three observations, ours.
None of them changed a conclusion. That is the point. Individually trivial, collectively they describe a citation layer noisier than the literature it points to.
Two are worth remembering. In the forty-second, a peer-reviewed reference list rendered a title as "The trouble with overconfidece". In the forty-sixth, a source attributed the availability heuristic to a 1991 paper whose subtitle announces it as a reconsideration of the 1973 original, and misdefined it as taking the first explanation that comes to mind. An author reached for the first attribution that came to mind, in a sentence about what comes to mind.
And the practical lesson is small and real: if you are citing something you have not read, you are probably copying someone else's error.
What We Would Actually Rely On
The findings we would be comfortable acting on, after forty-nine articles. Ours.
Independence in aggregation. The thirty-sixth article's guarantee is mathematical rather than empirical: averaging estimates that bracket the truth cannot lose. The forty-first showed what its absence costs, at 95 percent accuracy pooled against 84.5 in a queue from identical inputs[4].
Structured processes over improved people. Process-level evidence is measured at the outcome, which removes the intermediate link where the chains above kept breaking.
Contemporaneous records. The forty-third's finding that memory reconstructs toward outcomes makes a written prediction the only version not subject to it.
Reciprocity, with the cost question open. The forty-ninth found a randomised field experiment of 10,000 letters, which is among the best evidence in the series, and no source reporting what the gift cost.
One observation. Notice how few of these are biases. Three of the four are about designing a process, and the series arrived there without intending to.
What We Got Wrong
The corrections issued so far, collected. Ours.
Two articles have corrected this publication.
The fortieth withdrew a claim made elsewhere on this site that loss aversion is "among the most robustly replicated results in behavioural economics", after a 2018 review disputed it and drew a formal published dialogue[3]. The error was specific: we repeated a claim about the state of the evidence rather than a claim from the evidence.
The thirty-ninth corrected an inherited claim in a draft, which asserted that two sources gave a variant title where only one did.
And this article makes a third, above, to a chart we built ourselves.
Two observations. Three corrections in fifty articles is not evidence of unusual care; it is the number we happened to catch. And all three were caught by checking a specific number against what it referred to, which is the only method that has reliably worked.
The Honest Verdict At Halfway
What we would say to someone deciding whether any of this is useful. Ours.
Four statements.
The field is in better shape than its critics suggest and worse than its popularisers do. Sixty percent of graded claims at A or B is not a discipline in ruins. Forty percent at C or D is not a body of settled knowledge either.
Magnitude is where the damage happens. Directions survive popularisation; sizes do not. The single most common structure in this series is an effect graded A for existence and C for size.
The corrections are usually already published. In nearly every case where a famous finding needed qualifying, someone had qualified it in print, sometimes decades earlier, and the qualification simply had not travelled.
And the useful residue is procedural. After fifty articles the advice that has survived is mostly about how to structure a decision rather than which bias to avoid, and that was not the answer we expected to arrive at.
What To Do
Ask for the size, not the direction. The most common finding in this series is an effect that exists and whose magnitude is unestablished.
Check whether a number is an effect or a threshold. We made that error in this article's own chart and caught it in ten minutes.
Treat a shared label as a warning. Four times a single term covered distinguishable phenomena, and each time the transferred advice was wrong for some of them.
Ask what the second link is. Most evaluations demonstrate that something moved, not that the thing you are buying moved.
Look for the published correction before commissioning new work. It usually exists.
Prefer process design to bias training. Three of the four findings we would rely on are about structuring a decision.
Do not cite what you have not read. Eighteen variants in fifty articles suggests most citation errors are inherited.
Write the number down. It is the only defence against a memory that reconstructs toward what happened.
The Limits Of This Analysis
Several caveats matter, and they are unusually important here. Every figure in this article describes our own published corpus and nothing else. It is not a survey of behavioural science, not a sample of any literature, and not evidence about the field; it is a count of what this publication wrote. The extraction was done by regular expression and counts text patterns rather than meaning, so a grade discussed counts identically to a grade awarded, and effect sizes reported in formats other than the one searched are absent from the tally entirely. The set of 21 effect sizes is therefore incomplete and its summary statistics should be read accordingly. Three of those values were misclassified in our first pass, being detection thresholds and a group-performance penalty rather than measured effects, which we corrected in the body; other misclassifications may remain. The topics covered were selected by us, not sampled, so the grade distribution reflects our choices about what to write about as much as anything about the research. Our grades are our own judgments applied by a single publication with no external check. And the count of 488 unobtained sources is a statement about our access, largely reflecting subscription paywalls, rather than about the quality of the work behind them.
Frequently Asked Questions
What did the audit actually find?
Is behavioural science in trouble?
What is the single most common problem?
What error did you make in this article?
What would you actually rely on?
Why do you keep doing your own arithmetic?
References
- The 49 preceding articles in this series, published on this site between silo 14 numbers 1 and 49, from which all counts in this article were extracted by script. The extraction matched the strings "Grade A" through "Grade D", the phrase "did not obtain", the stem "truncat", and effect sizes in the specific format "d = " followed by a decimal value. Note: this is a self-referential source. Every figure in this article describes this publication's own output and is not evidence about behavioural science, any literature, or any sample of research. The method counts text patterns rather than meaning. gshfinancial.com/insights
- The first article in this series, Thirty-Six Percent: What Survived In Behavioural Finance, which established the four-level grading scheme applied throughout and audited in this article. Note: the source of the grading scheme whose output is counted here. The scheme is this publication's own construction and has no external validation. gshfinancial.com
- The fortieth article in this series, A Correction To Our Own Back Catalogue, recording the withdrawal of a claim that loss aversion is among the most robustly replicated results in behavioural economics, following a 2018 review in the Journal of Consumer Psychology that disputed it and drew a formal published research dialogue. Note: a self-reference, cited here as one of the three corrections this series has published about itself. gshfinancial.com
- The thirty-sixth and forty-first articles in this series, on the averaging principle and on information cascades, from which the pooled and sequential accuracy figures of 95.2 percent and 84.5 percent are drawn. Note: self-references. Those figures come from a simulation this publication built, using invented parameters, and demonstrate a mechanism rather than reproducing anyone's data. gshfinancial.com
Every figure in this article was extracted by script from this publication's own 49 preceding articles. It describes what this publication wrote and is not a survey of behavioural science, a sample of any literature, or evidence about the state of the field. The extraction counts text patterns rather than meaning, the set of effect sizes captured is incomplete, and three values in it were misclassified on the first pass and corrected in the body. The topics covered were selected rather than sampled, and the grades are this publication's own judgments with no external check.