A finance article cites a psychology study. The study is real, the journal is prestigious, the finding is memorable, and the citation is accurate. None of that tells you whether the effect exists.
Key Takeaway
The Reproducibility Project: Psychology replicated 100 studies from three leading journals. Ninety-seven percent of the originals reported statistically significant results; thirty-six percent of the replications did, and replication effects were half the magnitude of the originals[1]. Parallel projects in experimental economics and in social science published by Nature and Science found 61 percent and 62 percent respectively[2][3]. Ego depletion, one of the most cited ideas in self-control research, was tested across 23 laboratories with 2,141 participants and produced an effect close to zero[5]. This article sets out how the rest of this series will decide what to rely on.
A Note On Method
Everything in this article is taken from the primary literature or from the published abstracts of the studies concerned, verified in August 2026.
Where we give a figure, it comes from the paper reporting it rather than from a secondary summary. Where two sources describe the same project differently, we say so.
One such difference appears immediately. The Experimental Economics Replication Project is described by one source as covering studies published between 2011 and 2014[4] and by another as between 2011 and 2015[7]. We have not resolved it and do not rely on the date range for anything.
We did not obtain the full text of any of the replication projects. We rely on published abstracts, on the journals' own summaries, and on subsequent papers that reanalyse the datasets.
We were unable to verify several claims commonly made in this area, including the frequently quoted loss aversion coefficient and its published critiques. Those are not asserted here.
This article is about research quality. It is not investment, tax or psychological advice.
Why This Article Comes First
A statement of intent for the series that follows. This section is our own.
Behavioural finance is unusually easy to write badly. The findings are vivid, they explain things people already suspect about themselves, and they travel well as anecdotes. A writer can assemble a persuasive article entirely from studies that no longer hold.
Three features of the field make this worse than in most disciplines.
The canonical results are old. Much of what gets cited was published between 1974 and 2000, before preregistration, before power norms tightened, and before anyone was checking.
The secondary literature is enormous. Popular books, practitioner articles and training materials cite each other rather than the papers, so a finding can be repeated for two decades after the underlying evidence has been withdrawn from under it.
And the failures are quieter than the findings. A striking result gets a book chapter. Its failure to replicate gets a technical paper in a methods journal.
So the first useful thing this series can do is not to describe a bias. It is to establish what we will accept as evidence, and to be checkable on it.
The Three Projects
Three large coordinated replication efforts define what is known about the reliability of this literature. They are the factual base for everything below.
The Reproducibility Project: Psychology, published by the Open Science Collaboration in Science in 2015[1].
The Experimental Economics Replication Project, Camerer and colleagues, Science, 2016[2].
The Social Sciences Replication Project, Camerer and colleagues, Nature Human Behaviour, 2018[3].
Two features they share are worth noting before the numbers, and this observation is ours.
All three selected their targets systematically rather than choosing studies they suspected were weak. The Social Sciences Replication Project describes its 21 studies as systematically selected[3], and a subsequent methodological paper notes that this systematic selection is itself an analytical advantage, because it avoids the bias that would arise from cherry-picking suspicious results[8].
And all three were high-powered by design. These were not underpowered attempts that failed for lack of participants. The Social Sciences project ran samples on average about five times larger than the originals[3].
Reproducibility Project: Psychology
The largest of the three, and the one that changed the conversation.
The Open Science Collaboration conducted replications of 100 experimental and correlational studies published in three psychology journals, using high-powered designs and original materials where available[1]. The journals were Psychological Science, the Journal of Personality and Social Psychology, and the Journal of Experimental Psychology: Learning, Memory, and Cognition, and the target year was 2008[9].
The reported results, in the authors' own terms[1]:
Ninety-seven percent of original studies had statistically significant results.
Thirty-six percent of replications had statistically significant results.
Forty-seven percent of original effect sizes fell within the 95 percent confidence interval of the replication effect size.
Thirty-nine percent of effects were subjectively rated to have replicated the original result.
And, assuming no bias in the original results, combining originals with replications left 68 percent with statistically significant effects[1].
A methodological paper describing the project's mechanics adds useful detail: 158 articles were made available for replication, 111 were assigned, and 100 were completed in time for inclusion, of which 84 replicated the final result of the original paper[9].
Two Numbers That Get Conflated
A distinction that matters and is routinely lost. This section is our own analysis.
The project reported 36 percent and 39 percent, and they are different measurements.
Thirty-six percent is the proportion of replications that reached statistical significance. It is a mechanical threshold applied to a p-value.
Thirty-nine percent is the proportion of effects that replication teams subjectively rated as having replicated[1].
Three consequences.
A writer quoting either number as the replication rate is quoting one of several, and the paper reports at least four.
The 47 percent figure, being originals falling inside the replication's confidence interval, is a third and more forgiving criterion, and it produces the highest of the three.
And the fact that a single project yields 36, 39, 47 and 68 percent depending on the criterion is not an embarrassment. It is what happens when a genuinely difficult question is measured carefully rather than reduced to a headline.
Our own view is that anyone citing a single replication rate without saying which criterion produced it has not read the abstract.
The Finding That Matters More Than The Rate
The result that should have led the coverage, and generally did not.
The Open Science Collaboration reports that replication effects were half the magnitude of original effects, representing a substantial decline[1].
The Social Sciences Replication Project found the same thing independently: the effect size of the replications was on average about 50 percent of the original effect size[3].
Two projects, different fields, different journals, different teams, same halving.
Three consequences, ours.
This is a different and more useful claim than a binary replication rate. It says that even where an effect is real, the published estimate of its size is likely to be roughly double the truth.
For a practitioner, that is the operative number. Someone deciding whether a behavioural intervention is worth implementing does not principally need to know whether the effect exists. They need to know how big it is, because that determines whether it clears the cost of doing anything about it.
And it means the correct mental adjustment when reading an older behavioural paper is not to disbelieve it. It is to halve it, and then ask whether what remains would still change a decision.
The Social Sciences project offers a more refined version. Its Bayesian analysis estimated a true-positive rate of 67 percent, with the relative effect size of those true positives estimated at 71 percent of the original, which the authors read as evidence that both false positives and inflated effect sizes of true positives contribute to the gap[3].
So there are two distinct problems running at once: some findings are not real, and many of the real ones are smaller than reported.
The Economics Project
The second project, and the one closest to this publication's subject matter.
Camerer and colleagues conducted replications of experimental studies published in the American Economic Review and the Quarterly Journal of Economics, and found that 11 of 18 replications, being 61 percent, were successful[2][4].
As noted in our method section, our sources differ on whether the target window was 2011 to 2014[4] or 2011 to 2015[7], and we do not rely on it.
Two observations, ours.
Sixty-one percent is markedly better than 36 percent, and the difference invites an obvious explanation: that experimental economics is more rigorous than psychology. That explanation turns out to be at least partly wrong, for reasons set out two sections below.
And 61 percent is still four in ten failing. A discipline in which two-fifths of its published laboratory results do not reproduce under high-powered direct replication is not in a comfortable position, whatever the comparison.
The Social Sciences Project
The third, and methodologically the most demanding.
Camerer and colleagues replicated 21 systematically selected experimental studies in the social sciences published in Nature and Science between 2010 and 2015. The replications followed analysis plans reviewed by the original authors and were preregistered prior to the replications[3].
They found a significant effect in the same direction as the original for 13 of 21 studies, being 62 percent, with replicability varying between 12 (57 percent) and 14 (67 percent) depending on which complementary indicator was used[3].
Three features deserve emphasis, and these observations are ours.
The targets were published in Nature and Science, which is to say the most selective venues available. Prestige of outlet did not protect the finding.
The original authors reviewed the analysis plans. This forecloses the most common defence against a failed replication, that the replicators did it wrong.
And the replications were preregistered, which forecloses the second most common defence, that the replicators kept testing until they found nothing.
The journal's own editorial framing of the result is worth recording, because it addresses the power objection directly: despite increasing power substantially, with sample sizes on average about five times higher than the original studies, the failures persisted[6].
Power, Not Fraud
The most sophisticated point in this literature, and the one most often missed.
A natural reading of a 36 percent replication rate is that most of those findings were never real. A subsequent analysis suggests the picture is more mechanical than that.
A methodological paper modelling replication rates against publication bias reports that in experimental economics, the predicted replication rate is 60.1 percent against an observed rate of 61.1 percent, and concludes that issues with common power calculations can explain essentially the entire gap between observed and target replication rates in that field, even in a simple model without treatment effect heterogeneity, researcher manipulation, or measurement error[7].
For psychology, the same model predicts 53.9 percent against an observed 35.6 percent. That is well below the mean intended power of 92 percent but well above what was observed, and the authors report that the model accounts for two-thirds of the psychology replication gap[7].
Three consequences, ours.
In economics, the shortfall may be almost entirely a statistical artefact of how replication power was calculated, rather than evidence that the original findings were false.
In psychology, the same mechanism explains most but not all of it, leaving roughly a third of the gap requiring some other explanation.
And this materially changes what a reader should conclude. The honest summary is not most psychology is wrong. It is that a large share of the apparent failure is explained by underpowered replication design, and a residual is not.
We note that this analysis defines the replication rate as the share of original estimates whose replications produced statistically significant findings of the same sign, and that in both applications a small number of original results with p-values slightly above 0.05 were treated as positive results and included[7]. Definitional choices of that kind move these numbers.
Why Economics Looked Better
Our own reading of the comparison, offered as interpretation rather than finding.
The tempting conclusion from 61 percent against 36 percent is disciplinary: economists run tighter experiments.
The power analysis above undercuts that. If essentially the whole economics gap is explained by power calculation issues, then economics did not outperform psychology on the underlying reliability of its findings. It outperformed on a measurement that is sensitive to design choices.
Two more prosaic differences are worth naming.
Effect sizes in economics experiments are often larger by construction, because the manipulations are frequently monetary and unsubtle. Larger true effects are easier to detect on replication regardless of the field's rigour.
And the two projects sampled different things. One took a year's output from three psychology journals; the other took experimental papers from two economics journals. These are not matched populations.
We offer this as a caution against the disciplinary reading, not as a defence of either field.
The Case Study: Ego Depletion
One finding, followed through its whole arc, because the arc is instructive.
Ego depletion is the proposition that self-control draws on a limited resource which becomes depleted after exertion, so that a person who has just exercised self-control performs worse on a subsequent self-control task. The model has typically been tested using a sequential-task paradigm[10]. The original effect is attributed to Baumeister and colleagues in 1998[11].
It is difficult to overstate how widely this was adopted. It supplied the mechanism for a large popular literature on willpower, decision fatigue and depletion of judgment over a working day, much of which reached finance and management writing.
In 2016, Hagger, Chatzisarantis and a large group of collaborators published the first registered multilab replication of the effect, in Perspectives on Psychological Science[5].
The design: 23 participating laboratories, 2,141 participants, in both English-speaking and non-English-speaking countries, running a direct replication of the ego-depletion paradigm reported by Sripada, Kessler and Jonides in 2014[5][12].
The result, in the authors' own words: results across the 23 laboratories revealed small effect sizes on the primary and secondary dependent variables, and the 95 percent confidence intervals for the effect sizes for the majority of laboratories' replications included the value of zero. Their conclusion was that if there is any effect, it is close to zero[5].
What Bias Correction Did To It
The part of the story that explains why the replication result was not a surprise to specialists.
Before the multilab replication, Carter and colleagues had published a series of meta-analytic tests under a title that states its own conclusion: self-control does not seem to rely on a limited resource[13].
The numbers are the interesting part. In Carter and colleagues' revision of an earlier meta-analysis, 41 percent of the included studies were unpublished, and that revised analysis produced g = 0.43. Their trim-and-fill analysis, a standard correction for publication bias, produced g = 0.24. And a regression-based estimate using the precision effect estimation with standard error technique produced g = 0.003[5].
The multilab replication's own overall effect size closely mirrored that last figure[5].
Three observations, ours.
The estimate moves from a respectable 0.43 to a trivial 0.003 depending purely on which bias correction is applied. No new data was collected to produce that collapse.
The fact that 41 percent of the studies were unpublished tells you what was happening: null results existed in quantity and were not appearing in journals.
And the convergence between the most aggressive bias correction and the eventual 23-lab result is the strongest part of the story. The correction predicted the replication. That is the meta-analytic method working exactly as intended.
The authors of the replication put the implication carefully: estimates of the size of the depletion effect should, at the very least, be revised downwards from the effect size reported in bias-uncorrected meta-analyses[5].
The Objection That Is Still Live
The part a one-sided account would omit, and we are not going to omit it.
A commentary published in Frontiers in Psychology shortly after the replication accepts that the effect was not replicated across the 23 laboratories, but argues that cautious attention should be paid to the effectiveness of the depleting task used in the replication project, being the letter-e crossing task[14].
The argument is specific rather than rhetorical. The commentary notes that the task as invented by Baumeister and colleagues has particular features, including that the depletion condition involves more complex crossing rules than the control condition[14]. The implication is that if the manipulation did not actually deplete anyone, the study tested nothing.
The replication's own supplementary material engages with the same territory, reporting an examination of whether the depletion version of the letter-e task was more effortful and aversive than the no-depletion version[5].
Two observations, ours.
This is a real methodological objection, not special pleading, and it is the standard difficulty with any direct replication: a failed replication is ambiguous between the effect is absent and the manipulation did not work.
But it has to be weighed against the meta-analytic convergence described above. The bias-corrected estimate arrived at approximately zero before the replication ran, using an entirely different method on an entirely different body of data.
Our own assessment, offered as a judgment: the manipulation objection is legitimate and the accumulated evidence still points strongly against a practically meaningful depletion effect in this paradigm. A reader who wants to hold the contrary view has a principled basis for it, and should hold it explicitly rather than by default.
What That Sequence Teaches
Our own extraction of the general lesson.
Follow the ego depletion arc as a template and it contains five stages.
A striking original finding in a high-profile venue, in 1998.
A large secondary literature adopting it as mechanism, spreading well beyond the original field.
A bias-corrected meta-analysis in 2015 finding the effect collapses under correction.
A preregistered multilab replication in 2016 finding an effect close to zero.
And a continuing methodological dispute about whether the replication tested the right thing.
The practically important observation is about the gap between stages two and three. For roughly seventeen years, the finding was being applied confidently in adjacent fields while the evidence supporting it was, in retrospect, thin. Nobody in those adjacent fields did anything wrong. They cited a real paper accurately.
This is the structural reason a behavioural finance series needs an evidentiary policy rather than good intentions. Accuracy of citation is not the same as reliability of claim.
Researcher Degrees Of Freedom
The mechanism underneath, as the methodological literature describes it.
A survey of this literature notes that attention shifted from outright fraud to the subtler mechanics of researcher degrees of freedom: flexible decisions about exclusion rules, stopping, outcome definitions and model specifications that can produce significance even when evidence is weak, and the related problem of the garden of forking paths where many defensible analytic routes exist[15].
The same account records that early meta-research argued selective reporting, low statistical power, and strong incentives for novelty can produce a literature with inflated effect sizes and excess false positives, and that survey evidence indicated such practices were not rare edge cases but part of routine scientific workflow[15].
Three observations, ours.
The problem is not dishonesty. It is that a researcher making a series of individually reasonable analytic choices, each defensible, can arrive at significance without ever intending to.
Which means the failure is systemic and correctable, and the correction is procedural: preregistration, which the account identifies as the most widely endorsed reform to emerge[15].
And it gives a reader a usable heuristic. A finding from a preregistered, adequately powered study is in a different evidentiary class from a finding of similar vintage that was not, regardless of how interesting either is.
A Grading Scheme For Behavioural Claims
The operational output of this article, and it is ours. We will apply it throughout the series.
Grade A. Replicated in a preregistered, high-powered, multi-site design, or observed consistently in large field datasets rather than only in the laboratory. Safe to build an argument on.
Grade B. Supported by bias-corrected meta-analysis with a non-trivial residual effect, or by multiple independent laboratories without a formal multi-site replication. Usable, with the effect size treated as an upper bound.
Grade C. A well-known single-study or single-laboratory finding, not preregistered, with no replication either way. Describable as an idea; not usable as a basis for advice.
Grade D. Contradicted by a preregistered multilab replication or collapsing under standard bias correction. Reportable only as history, and only with its status stated.
Two rules attach to this scheme.
Halve reported effect sizes by default for anything below Grade A, on the basis that two independent large projects found replication effects at approximately half the original magnitude[1][3].
And state the grade rather than implying it. A reader should not have to infer from our hedging how confident we are.
Applying It To Ourselves
The scheme is worthless if it is not turned inward. This section is our own.
This publication already carries articles on loss aversion and the disposition effect, mental accounting, sunk cost escalation, the planning fallacy, home bias and regret. Each of those rests on a literature of the kind this article has just described.
We are not going to claim they were written under this scheme, because they were not. It did not exist when they were written.
What we will say is what the scheme implies for them.
Findings observed in large field datasets, such as trading records rather than laboratory tasks, are in the strongest position, because they do not depend on a manipulation working.
Findings that exist principally as laboratory demonstrations from the 1970s and 1980s need their current status checked before being relied on, however famous they are.
And any specific coefficient quoted from that era should be treated as an upper bound rather than a measurement.
Where a later article in this series revises something an earlier one asserted, we will say so in the later article rather than editing the earlier one silently.
What This Series Will Not Cite
A negative commitment, which is easier to check than a positive one.
We will not cite ego depletion as an established mechanism. On the evidence above it is Grade D in the sequential-task paradigm, and the popular derivations from it, including confident claims about willpower being consumed across a day, inherit that status.
We will not cite a replication rate without naming the criterion that produced it.
We will not cite an effect size from a pre-preregistration study as though it were a measurement.
We will not present a failed replication as settled where a specific methodological objection to it remains live, as with the depleting-task objection above.
And we will not cite anything whose primary source we have not seen, even where it is repeated everywhere. Our own experience of auditing this publication is that a reference which cannot be traced to a document is worse than no reference at all, because it looks like support.
What Tends To Survive
The constructive half. This section is our own analysis and is deliberately cautious about specifics.
The replication literature does not merely destroy. It tells you what kind of finding tends to hold, and three properties recur.
Large true effects survive. The power analysis above implies that much of the apparent failure is a detection problem, which by definition afflicts small effects more than large ones.
Findings in field data survive better than laboratory demonstrations, because they do not depend on a manipulation being effective and because the samples are frequently enormous.
And findings that were preregistered survive better, which is close to tautological but is the reason preregistration was adopted.
We are deliberately not listing specific behavioural finance findings as surviving. That requires checking each against the current literature, which is the work of the individual articles that follow rather than of this one, and we are not going to assert a list we have not verified.
What we will commit to is the procedure: each article in this series states the grade of its central claim, and where the grade is C or D, says so in the body rather than in a footnote.
The Practitioner's Actual Problem
Bringing this back to a person making decisions. This section is our own.
An owner or adviser reading behavioural finance is not conducting a literature review. They want to know whether to change something.
The grading scheme translates into three practical positions.
For a Grade A finding, the question is ordinary cost-benefit: does the effect, at half its reported size, still justify the intervention.
For Grade B or C, the finding is best used as a hypothesis about your own data rather than as a fact about people. If mental accounting predicts you treat a tax refund differently from operating cash, that is checkable in your own records, and the check is more informative than the literature.
For Grade D, the correct action is to stop citing it, including internally.
Our own observation is that the second of those is the most underused. A behavioural claim about a population is weak evidence about an individual business. But it is an excellent prompt for a question that individual business can answer from its own ledger, and the answer is not subject to any replication problem at all.
This Is Not An Argument For Nihilism
The closing position, stated plainly because articles of this kind are often read as debunking.
Three things follow from the evidence above, and a fourth does not.
Published effect sizes in this literature are, on average, roughly double what better-powered replication finds[1][3].
A substantial minority of high-profile findings do not reproduce under preregistered high-powered replication[1][2][3].
And at least one canonical mechanism, ego depletion in the sequential-task paradigm, sits at approximately zero under both bias correction and multilab replication[5][13].
What does not follow is that behavioural science is worthless. The projects described here are themselves behavioural science, conducted by the field on itself, published in its leading journals, and funded on the assumption that the answer mattered. A discipline that runs a 23-laboratory test of its own most popular idea and publishes a null result is not a discipline to write off.
The correct posture is neither credulity nor dismissal. It is grading, and then proportionate reliance.
That is what the remaining articles in this series will attempt, and this one exists so that a reader can hold us to it.
The Limits Of This Analysis
Several caveats matter. This article is about research quality and is not investment, tax or psychological advice. Everything is stated as verified in August 2026. We did not obtain the full text of any of the three replication projects and rely on published abstracts, journal summaries and subsequent papers reanalysing the datasets. Our sources differ on the target window of the Experimental Economics Replication Project, one giving 2011 to 2014 and another 2011 to 2015, and we do not rely on it. We did not obtain the full text of the Hagger and colleagues replication and rely on its abstract, its published supplementary discussion as reproduced by third parties, and commentaries on it. We did not obtain Carter and colleagues' meta-analysis directly and take its figures from the replication paper's discussion of it. We did not obtain Baumeister and colleagues 1998 or Sripada and colleagues 2014 and cite them only as the studies the replication targets. We were unable to verify the frequently quoted loss aversion coefficient or its published critiques, and have asserted nothing about them. We have deliberately not listed specific behavioural finance findings as having survived replication, because we have not checked them individually. The power-calculation analysis is a modelling result from a working paper and is presented as such. The grading scheme, the halving rule, the five-stage arc, the practitioner translation and the disciplinary-comparison caution are our own and are not drawn from any source.
Frequently Asked Questions
What is the replication rate in psychology?
Does that mean most behavioural findings are false?
What is the single most useful number here?
Is ego depletion real?
How should I use a behavioural finding in my own business?
Why publish this before the rest of the series?
References
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716, on replications of 100 experimental and correlational studies published in three psychology journals using high-powered designs and original materials where available; on replication effects being half the magnitude of original effects, representing a substantial decline; on 97 percent of original studies having statistically significant results while 36 percent of replications did; on 47 percent of original effect sizes falling within the 95 percent confidence interval of the replication effect size; on 39 percent of effects being subjectively rated to have replicated the original result; and on combining original and replication results leaving 68 percent with statistically significant effects if no bias in original results is assumed. Note: cited from the published abstract; we did not obtain the full text. pubmed.ncbi.nlm.nih.gov
- Camerer, C. F., et al. (2016). Evaluating replicability of laboratory experiments in economics. Science, 351(6280), 1433–1436, on replications of 18 experimental economics studies published in the American Economic Review and the Quarterly Journal of Economics, of which 11 (61 percent) were successful. Note: cited through subsequent papers that reanalyse the dataset; we did not obtain the full text. nature.com
- Camerer, C. F., Dreber, A., Holzmeister, F., Ho, T.-H., Huber, J., Johannesson, M., Kirchler, M., Nave, G., Nosek, B. A., Pfeiffer, T., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637–644, on replicating 21 systematically selected experimental studies published in Nature and Science between 2010 and 2015; on the replications following analysis plans reviewed by the original authors and being preregistered prior to the replications; on sample sizes averaging about five times higher than the original studies; on finding a significant effect in the same direction as the original for 13 (62 percent) studies; on the effect size of replications averaging about 50 percent of the original effect size; on replicability varying between 12 (57 percent) and 14 (67 percent) on complementary indicators; and on a Bayesian estimated true-positive rate of 67 percent with the relative effect size of true positives estimated at 71 percent, suggesting both false positives and inflated effect sizes contribute. Note: cited from the published abstract. nature.com
- Chirco, A., et al. Classroom experiments as a replication device. Journal of Behavioral and Experimental Finance, on the Open Science Collaboration (2015) finding that of attempts to replicate 100 studies published in top psychology journals only 39 percent were successful with mean replication effect sizes half the magnitude of the originals; on Camerer et al. (2018) finding 13 of 21 (62 percent) social science experiments replicable; and on Camerer et al. (2016) replicating experimental studies published in the American Economic Review and the Quarterly Journal of Economics from 2011 to 2014, showing 11 of 18 (61 percent) successful. Note: a secondary source used for the economics project's target window; its stated window differs from reference 7 and we do not rely on it. sciencedirect.com
- Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., Anggono, C. O., Batailler, C., Birt, A. R., et al. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science, 11(4), 546–573, on presenting the first registered multilab replication of the ego-depletion effect; on results across 23 participating laboratories (N = 2,141) revealing small effect sizes on the primary and secondary dependent variables; on the 95 percent confidence intervals for the majority of laboratories' replications including zero; on the conclusion that if there is any effect it is close to zero; on the overall effect size closely mirroring the precision effect estimation with standard error estimate reported by Carter and colleagues (g = 0.003), against Carter's revised meta-analysis (g = 0.43, with 41 percent of included studies unpublished) and trim-and-fill analysis (g = 0.24); on substantial heterogeneity in effect size across laboratories; on examination of whether the depletion version of the letter-e task was more effortful and aversive than the no-depletion version; and on the conclusion that estimates of the depletion effect should at the very least be revised downwards from bias-uncorrected meta-analytic figures. Note: cited from the published abstract and from the discussion as reproduced by the publisher and by repository copies. journals.sagepub.com
- Nature Human Behaviour editorial. (2018). Learning from replication. Nature Human Behaviour, on the Reproducibility Project: Psychology and the Experimental Economics Replication Project having successfully replicated 36 percent and 61 percent of their target studies respectively; and on Camerer and colleagues increasing power substantially, with sample sizes on average approximately five times higher than the original studies, and the failures nonetheless persisting. Note: a journal editorial accompanying reference 3. nature.com
- Working paper. Can the replication rate tell us about publication bias?, on the predicted replication rate in experimental economics being 60.1 percent against an observed rate of 61.1 percent, suggesting that issues with common power calculations can explain essentially the entire gap even in a simple model without treatment effect heterogeneity, researcher manipulation or measurement error; on the corresponding model for psychology predicting 53.9 percent against an observed 35.6 percent, accounting for two-thirds of the replication rate gap; on mean intended power of 92 percent in both applications; on the replication rate being defined as the share of original estimates whose replications have statistically significant findings of the same sign; and on a small number of original results with p-values slightly above 0.05 being treated as positive results in both applications. Note: a working paper rather than a peer-reviewed publication, presented here as a modelling result. arxiv.org
- Working paper. Identification of and correction for publication bias, on the systematic selection of results for replication in the Open Science Collaboration project being an analytical advantage for bias-correction purposes. Note: a working paper. arxiv.org
- Working paper. Identification of and correction for publication bias, on the Open Science Collaboration having considered studies published in Psychological Science, the Journal of Personality and Social Psychology, and the Journal of Experimental Psychology: Learning, Memory, and Cognition in 2008; on papers being assigned to replication teams on a rolling basis; and on 158 articles being made available for replication, 111 assigned, 100 completed in time for inclusion, and 84 of the 100 completed replications considering the final result of the original paper. Note: same working paper as reference 8, cited separately for the project's mechanics. arxiv.org
- Hagger, M. S., & Chatzisarantis, N. L. D., et al. (2016), as abstracted by the publishing repositories, on self-control being conceptualised as a limited resource which becomes depleted after a period of exertion resulting in self-control failure, and on the model typically being tested using a sequential-task experimental paradigm in which people completing an initial self-control task have reduced capacity and poorer performance on a subsequent task. Note: the same study as reference 5, cited separately for its statement of the theory. research.tilburguniversity.edu
- Baumeister, R. F., et al. (1998). Journal of Personality and Social Psychology, 74, 1252–1265, cited by the replication literature as the origin of the ego-depletion effect and of the letter-e crossing task. Note: we did not obtain this paper and cite it only as the study the later work targets. ncbi.nlm.nih.gov
- Sripada, C., Kessler, D., & Jonides, J. (2014). Methylphenidate blocks effort-induced depletion of regulatory control in healthy volunteers. Psychological Science, 25(6), 1227–1234, identified as the specific ego-depletion paradigm directly replicated by the multilab project. Note: we did not obtain this paper and cite it only as the replication target. journals.sagepub.com
- Carter, E. C., Kofler, L. M., Forster, D. E., & McCullough, M. E. (2015). A series of meta-analytic tests of the depletion effect: Self-control does not seem to rely on a limited resource. Journal of Experimental Psychology: General, 144, 796–815. Note: we did not obtain this paper; its figures are taken from the discussion in reference 5. ncbi.nlm.nih.gov
- Dang, J. (2016). Commentary: A multilab preregistered replication of the ego-depletion effect. Frontiers in Psychology, 7, 1155, on the ego depletion effect not having been replicated by a project including 23 laboratories (N = 2,141) in both English-speaking and non-English-speaking countries; and on cautious attention nonetheless being due to the effectiveness of the depleting letter-e crossing task used in the replicating project, noting that the depletion condition includes more complex rules of crossing than the control condition. Note: a published commentary presenting a live methodological objection. ncbi.nlm.nih.gov
- Working paper surveying research-methods literature, on psychology, behavioral economics and adjacent social sciences having confronted a replication crisis; on early meta-research arguing that selective reporting, low statistical power and strong incentives for novelty can yield a literature with inflated effect sizes and excess false positives; on attention shifting from outright fraud to researcher degrees of freedom, being flexible decisions about exclusion rules, stopping, outcome definitions and model specifications, and the garden of forking paths where many defensible analytic routes exist; on survey evidence reinforcing that such questionable research practices were part of routine scientific workflow; and on preregistration emerging as one of the most widely endorsed methodological reforms. Note: a working paper, used for its survey of the methods literature rather than as primary authority for any empirical claim. arxiv.org
This article concerns research quality and is not investment, tax or psychological advice. The full texts of the three replication projects and of the ego-depletion replication were not obtained; the article relies on published abstracts, journal summaries and subsequent reanalyses. Three references are working papers rather than peer-reviewed publications and are identified as such. Two sources give different target windows for the economics replication project and the difference is not resolved. The grading scheme and the halving rule are the authors' own proposals.