Someone has advised you to change a default, reorder a menu, or add a line about what most customers choose. The evidence base for that advice had a difficult year in 2022, and the dispute is unusually worth understanding because both sides published their data.

Key Takeaway

Mertens and colleagues meta-analysed 447 nudge experiments and concluded that choice architecture is an effective and widely applicable behavior change tool, while also reporting moderate publication bias[1][2]. Maier and colleagues took the same dataset and argued the bias finding was the real headline, concluding the literature is characterised by severe publication bias and that after correcting for it, no evidence remains that nudges are effective[2][3]. A third commentary in the same issue argued there is no reason to expect large and consistent effects[4], and a fourth critique attacked both analyses for pooling heterogeneous studies[5].

Our Grades For These Claims

Applying the scheme from the first article in this series.

That the nudge literature contains publication bias is Grade A. This is the one point on which every party agrees. The original meta-analysis reported it; the critics agree it exists and differ only on its severity[2].

That nudging is an effective general-purpose tool is Grade C. It rests on a large meta-analysis whose bias-corrected reanalysis, by a different team, found nothing left.

Any specific nudge effect size is Grade D. The headline figure is disputed at the level of whether the effect exists at all, and a separate critique argues the pooled average is not a meaningful quantity in the first place[5].

Our own summary, stated up front so a reader can weigh the rest against it: this is the most public and best-documented disagreement in behavioural science, and its resolution is not available. What is available is a clear view of why competent people disagree, which is more useful than a verdict.

A Note On Method

Everything here is verified to August 2026.

We obtained the full text of the Maier and colleagues letter through a public repository copy[3], and the abstract and quoted passages of the Mertens and colleagues meta-analysis as reproduced in that letter and in the surrounding commentary[1][2].

We did not obtain the Mertens meta-analysis directly, which matters, because we are describing a paper largely through the words of the team disputing it. We have flagged every place that occurs.

We did not obtain the per-category or per-domain results of the bias-corrected reanalysis, and therefore cannot say which kinds of nudge survived correction and which did not. This is the most significant gap in the article and we return to it.

We did not obtain the Szaszi and colleagues commentary beyond its title, nor the Datacolada post beyond its title and the rejoinder's characterisation of it.

This article reviews research methods and behavioural evidence. It is not marketing, legal or investment advice, and choice architecture in consumer-facing contexts carries obligations under consumer protection and competition law that this article does not address.

What Was At Stake

Why this dispute mattered beyond academia. This section is our own.

The letter's opening line captures the position: Thaler and Sunstein's "nudge" has spawned a revolution in behavioral science research, and despite its popularity, the nudge approach has been criticized for having a "limited evidence base"[3].

By 2022 the concept had travelled a long way from a book. Governments had established behavioural insights teams. Consultancies sold nudge programmes. Regulators used the framework. And a very large number of businesses had reordered menus, changed defaults and added social proof messages on the strength of it.

Two observations, ours.

The Mertens meta-analysis was explicitly a response to the evidence-base criticism. The letter describes it as seeking to address that limitation with a timely and comprehensive metaanalysis[3], which is a fair and generous characterisation from a team about to dismantle its conclusion.

And that framing raises the stakes. This was not a routine literature review. It was the field's attempt to settle the question of whether its most successful export actually works.

January: The Meta-Analysis

The first paper.

Mertens, Herberz, Hahnel and Brosch published The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains in the Proceedings of the National Academy of Sciences, 119(1), e2107346118[1]. A later methodological paper records that it compiled data from 447 nudge experiments[6].

Its headline finding, as quoted by the team that disputed it, is that choice architecture, meaning nudging, is an effective and widely applicable behavior change tool[3].

The paper also reported moderate publication bias[3], which becomes the pivot of everything that follows.

Two observations, ours.

Reporting the bias at all is to the authors' credit. They tested for it, found it, and published it. Nothing that follows involves anyone concealing anything.

And we should be explicit about our sourcing here: we are quoting the meta-analysis's own conclusions from the letter criticising it. The quotations are given with page references and the critics thank the original authors for their data, which is reassuring, but a reader wanting the original's full argument should obtain the original.

July: The Letter

The response, published six months later in the same journal.

Maier, Bartoš, Stanley, Shanks, Harris and Wagenmakers published No evidence for nudging after adjusting for publication bias in PNAS, 119(31), e2200300119, on 19 July 2022[2][7].

Their conclusion, stated in their own words: we conclude that the "nudge" literature is characterized by severe publication bias. Contrary to Mertens et al., our Bayesian analysis indicates that, after correcting for this bias, no evidence remains that nudges are effective as tools for behaviour change[3].

The method: model-averaged posterior mean effect size estimates with 95% credible intervals and Bayes factors for the absence of the effect, computed for the combined sample and also split by either the domain or intervention category[3].

Two observations, ours.

Note the phrase Bayes factors for the absence of the effect. This is not merely failing to find an effect. The analysis was constructed to quantify evidence for the null, which is a stronger claim than an inconclusive result.

And note that they split by domain and intervention category. That means per-category results exist. We did not obtain them, and they are precisely what a practitioner would want, because they would say whether any particular kind of nudge survived.

The Sentence That Did The Damage

One line in the letter does more work than the statistics.

It reads: we propose their finding of "moderate publication bias" (p. 1) is the real headline[3].

Three observations, ours, and this is the most instructive rhetorical move in the dispute.

The critics did not accuse the original authors of missing the bias. They accused them of reporting it and then not following it through.

Which means the disagreement is not about what is in the data. It is about which finding is the important one, and that is a judgment rather than a computation.

And it generalises well beyond nudging. In any analysis that reports both a headline result and a limitation, there is a decision about which one leads. That decision is rarely presented as a decision, and it determines what the reader takes away.

Our own view is that this is the single most transferable idea in this article. When you read a study, look at what it reports in its limitations, and ask what the paper would have said if that limitation had been the headline instead.

Moderate Or Severe

A disagreement about an adjective, which turns out to be the whole dispute.

The original team described the publication bias as moderate[3]. The critics described the same literature as characterised by severe publication bias[3].

Two observations, ours.

Both teams were looking at the same dataset. The critics state that their analysis is based on the corrected dataset supplied by the original authors, and thank them for sharing well-documented data and code[2].

So the difference is not in the evidence. It is in how much correction the evidence warrants, which depends on the statistical model chosen, and the choice of model is a defensible judgment on which specialists differ.

Our own summary of the state of play: everyone agrees the literature is biased. The question is by how much, and the answer determines whether the effect is 0.43 or approximately nothing. There is no neutral ground between those two positions, and no further data would resolve it, because the dispute is about how to treat the data that exists.

What Bias Correction Actually Does

The concept underneath, explained plainly, because the argument is unintelligible without it. This explanation is ours.

Suppose a hundred teams test the same nudge, and it has no real effect. By chance, a handful will produce a significant positive result. Those get published; the rest go in a drawer.

A meta-analyst later collects the published studies and finds a strong average effect. The average is correct as a description of the published literature and wrong as a description of reality.

Publication bias correction methods try to reconstruct what the whole distribution must have looked like, using the shape of what was published. If small studies report large effects and large studies report small ones, that pattern is a signature of selection, because a small study needs a large effect to reach significance.

Three consequences.

These corrections are inferences about missing data, and they necessarily rest on assumptions about how the selection worked.

Different methods make different assumptions and therefore produce different corrections. The critics used a model-averaging approach precisely to avoid committing to one[3].

And this is why two competent teams can reach opposite conclusions without either being incompetent or dishonest. The disagreement is about the invisible part of the literature, and reasonable people can model an absence differently.

A Ratio Worth Understanding

A figure from a later paper that illustrates what the correction methods are interrogating.

A methodological paper working with the same data notes that it restricted attention to the subset of 261 nudge experiments whose p-value was below the 5 percent two-sided significance level, out of the 447 compiled[6].

By our own arithmetic, that is 58.4 percent of the experiments reporting a statistically significant result.

Two observations, ours, and the first is a caution.

That figure is not evidence of bias by itself. A literature studying a real and reasonably large effect with adequately powered studies should produce a high proportion of significant results. That is what power means.

The bias question is whether the studies were powered well enough to justify that proportion. If the average study was small and the true effect modest, then 58 percent significant is more than the design should have delivered, and the excess has to come from somewhere.

That is the calculation, in essence, that separates the two teams. The first article in this series described the same logic being applied to psychology and economics replication rates, where power calculations explained most of the apparent failure. Here it is applied to a single literature, and the answer is contested.

And A Third Voice

The dispute was not two-sided.

Szaszi, Higney, Charlton, Gelman, Ziano, Aczel and colleagues, including Tipton, published No reason to expect large and consistent effects of nudge interventions in the same volume and issue, PNAS 119(31), e2200732119[4].

We obtained its title and citation only, and report none of its content.

Two observations, ours.

The title alone establishes that the January paper drew at least three separate published commentaries, from different teams, in the same issue. A further commentary by Bakdash and Marusich is also referenced[5].

And the framing of that title is different again from the Maier position. No reason to expect large and consistent effects is not the same claim as no evidence of any effect. It points at heterogeneity rather than at absence, which is the theme of the next section.

The Critique That Hit Both Sides

A fourth intervention, and the most uncomfortable one.

The Maier team's own rejoinder records that a post titled "Meaningless Means: The Average Effect of Nudging is d = 0.43" critiques both the PNAS meta-analysis and their own commentary for pooling studies that are very heterogeneous[5].

They add that the critique echoes many comments which raised the heterogeneity question immediately after our and the other commentaries were published[5].

Three observations, ours.

The argument is that the average is not a meaningful quantity, because the 447 experiments concern radically different interventions in radically different domains. Averaging a default on a pension form with a sign about towel reuse produces a number that describes nothing.

If that critique holds, it damages both published positions equally. The question does nudging work would then be malformed rather than merely unanswered.

And it is the same critique the third article in this series encountered in the choice overload literature: a pooled mean concealing populations that behave differently, where the useful work is identifying the conditions rather than computing the average.

The Rejoinder

How the criticised team responded, which is the most creditable moment in the whole episode.

They wrote: we thank Datacolada for writing up this post to re-iterate the important issue of heterogeneity and for inviting us to write a response to their post. We take this opportunity to reflect on our commentary and the critique of meta-analyses with heterogeneous studies[5].

Two observations, ours.

A team that had just published a high-profile null result accepted a critique of their own analysis, described the criticism as raising an important issue, and thanked the critics.

We did not obtain the substance of their response and cannot report how they resolved the heterogeneity objection. What we can report is the conduct.

Our own view: whatever the truth about nudging turns out to be, this exchange is a considerably better advertisement for the field than the finding itself would have been either way.

What To Make Of All This

Our own assessment, offered as reasoning rather than as a finding.

Four positions are on the table, and they are not mutually exclusive.

Nudging works, at around a small-to-medium effect, per the original meta-analysis.

Nudging shows no effect once bias is corrected, per the Maier letter.

There is no reason to expect large and consistent effects, per the Szaszi letter.

And the pooled average is not a meaningful quantity, per the heterogeneity critique.

Our reading is that the third and fourth are the most likely to survive, and that they are compatible with each other. A field studying hundreds of different interventions across dozens of domains should not expect a single answer, and the search for one produced a number that two teams could argue about indefinitely.

Which suggests the practically useful question is not whether nudging works. It is whether this specific change, in this specific context, moves this specific behaviour, which is a question no meta-analysis can answer for you.

The Part That Should Reassure You

A structural observation about how this dispute was possible at all. This section is our own.

The critics state that their analysis is based on the corrected dataset published by the original authors, cite the specific file and repository, and write: we thank Mertens et al. for sharing well-documented data and code[2].

Both teams' data and analysis scripts are deposited in a public repository, with dates[2].

Three consequences.

The reanalysis took six months, not the decade or more that this series has repeatedly documented for corrections to propagate. The ego depletion case took roughly seventeen years from finding to multilab null.

It happened because the data was public. Had the original authors not shared it, the dispute could not have occurred and the meta-analysis would have stood unchallenged in a very prominent journal.

And every reader can, in principle, check both teams' work, which is the opposite of the situation this series has described in almost every previous article.

Our own view: this is what a functioning correction mechanism looks like, and the discomfort of watching a headline finding dismantled in public is the price of having one.

A Deeper Problem Than Nudging

A commentary on the episode that generalises, quoted because it names something this series keeps encountering.

A later paper observes that a group may find a small to medium effect size for nudging behavioral interventions, while others can make different assumptions with the same data and find no effect whatsoever, and that while disagreement and debate are essential parts of the scientific process, existing norms often result in incommensurate theories and a set of findings which, while each interesting and defensible on their own, cannot be put together or made sense of in aggregate[8].

Two observations, ours.

The phrase each interesting and defensible on their own is the key. This is not a story about anyone being wrong. It is a story about a field producing results that cannot be combined.

And it is the fourth time in nine articles that this series has landed on the same structural problem: a literature whose aggregate says less than its parts, where the useful work is at the level of specific conditions and the headline number is an artefact of the aggregation.

So Can You Just Test It Yourself

The obvious response, and the arithmetic behind it. All calculations here are ours.

The first article in this series argued that where the literature is weak, a finding is best treated as a hypothesis about your own data. Nudges look ideal for that, because a default or a menu order can be varied by channel or period.

So: how many observations would a business need to detect an effect of a given size? Using standard power arithmetic for a two-group comparison at 80 percent power and a two-sided 5 percent threshold:

At the disputed headline effect, roughly 85 per group, 170 total.

At a small but commercially useful effect, roughly 393 per group, 786 total.

At a marginal effect, roughly 1,570 per group, 3,140 total.

At a trivial effect, roughly 6,280 per group, 12,560 total.

We flag that this is standard textbook power arithmetic computed by us, assumes a simple two-group comparison on a continuous outcome, and that real conversion or uptake tests need different and usually larger treatment.

The Trap In That Answer

Why the arithmetic above is less encouraging than it looks. This section is our own analysis.

Read the four rows again in light of the dispute.

If the effect is near the original meta-analysis estimate, a few hundred customers settles the question and any reasonably busy business can run the test.

If it is near the bias-corrected estimate, no small business will ever detect it, because the required sample runs into the thousands or is unbounded.

And you do not know which world you are in before running the test.

Three consequences.

A small business that runs a nudge test on two hundred customers and finds nothing has learned almost nothing. That result is equally consistent with no effect and with a real effect too small to detect at that sample size.

Worse, a business that runs several such tests will occasionally find a significant result by chance, and will have no way to distinguish it from a real one. It will then implement it.

Which means the honest position is uncomfortable: the smaller your business, the less able you are to resolve this question for yourself, and the more you are forced to rely on a literature that is currently in dispute.

What Survives For A Business

The constructive residue. This section is our own reasoning.

Three things survive the dispute intact.

Choice architecture is unavoidable. Every menu has an order, every form has a default, every price list has a first item. You are not choosing whether to have choice architecture; you are choosing whether yours was designed or accidental. That is true regardless of effect sizes.

The cost of most changes is near zero. Reordering a list or changing a default is not a capital project. An intervention with a contested effect and no cost is in a different decision category from one with a contested effect and a large cost.

And large-sample settings can still test. A business with tens of thousands of transactions is in the world where these questions are answerable, and should test rather than read.

Two things do not survive.

Confident effect size claims, whether from a vendor, a consultant or this publication.

And the assumption that a nudge which worked elsewhere will work for you, which the heterogeneity critique undermines more thoroughly than the bias critique does.

A Note On Defaults Specifically

A gap we are declaring rather than filling.

Defaults are the intervention most people mean when they say nudge, and they are the case usually cited as strongest.

The bias-corrected reanalysis split its results by intervention category[3], so a per-category answer exists.

We did not obtain it.

Consequently this article says nothing about whether defaults specifically survived correction, in either direction. We note that one commentary in our sources remarks that default nudges can have a large effect on people's decisions without restricting their freedom of choice, while immediately directing the reader to a contrary citation[8], which is about as unresolved as a statement can be.

Anyone whose decision turns on defaults should obtain the figure from the letter itself rather than rely on this article, and we would say the same about any other specific intervention category.

What To Do

Stop citing a nudge effect size. The headline figure is disputed at the level of whether the effect exists, and separately criticised as a meaningless average over heterogeneous studies.

Ask any vendor which of the four positions they are relying on. Anyone selling behavioural design in 2026 who has not encountered this dispute is not current in their own field.

Read the limitations section first. The critics' central move was to argue that the original paper's reported limitation was the real headline. That reading strategy transfers to everything.

Design your defaults anyway. Every form has one. The choice is between designed and accidental, and that is true whatever the effect size turns out to be.

Match the intervention's cost to the evidence's strength. A free change under contested evidence is a reasonable bet. A paid programme under the same evidence is not.

Know whether you have the sample. If a plausible effect needs thousands of observations and you have two hundred, testing will not tell you anything and may mislead you.

Distrust your own successful tests. Run enough small experiments and one will come back positive by chance. That is the mechanism that produced the publication bias in the first place, operating inside your business.

Prefer specific evidence to general evidence. The heterogeneity critique implies that a study of your exact intervention in your exact domain is worth more than a meta-analysis of 447 experiments in others.

The Limits Of This Analysis

Several caveats matter. This article reviews research methods and behavioural evidence. It is not marketing, legal or investment advice; choice architecture in consumer-facing contexts carries obligations under consumer protection and competition law not addressed here. Everything is verified to August 2026. We did not obtain the Mertens and colleagues meta-analysis directly, and describe it largely through the words of the team disputing it, which is flagged at each occurrence. We did not obtain the per-domain or per-category results of the bias-corrected reanalysis, which means this article cannot say whether any specific intervention type, including defaults, survived correction; that is the most significant gap here. We obtained the Szaszi and colleagues commentary as a title and citation only and report none of its content. We obtained the Datacolada critique only through the title and the criticised team's own characterisation of it, and did not obtain the substance of their rejoinder. We state no effect size of our own for any nudge. The figure of 261 significant results out of 447 comes from a later paper applying its own selection procedure, and our percentage is arithmetic on those two numbers. Our power calculations are standard textbook arithmetic computed by us, assume a two-group comparison on a continuous outcome at 80 percent power and a two-sided 5 percent threshold, and are illustrative; real uptake and conversion tests require different and usually larger samples. The four-position summary, the reading strategy, the analysis of what survives for a business and the small-sample trap are our own reasoning, not findings from the literature. Roughly four years of subsequent literature was not reviewed.

Frequently Asked Questions

Does nudging work?
It is genuinely contested. A meta-analysis of 447 experiments concluded it is an effective and widely applicable tool. A second team, using the same dataset corrected for publication bias, concluded no evidence of effectiveness remains. Both are published in the same journal, six months apart.
Who is right?
The dispute is not resolvable by more data, because it concerns how to treat the data that exists. Both teams agree the literature contains publication bias; they disagree on whether it is moderate or severe, and that adjective determines whether the effect is around 0.43 or approximately nothing.
What is the most useful thing to take from this?
Probably the critics' central move: they argued that the original paper's reported limitation was the real headline. When reading any study, look at what it concedes in its limitations and ask what the paper would have concluded had that led instead.
Should I still design my defaults deliberately?
Yes. Every form has a default and every menu has an order, so the choice is between designed and accidental rather than between nudging and not nudging. That holds regardless of what the effect size turns out to be, and most such changes cost nothing.
Can I just test it in my own business?
Only if you have the sample. On standard power arithmetic, detecting the disputed headline effect needs roughly 170 observations, but a marginal effect needs over 3,000. You do not know which you are looking for beforehand, so a null result on a small test tells you very little.
Is there anything good about this episode?
A great deal. The reanalysis happened within six months because the original authors published their data and code, and the critics thanked them for it. Compare that with the seventeen years this series documented between a famous finding and its multilab null. This is a correction mechanism working.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article states no nudge effect size of its own, and declares that it could not obtain the per-category results that would answer the question most readers actually have.

References

  1. Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences, 119(1), e2107346118. Note: we did not obtain this paper. Its conclusions are quoted in this article as reproduced, with page references, by the team disputing it at reference 3. pubmed.ncbi.nlm.nih.gov
  2. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences, 119(31), e2200300119, publisher record, on the analysis being based on the corrected dataset supplied by the original authors; on the authors thanking Mertens and colleagues for sharing well-documented data and code; and on both teams' data and analysis scripts being deposited in a public repository with dates. pnas.org
  3. Maier and colleagues (2022), full text via public repository copy, on Thaler and Sunstein's nudge having spawned a revolution in behavioral science research; on the nudge approach having been criticised for a limited evidence base; on Mertens and colleagues seeking to address that limitation with a timely and comprehensive meta-analysis; on their headline finding being that choice architecture is an effective and widely applicable behavior change tool (p. 8); on the authors proposing that the finding of moderate publication bias (p. 1) is the real headline; on their conclusion that the nudge literature is characterised by severe publication bias and that after correcting for it no evidence remains that nudges are effective; and on the analysis reporting model-averaged posterior mean effect size estimates with 95 percent credible intervals and Bayes factors for the absence of the effect, for the combined sample and split by domain or intervention category. Note: we did not obtain the per-domain or per-category figures themselves. ncbi.nlm.nih.gov
  4. Szaszi, B., Higney, A., Charlton, A., Gelman, A., Ziano, I., Aczel, B., et al., & Tipton, E. (2022). No reason to expect large and consistent effects of nudge interventions. Proceedings of the National Academy of Sciences, 119(31), e2200732119. Note: we obtained the title and citation only and report none of its content. bayesianspectacles.org
  5. Maier, M., and Bartoš, F., Rejoinder: No Evidence for Nudging After Adjusting for Publication Bias, on a post titled "Meaningless Means: The Average Effect of Nudging is d = 0.43" critiquing both the PNAS meta-analysis and the authors' own commentary for pooling studies that are very heterogeneous; on that critique echoing many comments raising the heterogeneity question immediately after publication, citing Bakdash and Marusich (2022) and Szaszi and colleagues (2022); and on the authors thanking the critics for re-iterating the important issue of heterogeneity and taking the opportunity to reflect on their own commentary. Note: an academic blog post by the criticised authors; we did not obtain the substance of their response or the original critique. bayesianspectacles.org
  6. Working paper on adaptive procedures for boundary false discovery rate control, on Mertens and colleagues (2022) having compiled data from 447 nudge experiments; and on the authors restricting attention, due to concerns of publication bias, to the subset of 261 nudge experiments whose p-value was below the 5 percent two-sided significance level. Note: a working paper, used only for the two counts; the percentage derived from them is our own arithmetic. arxiv.org
  7. University College London Discovery repository record for Maier and colleagues (2022), confirming the journal, volume, issue, article number and DOI. Note: an institutional repository record, used to confirm the citation independently. discovery.ucl.ac.uk
  8. ResearchGate record and associated commentary for Maier and colleagues (2022), on a group finding a small to medium effect size for nudging while others make different assumptions with the same data and find no effect whatsoever; on existing norms often resulting in incommensurate theories and findings which, while each interesting and defensible on their own, cannot be put together or made sense of in aggregate; and on default nudges being able to have a large effect on people's decisions without restricting freedom of choice, immediately qualified by a contrary citation. Note: a publisher record with third-party commentary, not peer-reviewed. researchgate.net

This article reviews research methods and is not marketing, legal or investment advice. The meta-analysis under discussion was not obtained and is described largely through the words of the team disputing it. The per-category results that would answer whether any specific nudge type survived correction were not obtained. Two of the four published commentaries were obtained as titles only. No nudge effect size is stated by this publication. Power calculations are the authors' own standard arithmetic and are illustrative.