The fiftieth article reported that eleven of forty-nine had found a famous result reproducible from something other than its stated mechanism. This is the strongest instance we know of, and it was missing from the series entirely.

Key Takeaway

From the abstract: "Despite the use of automated timing methods and a larger sample, our first experiment failed to show priming. Our second experiment was aimed at manipulating the beliefs of the experimenters... Strikingly, we obtained a walking speed effect, but only when experimenters believed participants would indeed walk slower."[1] The authors conclude that "both priming and experimenters' expectations are instrumental in explaining the walking speed effect."[1]

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the replication result. Published in a peer-reviewed open-access journal, and we obtained the paper itself rather than an abstract, which has been true of very few articles here.

Grade A that experimenter belief produced the effect in their second experiment, which is the paper's own reported result and the reason it matters.

Grade B that further stereotype-priming effects failed to replicate, reported from a citing source naming two studies we did not obtain.

Ungraded for the status of behavioural priming generally, which is a large literature we have not surveyed and about which this article makes no claim.

Our position: this is the best-sourced article in the series so far, and its scope is narrower than the topic's reputation suggests.

A Note On Method

Everything here is verified to August 2026.

Unusually, we obtained the replication paper itself, which is published open access under a Creative Commons licence permitting reproduction with attribution[1][2]. We have its abstract, its introduction and its stated motivations.

We did not obtain the 1996 original. Everything about it here comes from the replication paper's description and from citing sources.

We did not obtain the two later failed replications and report them from a citing source naming them.

We did not obtain any statement by any third party about this episode, including several widely quoted ones, and report none.

All simulation is ours and uses invented parameters.

This article discusses research methods. It is not advice on any commercial or professional matter, though its final sections draw our own analogies.

The Original Finding

What was claimed, described by the people who tried to reproduce it.

The replication describes it: "Bargh, Chen, and Burrows' (1996) famous study, in which participants unwittingly exposed to the stereotype of age walked slower when exiting the laboratory, was instrumental in defining this perspective."[1]

And its significance: "This striking finding, now widely cited, established that priming may occur automatically and influence behavior with little or no awareness. It subsequently generated considerable further research in social psychology."[3]

Two observations, ours.

The finding is remarkable if true, and that is not a criticism. Unscrambling sentences containing age-related words, then walking measurably slower without knowing why, would be a striking demonstration of unconscious influence on behaviour.

And note the framing the replicators use: the study defined a perspective. Its influence ran well beyond its own result, which is what makes its status consequential.

How The Priming Worked

The mechanism, because the design is elegant and worth understanding before it is questioned.

A description quoting the replication paper: the original "involved asking participants to indicate which word was the odd one out amongst an ensemble of scrambled words a number of which, when rearranged, form a sentence. Unbeknownst to participants, the word left out of the sentence was systematically related to the concept of 'being old'."[4]

Two observations, ours.

The manipulation is genuinely subtle. Participants are doing a language puzzle, and the primed content is the word they are discarding rather than the one they are using.

And the outcome measure, walking speed on leaving, is one participants have no reason to connect to the task. As a design for demonstrating unconscious influence it is close to ideal, which is part of why it travelled.

Why They Replicated It

The motivation, which the paper states plainly.

"This was motivated by three main reasons. The first is simply that the finding, influential as it is, was only replicated twice so far, with neither replication being exact."[3]

On the first of those replications: it "required participants to rate the walking speed of a character drawn on a sheet of paper after they had been primed with extreme exemplars associated to the concept of 'speed'", and found the expected effect. But "the latter only tells us of a bias in judgments of speed, as it did not require participants to actually perform any behavior."[3]

The paper adds that "our third motivation is methodological."[2]

Three observations, ours.

Twice, neither exact, for a finding that defined a research perspective. That is the pattern the fiftieth article identified as the most common in this series.

The distinction drawn about the earlier replication is exactly right and rarely made. Rating how fast a drawn figure walks is a judgment. Walking slowly is a behaviour. A study that establishes the first does not establish the second, however similar the theory.

And we did not obtain either prior replication and report only the replicators' characterisation of them.

Experiment One: Nothing

The straightforward result.

"Despite the use of automated timing methods and a larger sample, our first experiment failed to show priming."[1]

Three observations, ours.

Two improvements are named, and both matter. Automated timing removes the human with the stopwatch. A larger sample addresses power.

On its own this is an ordinary failed replication, of the kind this series has reported many times. It would have been a modest result.

And the paper does not stop there, which is what makes it the article it is. Failing to reproduce a finding tells you something did not happen. It does not tell you what happened originally.

Experiment Two: The Manipulation

The design that answers the second question.

"Our second experiment was aimed at manipulating the beliefs of the experimenters: Half were led to think that participants would walk slower when primed congruently, and the other half was led to expect the opposite."[1]

Three observations, ours.

The subject of the experiment has changed. The participants are no longer what is being studied; the experimenters are.

And the manipulation is two-sided. Half were told to expect the opposite result, which is a stronger design than simply comparing informed to uninformed, because it can detect an effect running in either direction.

This is the same structure this series praised in the twentieth article's adversarial collaboration and the forty-sixth's opposite-predictions design. Two conditions that predict contrary outcomes are worth far more than one condition and a hope.

Strikingly

The result, including the authors' own adverb.

"Strikingly, we obtained a walking speed effect, but only when experimenters believed participants would indeed walk slower."[1]

Three observations, ours.

The effect came back. Experiment one found nothing with automated timing; experiment two found the effect, and what predicted it was not the participants' condition but the experimenter's belief.

Note that this is not merely a measurement error story. The experimenters were not simply misreading a stopwatch, because the timing method was under the authors' control; the paper's framing is about beliefs shaping the interaction.

And we did not obtain the effect sizes for either experiment and report none.

The Conclusion They Drew

What the authors concluded, which is more careful than the result invites.

"This suggests that both priming and experimenters' expectations are instrumental in explaining the walking speed effect."[1]

And their closing sentence, which we quote exactly as published: "We conclude that unconscious behavioral priming is real, while real, involves mechanisms different from those typically assumed to cause the effect."[1]

Three observations, ours.

That closing sentence appears identically across every source we checked, including the publisher's own page and the paper itself, so the doubled construction is the published text rather than a transcription error on our part. Read past it and the claim is clear.

The claim itself is conservative. The authors do not say priming is fake. They say it is real and works differently than assumed, having just demonstrated that experimenter belief can generate the signature result.

And the word both is doing real work. Their conclusion is that expectations are also instrumental, not that they are the whole story.

And They Noticed The Primes

A smaller finding that undercuts the theory from another direction.

"Further, debriefing was suggestive of awareness of the primes."[1]

Two observations, ours.

The entire theoretical interest of the original rests on the influence being unconscious. If participants noticed the age-related words, the finding becomes considerably less remarkable even where it appears.

Suggestive is the authors' own hedge and we report it at that strength. This is a debriefing observation, not a measured variable.

What An Unblinded Measurer Can Produce

Because the general lesson is quantifiable. Our own simulation, invented parameters, standard two-sample comparison; this demonstrates a mechanism and reproduces nobody's data.

We set the true effect to zero, gave thirty subjects per group, and let the person recording the measurement shade each reading very slightly in the direction they expected.

With no shading: a false positive rate of about 5.5 percent, which is the expected baseline.

Shading by a tenth of a standard deviation: 7.2 percent.

By two tenths: 12.6 percent.

By three tenths: 21.7 percent.

By half: 49.1 percent.

Two observations.

At half a standard deviation of unconscious shading, a study of a completely non-existent effect reaches conventional significance about half the time.

And the shading required is small. Three tenths of a standard deviation is not fraud; it is the amount by which a person who expects an outcome might round a judgment call.

We Overstated It And Checked

A note on our own process, in the manner of the fiftieth article.

Our first description of the simulation output said that a shading of a fifth of a standard deviation made a null effect significant about a quarter of the time.

The actual figures are 12.6 percent at two tenths and 21.7 percent at three tenths. We had compressed two rows into one claim in the wrong direction.

One observation, ours. It was caught by reading the table rather than our own summary of it, which is the same method that caught the error in the previous article, and it is the only method that has reliably worked.

The General Form

Why this is not a story about psychology. Ours.

Three points.

The mechanism requires only that the person recording a measurement knows which condition produced it, and has a view about how it should come out. Neither condition is unusual.

Which puts it in the same family as two earlier articles. The forty-third found outcome knowledge reorganising the perceived relevance of evidence; the forty-fourth found evaluators rating identical reasoning differently by result. All three are about a measurement contaminated by knowing what it should show.

And the remedy is the same in all three and is structural rather than motivational: separate the person who knows the expected answer from the person who records the result.

It Was Not The Only One

Further failures, reported at one remove.

A citing source records that "Shanks et al. (2013) and O'Donnell et al. (2018) were not able to replicate priming effects of stereotypes associated with intelligence (professor) vs. stereotypes associated with lower intelligence (soccer hooligans) on a knowledge test."[5]

We did not obtain either study and report the citing source's characterisation.

Two observations, ours.

That is a different paradigm from the walking speed study, testing whether priming a professor stereotype improves performance on a knowledge test.

Two independent failures on a second paradigm makes this a pattern rather than a single disputed study, which is why we grade the general claim B rather than declining to make it.

What It Triggered

The consequence, which is the constructive part of the story.

A citing source records: "Given growing questions about the reliability of published research, the journal Perspectives on Psychological Science then published a special section on replicability."[6]

We did not obtain that special section and report only that it followed.

Two observations, ours.

A single open-access replication in 2012 sits near the start of a chain that produced large multi-laboratory replication projects, several of which this series has reported.

And it is worth noticing what the replicators actually did. They did not write a commentary. They ran the study, then ran a second study to find out what the original had measured, which is more work than criticism requires and considerably more useful.

What This Does Not Say

Four limits, stated explicitly. Ours.

It does not say the original authors did anything improper. Unconscious experimenter expectancy is, by definition, not a choice, and the replicators' own conclusion assigns a role to priming as well.

It does not say all priming research is unsound. The authors conclude unconscious behavioural priming is real. We have not surveyed that literature and this article makes no claim about it.

It does not establish that experimenter expectancy explains the 1996 result. It establishes that expectancy can produce the same signature, which is a claim about what the evidence can distinguish rather than about what happened.

And it does not license the conclusion that unconscious influence on behaviour is impossible. That is a much larger question and nothing here bears on it.

Where This Lives In A Firm

The application. Ours, untested, and offered as recognisable situations.

Four settings where the person measuring knows the expected answer.

A pilot evaluated by the person who proposed it. The judgment calls in deciding whether it worked are made by someone with a view.

A file reviewed by the partner who staffed it. Every borderline call has an expected direction.

A new process assessed against the old one, where the assessor knows which period is which.

And any client satisfaction measure collected by the person who served the client, where the collection itself is an interaction.

One observation. In none of these is anyone behaving badly, which is precisely the point. The effect in the replication was produced by experimenters who were told what to expect and presumably tried to be accurate.

What To Do

Separate measurement from expectation. The person recording a result should not be the person who knows how it is supposed to come out.

Automate the measurement where you can. The replication's first experiment used automated timing and found nothing.

Design conditions that predict opposite outcomes. The second experiment told half the experimenters to expect the reverse, which is far stronger than informed against uninformed.

Ask what a failed replication has and has not shown. Failing to reproduce something says an effect did not appear; it does not say what the original measured, and finding that out took a second experiment.

Distrust small unblinded comparisons. On our own simulation, half a standard deviation of shading makes a non-existent effect significant about half the time at thirty per group.

Treat this as the same problem as hindsight and outcome bias. All three are measurements contaminated by knowing what they should show, and all three have the same structural remedy.

Do not over-read it. The replicators conclude unconscious priming is real and works differently than assumed, which is narrower than the way this episode is usually summarised.

The Limits Of This Analysis

Several caveats matter. This article discusses research methods and is not advice on any commercial or professional matter; its applications are our own analogies. Everything is verified to August 2026. We obtained the replication paper's abstract and introduction, which is unusually good sourcing for this series, but not its methods, statistics or effect sizes, and we report none. We did not obtain the 1996 original, and everything said about it here comes from the replicators' description and citing sources. We did not obtain either of the two earlier replications the paper discusses, nor the two later failed replications on a different paradigm, nor the journal special section, and report all of these from citing sources or from the replication's own characterisation. We did not obtain any third-party commentary on this episode and report none, including statements widely attributed to prominent figures. All simulation is ours, uses invented parameters including a sample of thirty per group, and demonstrates a mechanism rather than reproducing anyone's data; our first description of its output was wrong and is corrected in the body. The closing sentence of the paper's abstract is quoted exactly as published, including a doubled construction that appears identically across every source we checked. This article makes no claim about the status of behavioural priming as a field, which we have not surveyed.

Frequently Asked Questions

What was the original finding?
That participants unwittingly exposed to the stereotype of age walked slower when leaving the laboratory. The priming came from a scrambled-sentence task in which the word discarded from each sentence related to being old, and the outcome was walking speed on exit, which participants had no reason to connect to the task.
What did the replication find?
Two things. With automated timing and a larger sample, nothing. Then, having told half the experimenters to expect slower walking and half to expect the opposite, a walking speed effect appeared, but only among experimenters who expected it.
Does that mean the original was wrong?
It means expectancy can produce the same signature, which is a claim about what the evidence can distinguish rather than about what happened in 1996. The replicators themselves conclude that both priming and experimenters' expectations are instrumental, and that unconscious priming is real but works differently than assumed.
Why is the second experiment the important one?
Because a failed replication tells you an effect did not appear, not what the original measured. Answering the second question required running another study, which is more work than criticism demands and considerably more useful.
How much unconscious shading would it take?
Less than you would expect. On our own simulation with thirty per group and a true effect of zero, shading each reading by three tenths of a standard deviation makes the null result significant about 22 percent of the time, and half a standard deviation about half the time.
What does this mean outside a laboratory?
The mechanism needs only that the person recording a measurement knows which condition produced it and has a view about how it should come out. That describes a pilot evaluated by whoever proposed it, a file reviewed by whoever staffed it, and most satisfaction scores collected by whoever did the work.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This is the best-sourced article in the series so far, because the paper it covers is open access, and it corrects an error in its own simulation summary in the body.

References

  1. Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral Priming: It's All in the Mind, but Whose Mind? PLOS ONE, 7(1), e29081, DOI 10.1371/journal.pone.0029081, published 18 January 2012, publisher's page reproducing the abstract, on the perspective that behavior is often driven by unconscious determinants having become widespread in social psychology; on Bargh, Chen and Burrows' 1996 study, in which participants unwittingly exposed to the stereotype of age walked slower when exiting the laboratory, having been instrumental in defining that perspective; on the authors presenting two experiments aimed at replicating the original study; on their first experiment having failed to show priming despite the use of automated timing methods and a larger sample; on the second experiment having manipulated the beliefs of the experimenters, with half led to think participants would walk slower when primed congruently and half led to expect the opposite; on the authors strikingly having obtained a walking speed effect but only when experimenters believed participants would indeed walk slower; on this suggesting that both priming and experimenters' expectations are instrumental in explaining the walking speed effect; on debriefing having been suggestive of awareness of the primes; and on the authors concluding that unconscious behavioral priming is real, while real, involves mechanisms different from those typically assumed to cause the effect. Note: the publisher's page for an open-access article. We obtained the abstract in full; the closing sentence is quoted exactly as published, including a doubled construction appearing identically in every source checked. journals.plos.org
  2. Full text of the same paper as published, reproducing the abstract identically, recording the authors' affiliations at the Consciousness, Cognition and Computation Group and Social Psychology Unit of the Université Libre de Bruxelles and the Social and Developmental Psychology Department at the University of Cambridge; noting that the work was supported by named research foundations; recording the article's distribution under a Creative Commons Attribution License permitting unrestricted use with attribution; and stating that the authors' third motivation was methodological. Note: the paper itself, open access. We obtained the abstract, front matter and portions of the introduction, and not the methods, statistics or effect sizes. journals.plos.org
  3. Repository copy of the same paper, reproducing the abstract and introduction, on the striking and now widely cited finding having established that priming may occur automatically and influence behavior with little or no awareness and having generated considerable further research; on the replication having been motivated by three main reasons, the first being that the finding, influential as it is, was only replicated twice so far with neither replication being exact; and on the first of those replications having required participants to rate the walking speed of a character drawn on a sheet of paper after priming with extreme exemplars associated with the concept of speed, finding the expected priming effect, while telling us only of a bias in judgments of speed as it did not require participants to actually perform any behavior. Note: a repository copy; our source for the paper's stated motivations and its account of prior replications, neither of which we obtained. ncbi.nlm.nih.gov
  4. Research institute commentary quoting the replication paper's opening, on Bargh, Chen and Burrows having demonstrated in their seminal series of experiments that activating a trait construct such as being old is sufficient to elicit behavioral effects in the absence of awareness; and describing the demonstration as involving asking participants to indicate which word was the odd one out amongst an ensemble of scrambled words, a number of which when rearranged form a sentence, with the word left out of the sentence being systematically related to the concept of being old. Note: a research institute's commentary quoting the paper. Our source for the description of the priming task; we did not obtain the 1996 original. cognitionandculture.net
  5. Repository page for the replication carrying scholarly citing text, on the results of the highly cited 1996 study finding a priming effect of elderly stereotypes on participants' walking speed not having been replicated by Doyen and colleagues; and on Shanks and colleagues (2013) and O'Donnell and colleagues (2018) not having been able to replicate priming effects of stereotypes associated with intelligence, being professor, against stereotypes associated with lower intelligence, being soccer hooligans, on a knowledge test. Note: a repository page reproducing text from citing papers. We did not obtain either of the two studies named and report this characterisation only. researchgate.net
  6. Repository page carrying further citing text, on Bargh, Chen and Burrows (1996) having found that priming elderly stereotypes caused subjects to walk more slowly and having replicated the finding in their own paper; on Doyen, Klein, Pichon and Cleeremans (2012) not having been able to replicate the finding unless the experimenters were aware of the hypotheses being tested; and on the journal Perspectives on Psychological Science having then published a special section on replicability, given growing questions about the reliability of published research. Note: a repository page reproducing citing text. We did not obtain the special section referred to and report only that it followed. researchgate.net

This article discusses research methods and is not advice on any commercial or professional matter. The replication paper is open access and its abstract and introduction were obtained; its methods, statistics and effect sizes were not, and none are reported. The 1996 original was not obtained. No third-party commentary on this episode is reported. All simulation is the authors' own, uses invented parameters, and demonstrates a mechanism rather than reproducing anyone's data.