Many firms buy training built on a specific causal claim: that shifting something measured by a computerised task will change how people behave. That claim is testable, it has been tested at scale, and this article reports what the largest test found.
Key Takeaway
From 492 studies and 87,418 participants: "We found that implicit measures can be changed, but effects are often relatively weak (|ds| < .30). Most studies focused on producing short-term changes with brief, single-session manipulations."[1] Two independent summaries of the paper add the load-bearing result: "changes in implicit measures did not mediate changes in explicit measures or behavior."[2][3]
A Note On Scope
Stated first, because the scope of this article is narrower than the topic and the difference matters.
This article is about one empirical question: whether procedures that change scores on implicit tasks also change behaviour. That is a question about measurement and intervention efficacy.
It is not about whether discrimination exists, and nothing in it bears on that. It makes no claim about any group, any workplace, or any individual.
It is not employment, human resources, legal or equality law advice. Obligations in this area are set by legislation and by regulators, differ by jurisdiction, and are not addressed here.
And it does not examine whether implicit measures predict behaviour, which is a separate literature we did not survey. We say where that gap sits in its own section.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A that implicit measures can be changed but weakly. A network meta-analysis of 492 studies, abstract verified verbatim across six independent sources including the authors' own repository.
Grade B for the mediation failure, which is the article's most consequential claim and which reaches us through two independent summaries of the paper rather than the abstract text we obtained.
Ungraded for anything about predictive validity, which we did not examine.
Our position: the first finding is well established and the second is the one that matters, and we are explicit that our sourcing for the second is one step removed.
A Note On Method
Everything here is verified to August 2026.
We obtained the published abstract verbatim from six independent sources, including the lead author's institutional repository, an open preprint archive, the publisher's record, and a hosted copy of the paper itself, all agreeing word for word[1][4][5][6].
Our copy of the abstract truncates after its statement about which procedures worked best. The mediation finding therefore comes from two independent third-party summaries of the paper[2][3], and we flag that at each use.
We did not obtain the paper's body, its effect estimates by procedure, or its moderator analyses.
We did not survey the predictive validity literature and report nothing from it.
All arithmetic is ours.
The Study
The paper.
Forscher, Lai, Axt, Ebersole, Herman, Devine and Nosek published A Meta-Analysis of Procedures to Change Implicit Measures in the Journal of Personality and Social Psychology, 117(3), 522–559, in September 2019, DOI 10.1037/pspa0000160[1][4].
Its method and scale, quoted: "Using a novel technique known as network meta-analysis, we synthesized evidence from 492 studies (87,418 participants) to investigate the effectiveness of procedures in changing implicit measures, which we define as response biases on implicit tasks. We also evaluated these procedures' effects on explicit and behavioral measures."[1]
Three observations, ours.
The definition is careful and worth noticing: implicit measures are defined as response biases on implicit tasks. That is a description of what the instrument records, not a claim about what it reveals.
The scale is unusually large. Four hundred and ninety-two studies is more than most literatures this series has covered contain in total.
And the design tested three outcome families: implicit measures, explicit measures, and behaviour. Most intervention research stops at the first.
Who Wrote It
The detail that determines how to read this. Ours.
The seven authors are recorded at the University of Arkansas, Washington University in St. Louis, the University of Virginia, the University of North Carolina at Chapel Hill, the University of Wisconsin-Madison, and the Center for Open Science[6][7].
Three observations.
This is not an external critique. The author list includes figures central to this research area, including a researcher who developed a well-known prejudice intervention and a co-founder of the Center for Open Science whose name appears throughout the implicit cognition literature.
Which means the paper constrains its own authors' prior work. Whatever else this is, it is not a hostile outside attack on a field, and it should not be read as one.
A small detail from the preprint worth recording: the first two authors "contributed equally to this manuscript. Order was determined by coin flip."[7] That is a research-practice choice this publication approves of and rarely sees stated.
The First Finding
The good news for the intervention, and it is real.
"We found that implicit measures can be changed, but effects are often relatively weak (|ds| < .30)."[1]
Two observations, ours.
Can be changed is a positive result. The interventions are not inert; they move the thing they target.
And the bound is stated as an upper one. The abstract says effects are often weaker than 0.30, not that 0.30 is typical, which means the typical effect is somewhere below that and we do not know where.
How Large Is Weak
Placing the number, with our own conversions.
At d = 0.30, the distributions of treated and untreated overlap by roughly 88 percent. Most people in one group are indistinguishable from most people in the other.
Against what this series has reported: brainstorming penalty 1.39, the disputed nudge headline 0.43, feedback interventions 0.41, framing 0.31, implicit measure change below 0.30, growth mindset 0.08, choice overload 0.02.
Two observations.
It sits at the lower end of the middle, comparable to framing and well below the strongest effects covered here.
And a d of 0.30 on the target measure would still be worth having if it carried through. The next section is about whether it did.
The Second Finding
The result that determines whether any of this matters commercially, with its sourcing stated plainly.
Two independent third-party summaries of the paper state its finding as: "It is found that implicit measures can be changed, but effects are often relatively weak (|ds| < .30), and changes in implicit measures did not mediate changes in explicit measures or behavior."[2][3]
Our copy of the published abstract truncates before this sentence. We report it because two independent sources give it identically and because it is the paper's central practical conclusion, and we grade it B rather than A on exactly that basis.
Three observations, ours.
Did not mediate is a specific technical claim. It says the change in the implicit measure was not the pathway through which any behavioural change occurred.
That is stronger than saying behaviour did not change. It says that even where behaviour moved, the implicit measure was not carrying the effect.
And it targets the exact assumption the intervention rests on. The theory of change is that moving the measure moves the person, and this reports that the link was not found.
The Chain That Has To Hold
Why that single word matters so much. Ours.
The commercial claim behind this kind of training is a chain with two links.
First: the training changes the implicit measure. The meta-analysis says yes, weakly.
Second: the changed measure produces changed behaviour. This is the link the mediation analysis addresses.
Two consequences.
A chain is only as good as its weakest link, and a mediation failure is not a weak link but an absent one.
Which means the first finding, on its own, cannot justify the expenditure. Demonstrating that a training moves a score is demonstrating that the training does something, not that it does the thing anyone is buying.
Why Weak Stages Compound
A general point about chains, worked. Our own arithmetic, illustrative only; the meta-analysis reports a mediation failure, which is a stronger claim than a weak second link, and we use the chain here only to show why weak stages multiply.
If an effect must pass through two stages, the end-to-end association is roughly the product of the two.
With both links at r = .30, the end-to-end correlation is .090, equivalent to d = 0.18.
With links of .30 and .20: .060, or d = 0.12.
With links of .15 and .15: .022, or d = 0.05.
Two observations.
Even in the most generous case, with both links at the stated upper bound, the end-to-end effect lands at d = 0.18, which is near the growth mindset figure this series graded harshly in its thirteenth article.
And this is the general lesson, transferable beyond the topic: an intervention that works through an intermediate step inherits the product of both stages, so a two-stage theory of change needs both stages to be strong, and most are not.
What Moved The Measure Most
The paper's own breakdown, which is more useful than the headline.
"Procedures that associate sets of concepts, invoke goals or motivations, or tax mental resources changed implicit measures the most, whereas procedures that induced threat, affirmation, or specific moods/emotions changed implicit measures the least."[1]
We did not obtain the effect estimates for any individual procedure and report only this ordering.
Two observations, ours.
Note what changed the measure least: threat, affirmation, and mood induction. Several of those describe the emotional register of a good deal of workplace training.
And taxing mental resources appearing among the most effective is interesting rather than encouraging. A procedure that shifts a score by loading someone cognitively is telling you something about the instrument as much as about the person.
Almost Everything Was Short-Term
The methodological limitation the authors state about their own field.
"Most studies focused on producing short-term changes with brief, single-session manipulations."[1]
Three observations, ours.
This is a statement about what the literature contains, not about what is possible. The authors are reporting that the evidence base is dominated by brief single sessions.
Which means the question a firm actually has, whether anything persists, is largely unaddressed by the studies available.
And it echoes the twenty-ninth article in this series, where a review found expectancy effects may be more likely to dissipate than accumulate. Two separate literatures, the same problem: short-horizon evidence used to justify long-horizon claims.
A Name Conflict
A small sourcing note, recorded for consistency with the rest of this series.
The third author's surname appears as Axt in five of our six sources and as Ast in one institutional repository record[1].
The same record also gives that author's affiliation as Duke University, while the paper's own front matter gives the University of Virginia[1][7].
One observation, ours. This is the twelfth instance in this series of a bibliographic detail being reproduced inconsistently by careful sources. It changes nothing here, and the accumulation continues to be the point.
What This Article Does Not Examine
The gap, stated in the body rather than buried. Ours.
There is a separate and substantial literature on whether implicit measures predict behaviour, which is a different question from whether changing them changes behaviour.
Three points.
We did not survey it and report nothing from it. A reader forming a complete view needs it and will not get it here.
One of our sources reproduces a third-party summary asserting insufficient evidence for the claim that the Implicit Association Test measures individual differences in implicit social cognition[3]. We could not identify which paper that summarises, did not obtain it, and are recording it as an unverified pointer rather than a finding.
And the two questions are logically independent. An instrument could predict behaviour well while being hard to shift usefully, or be easy to shift while predicting nothing. Conflating them would be an error, and this article addresses only the second.
What This Does Not Mean
Four things this article is not saying, stated explicitly because the topic invites overreach in both directions. Ours.
It does not say that discrimination does not occur. Nothing in a meta-analysis of intervention procedures bears on that question, and this article makes no claim about it.
It does not say the researchers were wrong to try. Four hundred and ninety-two studies represent an enormous amount of serious work, and the meta-analysis exists because the field examined itself.
It does not say nothing works. It says that the specific pathway of changing an implicit measure was not found to carry through, which leaves every other pathway untouched.
And it does not settle anything about obligations. What an employer must do is set by law and regulators, differs by jurisdiction, and is outside this article entirely.
The Spending Question
The decision a firm owner actually faces. Ours, and framed as a budgeting question rather than anything else.
Three questions to ask of any programme sold on this basis.
What outcome is being measured? If the evaluation ends at a score on a task, it has measured the first link of a two-link chain and stopped.
Over what horizon? The meta-analysis reports the literature as dominated by brief single-session manipulations, so persistence claims need their own evidence.
And through what mechanism? If the answer is that shifting an implicit measure shifts behaviour, that is the specific claim the mediation analysis addresses.
One observation. None of those questions is hostile. They are the questions any firm would ask about any purchase whose benefit is asserted rather than demonstrated, and this series has applied the same three to nudging, growth mindset and feedback.
Where The Evidence Is Better
The constructive half, and it is not our own invention. Ours in application.
The sixth article in this series covered a large 2022 revision of the hiring validity literature, in which structured approaches to selection were central. That literature is about changing the procedure rather than changing the person.
Three observations.
A structured process is evaluated on the decision it produces, not on a score measured beforehand, which removes the mediation problem entirely because there is no intermediate step.
The thirty-second article's noise audit is the same shape: measure the decisions, not the decision-makers, and a firm can run it without anyone's cooperation or belief.
And the sixth article's own finding was that the corrections in that literature had been applied too generously, so we are not offering it as settled either. We are observing that process-level evidence is measured at the outcome, which is a structural advantage over any intervention that has to travel through a person's measured state.
What To Do
Ask what the evaluation measured. A demonstrated change in a task score is the first link of a chain, and the meta-analysis reports that link as weak.
Ask about the mechanism. Two independent summaries of the paper state that changes in implicit measures did not mediate changes in behaviour, which addresses the assumed pathway directly.
Ask about the horizon. The authors describe the literature as dominated by brief, single-session manipulations, so any persistence claim needs separate support.
Note who wrote it. The author list includes leading figures in this research area, so the paper is a field examining itself rather than an outside attack.
Remember chains multiply. On our own arithmetic, even two links at the stated upper bound give an end-to-end effect near d = 0.18.
Prefer measuring the decision to measuring the decider. Process-level evidence is evaluated at the outcome, which removes the intermediate step where this chain broke.
Keep the two questions separate. Whether an instrument predicts behaviour and whether changing it changes behaviour are logically independent, and this article addresses only the second.
Do not extend any of this to the underlying social question. A meta-analysis of intervention procedures says nothing about whether discrimination occurs.
The Limits Of This Analysis
Several caveats matter. This article concerns measurement and intervention efficacy only. It is not employment, human resources, legal or equality law advice, makes no claim about whether discrimination exists, and makes no claim about any group, workplace or individual; obligations in this area are set by legislation and regulators, differ by jurisdiction, and are not addressed. Everything is verified to August 2026. We obtained the published abstract verbatim from six independent agreeing sources but not the paper itself, and report none of its effect estimates by procedure, its moderator analyses, or its confidence intervals. Our copy of the abstract truncates before the mediation finding, which is this article's most consequential claim and which reaches us through two independent third-party summaries; it is graded B on that basis. We did not survey the literature on whether implicit measures predict behaviour, which is a separate question, and one unverified pointer to that literature found in our sources is reported as such rather than as a finding. We found a name and affiliation conflict for one author across sources and report it without resolving it. All arithmetic is ours; the chain calculation is illustrative and the meta-analysis reports a mediation failure, which is a stronger claim than a weak second link. The sections on spending questions and on process-level alternatives are our own reasoning, and the hiring literature referenced there was itself reported in this series as having been over-corrected.
Frequently Asked Questions
What did the meta-analysis find?
Why does mediation matter so much?
Is this a hostile critique of the field?
Does this mean discrimination is not real?
Does the test predict behaviour?
What should a firm do instead?
References
- Forscher, P. S., Lai, C. K., Axt, J. R., Ebersole, C. R., Herman, M., Devine, P. G., & Nosek, B. A. (2019). A Meta-Analysis of Procedures to Change Implicit Measures. Journal of Personality and Social Psychology, 117(3), 522–559, institutional repository record reproducing the abstract, on the authors using a network meta-analysis to synthesize evidence from 492 studies and 87,418 participants to investigate the effectiveness of procedures in changing implicit measures, defined as response biases on implicit tasks; on the authors also evaluating these procedures' effects on explicit and behavioral measures; on the finding that implicit measures can be changed but effects are often relatively weak, with absolute values of d below .30; on most studies having focused on producing short-term changes with brief, single-session manipulations; and on procedures that associate sets of concepts, invoke goals or motivations, or tax mental resources having changed implicit measures the most, whereas procedures that induced threat, affirmation, or specific moods and emotions changed them the least. Note: the lead author's institutional repository. Our copy of the abstract truncates after the statement about which procedures worked best. This record gives the third author's surname as "Ast" and affiliation as Duke University, both differing from other sources. scholarworks.uark.edu
- Bibliographic record for Forscher and colleagues (2019), reproducing the abstract and stating the paper's finding as: it is found that implicit measures can be changed, but effects are often relatively weak with absolute values of d below .30, and changes in implicit measures did not mediate changes in explicit measures or behavior. Note: a bibliographic database's summary of the paper, not the abstract text we obtained elsewhere. This is one of two independent sources for the mediation finding, which is graded B on that basis. semanticscholar.org
- Academic paper indexing service, reproducing the abstract and independently stating the paper's finding as: it is found that implicit measures can be changed, but effects are often relatively weak, and changes in implicit measures did not mediate changes in explicit measures or behavior; recording the authors' institutional affiliations; and separately reproducing a third-party summary asserting that there is insufficient evidence for the claim that the Implicit Association Test measures individual differences in implicit social cognition. Note: an indexing service's summary. This is the second independent source for the mediation finding. The assertion about individual differences summarises a paper we could not identify and did not obtain; it is recorded in this article as an unverified pointer, not as a finding. scispace.com
- Publisher record for Forscher and colleagues, Journal of Personality and Social Psychology, 117(3), 522–559, September 2019, DOI 10.1037/pspa0000160, reproducing the abstract identically. Note: the publisher's own record; independent corroboration of the abstract text and the source for the volume, issue, pages and date. ovid.com
- Open preprint archive record for Forscher and colleagues, reproducing the abstract identically and recording the preprint deposit date of 15 August 2016 with the same DOI as the published article. Note: an open preprint archive; a further independent confirmation of the abstract text and evidence that the work was publicly available roughly three years before journal publication. osf.io
- Hosted copy of the published paper's front matter, reproducing the abstract and recording author contributions, institutional affiliations including the University of Arkansas, Washington University in St. Louis, the University of Virginia, the University of Wisconsin-Madison and the Center for Open Science, and acknowledging support from a National Institutes of Health grant awarded to Patricia G. Devine. Note: a hosted copy of the paper; we obtained the front matter and abstract only, not the body, its effect estimates or its moderator analyses. gwern.net
- Preprint record for Forscher and colleagues, reproducing the abstract and recording the seven authors' departmental affiliations at the University of Arkansas, Washington University in Saint Louis, the University of Virginia, the University of North Carolina at Chapel Hill, the University of Wisconsin-Madison and the Center for Open Science; and noting that the first two authors contributed equally to the manuscript, with order determined by coin flip. Note: a preprint record; the source for the author affiliations and for the coin-flip authorship note. researchgate.net
This article concerns measurement and intervention efficacy only. It is not employment, human resources, legal or equality law advice, makes no claim about whether discrimination exists, and makes no claim about any group, workplace or individual. The paper was not obtained in full and none of its effect estimates by procedure are reported. The mediation finding, which is this article's most consequential claim, reaches it through two independent third-party summaries rather than the abstract text obtained, and is graded accordingly. The separate literature on whether implicit measures predict behaviour was not surveyed.