The eighty-second article documented a correction we had to make to our own work. This one is about a correction the research literature made to itself, twice, and about whether it took.

Key Takeaway

A 2004 paper in the discipline's flagship journal reported that a famous 1995 study was "widely misinterpreted in both popular and scholarly publications" because "scores were statistically adjusted for differences in students' prior SAT performance."[2] On our own arithmetic, adjusting for a covariate on which two groups differ removes that difference by construction, so an adjusted gap of zero says nothing whatever about the raw gap.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The original finding is real and is not what is at issue. The 1995 abstract reports a genuine experimental effect, and nothing in this article disputes it.

Two. The misinterpretation is documented rather than alleged. A 2004 paper in American Psychologist reviewed how the study was characterised in media, textbooks and scholarly reports and found a specific error repeated across all three.

Three. The original authors replied and did not dispute the central point. The 2004 authors' response to that reply states they saw no disagreement on the key issues.

Four. The statistical point is not a matter of opinion. On our own arithmetic, an adjustment for a covariate removes the covariate-predicted difference exactly, so a zero adjusted gap is a property of the adjustment.

Five. Somebody checked fifteen years later whether the correction had worked, which is the most interesting thing in this literature and the finding we could not obtain.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the 1995 finding, from an abstract obtained verbatim from a national medical library database.

Grade A for the documented misinterpretation, from a 2004 abstract obtained verbatim from the publisher's own record and from a second reproduction.

Grade A for the exchange, from a passage of the 2004 authors' reply obtained verbatim.

Grade B for the fifteen-year follow-up, whose abstract we obtained and whose result we did not.

Grade D for the 2019 meta-analysis, which reaches us as a title in a reference list and nothing more.

Grade A for our own arithmetic, which demonstrates a property of covariate adjustment on entirely invented figures.

A Note On Method

Everything here is verified to August 2026.

We obtained the 1995 abstract verbatim from a national medical library database[1].

We obtained the 2004 abstract verbatim from the publisher's record[2], and a fuller reproduction including its proposed corrected summary from a second source[3].

We obtained a passage of the 2004 authors' reply to the original authors' response[3].

We obtained the fifteen-year follow-up's abstract and part of its discussion[4], but not its result, which we flag wherever it arises.

We did not obtain the 1995 paper itself, the original authors' 2004 reply in full, or any meta-analysis, and we report those through other sources or as citations[5].

All arithmetic is ours and every figure in it is invented.

What This Article Is And Is Not About

Because the underlying research concerns a socially charged subject, we state the boundaries before going further.

Four points.

This article is about what a statistical control variable removes from a comparison. That is a technical question with a definite answer, and it is the question our arithmetic addresses.

This article takes no position on any question about groups of people, makes no claim about the causes of any measured difference in the world, and does not treat the research it describes as evidence about those questions.

The 2004 paper we rely on affirms the original finding while correcting how it was described, and we follow it in that. Its authors state plainly that the effect was demonstrated.

And the reason this belongs in a series read by business owners is general rather than specific. Adjusted comparisons appear in benchmarking reports, salary surveys, industry studies and any regression output, and the error documented here is available to anyone reading any of them.

One further reason for choosing this episode over a less charged one, since we considered several. It is the best-documented case we could find of a specific misreading being counted, published, and then re-counted years later.

Most of the transmission failures this series has described are inferred from comparing a paper with its popular account. Here somebody actually went and tallied where the error appeared, in a flagship journal, and that evidentiary quality is rare enough to be worth using.

The Finding

What the 1995 study reported, in its own abstract.

It defines the concept: "Stereotype threat is being at risk of confirming, as self-characteristic, a negative stereotype about one's group."[1]

The citation: Steele, C. M., and Aronson, J. (1995), Stereotype threat and the intellectual test performance of African Americans, Journal of Personality and Social Psychology, 69(5), 797–811, November, DOI 10.1037//0022-3514.69.5.797[1].

Three observations, ours.

The definition is about a situation rather than a person. It describes a risk present in a context, which is what makes it experimentally manipulable.

The journal and the citation are among the most heavily cited in social psychology, and the finding has generated a literature of its own[5].

And nothing in this article questions the finding. What follows concerns how it was reported.

Two further notes on scale, since they bear on why a misreading of this particular study mattered enough to correct formally.

A reference list we obtained shows the concept generating a review chapter, a further American Psychologist article, a meta-analysis and a process model within thirteen years[5], and being applied in fields as distant as medical training.

Which is the context for the 2004 paper. A misdescription of a heavily cited finding propagates into every field that borrows it, and correcting it at the source is the only intervention available.

The Design

What was done, from the abstract.

It records: "Studies 1 and 2 varied the stereotype vulnerability of Black participants taking a difficult verbal test by varying whether or not their performance was ostensibly diagnostic of ability, and thus, whether or not they were at risk of fulfilling the racial stereotype about their intellectual ability."[1]

And a further study: "Study 3 validated that ability-diagnosticity cognitively activated the racial stereotype in these participants and motivated them not to conform to it, or to be judge", at which point our reproduction is cut off.[1]

Four observations, ours.

The manipulation is a description rather than a treatment. The same test was presented as diagnostic of ability or not, which is a change in framing and nothing else.

That makes it a clean experiment in the sense this series has praised elsewhere. Everything is held constant except the sentence the participant reads about what the test measures.

The test is described as difficult, and the paper's own body notes the assumption that threat is most likely when a test is frustrating[6], which is a stated boundary condition rather than a hidden one.

And Study 3 is a mechanism check, testing whether the manipulation did what it was supposed to rather than only whether scores moved, which is more than many studies in this series attempted.

The Parenthetical

The four words this article is named for.

The abstract's result sentence: "Reflecting the pressure of this vulnerability, Blacks underperformed in relation to Whites in the ability-diagnostic condition but not in the nondiagnostic condition (with Scholastic Aptitude Tests controlled)."[1]

Four observations, ours.

The parenthetical is in the abstract, at the end of the sentence reporting the result. It was never hidden, and the authors put it exactly where a careful reader would need it.

What it means is that the comparison being reported is between scores that have been adjusted for participants' prior test performance, not between raw scores.

Which changes the meaning of the phrase "but not in the nondiagnostic condition" completely, and the next several sections are about how.

And the placement is worth pausing on. The qualification most necessary to understanding the sentence appears after the sentence has already landed, which is a fact about how abstracts are written rather than about these authors.

Two things follow that are worth separating carefully, because the distinction governs the rest of this article.

The finding about the manipulation is unaffected. Whether framing a test as diagnostic changes performance is a within-experiment question, and adjusting for prior scores makes that comparison cleaner rather than weaker.

The claim about the size of a real-world difference is a different matter entirely, because that difference is exactly what the adjustment removed, and a study cannot report on a quantity it has subtracted out.

The 2004 Paper

The source that documented what happened next.

Sackett, P. R., Hardison, C. M., and Cullen, M. J. (2004), On Interpreting Stereotype Threat as Accounting for African American-White Differences on Cognitive Tests, American Psychologist, 59(1), 7–13, DOI 10.1037/0003-066X.59.1.7[2].

It opens by affirming the finding: "C. M. Steele and J. Aronson (1995) showed that making race salient when taking a difficult test affected the performance of high-ability African American students, a phenomenon they termed stereotype threat."[2]

Three observations, ours.

The venue matters. American Psychologist is the American Psychological Association's flagship journal, which is the most visible place in the discipline to publish a correction of this kind.

The opening sentence affirms the original finding before criticising anything, and the word used is "showed" rather than "claimed."

And that framing is what makes the paper usable. It is a correction about description, not an attack on a result, and this article follows it in that distinction.

One further point about who wrote it, since it bears on how to read the exchange. The 2004 authors work in personnel selection and employment testing, which is the applied field where claims about what tests measure have direct consequences for hiring practice.

That is a reason they would care about the precision of the claim rather than a reason to discount them. A finding used to justify decisions about real employment tests needs its scope stated accurately, and that is the interest the paper represents.

A Documented Misinterpretation

What the 2004 authors found.

Their abstract continues: "The authors document that this research is widely misinterpreted in both popular and scholarly publications as showing that eliminating stereotype threat eliminates the African American-White difference in test performance."[2]

And gives the reason: "In fact, scores were statistically adjusted for differences in students' prior SAT performance, and thus, Steele and Aronson's findings actually showed that absent stereotype threat, the two groups differ to the degree that would" be expected from those prior differences[3].

Four observations, ours.

The verb is "document." They did not assert that the misreading was likely; they went and counted where it appeared.

The scope is what makes it notable. Popular and scholarly publications both, and a separate summary specifies media, textbooks and scholarly reports[3].

The misreading described is not subtle. "Eliminating stereotype threat eliminates the difference" is a far stronger claim than anything the study tested, and it is the version that travelled.

And the correction rests entirely on the parenthetical. The whole dispute is about what four words in an abstract imply, which is why this article treats it as a lesson about reading rather than about any subject matter.

One clarification on what is and is not being alleged, because it is easy to blur. Nobody is accused of misreporting anything. The study reported what it did accurately and in the standard format.

The claim is about the chain from paper to summary to textbook to public account, and about a qualification dropping out somewhere along it. That is a failure of transmission rather than of research, and it is the failure this series has documented more often than any other.

They Wrote The Corrected Version

The most useful thing in the 2004 paper, and an unusual thing for a critique to include.

They supply a replacement summary: "An improved blurb would read: When their race is made salient to them and they are given a difficult test, high-ability African American students perform more poorly than otherwise. This previously demonstrated 'stereotype threat' effect may not, however, be the complete explanation for large African American-White differences on tests. This fact, the authors show, has been missed by a number of writers of media, textbooks, and scholarly reports."[3]

Four observations, ours.

The corrected version keeps the finding entirely. It says the effect is demonstrated and that students perform more poorly, which is the original result.

What it removes is a single inference: that the effect is the complete explanation for a difference measured outside the experiment.

Writing the replacement is a generous move and a rare one. Most critiques stop at saying what is wrong, and this one supplies the sentence a textbook could use instead.

And the hedge in it is careful. "May not, however, be the complete explanation" claims only that the study does not establish completeness, which is exactly as far as the argument goes.

The Exchange That Followed

What happened when the original authors replied, which matters for how much weight to give the correction.

The original authors published a response in the same issue[3], and the 2004 authors replied to it. Their reply opens: "We see no disagreement by Steele and Aronson (2004, this issue) with the key issues that prompted our article."[3]

The original authors' response does contest other matters, including that "Sackett et al. (2004) assessed neither" and "Sackett et al. (2004) offered no systematic" certain things, in passages our reproduction gives only in fragments[3].

Four observations, ours.

The claim of no disagreement is the 2004 authors' characterisation of the reply, not a concession quoted from the original authors, and we flag that distinction because it matters.

The fragments we have show the reply did contest things, including the scope of the 2004 review, so this was not an unqualified acceptance.

We did not obtain the original authors' reply in full, and a reader who wants to weigh this exchange properly should read both, which we could not.

But the statistical point at the centre is not the kind of thing that admits of disagreement. Whether an adjusted comparison controls for prior scores is a matter of what the analysis did, and the next section shows why that settles the narrow question.

What Adjusting Actually Does

Our own arithmetic, on entirely invented figures, demonstrating a property of covariate adjustment. This is a technical illustration and concerns no real group of people.

Take two groups, A and B, taking a test. They differ on a prior measure: A averages 600 and B averages 500. The prior measure predicts test performance with a slope of 0.6. In one condition a manipulation costs group B 15 points.

The raw scores:

In the non-diagnostic condition, A scores 360 and B scores 300, a raw gap of 60.

In the diagnostic condition, A scores 360 and B scores 285, a raw gap of 75.

Three observations.

The manipulation did what it was supposed to. It widened the gap by exactly the 15 points we put in, which is the experimental effect and it is real.

The gap in the non-diagnostic condition is 60 points and does not disappear, because it comes from the prior measure rather than from the manipulation.

And now we adjust.

Zero By Construction

The result. Ours, same invented figures.

Adjusting for the prior measure removes the slope multiplied by the prior difference, which is 0.6 times 100, or exactly 60 points.

The adjusted gaps:

In the non-diagnostic condition: zero. In the diagnostic condition: 15, being the manipulation and only the manipulation.

Four observations.

The adjusted gap in the non-diagnostic condition is zero because the adjustment removed it, not because the manipulation eliminated anything. We put 60 points of prior difference in and subtracted 60 points of prior difference out.

The raw gap in that same condition is still 60 points, unchanged, sitting in the data untouched by the analysis that reported no difference.

So the sentence "the groups did not differ in the nondiagnostic condition" is true of the adjusted comparison and false of the raw one, and both are accurate descriptions of the same experiment.

And that is the entire content of the 2004 correction. An adjusted zero cannot tell you about a raw difference, because the adjustment is what produced the zero, and this holds for every covariate in every study regardless of subject matter.

Two ways to state the same point, because it is worth having in more than one form.

Mechanically: the adjustment subtracts a fixed quantity from the comparison in every condition, so it cannot create or destroy a treatment effect, and it can absolutely create or destroy an apparent group difference.

In plain terms: if you take out the head start before the race, the finish looks level, and reporting a level finish is accurate as long as nobody forgets what was taken out.

What Share An Effect Explains

The quantitative version. Ours, invented figures throughout.

Given a prior-measure gap and a treatment effect, what share of the raw diagnostic-condition gap does the treatment account for?

Prior gap 100, effect 15: the effect explains 20 percent.

Prior gap 100, effect 25: 29 percent. Prior gap 100, effect 40: 40 percent.

Prior gap 60, effect 15: 29. Prior gap 60, effect 25: 41.

Prior gap 150, effect 15: 14. Prior gap 150, effect 40: 31 percent.

Four observations.

The effect is real in every row, and in no row is it the whole difference, because the rest of the difference was present before the manipulation was applied.

The share rises with the size of the effect and falls with the size of the prior gap, which is arithmetic rather than a finding.

These are invented numbers illustrating a structure, and we are not asserting any of them describes anything real.

And the reason we can say nothing more specific is itself the point. A study that controls for the prior gap has removed the information needed to compute this share, and no amount of re-reading it will produce the number.

The General Rule

The formula, worth carrying. Ours.

The share of a raw difference explained by a treatment equals the effect divided by the effect plus the slope times the prior gap.

Four observations.

It approaches 100 percent only as the prior gap approaches zero, which is the case where the covariate was not doing anything and the adjustment was unnecessary.

Which yields a clean statement of when the misreading is harmless. If the groups did not differ on the covariate, adjusted and raw comparisons agree, and nobody can be misled.

And it yields the condition under which it is most dangerous. The larger the group difference on the control variable, the more the adjustment removes and the more misleading an adjusted zero becomes.

That is a general property of analysis of covariance, is taught in every methods course, and was nonetheless missed widely enough to require a paper in a flagship journal.

Which is the durable lesson and the reason we would not treat this as a story about one literature. Knowing a rule and applying it while reading are separate abilities, and the second one fails under time pressure in a way the first does not.

Fifteen Years Later

The part of this story we find most interesting and can report least completely.

A later paper states: "Steele and Aronson (1995) showed that stereotype threat affects the test performance of stereotyped groups. A careful reading shows that threat affects test performance but does not eliminate Black-White mean score gaps. Sackett et al. (2004) reviewed characterization of this research in scholarly articles, textbooks, and popular press, and found that many mistakenly inferred that removing stereotype threats eliminated the Black-White performance gap."[4]

And states its own question: "We examined whether the rate of mischaracterization of Steele and Aronson had decreased in the 15 years since Sackett et al. highlighted the common misinterpretation."[4]

We did not obtain the answer.

Four observations, ours.

That anyone thought to check is the encouraging part. A correction published in a flagship journal is an intervention, and somebody measured whether it worked.

The design is straightforward and the right one. Count the mischaracterisations before and after, which is exactly what the original 2004 paper did and therefore directly comparable.

We would very much like the number, and its absence is the largest gap in this article. A reader who obtains it will know something we do not.

And the fact that a paper on this was published at all in 2019 or later carries a weak implication. Nobody writes a follow-up to report that a problem went away, though we would not put much weight on that inference.

Two reasons that inference is weak enough that we would not rest anything on it, ours. A paper reporting improvement is publishable too, particularly in an applied journal where the practical question is whether a correction strategy worked.

And the honest position is the one we are in. We know the question was asked and we do not know the answer, and a reader who wants it can obtain the paper in a few minutes, which is more than we managed.

Why It Was Missed

The follow-up paper offers a reason, and it is disarming.

It records that among the explanations, "The first is that some people who discuss Steele and Aronson (1995) might not have noticed that the performance of the participants was adjusted for SAT."[4]

Four observations, ours.

The proposed explanation is not motivated reasoning or ideology. It is that people did not notice four words in a bracket.

That is entirely plausible and it is the reason this article exists. The parenthetical is easy to skim past, and a reader who skims it takes away a sentence that means something different.

It also suggests the failure is reproducible by anyone, including us, and including on subjects nobody has strong feelings about.

And it points at where the responsibility sits, which is not obvious. The authors put the qualification in the abstract; the readers did not read it, and no amount of author diligence fixes that.

Laboratory And Operational Settings

A further line of work, which we can name and not describe.

A reference list records: Shewach, O. R., Sackett, P. R., and Quint, S. (2019), Stereotype threat effects in settings with features likely versus unlikely in operational test settings: A meta-analysis, Journal of Applied Psychology, 104(12), 1514–1534[4].

We obtained the title and nothing else.

Three observations, ours.

The title states the question being asked, which is whether effects found in laboratory conditions appear in real testing situations.

That is the same question this series has asked repeatedly, and it is the right one. A laboratory effect is a demonstration that something can happen, not evidence about how often it does.

And we will not speculate about the answer. A title tells you the question and not the finding, and we have graded this accordingly.

Two things the title does establish, which are worth having even without the result. The distinction between laboratory and operational conditions was considered important enough to build a meta-analysis around, by authors including one of the 2004 paper's authors.

And the question is the applied one: not whether the effect exists, which the same authors affirmed in 2004, but whether the conditions that produce it are present when real tests are administered. That is the question an employer would have.

What Actually Survives

Our reading, stated directly.

Five statements.

The experimental finding stands and is not in dispute here. The 2004 correction opens by affirming it.

A specific misreading was documented across media, textbooks and scholarly reports, and the correction was published in the discipline's flagship journal.

The statistical point is not contestable. Adjusting for a covariate removes the covariate-predicted difference by construction, so an adjusted zero is a property of the analysis.

The exchange that followed was not an unqualified concession, and we obtained only one side of it in any detail.

And the correction's effectiveness was measured fifteen years later by somebody whose answer we do not have, which is the open question in this article.

When The Adjustment Is The Right Analysis

Because this article could be misread as an argument against controlling for things, which would be wrong. Ours.

Four observations.

The adjustment in the study we have described was the correct choice for the question the study asked. If you want to know whether a framing manipulation changes performance, you need to remove pre-existing differences between participants, or the manipulation's effect is buried in noise.

Controlling for a covariate does two useful things at once. It increases the precision of the estimate and it removes a confound, and a study that skipped it would be worse, not more honest.

So the fault in the episode this article describes lies nowhere near the analysis. The design was right, the reporting was accurate, and the parenthetical was in the abstract.

And that is what makes it instructive rather than merely embarrassing for somebody. Every party did their job and the wrong claim spread anyway, which means the failure is in the space between a correct result and a reader in a hurry.

Which is not a comfortable conclusion, because it offers nobody to hold responsible and no procedure that would have prevented it. The only defence available operates at the point of reading, one reader at a time, and that is why this article ends in a checklist rather than a recommendation.

Two Questions That Sound Identical

The distinction underneath everything above, stated as plainly as we can. Ours.

Four observations.

Does this factor affect performance is a question about a causal influence, and an adjusted experimental comparison is the right way to answer it.

How much of the observed difference does this factor account for is a question about a decomposition, and it requires knowing the size of everything else, which the adjustment removed.

In ordinary speech those two questions sound like the same question, which is most of why this happens. An answer to the first gets read as an answer to the second, and nothing in the wording flags the substitution.

And the same trap sits in every business conversation about drivers of a number. Showing that price affects volume is not showing how much of last quarter's decline was price, and the second is almost always the question being asked.

One test that separates them reliably, ours. Ask whether the answer would change if the other factors were larger. A causal claim survives that question and a decomposition does not, because a decomposition is a share of a fixed total.

Which is why the second question is so much harder to answer and so much more often asked. It requires an inventory of everything else that mattered, and an experiment designed to isolate one factor has deliberately removed that inventory.

Reading Any Study With Controls

The transferable lesson, which is why this is in a business publication. Ours.

Four observations.

Every regression you will ever read reports adjusted relationships, and the adjustment always removes something.

Phrases to stop at: controlling for, adjusted for, holding constant, net of, after accounting for. Each announces that a quantity has been removed from the comparison being reported.

The question that follows is always the same one. Did the groups being compared differ on the thing that was removed? If they did, the adjusted comparison is not about the total difference and cannot be read as though it were.

And this is not a criticism of adjustment, which is usually the correct analysis. The adjusted comparison answers a real and often better question, and the error is reading its answer as though it addressed a different one.

Three Questions To Ask

A practical checklist. Ours.

What was controlled for? If the paper does not say prominently, look in the methods, and if the abstract states a result without stating the controls, treat the result as incomplete rather than as reported.

Did the groups differ on it? This is the question that determines whether the adjustment changed anything, and it is frequently not reported alongside the result.

Which question do I actually have? If you want to know whether a difference exists at all, you want raw figures. If you want to know whether a difference remains once something else is accounted for, you want the adjusted ones.

Three observations, ours.

The third question is the one people skip and it determines everything. Most readers arrive with a question about totals and are handed an answer about residuals.

The three take a minute and would have prevented the entire episode this article describes.

And they are not specialist skills. Nothing above requires statistical training, only the habit of asking what a number is a number of.

One reason to build the habit even if you rarely read research, ours. The same three questions apply to anything an adviser, a supplier or a software dashboard tells you, and dashboards in particular present adjusted figures without labelling them as such.

One addition for anyone reading a study rather than a report. Look at whether the abstract's result sentence and its method sentence agree about what was compared, since the abstract is usually the only part that gets read and is where the compression happens.

In the case this article describes, the abstract was accurate and complete, and the compression happened in the retelling rather than in the paper, which is the more common failure and the harder one to guard against.

Your Own Numbers Do This Too

Where a business owner meets this without reading any research. Ours.

Four points.

Same-store or like-for-like sales are an adjusted figure. They remove the contribution of new locations, which is usually the right analysis and is not a statement about total revenue.

Margin excluding one-off items is an adjusted figure, and whether the exclusions were genuinely one-off is exactly the question the adjustment assumes away.

Revenue per employee adjusted for headcount changes, cost per unit at constant volume, and growth on a constant-currency basis are all the same move: a quantity removed to make a comparison cleaner.

And every one of them can produce the error above. A metric that is flat after adjustment and down before it is telling you two true things, and which one matters depends on the decision in front of you.

Two decisions where the choice between them is not a matter of taste, ours. If you are paying rent, wages and suppliers, the total is what matters, because creditors are not paid in like-for-like figures.

If you are deciding whether a store is performing, the adjusted figure is the right one, because you want to know about the store rather than about how many stores you opened. Both numbers are correct and only one answers each question.

The Benchmarking Version

The place a small firm is most likely to be misled by this. Ours, and not advice on any particular report.

Four observations.

Industry benchmark reports routinely present figures adjusted for firm size, region or sector, and the adjustment is usually sensible.

The trouble is the comparison a reader then makes. You are a specific firm of a specific size in a specific place, and an adjusted benchmark has removed exactly the characteristics that make you comparable or not.

Which means an adjusted benchmark can tell you that you are typical while your raw numbers are nothing like the raw numbers of the firms in the sample.

And the practical response is unglamorous. Ask for the unadjusted figures for firms like yours, and if the sample is too small to provide them, that is itself the useful finding.

One thing worth checking that most readers never do, ours. Ask what the sample was before treating any benchmark as a standard, because a median computed over firms unlike yours is a number about them.

And a benchmark that cannot say how many firms of your size and sector are in it is not a benchmark. It is an average with a confident presentation, and the confidence is doing work the sample cannot support.

What To Do

Read the parenthetical. The qualification that changes a result's meaning is often at the end of the sentence reporting it, and a documented episode in the research literature turned entirely on four such words.

Ask what was controlled for before accepting any adjusted comparison, in a study or in a report about your own business.

Ask whether the groups differed on it. If they did, the adjustment removed that difference by construction and the adjusted result cannot speak to the total.

Decide which question you have. Raw figures answer whether a difference exists; adjusted figures answer whether it survives accounting for something else, and they are different questions.

Apply this to your own metrics. Like-for-like sales, adjusted margin and constant-currency growth are all the same move, and each can be flat while the total is not.

Ask benchmark providers for unadjusted figures for firms like yours, and treat an inability to supply them as information.

Do not treat a correction as a refutation. The 2004 paper affirmed the finding and corrected its description, and conflating those two is its own error.

Assume you are capable of this mistake. The explanation offered for how it spread is that readers did not notice four words, which is available to anyone on any subject.

The Limits Of This Analysis

Several caveats matter. This article is about what a statistical control variable removes from a comparison; it takes no position on any question about groups of people and treats none of the research it describes as evidence about such questions. Everything is verified to August 2026. We did not obtain the 1995 paper, only its abstract from a national medical library database and one body passage; our reproduction of its third study's description is cut off mid-sentence. We did not obtain the 2004 paper in full, only its abstract from the publisher's record and a fuller reproduction from a second source. We did not obtain the original authors' 2004 reply, only fragments and the other side's characterisation of it, and we flag that the statement of no disagreement is the 2004 authors' description rather than a concession quoted from the original authors; the fragments we do have show the reply contested several matters, so this was not an unqualified acceptance and a reader wanting to weigh the exchange should read both papers, which we could not. We did not obtain the result of the fifteen-year follow-up, only its abstract and part of its discussion, so we can report that the question was asked and not what was found; that is the largest gap here. We obtained the 2019 meta-analysis as a title in a reference list and nothing more, and have not speculated about its findings. All arithmetic is ours and every figure in it is invented: the prior scores, the slope, the effect size and the resulting gaps are illustrative constructions demonstrating a property of covariate adjustment, and none of them describes anything real or is offered as an estimate of anything. The demonstration assumes a linear relationship with a common slope across groups, which is the standard assumption of the analysis it illustrates and which is not always warranted.

Frequently Asked Questions

What was the documented misinterpretation?
A 2004 paper in American Psychologist reported that a 1995 study was widely misinterpreted in popular and scholarly publications as showing that eliminating stereotype threat eliminates a difference in test performance. In fact scores had been statistically adjusted for participants' prior test performance.
Does this mean the original finding was wrong?
No, and the 2004 paper says so explicitly in its first sentence, using the word "showed". The correction concerns how the finding was described, not whether it occurred. Treating a correction of description as a refutation of a result is its own error.
Why does adjusting for a covariate matter so much?
Because it removes the covariate-predicted difference by construction. On our own invented illustration, two groups differing by 100 points on a prior measure with a slope of 0.6 show an adjusted gap of exactly zero and a raw gap of 60 in the same condition. Both are accurate descriptions of the same data.
How was the error explained?
A later paper offers as its first explanation that some people discussing the study might not have noticed the performance was adjusted for prior scores. The qualification is four words in a bracket at the end of the abstract's result sentence, and it is easy to skim past.
Did anyone check whether the correction worked?
Yes. A later paper examined whether the rate of mischaracterisation had decreased in the fifteen years since the correction was published. We obtained its abstract and not its result, which is the largest gap in this article.
Where does this affect a business owner?
Anywhere an adjusted figure is reported. Like-for-like sales, margin excluding one-off items, constant-currency growth and size-adjusted industry benchmarks are all the same move, and each can be flat after adjustment while the total is not. Which figure matters depends on the decision you are making.
What should I ask about an adjusted number?
Three things: what was controlled for, whether the groups being compared differed on it, and which question you actually have. If the groups differed on the control, the adjusted result cannot speak to the total difference, because removing that difference is what the adjustment does.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article concerns what a statistical control variable removes from a comparison, and takes no position on any question about groups of people.

References

  1. National medical library database record for Steele, C. M., & Aronson, J. (1995), Stereotype threat and the intellectual test performance of African Americans, Journal of Personality and Social Psychology, 69(5), 797–811, November, DOI 10.1037//0022-3514.69.5.797, reproducing the abstract: that stereotype threat is being at risk of confirming, as self-characteristic, a negative stereotype about one's group; that Studies 1 and 2 varied the stereotype vulnerability of Black participants taking a difficult verbal test by varying whether their performance was ostensibly diagnostic of ability, and thus whether they were at risk of fulfilling the racial stereotype about their intellectual ability; that reflecting the pressure of this vulnerability, Blacks underperformed in relation to Whites in the ability-diagnostic condition but not in the nondiagnostic condition, with Scholastic Aptitude Tests controlled; and that Study 3 validated that ability-diagnosticity cognitively activated the racial stereotype in these participants and motivated them not to conform to it, at which point our reproduction is cut off. Note: a national medical library database record. Our source for the abstract, including the parenthetical this article concerns. We did not obtain the paper. pubmed.ncbi.nlm.nih.gov
  2. Publisher record for Sackett, P. R., Hardison, C. M., & Cullen, M. J. (2004), On Interpreting Stereotype Threat as Accounting for African American-White Differences on Cognitive Tests, American Psychologist, 59(1), 7–13, DOI 10.1037/0003-066X.59.1.7, reproducing the abstract and its proposed replacement summary: that Steele and Aronson (1995) showed that making race salient when taking a difficult test affected the performance of high-ability African American students, a phenomenon they termed stereotype threat; that the authors document this research is widely misinterpreted in both popular and scholarly publications as showing that eliminating stereotype threat eliminates the African American-White difference in test performance; and that an improved blurb would read that when their race is made salient to them and they are given a difficult test, high-ability African American students perform more poorly than otherwise, that this previously demonstrated stereotype threat effect may not however be the complete explanation for large African American-White differences on tests, and that this fact has been missed by a number of writers of media, textbooks and scholarly reports. Note: the publisher's record for a paper in the American Psychological Association's flagship journal. Our source for the documented misinterpretation. We obtained the abstract and not the paper. psycnet.apa.org
  3. Academic networking site records for the 2004 exchange, reproducing a fuller portion of the Sackett, Hardison and Cullen abstract, including that in fact scores were statistically adjusted for differences in students' prior SAT performance, and thus Steele and Aronson's findings actually showed that absent stereotype threat the two groups differ to the degree that would be expected; and reproducing the opening of the same authors' reply to the original authors' response, stating that they see no disagreement by Steele and Aronson (2004, this issue) with the key issues that prompted their article. The same records carry fragments of the original authors' response, including phrases indicating that it contested the scope of the 2004 review. Note: academic networking site records, not publisher records. Our source for the covariate explanation and for the exchange. The statement of no disagreement is the 2004 authors' characterisation of a reply we did not obtain, and the fragments show that reply contested several matters. researchgate.net
  4. University repository record for a later paper, On the Continued Misinterpretation of Stereotype Threat as Accounting for Black-White Differences on Cognitive Tests, in Personnel Assessment and Decisions, volume 8, issue 1, reproducing its abstract: that Steele and Aronson (1995) showed stereotype threat affects the test performance of stereotyped groups; that a careful reading shows threat affects test performance but does not eliminate Black-White mean score gaps; that Sackett and colleagues (2004) reviewed characterisation of this research in scholarly articles, textbooks and popular press and found that many mistakenly inferred that removing stereotype threats eliminated the performance gap; and that the authors examined whether the rate of mischaracterisation had decreased in the fifteen years since. The same record reproduces part of its discussion, giving as the first explanation that some people who discuss the 1995 study might not have noticed that the performance of the participants was adjusted for SAT, and carries a reference list confirming Shewach, O. R., Sackett, P. R., and Quint, S. (2019), Stereotype threat effects in settings with features likely versus unlikely in operational test settings: A meta-analysis, Journal of Applied Psychology, 104(12), 1514–1534. Note: a university repository record. Our source for the fifteen-year follow-up and for the proposed explanation. We obtained the abstract and part of the discussion and NOT the result, which is the largest gap in this article. The 2019 meta-analysis reaches us as a title only. scholarworks.bgsu.edu
  5. Reference list carried on a registered clinical trial protocol applying stereotype threat research to a medical training context, confirming Steele, C. M., and Aronson, J. (1995), Journal of Personality and Social Psychology, 69(5), 797; Steele, C. M., Spencer, S. J., and Aronson, J. (2002), Contending with group image: The psychology of stereotype and social identity threat, in Advances in Experimental Social Psychology, Vol. 34, 379–440; Steele, C. M. (1997), A threat in the air: How stereotypes shape intellectual identity and performance, American Psychologist, 52(6), 613; Nguyen, H. H. D., and Ryan, A. M. (2008), Does stereotype threat affect test performance of minorities and women? A meta-analysis of experimental evidence, Journal of Applied Psychology, 93(6), 1314; and Schmader, T., Johns, M., and Forbes, C. (2008), An integrated process model of stereotype threat effects on performance. Note: a reference list; citations only. Recorded to establish the scale of the subsequent literature. We obtained none of the works named here. cdn.clinicaltrials.gov
  6. Copy of the 1995 paper hosted by a research centre at the first author's own university, reproducing a body passage stating the authors' assumption that for African American students the act of taking a test purported to measure intellectual ability may be enough to induce the threat, but that this is assumed most likely to happen when the test is also frustrating. Note: a copy hosted by a centre at one author's own university. Recorded as our source for the stated boundary condition on test difficulty; we obtained this passage and not the paper's results. sparq.stanford.edu

This article concerns what a statistical control variable removes from a comparison. It takes no position on any question about groups of people and treats none of the research described as evidence about such questions. The 2004 paper relied on here affirms the original experimental finding while correcting how it was described, and this article follows it in that distinction. All arithmetic is the authors' own and every figure in it is invented.