The previous article was about decisions where nobody can supply a probability. This one is about the opposite failure: decisions where you have plenty of data, have looked at it carefully, and have extracted a relationship that is not there.

Key Takeaway

The founding definition is "the report by observers of a correlation between two classes of events which, in reality, (a) are not correlated, or (b) are correlated to a lesser extent than reported"[1]. In the standard experiment, evaluatively equivalent information is provided about both groups and people perceive them differently anyway[2]. On our own arithmetic the correlation in that design is exactly zero, and the cell driving the illusion contains four observations.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The effect is real and meta-analytically supported. A 1990 meta-analysis describes it as "highly significant, and of moderate strength."

Two. The mechanism is disputed and still was as of 2011. A recent paper states plainly that the theoretical explanation "is still a matter of debate," with at least two competing accounts.

Three. The meta-analysis contains a finding that complicates its own headline. Judgments were significantly predicted by a strategy the authors say reflects responsiveness to the information presented, which is a partly rational reading of what looks like a bias.

Four. The commercially decisive point is arithmetic and needs no psychology. On our own figures, five different contingency tables sharing the same count in the memorable cell produce verdicts ranging from the campaign works by 40 points to the campaign hurts by 20.

Five. The most memorable cell is the least reliable one. The distinctive cell in the standard design holds four observations, whose rate carries a 95 percent interval of roughly plus or minus 49 percentage points on our own calculation.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the effect existing, from a meta-analysis whose abstract we obtained verbatim from three independent sources.

Grade B for the founding clinical work, which we report from citations and one encyclopedia summary and did not obtain.

Grade A for the 1976 paradigm's structure, from its abstract obtained verbatim, though we did not verify the exact cell counts against the paper.

Grade C for the mechanism, because the literature itself describes it as unsettled and we obtained none of the competing papers.

Grade A for our own arithmetic, which is elementary contingency table mathematics.

Our position: the phenomenon is well supported, the explanation is open, and the useful part is a habit rather than a fact.

A Note On Method

Everything here is verified to August 2026.

We obtained the 1976 abstract verbatim from three independent sources, including the publisher and a bibliographic database[2][3]. We did not obtain the paper, so the cell counts used in our arithmetic are the standard ones reported in secondary accounts and we did not verify them.

We obtained the 1990 meta-analysis's abstract verbatim from three sources[4], and its final sentence is truncated in all of them, which we flag where it matters.

We obtained the 1967 definition verbatim with its page reference, quoted inside two later papers[1]. We did not obtain the 1967 paper, nor either of the two clinical papers.

We did not obtain the competing theoretical accounts, nor the field study discussed near the end, and report all from citation records and titles.

All arithmetic is ours. The contingency calculations are elementary; the business figures are invented throughout.

This article discusses research on judgment. It is not analytics, hiring or legal advice, and nothing here bears on the lawfulness of any employment practice.

The Definition

The source and the wording.

Chapman, L. J. (1967), Illusory correlation in observational report, Journal of Verbal Learning and Verbal Behavior, 6(1), 151–155, DOI 10.1016/S0022-5371(67)80066-5[5].

The 1976 paper quotes it with a page reference: illusory correlation refers to "the report by observers of a correlation between two classes of events which, in reality, (a) are not correlated, or (b) are correlated to a lesser extent than reported" (p. 151)[1].

Three observations, ours.

The definition is about reports, not beliefs. It concerns what observers say they saw, which is measurable in a way that belief is not.

It concerns two classes of events, which in practice means a two-by-two table: something present or absent, something else present or absent. Everything below follows from that structure.

And we did not obtain the 1967 paper. We have the definition verbatim with its page number because two later papers quote it, which is better than a paraphrase and worse than the original.

The Clause Everyone Drops

Part (b) of that definition, which almost every summary omits and which is the commercially important half. Ours.

Four observations.

Clause (a) covers seeing a relationship where none exists. That is the dramatic version and the one the name suggests.

Clause (b) covers seeing a relationship that exists but is weaker than reported. That is the far more common case and it is much harder to notice.

The difference matters because clause (b) cannot be refuted by finding some relationship. A business owner who says referrals produce better customers may be right about the direction and badly wrong about the size, and confirming the direction will feel like vindication.

And this is the same distinction the fiftieth article found running through the whole series. Direction survives popularisation and magnitude does not, and here it is written into the founding definition.

The Clinicians

Where the concept came from, which is more pointed than the laboratory work that followed.

Reference lists identify Chapman, J. L., and Chapman, J. P. (1967), Genesis of popular but erroneous psychodiagnostic observations, Journal of Abnormal Psychology, 72, 193–204, and Chapman, L. J., and Chapman, J. P. (1969), Illusory correlation as an obstacle to the use of valid psychodiagnostic signs, Journal of Abnormal Psychology, 74(3), 271–280[6].

An encyclopedia entry records that the concept was used "to question claims about objective knowledge in clinical psychology" through a refutation of a set of purported diagnostic signs that practitioners widely believed in[7]. This is an encyclopedia, not an academic source, flagged here and at every use, and we obtained neither clinical paper.

Four observations, ours.

The subjects were trained, experienced professionals reporting on material within their own expertise. This was not undergraduates guessing.

The title of the earlier paper is worth reading twice. Genesis of popular but erroneous observations claims not merely that the observations were wrong but that the research explains how they came to be widely held.

The commercial parallel is exact and uncomfortable. An experienced owner reporting what their customers are like is doing the same task, from memory, on unrecorded data, with a strong prior.

And we are deliberately not describing the specific diagnostic claim at issue, which concerned a historical classification now abandoned. The methodological point stands without it and the substance is not ours to relitigate.

An Obstacle To The Valid Ones

The 1969 title, which contains a claim the laboratory literature mostly leaves out. Ours.

Three observations.

"An obstacle to the use of valid psychodiagnostic signs" asserts that the false patterns did not merely coexist with the real ones. They crowded them out.

That is a stronger and more useful claim than the existence of the illusion. It says attention is finite, and a confident false rule occupies the space a real one would have used.

And we did not obtain the paper, so we report the claim as its title states it and cannot tell you how it was demonstrated or how large the displacement was.

The 1976 Paradigm

The experimental design that made the effect studyable at scale.

Hamilton, D. L., and Gifford, R. K. (1976), Illusory correlation in interpersonal perception: A cognitive basis of stereotypic judgments, Journal of Experimental Social Psychology, 12(4), 392–407, July, DOI 10.1016/S0022-1031(76)80006-6[2].

Its abstract: "Illusory correlation refers to an erroneous inference about the relationship between two categories of events. One postulated basis for illusory correlation is the co-occurrence of events which are statistically infrequent; i.e., observers overestimate the frequency of co-occurrence of distinctive events. If one group of persons 'occurs' less frequently than another and one type of behavior occurs infrequently, then the above hypothesis predicts that observers would overestimate the frequency that that type of behavior was performed by members of that group."[2]

And its result: "Results of two experiments testing this line of reasoning provided strong support for the hypothesis."[8]

Three observations, ours.

The design has two things that are rare: a smaller group and a less common behaviour. The prediction concerns the combination of the two.

The published conclusion is carefully bounded: distortions can result from these mechanisms "at least when the various events co-occur with differential frequencies"[3]. That conditional is in the paper's own abstract and is usually dropped.

And the reason this design matters commercially has nothing to do with its social psychology framing. It is a general model of what happens when you observe two binary variables at unequal frequencies, which describes most business record-keeping.

Identical Ratios

The feature that makes the demonstration airtight.

A summary records: "Although evaluatively equivalent information was provided about both groups, people perceived the groups differently because of the effect of distinctive information."[9]

Four observations, ours.

The proportion of desirable to undesirable behaviour is the same for both groups. Only the totals differ, because one group appears more often than the other.

So there is no relationship to find. Knowing which group someone belongs to tells you nothing about their behaviour, by construction.

That is what separates this from a finding about prejudice. The participants had no prior information about either group, which were unlabelled, so nothing they brought with them can explain the result.

And it means the effect is a claim about arithmetic and memory rather than about attitudes, which is why it belongs in a series about business judgment.

The Correlation Is Exactly Zero

Our own arithmetic, using the cell counts reported in secondary accounts, which we did not verify against the paper.

The standard array gives the larger group 26 desirable and 13 undesirable behaviours, and the smaller group 8 desirable and 4 undesirable. Fifty-one observations in total.

Both ratios are exactly 2.00 to 1.

The phi coefficient, which is the correlation between two binary variables, comes to 0.0000.

Three observations.

Not approximately zero. Exactly zero, because equal ratios across both rows is precisely the condition for independence in a two-by-two table.

The distinctive combination, the smaller group with the less common behaviour, occurs four times out of fifty-one. That is under eight percent of the array and it is the most unusual thing in it.

And that is the entire finding in one sentence. Four observations, correctly counted, would tell you nothing. Four observations, vividly remembered, tell you something false.

Distinctiveness

The proposed mechanism, in the literature's own words.

A recent paper states it: "If the distributions of two binary variables are skewed, people erroneously perceive a correlation even if the variables are actually uncorrelated. Specifically, people perceive a correlation between the variables' infrequent (vs. frequent) levels." And: "As proposed in the distinctiveness-based account, ICs arise due to a memory advantage for infrequent events."[10]

Three observations, ours.

The first sentence is the cleanest general statement of the phenomenon we found, and it is purely statistical. Two skewed binary variables, no correlation, and an erroneous perception.

The prediction is specific in a way most bias claims are not. The perceived link is always between the infrequent levels, not just any spurious link, which makes it testable.

And the mechanism named is a memory advantage, which locates the failure in encoding and recall rather than in reasoning. That distinction matters for the remedies at the end.

The Meta-Analysis

The quantitative synthesis.

Mullen, B., and Johnson, C. (1990), Distinctiveness-based illusory correlations and stereotyping: A meta-analytic integration, British Journal of Social Psychology, 29, 11–28[4].

Its abstract: "This article reports the results of a meta-analytic integration of previous research on illusory correlation in stereotyping effects. The following patterns were observed. The basic distinctiveness-based illusory correlation effect is highly significant, and of moderate strength. Consistent with theoretical expectations, distinctiveness-based illusory correlation effects are stronger when the distinctive behaviour is negative. Effects are also stronger as a function of the number of exemplars presented in the stimulus array. This is consistent with the effects of memory load on covariation judgement demonstrated elsewhere."[4]

We obtained the abstract and not the paper, and it reports no numerical effect size that we could obtain.

Highly Significant And Of Moderate Strength

Reading that carefully, because the two halves say different things. Ours.

Four observations.

"Highly significant" is a statement about confidence that the effect is not zero. "Of moderate strength" is a statement about size. Those are independent and this series has repeatedly found writers collapsing them.

Moderate is the right word for a real effect that will not dominate a decision. It sits in the same territory as the median effect size the fiftieth article computed across this publication's first forty-nine articles, which was 0.41.

The second moderator is the one with commercial teeth. Effects are stronger when the distinctive behaviour is negative, so the illusion is worse for problems than for successes, which is exactly the direction that generates unfair reputations.

And the third is stranger and more useful. Effects are stronger with more exemplars, which the authors link to memory load. More data made the illusion worse, not better, which is the opposite of what anyone would assume.

The Finding Inside It That Cuts The Other Way

The last sentence of that abstract, which complicates the headline and which our sources truncate.

The abstract continues: judgments "are significantly predicted by the paired distinctive covariation judgement strategy. This indicates that subjects' judgements of covariation in the illusory correlation in stereotyping paradigm seem to reflect a responsiveness to the information being presented to" them, at which point all three of our sources cut off[4].

Four observations, ours.

The phrase "responsiveness to the information being presented" is doing something important. It says participants were not inventing; they were tracking a specific feature of the data.

That is a partly rational reading of what is presented as a bias. If people are systematically weighting one cell, they are applying a rule to real information rather than hallucinating.

Which would relocate the error. Not a failure of perception but a badly chosen statistic, and a badly chosen statistic is correctable by naming the right one, which is what the arithmetic sections below do.

And we cannot tell you how the sentence ends. It is truncated identically in every source we found, and it is the sentence that would tell you how far the rational reading goes.

The Mechanism Is Disputed

Where the explanation stands, stated by the literature rather than by us.

A recent paper: "While such systematic Illusory Correlations (ICs) can account for important phenomena, including erroneous stereotypes linking minority groups with infrequent attributes, the theoretical explanation is still a matter of debate."[10]

Three observations, ours.

"Still a matter of debate" about a phenomenon first named in 1967 is a long time for a mechanism to remain open, and it is the third consecutive article in this series to report that structure.

The same source describes the competing consideration: a smaller number of studies investigate "biased memory and systematic inferences within the same paradigm"[10], which names two rival accounts.

And the practical consequence is limited, which is worth saying plainly. The arithmetic below works regardless of why people get it wrong, so a reader can act on this article without the dispute being settled.

Information Loss

The principal alternative account, which we can name and not evaluate.

Reference lists identify Fiedler, K. (1991), The tricky nature of skewed frequency tables: An information loss account of distinctiveness-based illusory correlations, and Fiedler, K., Hemmeter, U., and Hofmann, C. (1984), On the origin of illusory correlations, European Journal of Social Psychology, 14, 191–201[11][6].

We obtained neither and report the titles.

Three observations, ours.

The phrase information loss suggests a different story from distinctiveness. Not that rare events are remembered better, but that information degrades as it is stored, and degradation hurts small samples more.

If that is right the prediction is similar and the remedy differs. Distinctiveness says pay less attention to vivid cases; information loss says write things down at the time.

And the same author appears in a 2011 paper on the topic[11], so this is a sustained programme rather than a single objection, which is why we name it rather than treating distinctiveness as settled.

The Four Cells

The structure everything in this article reduces to, stated once so the rest is readable. Ours.

Any claim of the form X goes with Y has four relevant counts, not one.

Cell A: X present, Y present. Cell B: X present, Y absent. Cell C: X absent, Y present. Cell D: X absent, Y absent.

Three observations.

Only cell A produces memorable events. The customer who came from a referral and turned out excellent is a story; the customer who came from a referral and turned out ordinary is not.

Cell D is almost never recorded at all. Nobody logs the prospects who did not come from a referral and did not become customers, and there are usually more of them than everything else combined.

And a relationship exists only if the rate in the first row differs from the rate in the second, which requires all four. That is the whole of the statistics and it is what the next two sections demonstrate.

Cell A Alone

Our own arithmetic, invented figures throughout.

A campaign ran. Afterwards you count: 40 sales followed the campaign. Forty is real, countable and memorable.

Now the full table. Campaign run: 40 sales, 60 no sales, a rate of 40.0 percent. No campaign: 40 sales, 60 no sales, a rate of 40.0 percent.

Three observations.

The rates are identical, so the campaign did nothing whatever.

Cell A is unchanged by that fact. Forty sales did follow the campaign, and every one of them is a real sale to a real customer.

And this is the shape of most commercial pattern-finding. The thing you notice is cell A, the thing that matters is the difference between two rates, and cell A is one of the four numbers you need.

Same Cell A, Five Different Answers

Holding cell A fixed at forty and varying the rest. Our own arithmetic, invented figures.

40, 60, 40, 60: true lift 0.0 points. No effect.

40, 60, 20, 80: true lift +20.0 points. The campaign works.

40, 60, 60, 40: true lift negative 20.0 points. The campaign hurts.

40, 10, 40, 60: true lift +40.0 points. The campaign works very well.

40, 160, 40, 60: true lift negative 20.0 points. The campaign hurts.

Four observations.

Cell A is forty in every row. Anyone reasoning from it alone has the same evidence in all five cases and the truth ranges from plus forty points to minus twenty.

Two rows produce identical verdicts from completely different tables, which is the point restated: cell A carries no information about the answer at all.

Note that rows three and five differ only in how many times the campaign ran without producing a sale, which is the number nobody counts and which flips a strong positive impression into a negative result.

And the practical form is a question rather than a calculation. How often did it happen when we did not do the thing? That single question supplies cells C and D and does most of the work.

One further point about row four, which is the one worth arguing with. It reports the campaign working by forty points, and it does so because the campaign ran only fifty times rather than a hundred. Nothing about the successes changed; the denominator did.

That is the most common way a real business overstates a result. You remember the times you ran the campaign and it worked, and you forget how many times you ran it at all, because the failures were not events. A quote that went nowhere leaves no trace in anyone's memory and no entry in most systems.

And it is why the fix is a counting discipline rather than an analytical one. The number that would settle these five rows apart is not a statistic, it is a tally, and it has to be kept at the time because it cannot be reconstructed afterwards.

The Rare Cell Is The Least Reliable One

The observation we think is the most useful thing in this article, and we have not seen it made in these terms. Our own arithmetic, standard sampling error on a proportion.

The half-width of a 95 percent interval on a rate, by the number of observations behind it:

4 observations: plus or minus 49.0 percentage points. 8: 34.6. 20: 21.9. 50: 13.9. 100: 9.8. 400: 4.9.

Four observations.

The distinctive cell in the standard experimental design contains four observations, whose rate is uncertain to nearly fifty percentage points. It is close to uninformative.

And it is simultaneously the most memorable cell in the array, because it is the rarest combination. Those two facts are the whole problem stated arithmetically.

So the finding can be restated without any psychology at all. Memorability and statistical reliability run in opposite directions, because what makes an observation stand out is that there are few like it.

That inverse relationship is general, and it is why the remedy is not to think harder. Thinking harder about a vivid case gives more weight to the least reliable evidence available, which makes the judgment worse rather than better.

And it explains the meta-analytic moderator that otherwise makes no sense. Effects were stronger with more exemplars presented, which sounds backwards until you notice that a larger array makes the rare combination rarer in relative terms and therefore more distinctive, while adding memory load. More data sharpens the thing you should ignore.

A Field Test We Could Not Obtain

The study we would most like to report, named only.

Reference lists identify Redelmeier, D. A., and Tversky, A. (1996), On the belief that arthritis pain is related to the weather[11].

We did not obtain it and report the title and authors only.

Three observations, ours.

The design implied by the title is the ideal test: a widely held belief about a relationship, real patients, and independently recorded weather data. Nothing hypothetical.

The senior author is the same Amos Tversky who appears throughout this series, and the study is Canadian in origin through its lead author's affiliation, which we mention because this publication's readers are.

And we will not tell you what it found. A title is not a finding, and the title is phrased neutrally enough that guessing would be inventing. We name it so a reader knows where to look.

Two Different Things Under One Name

A distinction our sources draw and most treatments blur, which matters because the two have different remedies. Ours.

A reference encyclopedia separates them: distinctiveness-based illusory correlation, where the relationship is believed because of attention given to infrequent information, and expectancy-based illusory correlation, which is "misperceptions of relationships due to people's preexisting expectations."[9]

Four observations.

These are opposite in their inputs. The first requires no prior belief at all and is generated purely by the frequency structure of what you observe. The second requires a prior belief and is generated by biased processing of what arrives.

The everything-in-this-article arithmetic applies to the first only. The 1976 design deliberately used unlabelled groups precisely so that no expectation could be at work.

And they call for different fixes. Distinctiveness is corrected by counting all four cells. Expectancy is corrected by pre-registering what you expect before you look, which is the confirmation bias remedy the forty-fifth article covered and a different instrument entirely.

This is the sixth literature in which this series has found one name covering more than one thing. The pattern is common enough now that we would treat it as the default expectation rather than a surprise, and the practical rule follows: before adopting a remedy, establish which version of the effect you are facing.

What Actually Survives

Our reading, stated directly.

Five statements.

The effect is real and moderate. A meta-analysis calls it highly significant and of moderate strength, which is two claims and both matter.

It gets worse with more data and with negative content. Both moderators are reported in that abstract and both cut against intuition.

The mechanism remains disputed after nearly sixty years, with distinctiveness and information loss as competing accounts and the literature itself calling the question open.

Part of it may be responsiveness to real information, on the meta-analysis's own final sentence, which our sources truncate before it finishes.

And the arithmetic does not depend on any of that. Five tables with the same memorable cell give verdicts from plus forty to minus twenty, whatever the psychology turns out to be.

Your Business Patterns

The first application. Ours, untested, and not analytics advice.

Four points.

Every rule of the form our best customers always come from X is a claim about a two-by-two table, and it is almost always being made from cell A.

The test is one question. What fraction of customers from X are good, against what fraction of customers not from X? If you cannot answer the second half, you do not have a finding.

The meta-analytic moderator makes this worse with scale rather than better. Effects are stronger with more exemplars, so a longer trading history does not correct the impression; it may entrench it.

And clause (b) of the definition means finding some relationship does not vindicate you. The rule may be directionally right and badly overstated, which is the more common failure and the harder one to see.

Hiring And Clients

The second application, and the one with the most at stake. Ours, and nothing here bears on the lawfulness of any employment practice.

Four points.

The founding clinical work is directly on point. Experienced professionals, working within their expertise, reported relationships their data did not contain, which is precisely the claim behind most confident hiring intuition.

The structural problem is that you observe cell A and cell B and almost never C or D. You see how the people you hired worked out. You do not see how the people you rejected would have.

That is the fifty-eighth article's survivorship problem in a two-by-two frame, and this publication already has an article on hiring validity that reports what the selection literature actually supports.

And the negative-content moderator is the one to watch here. The illusion is stronger when the distinctive behaviour is negative, so a rare bad outcome from an unusual candidate profile will generate a rule far out of proportion to its evidence.

The Data You Never Collected

The third application, and the cheapest fix in this article. Ours.

Four points.

Cell D is the largest cell in most business situations and the only one nobody records. The quotes that went nowhere, the leads that never converted, the months the problem did not occur.

Its absence is not neutral. A table missing cell D cannot produce a rate for the second row, which means no comparison is possible and any perceived pattern is cell A wearing a suit.

The practical version is a logging decision rather than an analysis one. Record the non-events for one quarter: enquiries that did not convert, deliveries that were not late, weeks the supplier did not disappoint.

And that is the remedy the information loss account would favour, if it is right. Writing things down at the time defeats degradation in storage, which is the same instrument this series has now recommended in six articles for six different failures.

Whatever You Have Less Of

A closing implication, and it is the one with the widest reach. Ours.

Four observations.

The prediction is specifically about the infrequent levels of two skewed variables. Whatever you have less of will attract a false association with whatever else you have less of.

In a business that means your newest product line, your smallest customer segment, your one supplier in a different country, your most recent hire. Each is a minority category in your own records.

And because the effect is stronger for negative content, the specific prediction is uncomfortable and testable. The rare bad outcome will attach itself to the rare category, and the resulting rule will feel like experience.

The check is the same one throughout. Count the rate, not the incidents, and demand the denominator before accepting any rule about a category you have few of.

What To Do

Ask for all four cells before accepting any pattern. On our own arithmetic, five tables sharing the same memorable cell give verdicts from plus forty points to minus twenty.

Ask the one question that supplies the missing half. How often did it happen when we did not do the thing? That fills cells C and D and does most of the work.

Distrust your most vivid cases specifically. Memorability and reliability run in opposite directions, because what makes an observation stand out is that there are few like it.

Treat rules about small categories with extra suspicion. The predicted illusion is between the infrequent levels of both variables, which in a business means new lines, small segments and recent hires.

Expect it to be worse for problems than for successes. The meta-analysis reports stronger effects when the distinctive behaviour is negative.

Do not expect more experience to fix it. The same meta-analysis reports stronger effects with more exemplars, which the authors link to memory load.

Log the non-events for one quarter. Cell D is the biggest cell and the only one nobody records, and without it no rate comparison is possible.

Remember that being directionally right is not vindication. The founding definition covers relationships that exist and are weaker than reported, which is the more common and less visible failure.

The Limits Of This Analysis

Several caveats matter. This article discusses research on judgment and is not analytics, hiring or legal advice; nothing here bears on the lawfulness of any employment practice, and the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 1967 paper that defines the term, and have its definition verbatim only because two later papers quote it with a page reference. We did not obtain either of the clinical papers, which are the origin of the concept and the most pointed evidence in this article, and report them through citations and one encyclopedia summary; we have deliberately not described the specific historical diagnostic claim at issue, which concerned a classification now abandoned. We did not obtain the 1976 paper, only its abstract from three independent sources, so the cell counts underlying our arithmetic are the standard ones from secondary accounts and we did not verify them; if they are wrong, our phi calculation is wrong with them, though the structural point that equal ratios give zero correlation holds regardless. We obtained the 1990 meta-analysis's abstract and not the paper, and it reports no numerical effect size we could obtain, so "moderate strength" is the authors' characterisation rather than a figure; its final sentence is truncated identically in all three of our sources, and it is the sentence bearing on how far a rational reading of the effect extends. We obtained none of the competing theoretical accounts and report the information loss alternative from titles alone. We did not obtain the field study on a real medical belief and report its title and authors without any indication of what it found. Two sources are an encyclopedia and a teaching site, flagged at every use. All arithmetic is ours; the contingency and sampling calculations are elementary and reproducible, and every business figure is invented. And the laboratory paradigm involves unlabelled groups and a single sitting, which is what makes it a clean demonstration and also what limits how far it transfers to a business owner reasoning over years of accumulated impressions.

Frequently Asked Questions

What is illusory correlation?
The founding 1967 definition is the report by observers of a correlation between two classes of events which in reality are either not correlated, or correlated to a lesser extent than reported. That second clause covers overstatement rather than fabrication and is the more common case.
How was it demonstrated?
By showing people behaviours attributed to a larger and a smaller group, with the ratio of desirable to undesirable identical for both. On our own arithmetic the correlation is exactly zero. People perceived the groups differently anyway, driven by the rarest combination, which occurred four times in fifty-one observations.
How strong is the effect?
A 1990 meta-analysis calls it highly significant and of moderate strength. It also reports the effect is stronger when the distinctive behaviour is negative, and stronger with more exemplars presented, which the authors link to memory load. More data made it worse.
Is the mechanism settled?
No. A recent paper states the theoretical explanation is still a matter of debate. The distinctiveness account says rare events get a memory advantage; an information loss account says information degrades in storage and degradation hurts small samples most. We obtained neither of the competing papers.
What is the practical test?
One question: how often did it happen when we did not do the thing? On our own arithmetic, five tables sharing the same count in the memorable cell produce verdicts from plus forty points to minus twenty. The memorable cell carries no information about the answer.
Why is the vivid case the wrong one to trust?
Because memorability and statistical reliability run in opposite directions. What makes an observation stand out is that there are few like it, and on our own calculation a rate based on four observations carries a 95 percent interval of roughly plus or minus 49 percentage points.
Which of my business categories are most at risk?
Whatever you have least of. The prediction concerns the infrequent levels of two skewed variables, so newest product lines, smallest segments, single foreign suppliers and recent hires. And because the effect is stronger for negative content, the rare bad outcome will attach itself to the rare category.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article reports a finding whose mechanism has been disputed for nearly sixty years, and an arithmetic point that holds whichever side turns out to be right.

References

  1. Hosted copy of the 1976 paper reproducing its opening text, which quotes the founding definition with a page reference: that Chapman (1967) introduced the concept of illusory correlation to refer to the report by observers of a correlation between two classes of events which, in reality, are either not correlated or are correlated to a lesser extent than reported, at page 151; and recording that most research on the topic had been concerned with illusory correlation as a basis for erroneous reports of relationships between symptoms of patients and performance on psychodiagnostic tests, citing Chapman and Chapman (1967, 1969), Golding and Rorer (1971) and Starr and Katkin (1969). Together with a second hosted teaching copy reproducing the same quotation, the paper's affiliation at Yale University, and its summary that results of two experiments testing this line of reasoning provided strong support for the hypothesis. Note: hosted copies of the paper's opening text, including a teaching site. Our source for the 1967 definition verbatim with its page reference; we did not obtain the 1967 paper. cliffsnotes.com
  2. Hamilton, D. L., & Gifford, R. K. (1976). Illusory correlation in interpersonal perception: A cognitive basis of stereotypic judgments. Journal of Experimental Social Psychology, 12(4), 392–407, July, DOI 10.1016/S0022-1031(76)80006-6. Publisher record reproducing the abstract: on illusory correlation referring to an erroneous inference about the relationship between two categories of events; on one postulated basis being the co-occurrence of events which are statistically infrequent, such that observers overestimate the frequency of co-occurrence of distinctive events; and on the prediction that if one group of persons occurs less frequently than another and one type of behavior occurs infrequently, observers would overestimate the frequency with which that type of behavior was performed by members of that group. Note: the publisher's record. We obtained the abstract only and not the paper, so the cell counts used in our arithmetic come from secondary accounts and are unverified. sciencedirect.com
  3. Bibliographic database record for the same 1976 paper, giving the citation as Hamilton, David L. and Gifford, Robert K., Journal of Experimental Social Psychology, 12, 4, 392–407, July 1976, and reproducing the author abstract: on illusory correlation referring to an erroneous inference a person makes about the relationship between two categories of events; and on the present findings demonstrating that distortions in judgment can result from the cognitive mechanisms involved in processing information about co-occurring events, at least when the various events co-occur with differential frequencies. Note: an educational resources database. Our source for the paper's own conditional qualification, which is usually dropped from summaries. eric.ed.gov
  4. Mullen, B., & Johnson, C. (1990). Distinctiveness-based illusory correlations and stereotyping: A meta-analytic integration. British Journal of Social Psychology, 29, 11–28. Bibliographic service record reproducing the abstract: on the article reporting the results of a meta-analytic integration of previous research on illusory correlation in stereotyping effects; on the basic distinctiveness-based illusory correlation effect being highly significant and of moderate strength; on effects being stronger when the distinctive behaviour is negative, consistent with theoretical expectations; on effects also being stronger as a function of the number of exemplars presented in the stimulus array, consistent with the effects of memory load on covariation judgement demonstrated elsewhere; and on subjects' judgements of covariation being significantly predicted by the paired distinctive covariation judgement strategy, indicating that those judgements seem to reflect a responsiveness to the information being presented to them, at which point the reproduction is cut off. The publisher's record and a repository record reproduce the same abstract and are cut off at the same point. Note: a bibliographic service record, corroborated by two others. The abstract reports no numerical effect size we could obtain, and its final sentence is truncated identically in all three sources. semanticscholar.org
  5. Reference list carried on a hosted academic reading, confirming Chapman, L. (1967), Illusory correlation in observational report, Journal of Verbal Learning and Verbal Behavior, 6(1), 151–155, DOI 10.1016/S0022-5371(67)80066-5; Chapman, Loren J. and Jean P. (1969), Illusory Correlation as an Obstacle to the Use of Valid Psychodiagnostic Signs, Journal of Abnormal Psychology, 74(3), 271–280, DOI 10.1037/h0027592, PMID 4896551; and Hamilton, D. and Gifford, R. (1976), Journal of Experimental Social Psychology, 12(4), 392–407, DOI 10.1016/S0022-1031(76)80006-6. Note: a hosted reference list, used to confirm citations, DOIs and the PubMed identifier independently. We obtained none of the papers named. facultypsy.hope.edu
  6. Publisher record for a book chapter on illusory correlations and stereotype theory, carrying a reference list confirming Chapman, J. L. and Chapman, J. P. (1967), Genesis of popular but erroneous psychodiagnostic observations, Journal of Abnormal Psychology, 72, 193–204; Chapman, L. J. and Chapman, J. P. (1969), Journal of Abnormal Psychology, 74, 271–280; Crocker, J. (1981), Judgment of covariation by social perceivers, Psychological Bulletin, 90, 272–292; Fiedler, K., Hemmeter, U. and Hofmann, C. (1984), On the origin of illusory correlations, European Journal of Social Psychology, 14, 191–201; Golding, S. L. and Rorer, L. G. (1972), Illusory correlation and subjective judgment, Journal of Abnormal Psychology, 80, 249–260; and Hamilton, D. L. and Rose, T. L. (1980), Journal of Personality and Social Psychology, 39, 832–845. Note: a publisher's reference list; citations only. We obtained none of the works named, including both clinical papers. link.springer.com
  7. Encyclopedia entry on illusory correlation, recording that the term was originally coined by Chapman (1967) to describe people's tendencies to overestimate relationships between two groups when distinctive and unusual information is presented; that the concept was used to question claims about objective knowledge in clinical psychology through the Chapmans' refutation of a set of purported diagnostic signs then widely used by clinicians; and that Hamilton and Gifford (1976) conducted a series of experiments in which participants read sentences describing either desirable or undesirable behaviours attributed to either a majority or a minority group. Note: an encyclopedia, not an academic source, flagged at every use. We have deliberately not described the specific historical diagnostic claim at issue. en.wikipedia.org
  8. Hosted student copy of the 1976 paper reproducing its abstract and front matter, including the authors' Yale University affiliation, the acknowledgement of NSF and NIMH grant support, and the statement that the differential perception of majority and minority groups could result solely from the cognitive mechanisms involved in processing information about stimulus events differing in their frequencies, with results of two experiments providing strong support for the hypothesis. Note: a student document-sharing site, not an academic source, flagged at every use. Used only to corroborate abstract wording already obtained from the publisher and a bibliographic database. studocu.vn
  9. Psychology reference encyclopedia entry on illusory correlation, recording that although evaluatively equivalent information was provided about both groups, people perceived the groups differently because of the effect of distinctive information; that this is known as a distinctiveness-based illusory correlation because a relationship is believed to exist between two variables as a result of the special attention given to distinctive, meaning infrequent, information; and that expectancy-based illusory correlations are a separate category, being misperceptions of relationships due to people's preexisting expectations. Note: a reference encyclopedia, not a peer-reviewed source, flagged at every use. Our source for the statement that the information provided about both groups was evaluatively equivalent. psychology.iresearchnet.com
  10. Repository page for the 1990 meta-analysis carrying scholarly citing text, recording that if the distributions of two binary variables are skewed people erroneously perceive a correlation even if the variables are actually uncorrelated, and specifically perceive a correlation between the variables' infrequent rather than frequent levels; that while such systematic illusory correlations can account for important phenomena the theoretical explanation is still a matter of debate; that as proposed in the distinctiveness-based account they arise due to a memory advantage for infrequent events; and that besides many studies demonstrating illusory correlations in contingency judgments, a smaller number investigate the key explanatory variables of biased memory and systematic inferences within the same paradigm. Note: a repository page reproducing text from citing papers. Our source for the statement that the mechanism remains disputed and for the general statistical formulation of the effect. researchgate.net
  11. Publisher record for a 2011 paper revisiting illusory correlations, carrying a reference list confirming Mullen, B. and Johnson, C. (1990), British Journal of Social Psychology, 29, 11–28; Fiedler, K. (1991), The tricky nature of skewed frequency tables: An information loss account of distinctiveness-based illusory correlations; Sanbonmatsu, D. M., Sherman, S. J. and Hamilton, D. L. (1987), Illusory correlation in the perception of individuals and groups, Social Cognition, 5, 1–25; and Pryor, J. B. (1986), Personality and Social Psychology Bulletin, 12, 216–226. Together with a separate hosted reference list confirming Redelmeier, D. A. and Tversky, A. (1996), On the belief that arthritis pain is related to the weather, and Kurtz, R. M. and Garfield, S. L. (1978), Illusory correlation: a further exploration of Chapman's paradigm, Journal of Consulting and Clinical Psychology, 46, 1009–1015. Note: reference lists; titles and citations only. Our source for the competing information loss account and for the field study, neither of which we obtained; we report the field study's title without any indication of what it found. journals.sagepub.com

This article discusses research on judgment and is not analytics, hiring or legal advice; nothing here bears on the lawfulness of any employment practice. The 1967 paper and both clinical papers were not obtained. The 1976 paper was obtained as an abstract only, so the cell counts underlying the arithmetic come from secondary accounts and are unverified. The 1990 meta-analysis reports no numerical effect size that could be obtained, and its final sentence is truncated identically in all three sources consulted. The competing theoretical accounts and the field study are reported from titles alone. Three sources are non-academic and flagged at every use. All arithmetic is the authors' own; business figures are invented throughout.