Ask a client to list ten reasons they value your service and they may end up valuing it less than if you had asked for three. That is the claim, and it comes from one of the most cited experiments in social psychology. This article covers what it found, and what happened when someone tried it again with sixty times the sample.

Key Takeaway

From the abstract: "Ss who had to recall 12 examples of assertive (unassertive) behaviors, which was difficult, rated themselves as less assertive (less unassertive) than subjects who had to recall 6 examples, which was easy. In fact, Ss reported higher assertiveness after recalling 12 unassertive rather than 12 assertive behaviors. Thus, self-assessments only reflected the implications of recalled content if recall was easy."[1] The paper's tables state "n = 9 or 10 per condition."[2] A later team "tried and failed to replicate" it with N = 661[3].

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade C for the ease-of-retrieval effect on judgment. An elegant original with very small cells, and a well-powered independent attempt that failed to reproduce it.

Grade A that the manipulation changes felt difficulty. The replication reports this held in three of its four studies, which is the strongest thing in this article.

Grade B for the underlying availability heuristic, which is older, broader, and not what the replication tested.

Our position: the mechanism is real at the first step and unproven at the second, which is the same shape as the thirty-eighth article's broken chain, arrived at from a completely different literature.

A Note On Method

Everything here is verified to August 2026.

We obtained the 1991 paper's abstract verbatim from four independent sources and, unusually for this series, a hosted copy of the paper itself from a university teaching page, from which we took the introduction and the table notes[1][2][4][5].

We did not obtain the paper's statistics, effect sizes or full results, only its abstract, introduction and the sample-size notes beneath its tables.

The replication reaches us as a publisher abstract that truncates mid-sentence[3]. We did not obtain that paper either.

We report no effect size for anything in this article, from either the original or the replication.

All power calculations are ours and are illustrative of scale rather than exact.

This article discusses research on judgment and survey design. It is not market research, professional conduct or client engagement advice.

The Heuristic It Revisits

The older idea the paper takes another look at.

Quoting from the paper's own introduction: "One of the most widely shared assumptions in decision making as well as in social judgment research holds that people estimate the frequency of an event, or the likelihood of its occurrence, 'by the ease with which instances or associations come to mind' (Tversky & Kahneman, 1973, p. 208)."[2]

A later source restates the logic: "If an incident comes to mind easily, people believe there must be many such incidents in the population from which it is drawn. Conversely, the more difficult it is to remember an incident, the smaller one should perceive the overall population."[5]

Two observations, ours.

The 1973 formulation is about ease, not about count, and that distinction is the whole of what follows. The heuristic as originally stated is a claim about a subjective experience.

And we did not obtain the 1973 paper. We have its central phrase with a page number, quoted inside a paper we did obtain, which is better than most secondhand citation in this series but is still secondhand.

The Confound The Paper Attacks

Why the study was designed, which is the most impressive thing about it.

From the introduction: "Typically, the manipulations that are introduced to increase the subjectively experienced ease of recall are also likely to affect the amount of subjects' recall. As a result, it is difficult to evaluate if the obtained estimates of frequency likelihood, or typicality are based on subjects' subjective experiences or on a biased" sample of what was recalled, with our source truncating[2].

Three observations, ours.

That is a genuinely sharp identification problem, and the authors state it themselves in their opening. Anything that makes recall feel easier usually also produces more recalled material, so the two explanations are tangled.

The design separates them by making the harder condition produce more content. If content drove the judgment, recalling twelve examples of your own assertiveness should make you feel more assertive.

And this is what makes the experiment elegant rather than merely surprising. The two accounts predict opposite results, which is the structure this series praised in the twentieth article's adversarial collaboration.

Six Versus Twelve

The result.

"Ss who had to recall 12 examples of assertive (unassertive) behaviors, which was difficult, rated themselves as less assertive (less unassertive) than subjects who had to recall 6 examples, which was easy."[1]

Two observations, ours.

The participants who produced twice as much evidence of their own assertiveness came away thinking they were less assertive.

And the parenthetical matters: it worked in both valences. Recalling twelve unassertive behaviours made people rate themselves as less unassertive, which is the same effect running in the opposite direction and is the harder version to explain away.

The Reversal

The sentence that makes the content account untenable.

"In fact, Ss reported higher assertiveness after recalling 12 unassertive rather than 12 assertive behaviors."[1]

And the conclusion the authors draw: "Thus, self-assessments only reflected the implications of recalled content if recall was easy."[1]

Two observations, ours.

Read the first sentence carefully. People who spent the task listing twelve ways they had been unassertive then rated themselves as more assertive than people who listed twelve ways they had been assertive.

If that holds, the content of what you recall is not merely diluted by the difficulty of recalling it. It can be overwhelmed and reversed by it.

The Discrediting Manipulation

The control that tests the mechanism rather than the effect.

"The impact of ease of recall was eliminated when its informational value was discredited by a misattribution manipulation."[1]

A separate description adds that recalling six produced higher assertiveness ratings than twelve "except when recall difficulty was discredited, which reversed typical patterns."[5]

We did not obtain the details of the manipulation and report only that it existed and what it did.

Three observations, ours.

A misattribution manipulation gives participants an alternative explanation for why the task felt hard, so the difficulty stops being informative about them.

That the effect disappeared under it is strong evidence about mechanism. It shows the difficulty was being read as a signal rather than simply interfering with the task.

And it is the most practically suggestive part of the paper, because the remedy it implies is telling people why the task is hard. We flag that we are inferring the application; no source proposes it.

Nine Or Ten Per Cell

The number that changes how to read all of the above.

The notes beneath the paper's own tables read "n = 9 or 10 per condition" for one experiment, "n = 18 to 20 per condition" for another, and "n is 9 or 10 per condition" for a third[2].

Three observations, ours.

Nine to ten people per cell is very small, and we report it because the paper reports it. It is stated plainly in the published tables rather than concealed.

The 1991 standard was different from today's, and this is not a charge of misconduct. It was normal practice.

But it does bear directly on how much weight the result can carry now, and that is a question of arithmetic rather than of blame.

What That Could Detect

Our own power calculation, a normal approximation to a two-sample comparison at conventional thresholds; the original used a factorial design so per-cell counts do not map exactly onto every contrast, and these figures are illustrative of scale.

With 9 per group, the smallest effect reliably detectable is about d = 1.32.

With 10: about d = 1.25.

With 20: about d = 0.89.

For scale, this series has reported framing at d = 0.31, feedback interventions at 0.41, and growth mindset at 0.08.

Two observations.

A design powered only for effects above d = 1.25 will detect a real effect only if that effect is enormous, or if the sample happened to fall favourably.

Which does not mean the finding is wrong. It means the study could not have distinguished a very large effect from a lucky draw, and that is why what happened next matters more than the original.

Six Hundred And Sixty-One

The attempt.

A study published in the journal Memory, 29(2), examined the effect in an eyewitness context. Its authors report: "We manipulated the number of reasons participants gave to justify their identification (Study 1; N = 343), and also the number of instances they provided of a weak or strong memory (Studies 2a & 2b; Ns = 350 & 312, respectively). Across the three studies, ease-of-retrieval did not affect eyewitnesses' confidence or other testimony-relevant judgements."[3]

And then: "We then tried, and failed, to replicate" the original ease-of-retrieval finding, "(Study 3; N = 661)."[3]

We did not obtain this paper and have its abstract, which truncates.

Two observations, ours.

On our own calculation, a study with 661 participants can detect an effect around d = 0.22. That is roughly six times more sensitive than the original.

And the authors did not merely study a different context and find nothing. They went back and attempted the original paradigm directly, which is the harder and more useful thing to do.

The detail that makes this a diagnosis rather than a demolition.

"In three of the four studies, ease-of-retrieval had the expected effect on participants' perceived task difficulty; however, frequentist and Bayesian testing showed no evidence for an effect on" the judgments, with our source truncating[3].

Three observations, ours.

The manipulation worked. Asking for more instances did make the task feel harder, and that was measured.

What did not appear was the second step, from felt difficulty to judgment. So the chain is: manipulation, then experienced difficulty, then evaluation. Link one held and link two did not.

And that is exactly the structure of the thirty-eighth article in this series, where a network meta-analysis found interventions changing a measure without the change carrying through to behaviour. Two unrelated literatures, the same failure at the same joint.

What Survives

Our honest residue. Ours.

Four statements.

Asking for more examples makes the task harder. This is measured and it held in the larger study.

Whether that difficulty changes the resulting judgment is unresolved, with an elegant small original on one side and a well-powered failure on the other.

The mechanism evidence is suggestive. That discrediting the difficulty eliminated the effect is the kind of result that is hard to get by chance, though it too rests on the small cells.

And the underlying availability heuristic is not what failed here. The replication tested one specific paradigm built on it, not the 1973 proposal itself.

Self Versus Other

A moderator worth recording, from a separate paper.

A publisher record states: "Four studies demonstrate that people rely relatively more on the experienced ease of recall when making judgments about the self compared to judgments about others." And: "Subjective retrieval ease was less informative when people were relatively less familiar with the specific other person."[6]

We did not obtain this paper and report the abstract's claims.

Two observations, ours.

If it holds, the effect is strongest exactly where the original tested it, which is self-assessment, and weaker where a firm might most want to use it, which is judgments about other people or things.

And it narrows the practical scope considerably. A questionnaire asking a client about themselves is the vulnerable case. One asking about your service may not be.

A Seventeenth Variant, And A Fitting One

The running sourcing note, and this one is the most self-illustrating yet.

A scholarly source citing this literature writes: "One of them is the availability heuristic (Schwarz et al., 1991), our tendency to take the first explanation for a phenomenon that comes to us."[5]

Two observations, ours.

That sentence contains two errors. The availability heuristic is Tversky and Kahneman's, from 1973; the 1991 paper is explicitly a reconsideration of it, as its subtitle says. And the definition given, taking the first explanation that comes to mind, is not the availability heuristic either.

The irony is worth stating plainly. An author reached for the first attribution that came to mind, and got it wrong, in a sentence about a heuristic concerning what comes to mind. That is the seventeenth bibliographic variant recorded in this series and the only one that demonstrates its own subject.

What This Means For Questionnaires

The application, ours and untested, offered with the grade attached.

Four points.

The number of examples you request is a design decision, not a neutral one. On the original finding it changes the answer; on the replication it changes the felt difficulty but perhaps not the answer.

Because the evidence is Grade C, we would not build anything important on it. But it costs nothing to ask for three examples rather than ten, and the downside of the smaller number is only less material.

The self-assessment case is the one to watch. If the moderator holds, questions asking people to rate themselves after listing evidence about themselves are where the risk concentrates.

And a long questionnaire is a difficulty manipulation whether or not you intended one. Question forty is answered by someone for whom the task has become hard, and that is true regardless of whether the 1991 judgment effect is real.

The Honest Position

Where we would leave it. Ours.

Three points.

This is a beautifully designed experiment. It identifies a real confound, resolves it with a design where the two accounts predict opposite outcomes, and includes a manipulation that tests the mechanism rather than just the effect. The thinking is better than most of what this series has covered.

It also has nine to ten people per cell, and a well-powered independent attempt did not reproduce its judgment effect.

And both of those things are true at once. Elegance of design and strength of evidence are separate properties, and this series has now found several cases where a study is admired for the first and cited as though it had the second.

What To Do

Ask for fewer examples than feels thorough. The downside is only less material, and on the original finding the larger request can move the answer against you.

Treat questionnaire length as a variable. A long form is a difficulty manipulation whether you intended one or not.

Watch self-assessment questions specifically. A separate paper reports the effect is stronger for judgments about the self than about others.

Do not build anything important on this. We grade the judgment effect C, and a study six times more sensitive than the original found no evidence for it.

Separate the two links. Asking for more instances does make the task harder, which is well supported. Whether that difficulty changes the judgment is what is unresolved.

Note that naming the difficulty may defuse it. The effect disappeared when its informational value was discredited, though we are inferring the application and no source proposes it.

Do not confuse an elegant design with strong evidence. This study has the first and, on its own, not the second.

The Limits Of This Analysis

Several caveats matter. This article discusses research on judgment and survey design and is not market research, professional conduct or client engagement advice. Everything is verified to August 2026. We obtained the 1991 abstract verbatim from four independent sources and a hosted copy of the paper's introduction and table notes, but not its statistics, effect sizes or results. We did not obtain the 1973 paper, and quote its central phrase secondhand from inside the 1991 paper. We did not obtain the replication, whose abstract truncates mid-sentence, nor the paper on self versus other judgments, nor the details of the misattribution manipulation. We report no effect size anywhere in this article, from either the original or the replication. All power calculations are ours, use a normal approximation at conventional thresholds, and are illustrative of scale rather than exact; the original used a factorial design, so per-cell counts do not map onto every contrast. The suggestion that naming a task's difficulty might defuse the effect is our inference from the misattribution result and is proposed by no source. The questionnaire applications are our own reasoning, untested, and are offered with the Grade C attached.

Frequently Asked Questions

What did the original find?
That people asked to recall twelve examples of their own assertive behaviour, which was difficult, rated themselves as less assertive than those asked for six. More evidence of assertiveness produced a lower self-rating, because the felt difficulty of producing it was read as a signal.
What is the strongest part of it?
The design. It identifies a real confound, which is that anything making recall feel easier usually also changes what gets recalled, and resolves it with a setup where the two explanations predict opposite outcomes. Plus a manipulation that discredited the difficulty and eliminated the effect.
What is the weakest part?
The sample. The paper's own tables state nine or ten people per condition. On our own calculation that detects only effects above about d = 1.25, where this series has reported framing at 0.31 and growth mindset at 0.08.
Did it replicate?
A team studying eyewitness confidence found no effect across three studies, then attempted the original paradigm directly with 661 participants and reports that they tried and failed to replicate it. On our own figures that study could detect an effect around d = 0.22.
So is the whole thing wrong?
No, and the replication is more informative than that. In three of four studies the manipulation did make the task feel harder, as measured. What did not appear was the step from felt difficulty to judgment. Link one held, link two did not.
Should I change my questionnaires?
Modestly. Asking for three examples rather than ten costs nothing but material, so it is a cheap precaution. We would not build anything important on a Grade C finding, but a long form is a difficulty manipulation whether or not the judgment effect is real.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article admires an experiment's design and reports its sample size in the same breath, because those are separate properties and the literature routinely treats them as one.

References

  1. Bibliographic service record for Schwarz, N., Bless, H., Strack, F., Klumpp, G., Rittenauer-Schatka, H., & Simons, A. (1991), Ease of retrieval as information: Another look at the availability heuristic, Journal of Personality and Social Psychology, 61, 195–202, reproducing the abstract, on experienced ease of recall having been found to qualify the implications of recalled content; on subjects who had to recall 12 examples of assertive or unassertive behaviors, which was difficult, rating themselves as less assertive or less unassertive than subjects who had to recall 6 examples, which was easy; on subjects in fact having reported higher assertiveness after recalling 12 unassertive rather than 12 assertive behaviors; on self-assessments thus only reflecting the implications of recalled content if recall was easy; on the impact of ease of recall having been eliminated when its informational value was discredited by a misattribution manipulation; and on the informative functions of subjective experiences being discussed. Note: a bibliographic service reproducing the abstract. We did not obtain the paper's statistics or effect sizes from this source. semanticscholar.org
  2. Hosted copy of the 1991 paper on a university teaching page, reproducing its front matter, abstract and introduction, on one of the most widely shared assumptions in decision making and social judgment research holding that people estimate the frequency of an event, or the likelihood of its occurrence, by the ease with which instances or associations come to mind, quoting Tversky and Kahneman (1973), page 208; on the manipulations typically introduced to increase subjectively experienced ease of recall also being likely to affect the amount of subjects' recall, making it difficult to evaluate whether obtained estimates are based on subjective experiences or on a biased sample, our source truncating; on the studies having been conducted at the Universität Heidelberg under the direction of the first three authors and supported by named research grants; and reproducing table notes stating that n = 9 or 10 per condition, that n = 18 to 20 per condition, and that n is 9 or 10 per condition. Note: a hosted copy on a university teaching page. Our source for the introduction and for the sample sizes, which appear in the paper's own published tables. We did not obtain its statistics or results. uni-muenster.de
  3. Publisher record for a study investigating the ease-of-retrieval effect in an eyewitness context, Memory, 29(2), on the authors having manipulated the number of reasons participants gave to justify their identification in Study 1 with 343 participants, and the number of instances they provided of a weak or strong memory in Studies 2a and 2b with 350 and 312 participants respectively; on ease-of-retrieval not having affected eyewitnesses' confidence or other testimony-relevant judgements across the three studies; on the authors having then tried, and failed, to replicate the original 1991 ease-of-retrieval finding in Study 3 with 661 participants; and on ease-of-retrieval having had the expected effect on participants' perceived task difficulty in three of the four studies, while frequentist and Bayesian testing showed no evidence for an effect on judgments, our source truncating. Note: a publisher record; the abstract truncates mid-sentence and we did not obtain the paper, its statistics or its effect sizes. tandfonline.com
  4. Repository record for the 1991 paper reproducing its front matter and abstract, recording the authors' institutional affiliations at ZUMA Mannheim, the Universität Mannheim, the Max Planck Institut für psychologische Forschung in Munich and the Universität Heidelberg. Note: a repository record; a further independent reproduction of the abstract text and the source for author affiliations. researchgate.net
  5. Repository page for the 1991 paper carrying citing text, on the findings having indicated that recalling six examples resulted in higher assertiveness ratings than recalling twelve, except when recall difficulty was discredited, which reversed typical patterns; on the availability heuristic stating that people tend to estimate the frequency of an event as a function of the ease with which it comes to mind, citing Tversky and Kahneman (1973), with an incident coming to mind easily leading people to believe there must be many such incidents and greater difficulty leading to a smaller perceived population; and separately containing a citing passage describing the availability heuristic as attributable to Schwarz and colleagues (1991) and defining it as the tendency to take the first explanation for a phenomenon that comes to us. Note: a repository page reproducing text from papers citing the 1991 article. The final passage misattributes the availability heuristic to the 1991 paper rather than to Tversky and Kahneman (1973) and misdefines it; both errors are reported in the body rather than repeated as fact. academia.edu
  6. Publisher record for a paper on the use of experienced retrieval ease in self and social judgments, on people basing judgments from memory on both the content of the information they retrieve and the ease they experience in retrieving it, citing Schwarz and colleagues (1991); on four studies demonstrating that people rely relatively more on the experienced ease of recall when making judgments about the self compared to judgments about others; on this pattern having been found for judgments of an average other and of a specific other; and on subjective retrieval ease having been less informative when people were relatively less familiar with the specific other person. Note: a publisher record; we obtained the abstract only and not the paper or its effect sizes. sciencedirect.com

This article discusses research on judgment and survey design and is not market research, professional conduct or client engagement advice. The 1991 paper's abstract, introduction and table notes were obtained; its statistics and results were not. The replication was not obtained and its abstract truncates. No effect size is reported anywhere in this article. All power calculations are the authors' own, are illustrative of scale rather than exact, and do not map precisely onto a factorial design.