The ninetieth article was about a study design that was almost never run. This one is about a study that was run properly, whose results are more interesting than either the original story or the debunking of it.

Key Takeaway

The 2018 abstract reports that an extra minute waited at age four predicted about one tenth of a standard deviation in achievement at fifteen, that this was "only half the size of those reported in the original studies" and "reduced by two thirds in the presence of controls," and that "most of the variation in adolescent achievement came from being able to wait at least 20 s."[1] On our own arithmetic, the first and last of those can only be reconciled if the relationship is sharply nonlinear.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The effect is real and smaller than advertised. The replication found a positive association at roughly half the original magnitude before controls.

Two. Most of the signal is in the first twenty seconds, which is stated in the abstract and is almost never repeated, and which changes what the test is measuring.

Three. Controls for background and early ability removed two thirds of what remained, which is the finding that generated the subsequent argument.

Four. The replicating authors explicitly reject the simple debunking reading, and we quote them saying so, because that sentence is dropped as reliably as the original's qualifications were.

Five. On our own arithmetic an association of the surviving size predicts very little about any individual, which is the point a business reader should take rather than any verdict about children.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the 2018 findings, from an abstract obtained verbatim from the publisher and confirmed against an open-access repository copy.

Grade A for the existence and shape of the subsequent debate, from a national database record listing the commentaries and responses with their citations.

Grade B for the replicating authors' concession, obtained verbatim from an academic networking site's reproduction of their later commentary.

Grade C for the public framing, from a professional association's news item.

Grade D for the original 1990 study, which reaches us only as a citation and through the replication's description of it.

Grade A for our own arithmetic, with one interpretive step in it that we flag explicitly rather than bury.

A Note On Method

Everything here is verified to August 2026.

We obtained the 2018 abstract verbatim from the publisher[1] and confirmed it against an open-access university repository copy[2].

We obtained a national database record listing the commentaries and responses[3], several of which carry no abstract at all.

We obtained the replicating authors' later concession verbatim from an academic networking site[4].

One source is a professional association's news item[5] and one a bibliographic service[6], both flagged where used.

We did not obtain the 1990 original, the 2018 paper itself, or any of the commentaries, so we report no figures beyond those in the abstract.

All arithmetic is ours. This article discusses research on child development and prediction. It is not parenting, educational or hiring advice.

The Original Study

What is being replicated, as the replicating authors describe it.

Their abstract records: "We replicated and extended Shoda, Mischel, and Peake's (1990) famous marshmallow study, which showed strong bivariate correlations between a child's ability to delay gratification just before entering school and both adolescent achievement and socioemotional behaviors."[1]

The citation: Shoda, Y., Mischel, W., and Peake, P. K. (1990), Predicting adolescent cognitive and self-regulatory competencies from preschool delay of gratification: Identifying diagnostic conditions, Developmental Psychology, 26, 978–986[1]. We did not obtain it.

Four observations, ours.

The description is generous rather than dismissive. "Strong bivariate correlations" is the replicating authors' own characterisation of the work they are testing.

The original title contains the phrase "Identifying diagnostic conditions", which suggests the original authors were themselves concerned with when the measure works rather than claiming it always does.

The outcomes are achievement and socioemotional behaviours in adolescence, measured a decade or more after the waiting task, which is a genuinely impressive study design whatever its results.

Two things that design buys, ours. A decade between measure and outcome removes any possibility that the outcome caused the measure, which is a real advantage over most correlational work.

And it makes the study expensive and rare, which is part of why the original stood unchallenged for so long. Following children for eleven years is not something anyone does casually.

And we cannot tell you the original's sample size, which matters and which we could not obtain, though the replication's framing implies it was considerably smaller.

Two reasons that omission matters more than usual here, ours. A small original sample would explain part of the shrinkage by itself, since estimates from small samples that reach publication tend to be the larger ones.

And the popular retelling of the original rests on vivid individual cases, which is a mode of persuasion that works identically at any sample size and tells a reader nothing about it.

The 2018 Replication

The source.

Watts, T. W., Duncan, G. J., and Quan, H. (2018), Revisiting the Marshmallow Test: A Conceptual Replication Investigating Links Between Early Delay of Gratification and Later Outcomes, Psychological Science, 29(7), 1159–1177, July, DOI 10.1177/0956797618761661[1][3].

Four observations, ours.

The subtitle says "conceptual replication", which is a specific term and an honest one. It means testing the same hypothesis with a different sample and different measures rather than repeating the original procedure exactly.

That has a cost worth stating. A conceptual replication that fails can always be answered by saying the conditions differed, which is the seventy-fifth article's dispute in a new setting, and the commentaries below make exactly that argument.

The venue is the same journal that later published the commentaries and the authors' response, all in one issue[3], which is how this kind of disagreement should be handled.

And the paper is open access, on the database record's own designation[3], so a reader can check everything here directly, which we could not.

What It Found

The results, in the abstract's own words.

It records: "Concentrating on children whose mothers had not completed college, we found that an additional minute waited at age 4 predicted a gain of approximately one tenth of a standard deviation in achievement at age 15."[1]

And: "But this bivariate correlation was only half the size of those reported in the original studies and was reduced by two thirds in the presence of controls for family background, early cognitive ability, and the home environment."[1]

Four observations, ours.

The first sentence reports a real and positive finding. The waiting task predicted something a decade later, which is not nothing and is what a debunking account tends to skip.

The word "But" opening the second sentence is doing the work, and the two reductions are separate. Half, and then two thirds of what was left.

The controls named are the interesting part. Family background, early cognitive ability, and the home environment, which are the things a four-year-old's capacity to wait might plausibly be a symptom of rather than a cause.

And the subsample restriction is stated up front. Children whose mothers had not completed college, which is a deliberate choice we return to below.

Two things worth noting about how that sentence is built, ours. The restriction appears before the finding, in the same sentence, which is about as prominent a disclosure as an abstract allows.

And it is the sort of clause that vanishes in a headline. A restriction stated in a subordinate clause is a qualification, and by now this series has documented what happens to those.

The Sentence Nobody Quotes

The last line of the abstract, and the most interesting thing in this article.

It reads: "Most of the variation in adolescent achievement came from being able to wait at least 20 s."[1]

Four observations, ours.

Twenty seconds. Not fifteen minutes, not the heroic feat of self-denial the story is built around, but roughly the time it takes to settle into a chair.

If that holds, the predictive content of the task is almost entirely in a threshold near the beginning, and everything after it adds very little.

Which changes what the measure plausibly captures. A child who cannot manage twenty seconds may be telling you something quite different from a child who manages three minutes rather than fifteen, and only the first distinction is doing work.

And it is a sentence that survives in no popular account we have seen, including the accounts written to debunk the study. Both the celebratory and the sceptical versions drop it, which is unusual and worth noticing.

Two reasons it is inconvenient to both, ours. It weakens the heroic version, because twenty seconds is not a feat of character.

And it weakens the dismissive version too, because a sharp threshold effect is a real and interesting finding rather than an absence of one.

An Internal Consistency Check

Our own arithmetic, using only the paper's own two figures, and it is the sort of check a reader can run on any abstract.

Take the first finding literally and extend it. If an extra minute predicts 0.1 standard deviations, then waiting:

1 minute predicts 0.1 SD. 2: 0.2. 5: 0.5. 7: 0.7. 10: 1.0. 15 minutes: 1.5 standard deviations.

Four observations.

Fifteen minutes would predict a gain of 1.5 standard deviations, which moves a child from the fiftieth percentile to roughly the ninety-third.

That is an implausibly large effect for a preschool waiting game, and the relationship therefore cannot be linear.

Which is precisely what the abstract's last sentence says. The two findings are consistent only if almost all the signal sits near the start, and our arithmetic reaches the paper's own conclusion from the other direction.

And that is why we ran it. An abstract reporting a per-minute figure invites a linear reading, the linear reading is absurd, and the sentence correcting it is the one that gets dropped.

Two general uses for that habit, ours, since it is the transferable part of this section. Extend any per-unit figure to the top of its range and see whether the answer is sane, which takes one multiplication.

And when it is not sane, the figure is either wrong or nonlinear, and reading further usually tells you which. Here the abstract itself supplies the answer two sentences later.

What The Test Is Measuring

The consequence. Ours.

Four observations.

If the signal is concentrated at twenty seconds, the task is closer to a yes-or-no question than to a measure of endurance.

The popular version tracks the wrong variable. It asks whether the child waited fifteen minutes and the data cares whether they waited twenty seconds, and those are different children.

The candidate explanations for failing twenty seconds are also different from the ones for failing fifteen minutes. Not understanding the instruction, not trusting it, or being hungry are all live at twenty seconds and mostly resolved by three minutes.

And this is our reasoning rather than the paper's. We did not obtain the paper's own interpretation of its threshold finding, and would want it before putting weight on any particular explanation.

What The Controls Did

Our own arithmetic, indexing the original studies at 1.0 and applying the two reductions the abstract reports.

Original studies, indexed: 1.000.

Replication, bivariate, at half the size: 0.500.

Replication with controls, reduced by two thirds: 0.167.

Four observations.

The surviving association is roughly one sixth of the figure the original studies reported, on this chaining.

Neither reduction alone would be remarkable. Replications routinely find smaller effects, and controls routinely absorb some of an association, and this series has covered both patterns repeatedly.

It is the combination that produces the headline, and the second reduction is the substantively interesting one, because of what the controls were.

And the chaining is our own reading of one sentence, which the next section addresses rather than glossing over.

A Note On Our Own Chaining

An interpretive step we are flagging rather than burying.

The abstract reads that the bivariate correlation "was only half the size of those reported in the original studies and was reduced by two thirds in the presence of controls."[1]

We have read "reduced by two thirds" as applying to the replication's own bivariate figure, giving 0.5 times one third, or 0.167.

Three observations, ours.

That is the natural reading of the sentence and we are reasonably confident in it, since the subject of the sentence is the replication's correlation.

But it is a reading rather than a quoted number, and a different baseline would give a different answer. We did not obtain the paper, which reports the actual coefficients and would settle it in a line.

We flag it because the resulting figure is the most quotable thing in this article, and a reader repeating "one sixth" should know it is our arithmetic on a ratio rather than a number the authors published.

What An Association That Size Predicts

Turning it into something a business reader can use. Our own arithmetic, on the paper's per-minute figure and on smaller values reflecting the controlled estimate.

Take two children, one of whom waited five minutes longer than the other. How often does the longer waiter score higher at fifteen?

At 0.05 SD per minute: 57 percent of the time. At 0.10: 64 percent. At 0.20: 76 percent.

Four observations.

A coin flip is fifty percent, so at the paper's unadjusted figure a five-minute difference buys you fourteen points of predictive edge over guessing.

That is a real edge and it is not a destiny. Roughly one child in three with the longer wait will score lower, at the unadjusted estimate, and the controlled estimate is worse.

Which is the eighty-second article's arithmetic in a new setting. A real association at the group level tells you very little about any individual, and the marshmallow story is told entirely about individuals.

Two versions of the story illustrate the gap, ours. "Children who waited did better on average" is what the data supports and is unremarkable enough that nobody would repeat it.

"The child who waited became the successful adult" is what circulates, and it is a claim about a person that a group-level association cannot license, however large the association is.

And the second sentence is the one people remember, which is the whole difficulty. A story about a person is memorable and an average is not, so the version that survives is the one the evidence cannot support.

And the calculation assumes normally distributed outcomes and independence, which are our assumptions and not the paper's, and which we flag as the soft spot here.

One further caveat on this specific calculation, ours. It also assumes the per-minute figure applies across the whole range, which the abstract's own last sentence says it does not.

So the true comparison for two children differing by five minutes is probably weaker than our table shows if both cleared twenty seconds, and possibly stronger if only one did. The threshold finding cuts against a smooth reading in both directions.

Which is a limitation of our own table rather than of the study, and worth saying plainly. We computed a smooth relationship the paper says is not smooth, because the smooth version is what the per-minute figure supports and the threshold figure gives us nothing to compute with.

The Debate That Followed

What happened next, which is the healthiest thing in this article.

A national database record lists, in the same journal and the same issue: Doebel, S., Michaelson, L. E., and Munakata, Y. (2020), Good Things Come to Those Who Wait: Delaying Gratification Likely Does Matter for Later Achievement (A Commentary on Watts, Duncan, and Quan, 2018), Psychological Science, 31(1), 97–99; Falk, A., Kosse, F., and Pinger, P. (2020), Re-Revisiting the Marshmallow Test: A Direct Comparison of Studies by Shoda, Mischel, and Peake (1990) and Watts, Duncan, and Quan (2018), Psychological Science, 31(1); and a response by Watts, T. W., and Duncan, G. J. (2020), Psychological Science, 31(1), 105–108[3][4].

Several of these carry no abstract on the database record[3], and we obtained none of them.

Four observations, ours.

Two critical commentaries and a response, published together, is exactly what a contested finding should produce, and this series has covered several literatures that produced nothing of the kind.

The first commentary's title is an argument in itself. "Delaying Gratification Likely Does Matter for Later Achievement" states its position before the colon.

The second is a direct comparison of the two studies, which is the analysis a reader most wants and which we could not obtain.

And the replicating authors published a further commentary titled "Controlling, Confounding, and Construct Clarity"[4], which names the three things the argument is actually about.

Those three words are worth unpacking, ours. Controlling is whether the covariates were appropriate. Confounding is whether something else explains both the waiting and the achievement. Construct clarity is whether the task measures what anyone thinks it does.

The third is the deepest. If the task measures something other than self-control, both the original finding and the replication are about a different thing than either title suggests, which is where the twenty-second threshold points.

And construct clarity is the question a business reader should carry away, ours. Before asking whether a measure predicts, ask what it is a measure of, because a strong predictor of the wrong construct is worse than a weak predictor of the right one.

What The Replicating Authors Conceded

The passage that should accompany every citation of this replication.

In their response they write: "In general, we agree that our study did not lend itself to making simplistic conclusions about the replicability of the Shoda et al. (1990) study, as we clearly observed positive and substantively important correlations between early gratification delay and later achievement."[4]

Four observations, ours.

"Positive and substantively important" is the replicating authors' own description of their own findings, written after the debunking headlines had run.

They also disown the reading their paper is most often used for. "Did not lend itself to making simplistic conclusions about the replicability" is a direct rejection of the marshmallow-test-debunked framing.

Which puts anyone citing this paper as a refutation in an awkward position, ourselves included if we had. The authors say the correlations were important and that simple conclusions are not available, and both halves are theirs.

And this is the same structure as the eighty-sixth article's subtle-effects sentence and the ninetieth's caveat. Three consecutive articles in which the qualification is in the paper and absent from the reception, which is now less a pattern than a rule.

Two things follow that we would state as working practice, ours. Read the last two sentences of any abstract before the first two, because the qualifications cluster there and the headline is at the top.

And look for a published response by the original or replicating authors, which takes a minute and frequently contains the sentence that would have changed how you read the paper.

The Subsample Question

A design choice that drew comment, reported as such.

The abstract states the analysis concentrated on "children whose mothers had not completed college"[1]. A professional association's news item notes that much of the analysis focused on that subsample[5], and characterises the study as suggesting the link is heavily influenced by social and economic backgrounds[5]. That is a news item, not research, flagged.

Four observations, ours.

Restricting to a subsample is a defensible choice with a cost. It can improve comparability with the original and it narrows what the findings generalise to.

The direction of the effect on the results is not obvious to us, and we will not guess. A more homogeneous sample can either sharpen or attenuate an association depending on what it removes.

The replicating authors' own commentary title includes "Controlling, Confounding, and Construct Clarity"[4], which suggests this was among the criticisms they addressed, and we did not obtain their answer.

And it is the kind of detail that determines what a study means and never reaches a headline. "Marshmallow test fails to replicate" contains no room for a subsample restriction, and the restriction is in the second sentence of the abstract.

One thing we can say about the restriction without guessing, ours. It was disclosed, prominently, in the abstract, which is the standard this series has applied to everyone else and which the paper meets.

What Actually Survives

Our reading, stated directly.

Five statements.

The association is real and positive, on the replication's own finding, at roughly half the original bivariate size.

Controls for background, early ability and home environment removed about two thirds of what remained, on the abstract's own account.

Most of the predictive variation came from clearing twenty seconds, which reframes what the task measures and appears in no popular version.

The replicating authors describe their own correlations as positive and substantively important and reject simplistic conclusions about replicability.

And on our own arithmetic the surviving association predicts an individual's outcome barely better than chance, which is compatible with all of the above and is the part a business should act on.

Those five sit together without contradiction, which is worth saying because the public versions treat them as incompatible. A real effect, much smaller than reported, concentrated at a threshold, largely shared with circumstance, and nearly useless for predicting one person is a coherent picture and is what the abstract describes.

And it is a picture with a clear practical reading, ours. Treat it as a finding about populations and not as a fact about anybody, which is the only use the arithmetic supports and is how almost nobody uses it.

Self-Control Still Matters

Because this could be read as licence to dismiss the whole idea, which the evidence does not support. Ours.

Four observations.

Nothing here shows that self-control is unimportant. It shows that one measure of it, taken at age four, predicts adolescent achievement weakly once other things are accounted for.

A commentary published in the same journal argues that delaying gratification likely does matter for later achievement[3], and we did not obtain its evidence.

And the replication's own authors agree the correlations were substantively important[4], which is about as clear as a signal gets that the dismissive reading is wrong.

Two reasons that signal deserves weight, ours. It runs against their own interest, in the sense that a stronger negative finding would have been a bigger paper.

And it was published after the reception had already formed, so it is a correction of how their work was being used rather than a hedge written in advance.

What has been weakened is a specific and much stronger claim: that a brief early behavioural measure captures a stable trait that determines life outcomes. That claim was always doing more work than the data could carry.

One reason it was attractive enough to carry that weight, ours. It offers a simple, early, cheap intervention, and a claim promising large returns from a small change will always find an audience.

Against which the replication points somewhere expensive. If family background, early cognitive ability and home environment absorb most of the association, then the levers are the ones that are hard to move, which is a less appealing conclusion and possibly the right one.

One consequence for how a business should read this literature generally, ours. Findings that recommend cheap interventions are over-represented in what reaches you, not because anyone is dishonest but because expensive conclusions travel badly.

Conceptual And Direct Replication

A distinction that determines how much this study can settle. Ours.

Four observations.

A direct replication repeats the original procedure as closely as possible, and a failure is hard to explain away. A conceptual replication tests the same hypothesis with different measures and a different sample.

The 2018 paper describes itself as the second kind, in its own subtitle, which is an honest label rather than a concealed weakness.

But it means a defender always has an answer available. Different sample, different measures, different era, and the commentaries appear to have made exactly that argument.

Which is why the direct comparison paper is the one we most wanted. Somebody put the two studies side by side[4], and that analysis would settle what a conceptual replication cannot, and we did not obtain it.

One general point about the two kinds, ours, since the distinction recurs. A conceptual replication that succeeds is stronger evidence than a direct one, because it shows the finding survives a change of measures.

And a conceptual replication that fails is weaker evidence than a direct failure, for the same reason. The asymmetry is real and is why this paper's status is genuinely contested rather than merely disputed.

The Reception Went Wrong In Both Directions

An unusual feature of this case, worth naming. Ours.

Four observations.

The original reception overstated it, turning a correlational finding into a claim that a preschool behaviour determines a life.

The replication's reception overstated in the opposite direction, producing accounts of a famous study failing, which its authors explicitly reject.

And both receptions dropped the same sentence, the one about twenty seconds, which is the finding that would have complicated either story.

That symmetry is the useful observation. A qualification is not dropped because it inconveniences one side; it is dropped because it is a qualification, and both a celebratory and a sceptical account move faster without it.

Which suggests a cheap diagnostic, ours. When two opposed accounts of a study agree about what to omit, the omitted thing is usually the most informative part, because it is the part that does not serve either argument.

The Business Version Of This Claim

Where a reader of this publication meets the same structure. Ours.

Four observations.

The marshmallow story is used to support a family of business claims. That character predicts performance, that a brief signal reveals it, and that you can select on it.

Each step is separately doubtful and the chain multiplies the doubt. A trait may matter, be poorly measured by a short task, and still be swamped by circumstances, which is roughly what the replication reports.

The controls are the part that transfers most directly. Family background, early cognitive ability and home environment absorbed two thirds of the association, meaning much of what looked like character was standing in for circumstance.

And the equivalent controls in a business setting are rarely applied. Territory quality, account inheritance, tenure and manager are the family background of a sales figure, and a firm ranking people on raw output has run the uncontrolled version.

One thing that comparison does not license, ours. We are not claiming sales performance is mostly circumstance, which would need the study nobody has run.

The claim is narrower and still uncomfortable. A raw ranking has not established that it is not, and the replication is a demonstration of how much a ranking can change when circumstance is accounted for.

Two ways to run the controlled version cheaply, ours and untested. Compare people within the same territory and the same manager rather than across the whole team, which removes the two largest confounds at no analytical cost.

And look at the change in a person's numbers rather than the level, since a poor territory depresses the level and not the trend. Neither is a proper analysis and both are better than a raw league table.

Hiring For Character

The specific application, and the one where the arithmetic bites. Ours, and not hiring advice.

Four points.

Selection exercises intended to reveal character are short behavioural measures used to predict long-run outcomes, which is precisely the structure this literature examines.

On our own arithmetic, an association of the size that survived controls would let you rank two candidates correctly slightly more often than a coin, which is worth something and is not worth much.

The eighty-eighth article's problem compounds it. A firm making a handful of hires a year cannot detect whether its selection method works, so the practice never gets evaluated.

And the ninetieth article's requirement applies here too. Testing whether a selection method works needs the people you rejected, which almost nobody has, and which is why hiring practices persist without evidence.

One partial remedy exists and is rarely used, ours. Track the people you rejected who were hired elsewhere, where that is visible, which gives a rough and biased comparison rather than none at all.

One partial remedy exists and is rarely used, ours. Track the people you rejected who were hired elsewhere, where that is visible, which gives a rough and biased comparison rather than none at all.

It is a weak design by the ninetieth article's standard and it is the only one available to a small firm. A biased comparison you know is biased beats a confident number with no comparison behind it.

What The Controls Were Standing In For

The finding we would put most weight on. Ours.

Four observations.

The three controls are family background, early cognitive ability, and the home environment[1], and between them they absorbed two thirds of the association.

That does not mean the waiting task measured nothing. It means much of what it measured was also measured by those three things, which is what a control does.

Which is worth stating carefully because the word invites the wrong picture. Controlling does not remove an influence from the world; it removes it from the comparison, which is the eighty-third article's point and applies identically here.

The interpretation people reach for is that circumstance causes both the waiting and the achievement, and the analysis cannot establish that. Controls identify shared variance and not its direction.

Which is the honest position and it is uncomfortable. We know the association shrinks when circumstance is accounted for and not why, and the commentaries argue about exactly this.

Two positions a reader can hold responsibly here, ours. That the waiting task partly measures circumstance, in which case it is a symptom and intervening on it would do little.

Or that circumstance shapes a real capacity which then matters, in which case the capacity is causal and the controls have absorbed part of the causal path. The data in front of us does not separate them, and saying so is more useful than picking.

A Different Reading Of The Task

A line of work our sources point at, which we can name and not assess.

A bibliographic service's summary of related work records that "manipulations of social trust influence delaying gratification," and that this highlights "intriguing alternative reasons" for individual differences on the task[6].

This is a bibliographic service summarising a paper we did not obtain, flagged.

Four observations, ours.

If trust affects waiting, the task may partly measure whether a child expects an adult's promise to be kept, which is a reasonable thing to have learned from experience.

That reading fits the twenty-second threshold uncomfortably well, though this is our speculation rather than any finding. A child who does not expect the second marshmallow has no reason to wait even briefly.

It would also explain why background controls absorb so much, since whether adults keep promises is a feature of a household, though again we are reasoning rather than reporting.

And we flag all of that firmly. We obtained one clause from a summary of a paper we did not read, and have built a plausible story on it, which is the thing this series criticises when others do it.

We have left it in rather than cutting it because the speculation is testable and points somewhere, and because removing it would leave the article looking more certain than we are. But a reader should treat this section as the weakest thing here.

Bibliographic Note

The series keeps a count.

The 2018 paper's title is given as "Early Delay of Gratification" by the publisher and the database[1][3], and as "Early Gratification Delay" in one reproduction[4].

Three observations, ours.

The words are the same and the order is not, which would defeat an exact-phrase search while leaving a keyword search unaffected.

It is the mildest category we record, and we note it because the same source that carries the article's most valuable quotation also carries its title wrong, which is a useful reminder that a source can be reliable in one respect and not another.

Which is the practical reason we record these at all rather than a scolding one. Verifying a citation and trusting a quotation are separate decisions, and a source can pass one and fail the other.

That brings the running count of bibliographic variants across this series to thirty-seven.

One further note at this stage, ours. Nine of the last ten articles have produced at least one, which is a higher rate than the first fifty, and reflects that we are now checking every citation against multiple sources rather than one.

What To Do

Do not treat the marshmallow study as debunked. The replicating authors describe their own correlations as positive and substantively important and reject simplistic conclusions about replicability.

Do not treat it as destiny either. On our own arithmetic, at the paper's unadjusted figure a five-minute difference in waiting leaves the longer waiter scoring higher 64 percent of the time.

Extend a per-unit figure and see whether it stays sensible. Ours took 0.1 SD per minute to 1.5 SD over fifteen minutes, which is implausible, which is how we found the nonlinearity the abstract states.

Read the last sentence of an abstract. The twenty-second finding is there, changes what the study means, and appears in no popular account we have seen.

Ask what the controls were. Family background, early cognitive ability and home environment absorbed two thirds of the association, which is the substantive finding rather than a technical detail.

Apply the same controls to your own judgements of people. Territory, inherited accounts, tenure and manager are the equivalent, and a raw ranking has controlled for none of them.

Be sceptical of short signals used to predict long outcomes, which is the structure this literature examines and the structure of most selection exercises.

And notice when a qualification is in the paper and absent from the reception. That has now happened in three consecutive articles in this series.

The Limits Of This Analysis

Several caveats matter. This article discusses research on child development and prediction and is not parenting, educational or hiring advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 1990 original study, which reaches us only as a citation and through the replicating authors' description; we cannot report its sample size, which matters to the comparison. We did not obtain the 2018 paper itself, only its abstract, so every figure here comes from that abstract and we report no coefficients, confidence intervals or sample sizes from the study. We did not obtain any of the three commentaries or responses, several of which carry no abstract on the database record; we can therefore report that a substantive debate occurred and almost nothing about its content, which is the largest gap here since the direct comparison of the two studies is the analysis a reader most wants. The chaining that produces our one sixth figure is our own reading of a single sentence, flagged in its own section: we take "reduced by two thirds" to apply to the replication's bivariate figure, which is the natural reading and is not a number the authors published. All arithmetic is ours. The linear extension to fifteen minutes is a deliberate reductio and not a claim about what the study found; the probability calculations assume normally distributed outcomes and independence, which are our assumptions; and the per-minute values below the paper's 0.1 are illustrative of a controlled estimate rather than figures the paper reports. Our reading of the twenty-second threshold, and the social-trust speculation built on one clause from a summary of a paper we did not read, are explicitly our own reasoning and should be weighed as such.

Frequently Asked Questions

Was the marshmallow test debunked?
No, and the replicating authors say so themselves. In a later response they write that their study did not lend itself to simplistic conclusions about replicability, as they clearly observed positive and substantively important correlations between early gratification delay and later achievement.
What did the 2018 replication actually find?
That an additional minute waited at age four predicted about one tenth of a standard deviation in achievement at fifteen; that this was only half the size reported in the original studies; and that it was reduced by two thirds once family background, early cognitive ability and home environment were controlled for.
What is the twenty-second finding?
The abstract's last sentence states that most of the variation in adolescent achievement came from being able to wait at least twenty seconds. If that holds, the task is closer to a threshold than a measure of endurance, and the difference between waiting three minutes and fifteen carries little of the signal.
Why does the linear extension matter?
Because a per-minute figure invites a linear reading. On our own arithmetic, 0.1 standard deviations per minute over a fifteen-minute test would predict 1.5 standard deviations, moving a child from the fiftieth percentile to about the ninety-third. That is implausible, so the relationship cannot be linear, which is what the abstract's last sentence says.
How much does it predict about an individual?
On our own arithmetic, at the paper's unadjusted figure of 0.1 per minute, a child who waited five minutes longer scores higher at fifteen 64 percent of the time. A coin flip is 50. At smaller values reflecting the controlled estimate it falls toward 57 percent.
Does self-control not matter then?
That is not the conclusion. A commentary in the same journal argues that delaying gratification likely does matter for later achievement, and the replicating authors themselves call their correlations substantively important. What is weakened is the stronger claim that a brief early measure captures a trait that determines outcomes.
What is the business lesson?
That short signals used to predict long outcomes carry less than they appear to, and that much of what looks like character is standing in for circumstance. The equivalent controls in a business are territory, inherited accounts, tenure and manager, and a raw ranking of people has applied none of them.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article flags one interpretive step in its own arithmetic in a dedicated section rather than in the limits note.

References

  1. Publisher record for Watts, T. W., Duncan, G. J., & Quan, H. (2018), Revisiting the Marshmallow Test: A Conceptual Replication Investigating Links Between Early Delay of Gratification and Later Outcomes, Psychological Science, 29(7), 1159–1177, July, DOI 10.1177/0956797618761661, reproducing the abstract in full: that the authors replicated and extended Shoda, Mischel and Peake's (1990) famous marshmallow study, which showed strong bivariate correlations between a child's ability to delay gratification just before entering school and both adolescent achievement and socioemotional behaviours; that concentrating on children whose mothers had not completed college, they found an additional minute waited at age 4 predicted a gain of approximately one tenth of a standard deviation in achievement at age 15; that this bivariate correlation was only half the size of those reported in the original studies and was reduced by two thirds in the presence of controls for family background, early cognitive ability, and the home environment; and that most of the variation in adolescent achievement came from being able to wait at least 20 seconds. The record carries the paper's reference list, confirming Shoda, Y., Mischel, W., and Peake, P. K. (1990), Predicting adolescent cognitive and self-regulatory competencies from preschool delay of gratification: Identifying diagnostic conditions, Developmental Psychology, 26, 978–986. Note: the publisher's record. Our source for the abstract, which is the source of every figure in this article. We did not obtain the paper and report no coefficients, intervals or sample sizes from it. journals.sagepub.com
  2. Open-access university repository copy of the same article, reproducing the abstract identically, including the finding that an additional minute waited at age 4 predicted a gain of approximately one tenth of a standard deviation in age-15 achievement, and that the bivariate correlation was half the size of those in the original studies and reduced by two thirds in the presence of controls. Note: an open-access repository record, used to confirm the abstract independently of the publisher. The article is designated open access, so a reader can obtain the full text we could not. escholarship.org
  3. National database record listing the article and the exchange that followed it, confirming Watts, T. W., Duncan, G. J., and Quan, H. (2018), Psychological Science, 29(7), 1159–1177, July, epub 25 May 2018, PMID 29799765, designated a free full-text article; Doebel, S., Michaelson, L. E., and Munakata, Y. (2020), Good Things Come to Those Who Wait: Delaying Gratification Likely Does Matter for Later Achievement (A Commentary on Watts, Duncan, and Quan, 2018), Psychological Science, 31(1), 97–99, PMID 31850827, recorded as having no abstract available; and a further item by Watts, T. W., and Duncan, G. J. (2020), Psychological Science, 31(1), 105–108, PMID 31850825, also recorded as having no abstract available. Note: a national database record. Our source for the existence and citations of the commentaries and response. Several carry no abstract, and we obtained none of them, so we report that the debate occurred and almost nothing about its content. pubmed.ncbi.nlm.nih.gov
  4. Academic networking site record for Falk, A., Kosse, F., and Pinger, P., Re-Revisiting the Marshmallow Test: A Direct Comparison of Studies by Shoda, Mischel, and Peake (1990) and Watts, Duncan, and Quan (2018), reproducing the 2018 abstract and reproducing a passage from the replicating authors' later commentary, titled Controlling, Confounding, and Construct Clarity: A Response to Criticisms of "Revisiting the Marshmallow Test," stating that in general they agree their study did not lend itself to making simplistic conclusions about the replicability of the Shoda and colleagues (1990) study, as they clearly observed positive and substantively important correlations between early gratification delay and later achievement. Note: an academic networking site record. Our source for the replicating authors' own concession, which is the most important quotation in this article. Recorded also as a bibliographic variant: this source renders the 2018 title as "Early Gratification Delay" where the publisher gives "Early Delay of Gratification". researchgate.net
  5. Professional psychological association's news item on the 2018 study, characterising it as suggesting that the widely studied link between children's ability to delay gratification and their life outcomes is heavily influenced by social and economic backgrounds, and noting that much of the analysis focused on a subsample of children whose mothers had not completed college by the time the child was born. Note: an association news item, NOT a research source, flagged at every use. Cited for the public framing of the study and for the subsample point, which is also in the abstract. psychologicalscience.org
  6. Bibliographic service record for the 2018 paper, reproducing a condensed form of its findings and carrying summaries of related work, including one recording that manipulations of social trust influence delaying gratification and that this highlights intriguing alternative reasons to test for individual differences, and another describing a study extending the 2018 analytic approach to examine the long-term predictive validity of delay of gratification in a sample of 702 participants. Note: a bibliographic service record; summaries of papers we did not obtain. The social-trust clause is a single sentence about a study we have not read, and the speculation we build on it in this article is explicitly flagged as our own. semanticscholar.org

This article discusses research on child development and prediction and is not parenting, educational or hiring advice. Neither the 1990 original nor the 2018 paper was obtained, only the latter's abstract, which is the source of every figure here. None of the three commentaries or responses was obtained. All arithmetic is the authors' own, and the chaining producing the one-sixth figure is an explicitly flagged reading of a single sentence rather than a number the authors published.