A client's view of a year of work is not the average of that year. It is assembled from a small number of moments, and there is a substantial literature on which ones. There is also a recent finding that complicates the obvious conclusion.
Key Takeaway
In a randomised trial, 682 patients undergoing colonoscopy were assigned so that half had a short interval added to the end of their procedure, testing whether a memory failure observed in psychology experiments could be applied in a clinical setting to lessen patients' memories of the pain[1]. A 2022 meta-analysis of 174 independent samples reported strong support for the peak-end rule and the duration neglect phenomenon[2][3]. One source also reports that the average across the whole experience predicted overall evaluations just as well[3], which we could not verify and which sits awkwardly with the paper's own framing.
Our Grades For These Claims
Applying the scheme from the first article in this series.
That the peak and the end dominate retrospective evaluation is Grade A. A meta-analysis of 174 independent samples reporting strong support, a randomised controlled trial in a clinical setting, and a recent large experience-sampling study all point the same way.
That duration is neglected is Grade A, from the same meta-analysis.
That the peak-end rule outperforms a simple average is Grade D on our sourcing, and the section on that is the most important in this article.
That any of this is usable to design a client experience is Grade C, and it is our own extension.
A Note On Method
Everything here is verified to August 2026.
We obtained the published abstract of the 2003 randomised trial from PubMed, though it truncates before the results[1].
We obtained portions of the 2022 meta-analysis from the publisher, including its stated overall conclusion[2].
The qualification about the average performing equally well comes from an encyclopedia entry[3]. We did not verify it, and we flag its status at every mention because it is the single most consequential claim in this article.
The 1993 cold water study reaches us through a peer-reviewed paper describing it[4]; specific figures for it come from a commercial website and are flagged.
We noticed one secondary source conflating two different studies, and record that below.
This article reviews behavioural research. It is not marketing, service design or medical advice.
The Cold Water Study
Where it starts.
A peer-reviewed paper describes the original: participants underwent two painful treatments, being immersion of a hand in 14°C water for 60 seconds, and the same immersion for 60 seconds with an added 30 seconds in water raised to 15°C. Asked which they would rather repeat, most subjects chose to repeat the treatment including the additional time of milder pain, essentially preferring more pain over less[4].
The same source states these results are inconsistent with a purely hedonistic view of behaviour, which would entail minimizing negative events, and that the authors proposed the finding was due to participants giving more weight to the worst and final moments of painful experiences, being peak-end rather than a sum of negative moments[4].
The paper is Kahneman, Fredrickson, Schreiber and Redelmeier (1993), When more pain is preferred to less: Adding a better end, Psychological Science, 4(6), 401–405[4]. We did not obtain it.
Two observations, ours.
The second trial contained strictly more suffering. Everything in the first trial happened, and then more happened. There is no reading on which the longer trial was objectively better.
And a commercial source gives the sample as 32 subjects with 69 percent choosing the longer variant[5]. We could not verify either figure against the paper or against the peer-reviewed description, and we report them as that commercial source's, not as established.
The Randomised Trial
The strongest study, and the reason this is not a laboratory curiosity.
Redelmeier, Katz and Kahneman published Memories of colonoscopy: A randomized trial in Pain, 104(1–2), 187–194, in July 2003[1].
The abstract states the purpose plainly: "We tested whether a memory failure observed in psychology experiments could be applied in a clinical setting to lessen patients' memories of the pain of an unpleasant medical procedure."[1]
The design: consecutive outpatients undergoing colonoscopy who were medically stable, mentally competent, and able to speak English, n = 682. By random assignment, half the patients had a short interval added to the end of their procedure during which the tip of the colonoscope remained in the rectum. Pain was measured on a ten point intensity scale during the procedure, and memory afterwards using both a rating scale and a ranking task. Randomization resulted in two similar groups.[1]
Our copy of the abstract truncates before the results, and we therefore report the trial's own findings only as secondary sources describe them, below.
Three observations, ours.
This is a randomised controlled trial on real patients in a real clinical setting, which puts it in a different evidentiary category from almost everything else in this series.
The manipulation added discomfort. The extra interval was not pleasant; it was less unpleasant than what preceded it.
And note the framing: the authors describe what they are exploiting as a memory failure. That word is theirs, and it matters for the ethics section later.
What They Were Deliberately Doing
The reported results, from secondary sources, with our sourcing stated.
A design reference describes the trial as having randomly divided patients into one group that underwent a typical colonoscopy, and another that underwent the same procedure in addition to having the tip of the scope left in for three extra minutes without inflation or suction, with the result that patients who underwent the longer procedure experienced the final moments as less painful, rated their overall experience as less unpleasant, and ranked the procedure as less aversive in comparison to the other participants[6].
A commercial source adds specific figures, giving the final-moments ratings as 1.7 versus 2.5 on a ten-point scale and stating the trial had a median follow-up of 5.3 years[7]. We could not verify these against the paper and report them as that source's.
Two observations, ours.
A five-year follow-up, if accurate, would mean the trial tracked something considerably more consequential than a rating: whether patients returned for the screening they needed. That is a real behavioural outcome and it is the strongest possible version of this finding. Our source truncates before stating the result, so we do not report one.
And we should flag a sourcing failure we found. One commercial source attributes the sample of 682 patients to the 1996 study[5], when the PubMed abstract shows that figure belongs to the 2003 trial[1]. Two different studies have been merged. This is the seventh time in this series that a famous finding's numbers have been misreported by careful-looking secondary sources.
The Meta-Analysis
The aggregate test.
Alaybek, Dalal, Fyffe, Aitken, Zhou, Qu, Roman and Baines published All's well that ends (and peaks) well? A meta-analysis of the peak-end rule and duration neglect in Organizational Behavior and Human Decision Processes, 170, 104149, in 2022[2].
Its stated conclusion: "Overall, the results provided strong support for the peak-end rule and the duration neglect phenomenon. In addition, the results indicated that the peak-end rule was a useful heuristic in making summary evaluations, as" and our source truncates there[2].
The paper also examined conceptual and methodological moderators and compared the peak-end effect to the effects of other predictors that have been examined in the literature on memory biases[2].
Two observations, ours.
The acknowledgements thank researchers who provided additional information beyond what was publicly available[2]. Soliciting unpublished data is the standard defence against publication bias, and it is a good sign.
And the phrase compared the peak-end effect to the effects of other predictors is the important one. A meta-analysis that only asked whether the effect exists would be less useful than one that asks whether it beats the alternatives. The next section is about what that comparison apparently found.
The Qualification
The claim that changes how this should be used, reported with a clear statement of its status.
An encyclopedia entry describing the same meta-analysis states that it drew on 174 independent samples and found strong support for the peak-end rule across a wide range of contexts, but that the analysis also found that the average of how an experience felt across its entire duration predicted people's overall evaluations just as well as the peak-end score alone[3].
Four things about how we are treating this, all ours.
The source is an encyclopedia entry, not the paper. We flag that at every mention and we grade the claim D accordingly.
It sits awkwardly with the paper's own abstract, which says the results indicated the peak-end rule was a useful heuristic[2]. Those two statements are not flatly contradictory, since a heuristic can be useful without being uniquely predictive, but the tension is real and we did not resolve it.
The 174 sample figure also comes from this entry rather than from the publisher record we obtained.
And a reader whose decision depends on this should obtain the paper. We are reporting it because omitting a claim this consequential would be worse than reporting it with its provenance attached.
Why That Would Matter
The consequence, if the qualification is accurate. This section is our own reasoning.
The commercial appeal of the peak-end rule rests entirely on a comparison. It says: do not bother improving the average, engineer the peak and the ending instead. That is advice about where to spend money.
Three consequences.
If the simple average predicts evaluations just as well, then the rule has been confirmed as a description and simultaneously fails to beat the obvious alternative as a guide to action.
An effect can be real, robust and useless for a decision, and this series has now hit that structure twice. The thirteenth article distinguished a statistically significant effect from a meaningful one. The sixteenth found experts beating chance but losing to simple extrapolation. Here the question is the same: not whether the model works, but whether it beats the naive baseline.
And it inverts the practical advice. If the average does as well, then making the whole engagement better is not the inefficient option the peak-end framing suggests it is.
When The Two Models Disagree
Reconciling those two facts, with our own arithmetic. The profiles below are invented to demonstrate a mechanism, use one simple operationalisation of the peak-end score among several possible, and come from no study.
How can two different models predict equally well? Because on most real experiences they barely disagree.
Take an experience rated at eight points in time. A steady, unremarkable profile gives a peak-end score of 4.5 and an average of 4.3. A gradual improvement gives 4.5 and 4.5. The models agree.
Now the lopsided cases. One bad spike with an acceptable ending gives 6.0 against an average of 3.8. Fine throughout with a bad ending gives 9.0 against 3.8. Bad throughout with a good ending gives 5.0 against 7.3.
Two observations.
Across a large pool of mostly ordinary experiences, the two scores correlate highly and neither wins. That is entirely consistent with both the strong support and the qualification.
Which means the practical question is not which model is correct. It is whether your experience is shaped like the first two rows or the last three, and a professional engagement is usually shaped like the last three: long stretches of nothing, one difficult episode, and a delivery moment.
A Correction Exists
Something anyone citing this meta-analysis should know.
A corrigendum to the 2022 meta-analysis was published in the same journal, at Organizational Behavior and Human Decision Processes, 104278[8].
We could not access its content and do not know what was corrected.
Two observations, ours.
We record it because a reader relying on any figure from the 2022 paper should check the corrigendum first, and because we cannot tell them whether it affects anything in this article.
And a published correction is the system working rather than failing, in the same way the twentieth article in this series described an adversarial collaboration as the process functioning properly.
Duration Neglect, And A Way To Remove It
The companion phenomenon, and a boundary condition worth knowing.
Duration neglect is the second half of the finding: the length of an experience contributes little to how it is remembered. The meta-analysis reports strong support for it[2].
The same encyclopedia entry reports a boundary: presenting duration information visually rather than numerically was sufficient to eliminate the effect in controlled laboratory settings, suggesting the bias is partly a product of how information is formatted rather than a fixed feature of memory, and that simple changes to how patients or consumers record experiences may produce evaluations that are more reflective of actual lived experience[3].
We did not obtain the studies behind this and report only the entry's description.
Two observations, ours.
If duration neglect can be eliminated by a presentation change, it is not a fixed property of memory, which weakens any argument that treats it as an unavoidable feature of how clients think.
And it points somewhere useful and honest: showing a client what actually happened over time is a way of getting an evaluation closer to the work rather than further from it.
Recent Support
The most recent evidence we located.
A 2025 paper in a peer-reviewed personality journal reports that across four experience sampling samples, overall N = 1,889, total measurements = 131,575, retrospective well-being judgments were disproportionately influenced by the peak and end experiences from the assessment period[9].
We obtained only that statement and none of the paper's detail.
Two observations, ours.
The scale is substantial: over 130,000 measurements, in everyday life rather than a laboratory or a clinic.
And it concerns wellbeing over an extended period, not a discrete unpleasant episode, which extends the phenomenon well beyond its origins in pain research.
Where The Effect Has Been Found
The range, from a peer-reviewed source.
A paper lists the peak-end rule as having been observed in colonoscopies, lithotripsy, arthritis, headaches and childbirth, and in non-painful negative experiences including annoying sounds, poor air quality and a cognitively demanding task[4].
Two observations, ours.
That is a wide range of aversive experiences, which is a point in favour of generality.
But notice what is absent: almost everything on that list is unpleasant. The literature originates in pain and the extension to positive or mixed experiences is a separate question. One of our sources is a paper on evaluations of pleasurable experiences[10], which we did not obtain, and whose existence indicates the question was still being asked.
The Ethical Question
A question the medical literature has taken seriously and business writing generally has not.
A 2025 paper in a peer-reviewed clinical journal is titled Prolonging Medical Procedures to Exploit the Peak-End Rule: An Ethical Analysis[11]. We did not obtain its argument and report only that the question has been formally posed.
Three observations, ours.
The 2003 trial's own abstract describes what it is using as a memory failure[1]. Deliberately exploiting a failure in someone's memory is a different act from making their experience better, and the two are easy to confuse.
In the clinical case there is a defensible justification: if a better-remembered procedure means patients return for screening they need, the memory manipulation serves the patient. That justification is empirical, and it depends on the follow-up result we could not obtain.
And in a commercial setting the equivalent justification is much weaker. Engineering a client's memory of an engagement so they rate it above what it was does not serve them; it serves you. We think that distinction is worth holding onto, and it is the reason the practical sections below are framed around the ending being accurate rather than merely pleasant.
The Shape Of An Engagement
The application. This is our own extension to a setting none of the cited research examined.
A professional engagement has the lopsided profile described above. Long stretches where nothing is felt, one or two episodes of real difficulty, and a delivery.
Three consequences.
The quiet competent months contribute almost nothing to the remembered evaluation, because there is nothing to remember. The work that goes right is invisible, which is the ninth time this series has landed on that structure.
The peak is usually a problem, not a triumph. In advisory work the most intense moment is typically a bad one: the assessment, the query, the shortfall.
And there is generally a clear end, which is unusual and useful: filing, a completion, a signed set of accounts. Many services do not have one.
Which End Is The End
A point that determines whether any of this is actionable. Ours.
The research measures experiences with a defined end that the participant knows is the end.
In an ongoing client relationship, the end is not the completion of the work. It is whatever happened most recently, which is frequently an invoice.
Two consequences.
If the last contact before a long silence is a bill, then on this literature the bill is the ending, and it is being remembered as the final note of the whole engagement.
And that suggests something cheap and honest: do not let the invoice be the last thing. A short substantive contact after it, explaining what was done and what happens next, changes which moment occupies that position without manufacturing anything.
The Peak You Do Not Choose
The harder half. Ours.
You cannot choose whether a client has a bad moment. A reassessment arrives, a number is worse than expected, a deadline is missed by a third party.
Three observations.
What is controllable is the intensity of the worst moment, and intensity is partly a function of surprise. A client who was told in March that this was possible experiences a smaller peak in September than one who was not.
Which reframes proactive communication. It is not only a courtesy. On this literature it is the only available lever on the variable that dominates the memory.
And it is honest, unlike the alternative. Reducing a peak by warning someone in advance improves both the experience and the memory of it, which is precisely the distinction the ethics section drew.
What To Do
Expect the quiet months to count for nothing. There is nothing to remember about work that went smoothly, so a year of competence contributes little to how the year is recalled.
Identify what the ending actually is. For an ongoing relationship it is usually the most recent contact, and if that is an invoice, the invoice is the ending.
Do not let the bill be the last word. A short substantive contact afterwards costs nothing and changes which moment sits in that position.
Reduce the peak by removing surprise. Intensity is partly a function of not having seen it coming, and warning a client in March is the only honest lever on the variable that dominates.
Do not conclude the average is a waste of effort. One source reports the simple average predicting evaluations as well as the peak-end score, which if right undercuts the whole case for neglecting it.
Notice whether your engagement is lopsided. The two models only give different advice where the profile is uneven, and a professional engagement usually is.
Show clients what actually happened. Presenting duration visually is reported to eliminate duration neglect, which moves an evaluation closer to the work rather than further from it.
Keep the distinction between improving an experience and engineering a memory. The trial's own authors call what they exploited a memory failure, and a clinical justification for that does not transfer to a commercial one.
Check the corrigendum before citing the meta-analysis. A correction was published and we do not know what it changed.
The Limits Of This Analysis
Several caveats matter. This article reviews behavioural research and is not marketing, service design or medical advice. Everything is verified to August 2026. We obtained the 2003 trial's abstract but it truncates before the results, so its findings are reported only as secondary sources describe them; specific figures come from a commercial website and are unverified. We did not obtain the 1993 study, and its sample size and choice percentage come from a commercial source we could not verify. The claim that the average predicts as well as the peak-end score comes from an encyclopedia entry, is unverified, sits in unresolved tension with the meta-analysis's own abstract, and is graded D; the 174 sample figure comes from the same entry. A corrigendum to the meta-analysis exists and we could not access its content. We did not obtain the studies on visual duration presentation, the 2025 experience sampling paper beyond one sentence, the paper on pleasurable experiences, or the ethical analysis. We found and report one secondary source conflating the 1996 and 2003 studies. All arithmetic on peak-end versus average is ours, uses invented profiles and one operationalisation among several, and comes from no study. The literature originates in aversive experiences and its extension to positive or mixed ones is a separate question. Every application to client engagements is our own extension to a setting none of this research examined.
Frequently Asked Questions
What is the peak-end rule?
Did they really make a procedure worse on purpose?
So should I stop worrying about the middle of an engagement?
How can both be true?
What is the cheapest thing to change?
Is engineering a client's memory legitimate?
References
- Redelmeier, D. A., Katz, J., & Kahneman, D. (2003). Memories of colonoscopy: a randomized trial. Pain, 104(1–2), 187–194. DOI 10.1016/s0304-3959(03)00003-4, published abstract via PubMed, on patients' memories of the past potentially influencing their decisions about the future while being imperfect and susceptible to bias; on the authors testing whether a memory failure observed in psychology experiments could be applied in a clinical setting to lessen patients' memories of the pain of an unpleasant medical procedure; on the study covering consecutive outpatients undergoing colonoscopy who were medically stable, mentally competent and able to speak English, n = 682; on half the patients by random assignment having a short interval added to the end of their procedure during which the tip of the colonoscope remained in the rectum; on pain being measured with a ten point intensity scale and memory measured using both a rating scale and a ranking task; and on randomization resulting in two similar groups. Note: our copy of the abstract truncates before the results, which are therefore reported in this article only as secondary sources describe them. pubmed.ncbi.nlm.nih.gov
- Alaybek, B., Dalal, R. S., Fyffe, S., Aitken, J. A., Zhou, Y., Qu, X., Roman, A., & Baines, J. I. (2022). All's well that ends (and peaks) well? A meta-analysis of the peak-end rule and duration neglect. Organizational Behavior and Human Decision Processes, 170, 104149. DOI 10.1016/j.obhdp.2022.104149, publisher record, on the authors conducting a meta-analysis on the peak-end rule and duration neglect, examining conceptual and methodological moderators, and comparing the peak-end effect to the effects of other predictors examined in the literature on memory biases; on the results overall providing strong support for the peak-end rule and the duration neglect phenomenon; on the results indicating that the peak-end rule was a useful heuristic in making summary evaluations, our source truncating; on the authors thanking researchers who provided additional information beyond what was publicly available; and on main effect analyses having been conducted separately within subsets measuring experiential variables using unipolar versus bipolar scales. Note: the publisher's record; we obtained portions and the conclusion sentence is truncated in our source. sciencedirect.com
- Encyclopedia entry on duration neglect, on the meta-analysis by Alaybek and colleagues (2022) drawing on 174 independent samples and finding strong support for the peak-end rule across a wide range of contexts; on the analysis also finding that the average of how an experience felt across its entire duration predicted people's overall evaluations just as well as the peak-end score alone; on presenting duration information visually rather than numerically having been sufficient to eliminate the effect in controlled laboratory settings, suggesting the bias is partly a product of how information is formatted rather than a fixed feature of memory; and on the implication that simple changes to how patients or consumers record experiences may produce evaluations more reflective of actual lived experience. Note: an encyclopedia entry, not peer-reviewed. This is the sole source for the claim that the average predicts as well as the peak-end score, and for the 174 sample figure. We did not verify either, the first sits in unresolved tension with the meta-analysis's own abstract at reference 2, and we grade that claim D. en.wikipedia.org
- Peer-reviewed paper describing the original study, on Kahneman and colleagues having subjected participants to two painful treatments, being immersion of a hand in 14°C water for 60 seconds and the same immersion with an added 30 seconds in water raised to 15°C; on most subjects choosing to repeat the treatment including the additional time of milder pain, essentially preferring more pain over less; on these results being inconsistent with a purely hedonistic view of behaviour; on the authors proposing that the finding was due to participants giving more weight to the worst and final moments rather than a sum of negative moments; and on the peak-end rule having since been observed in colonoscopies, lithotripsy, arthritis, headaches and childbirth, and in non-painful negative experiences including annoying sounds, poor air quality and a cognitively demanding task. Note: a peer-reviewed paper's description; we did not obtain the 1993 study itself. ncbi.nlm.nih.gov
- Commercial customer-experience reference site, giving the 1993 study's sample as thirty-two subjects and stating that sixty-nine percent chose the longer variant; and separately attributing a sample of 682 patients to the 1996 study. Note: a commercial website, not peer-reviewed. Neither figure for the 1993 study could be verified. The attribution of 682 patients to the 1996 study conflicts with reference 1, which shows that figure belongs to the 2003 trial; we report this as a sourcing error we found rather than as a fact. ux-strategy.ch
- Design reference site, on the 1996 study finding that colonoscopy or lithotripsy patients consistently evaluated the discomfort of their experience based on the intensity of pain at the worst and final moments regardless of length or variation in intensity; and on the later study randomly dividing patients into one group undergoing a typical colonoscopy and another undergoing the same procedure with the tip of the scope left in for three extra minutes without inflation or suction, with patients in the longer procedure experiencing the final moments as less painful, rating their overall experience as less unpleasant, and ranking the procedure as less aversive. Note: a design reference website, not peer-reviewed; used because our copy of the trial's abstract truncates before its results. lawsofux.com
- Commercial blog, on the 2003 randomized controlled trial with 682 patients undergoing colonoscopy, on the extended-procedure group rating the final moments as less painful at 1.7 versus 2.5 on a ten-point scale, rating the entire experience as less unpleasant, ranking it as less aversive compared to seven other unpleasant experiences, and on the median follow-up being 5.3 years. Note: a commercial blog, not peer-reviewed. None of these figures could be verified against the paper, and the source truncates before stating the follow-up result. thelaunchpadincubator.com
- Publisher record for the corrigendum to Alaybek and colleagues (2022), Organizational Behavior and Human Decision Processes, 104278. Note: we could not access the corrigendum's content and do not know what was corrected; it is recorded here so that anyone citing the meta-analysis checks it. sciencedirect.com
- Scharbert, J., Utesch, K., Reiter, T., ter Horst, J., van Zalk, M., Back, M. D., & Rau, R. (2025). If you were happy and you know it, clap your hands! Testing the peak-end rule for retrospective judgments of well-being in everyday life, publisher record, on the authors finding across four experience sampling samples, overall N = 1,889 and total measurements = 131,575, that retrospective well-being judgments were disproportionately influenced by the peak and end experiences from the assessment period. Note: a publisher record; we obtained this statement only and none of the paper's detail. journals.sagepub.com
- Publisher record for a paper in Psychonomic Bulletin & Review on evaluations of pleasurable experiences and the peak-end rule, noting that prior research suggests the addition of mild pain to an aversive event may lead people to prefer and directly choose more pain over less; and its reference list confirming Fredrickson, B. L., & Kahneman, D. (1993), Duration neglect in retrospective evaluations of affective episodes, Journal of Personality and Social Psychology, 65, 44–55, and Schreiber, C. A., & Kahneman, D. (2000), Determinants of the remembered utility of aversive sounds, Journal of Experimental Psychology, 129, 27–42. Note: a publisher record; we did not obtain this paper or the works in its reference list. One other source gives the Fredrickson and Kahneman page range as 45–55. link.springer.com
- Peer-reviewed clinical journal article, Prolonging Medical Procedures to Exploit the Peak-End Rule: An Ethical Analysis, Journal of Evaluation in Clinical Practice, and its reference list confirming the citations for Kahneman and colleagues (1993), Redelmeier and Kahneman (1996), Redelmeier, Katz and Kahneman (2003), and Alaybek and colleagues (2022). Note: we did not obtain this paper's argument and report only that the ethical question has been formally posed in the clinical literature. onlinelibrary.wiley.com
This article reviews behavioural research and is not marketing, service design or medical advice. No paper discussed was obtained in full; the 2003 trial's abstract truncates before its results and those are reported only from secondary sources. The claim that a simple average predicts evaluations as well as the peak-end score comes from an encyclopedia entry, is unverified, and sits in unresolved tension with the meta-analysis's own abstract. A corrigendum to that meta-analysis exists whose content could not be accessed. All arithmetic is the authors' own and uses invented profiles. Every application to client engagements is an extension to a setting this research did not examine.