Overall evaluation is dominated by the peak and the end
Aliases: peak-end rule · duration neglect · retrospective evaluation · remembered utility
What it is
People's overall evaluation of an experience is not the average of its moment-by-moment feelings. Asked afterwards "how was it," the judgment is carried by two points: the moment of strongest feeling (the peak) and the feeling at the end — the peak-end rule. The first thing to keep straight is that two quantities are being distinguished: the total experienced in the moment (experienced utility) and the evaluation stored in memory (remembered utility) are not two measurements of one thing but two different things. "Longer suffering that ends mildly is remembered as better than shorter suffering" is not a paradox inside this framework — what gets evaluated is the memory, not the process.
Why it happens
The mechanism is compressive storage plus retrospective reconstruction. Experience arrives as a stream, but what is kept is not a continuous recording — it is snapshots of a few emotionally significant moments plus the final state; asked to rate the whole episode later, people reconstruct a judgment from those snapshots rather than replay it and average. Two consequences follow directly from this storage format. First, the peak is over-weighted: the most intense moment is precisely the one that received the most attention, the deepest encoding, and the easiest later retrieval, so it enters the answer at far more than its share of time. Second, duration neglect: duration must be encoded separately and preserved until retrieval, and snapshot storage discards it — two episodes with the same peak intensity score almost identically afterwards, no matter how different their lengths. The end is over-weighted for two reasons at once: it is the most recent and most vivid segment at retrieval, and it defines how the episode closed, setting the tone for the whole memory.
Studying it
- Colonoscopy paradigm: patients report their current pain at fixed intervals during the procedure, then give one retrospective rating of the whole experience; the retrospective rating tracks a combination of peak and end pain, while duration adds almost no explanatory power. Follow-up work found that the retrospective pain evaluation predicted willingness to return for the next procedure better than the real-time pain did — which of the two quantities predicts behavior is exactly the methodological point of this line of research.
- Cold-pressor paradigm: the same cold-water task in two versions — a short one that stops at the fixed duration, and a long one that appends, after the identical cold water, a slightly warmer and less painful tail. Asked which episode they would rather repeat, participants systematically choose the longer version; duration neglect and the better-end effect are demonstrated in one experiment.
- The interface and service version: sample experience during a task at fixed intervals (or with screen-recording replay to assist recall), collect a separate global rating afterwards, and regress the retrospective score on mean, peak, end, and duration terms to see how much variance each explains — the standard way to test the rule on real flows.
- Methodological cautions: real-time and retrospective evaluations are different constructs and must be measured and modeled separately; the timing of a survey alone decides which one you get — sampling during use yields a real-time curve, recall right after completion yields a retrospective judgment. Use retrospective measures to predict return visits, repurchase, and recommendation; use real-time state to predict mid-task abandonment. The two are not interchangeable.
Where it stops holding
- The "overall evaluation" in these experiments is a single score on a post-hoc questionnaire — the product of reconstruction. The evidence comes mostly from closed, single episodes of seconds to tens of minutes (medical procedures, cold water, film clips) in which participants had no goals or history beyond the episode — conditions far from a real product.
- Long-lived products slice experience into many sub-episodes, each with its own peak and end, while the user's attitude toward the product aggregates all of those remembered evaluations. Which episode's peak and end dominates the whole — the most recent, the worst, or the most often recalled — is itself an open research question; the single-episode result does not transfer directly to predicting long-term satisfaction.
- Duration neglect is bounded: when duration is explicitly priced or itself carries meaning (billed-by-the-minute waits, an announced delay), it re-enters the evaluation.
- Peak and end are not the whole input either: information received afterwards, the way the story gets told, and other people's evaluations can rewrite the memory's representation, and the retrospective evaluation keeps drifting after the fact.
Related
- Same group: P1.03.2 The ending has the highest return on investment · P1.03.3 Negative peaks are amplified in memory too
- Nearby: P1.01.3 The reflective layer is decided by meaning, memory, and self-image · P1.08.3 The reflective layer can rewrite the memory of the other two layers · P1.09.3 An emotional peak needs a concrete, attributable object
- Search terms:
peak-end rule·duration neglect·retrospective evaluation·experienced utility·remembered utility