A9.07.3Retrospective workload rating and the peak-end effectresearchdesign

Retrospective ratings taken after a task are dominated by the peak and the ending, diverging from the true average

Aliases: retrospective workload rating · peak-end effect · post-hoc overall rating

What it is

A retrospective rating asks the user for a single overall load score after the entire task is over, rather than sampling load in real time as it unfolds. Such ratings are dominated by the peak-end effect: what gets reported mostly reflects the moment of highest load during the task and the feeling right at the end, not the average level across the whole experience.

Why it happens

When people recall an experience, they don't integrate and average it moment by moment — they pull a few standout episodes from memory and reconstruct an overall impression from those. Load memory follows the same pattern: the intensity at the peak moment and the feeling near the end get overweighted in memory, while a longer but steadier middle stretch gets flattened or forgotten. This means two tasks with a similar average load across their full duration can produce wildly different retrospective ratings, simply because the peak lands in a different place or the ending is handled differently.

Studying it

A common way to test for this bias is to collect a real-time rating sequence during the task alongside a single retrospective rating afterward, then compare how strongly the retrospective score correlates with the sequence's mean, peak, and final value respectively — it typically correlates much more strongly with the peak and the final value than with the true average. The methodological point: if the goal is to find out exactly where in the process things got most demanding, collecting only a post-hoc overall rating loses that information — real-time sampling during the task is required to localize it.

Where it stops holding

When a task is short and doesn't fluctuate much, the peak-end effect has limited impact and the retrospective rating stays close to the true average. When a task is longer and has a clear load trajectory — easy-then-hard, or the reverse — relying solely on the post-hoc overall rating to judge "was this generally demanding" will be systematically dominated by the peak or the ending, and won't represent the typical or average operating state.

Applying it

  • When assessing the overall load of a longer flow — a multi-step form or an onboarding sequence, say — don't just ask "how did that feel overall" once at the end; insert brief real-time ratings at key checkpoints.
  • Deliberately check whether the last step of the flow is designed in a way that spikes load right at the end (requiring a bulk re-confirmation of information, for instance), since a poor ending will pull down the user's retrospective evaluation of the whole flow out of proportion to how much time it actually took.
  • Verification: compare the real-time rating curve against the post-hoc overall rating; when the two disagree sharply, prioritize the real-time curve to locate the stage that actually needs fixing rather than being misled by the overall score.

Related

  • Same group: A9.07.1 Multidimensional scales split load into mental, physical, temporal and other components, scored separately then weighted · A9.07.2 A single overall scale is simpler to administer but cannot reveal where load comes from · A9.07.4 Subjective scales lose discriminating power at very low or very high load, showing ceiling and floor effects
  • Adjacent: P1.03 Peak-End Rule
  • Search terms: peak-end effect · retrospective workload rating · in-task probe

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/A9.07.3