Engagement metrics absorb novelty more readily than task success does
Aliases: novelty-sensitive metrics · engagement versus success · exploratory clicks
What it is
Clicks, session length, page views, and feature-open counts tally exploration, so they move with novelty. Whether a task finishes to a pre-set standard measures an outcome, not how much wandering occurred, and usually does not spike just because someone wanted to inspect a new entry. After a redesign, rising engagement with flat or falling success looks like exploration, not a better experience. Making engagement the primary metric hands novelty a scale that both alarms and congratulates too easily.
Why it happens
Exploration creates extra low-intent events: opening a new tab, ducking in and out, clicking through a coach. Those events enter engagement’s numerator and need not enter success’s. Success’s denominator is people who attempted a task; its numerator is people who reached a defined result. An extra lap of browsing is neither a success nor automatically a failure. Time-based engagement can even score being lost as “more engaged.” After novelty fades, engagement often drops harder than success because the events being withdrawn were the aimless ones. Learning cost, by contrast, usually hits success and time: people are still attempting the task, just failing or going slowly, while engagement can look fine.
Studying it
Split metrics into engagement versus task outcomes in advance and park the decision on the latter. After release, report both families of curves: if only engagement rises then falls while success is flat, attribute the rise to novelty. Sensitivity-check engagement by collapsing it to task attempts and see whether the lift remains. Read session-level success, errors, and help-seeking alongside click volume; do not explain experience with total clicks. In lab or remote tests, give the same tasks and compare exploratory clicks with completion, as a reference for field engagement spikes.
Where it stops holding
For feeds and content products the “task” is browsing, so engagement sits closer to value—without turning aimless refresh into quality. Success near a ceiling is insensitive to novelty and also insensitive to real gains, so finer quality or cost measures are needed. Some novelty shows up as avoidance, and engagement falls first; task outcomes then decide whether people are afraid to click or the thing is actually harder. Ads and recs can move engagement and conversion together; calibrate against downstream funnel results rather than declaring engagement always untrustworthy.
Applying it
- Let critical-task completion, errors, and time carry redesign success; keep engagement as a companion observation.
- When a weekly report shows activity or duration up and task success unchanged, mark the lift as pending decay and keep it out of the outcome.
- Count “opened but did not finish a task” on new entries as a dedicated exploration monitor.
- Check: after novelty prompts come down, does engagement fall back while success stays. If only the former falls, do not credit the engagement lift to the design.
Related
- Same group: Q3.14.1 Early metric movement after a change can be transitory · Q3.14.2 Learning costs temporarily suppress metrics for existing users · Q3.14.3 Observation windows must be long enough · Q3.14.4 Novelty typically rises then falls; learning typically falls then rises · Q3.14.5 Segmented comparison is required to tell them apart · Q3.14.7 Major redesigns void historical baselines
- Adjacent: Q3.06 Task success rate · Q6.08 Experience metrics and business metrics
- Search terms:
engagement metrics·task success·novelty-sensitive outcome
Cards in the same group
- Q3.14.1Metric movement right after a change may not last
- Q3.14.2Relearning costs for existing users pull metrics down for a while
- Q3.14.3The observation window has to outlast the transient
- Q3.14.4Novelty usually rises then falls; learning usually falls then rises
- Q3.14.5Telling novelty from learning takes a split by people; the blended curve confuses them
- Q3.14.7After a major redesign the old baseline is gone; a new comparison period has to be built