L5.09.4typical failures lower overtrustdesignresearch

Showing typical failures can suppress overtrust, at the cost of short-term adoption

Aliases: adoption dip · failure samples · overtrust intervention

What it is

Onboarding includes an answer that looks sure of itself and is a fake citation. New users hand over more carefully the first time; short-term one-click adopt drops a notch. Lighting up typical failures can press overtrust down, at the cost of a short-term drop in adoption.

This is a calibration operation, not an incident. What is chosen is the kind of miss that really happens, not a crash invented to frighten.

Why it happens

Overtrust is missing downward samples. Put a typical failure on the same layer as success, and the estimate is pulled back toward the base rate. The behaviour of the pull-back is fewer “use all” clicks and more checking left in — short-term adoption therefore falls. The fall is part of the goal, not a side-effect incident. If the product takes adoption as the only KPI, it will read this calibration as a feature failure, pull the failure sample, and overtrust returns.

This is not “calibration requires exposing failures.” That is the principle: the estimate has to eat negative samples. This is a concrete intervention: use typical cases to press trust that is already high or about to be, and admit that adoption will dip first. On Lee and See’s picture, this is an intentional push down.

Studying it

Onboarding randomly inserts or does not insert a typical miss for this task (a fluent fake citation, a wrong function side-effect). Then measure first-handover scope, checking, and calibration a week later. Independent variables: whether the miss is typical, whether it sits on the same layer as success, timing of insertion (onboarding / after the first success). Dependent variables: short-term adoption, gap between estimate and base rate, whether still calibrated a week later.

A week is required. Looking only at how much adoption dipped that day will judge a successful calibration as a failure.

Where it stops holding

If the failure sample is unrepresentative (an unrelated crash, or an extreme incident), what is pressed down is willingness to use, not overtrust, and it may be pressed into under. On high-stakes features, onboarding failure samples need a checking act, or they only produce avoidance. Trust cannot be built from documents — failure cases in a white paper almost never move overtrust; they have to be met on the path of use. A real severe incident is not “a typical failure shown.”

Applying it

  • Before onboarding or the first widening of permission, put one miss that really occurs on this task on the same layer as one success, and tell people “check this class.”
  • Mark the next few days’ drop in adoption as expected, not as a rollback signal. Watch calibration error, not that day’s adopt control.
  • Rotate samples so they are not taken as a one-shot performance. Expired failures (a class already fixed) become false alarms.
  • Check: after insertion, whether the estimate sits nearer the base rate, whether one-click adopt falls as expected, and whether handover scope a week later still matches the base rate. Adoption not down and estimate still high means the sample was not typical enough, or not visible enough.

Related

  • Same group: L5.09.1 A reasonable trust level varies with task and situation; there is no globally correct trust · L5.09.2 A streak of successes pushes trust above the system's actual reliability · L5.09.3 Undertrust shows up as repeated manual checking, whose cost is often ignored · L5.09.5 Trust is formed from personal use, not from statements and documentation
  • Nearby: L5.03 Trust Calibration · L5.04 Collapse of Trust · L1.01 Mismatch Between Stochastic Output and Deterministic UI
  • Search terms: typical failure · overtrust intervention · adoption dip

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/L5.09.4