L5.03.1trust calibrationdesignresearch

The goal is for trust to match actual reliability

Aliases: calibrated trust · matching reliability · trust as attitude

What it is

A navigator states arrival “to the minute,” and people keep appointments to the minute; its real error is plus or minus eight minutes, and trust and reliability have come apart. The goal of calibration is for trust to sit against actual reliability — not to push trust as high as possible, and not to frighten people out of using the thing.

Lee and See treat trust in automation as an attitude that has to line up with capability. The lining-up is the goal, not a by-product.

Why it happens

Trust is a person’s internal estimate of “I can still hand this over next time.” Reliability is the system’s true frequency on that task, under those conditions. Both quantities can be written as probabilities; their sources differ. One comes from experience, interface promises, and vivid events; the other from repeated trials. If the interface feeds the first with fluency, badges, and successes only, and never names the scope of the second, the estimate drifts.

Calibration is not a feeling. The feeling can be excellent — people are at ease, or people are careful — while the estimate is still off. The objective is matching error, not satisfaction. Muir stresses that trust updates with experience; if the update only eats successes, the estimate walks one way.

Studying it

Measure the system’s reliability on the target task independently, then measure people’s estimate of “the next item will also be right,” and actual handover. Independent variables: precision the interface promises (to the minute / an error band), mix of successes and failures shown. Dependent variables: gap between estimate and base rate, whether handover sits in a range commensurate with the base rate.

Reliability must be sliced on the task in front of the user, not borrowed from another benchmark. Consecutive items in the lab will self-calibrate; sparse use once a day in a product drifts more easily, and should be reported separately.

Where it stops holding

On creative tasks with no measurable right or wrong, “reliability” as a quantity does not exist; the calibration goal has to be rewritten as consistency with the promise (draft, spark), not a hit rate. Experts in their own domain bring a base rate, and a quieter interface need not wreck the match. This entry only defines the goal. What kind of failure over and under each are, and whether failures have to be shown to calibrate, are later steps.

Applying it

  • Promised precision must not exceed error you have measured. Give an error band if you can; do not give a fake point estimate.
  • Before handover, show recent performance on that task slice, not a product-wide average.
  • Precision wording in marketing and empty states must use the same numbers as inside the product.
  • Check: ask “how many minutes’ slack will you leave for the next trip” or “will you hand over the next item.” Put the answer next to your real error or hit rate. A wide gap means the calibration goal was missed, whether or not people like the interface.

Related

  • Same group: L5.03.2 Overtrust and undertrust are both failures · L5.03.3 Calibration requires exposing failures, not hiding them
  • Nearby: L5.09 Overtrust and Trust Collapse · L5.04 Collapse of Trust · L5.10 Asymmetric Effect of First Failures on Trust
  • Search terms: trust calibration · trust and reliability · Lee and See

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/L5.03.1