Trust is formed from personal use, not from statements and documentation
Aliases: Muir experience · badges are not trust · trial cannot be skipped
What it is
An inbox classifier wears a “96% accurate” badge and a white paper. New users still look through dozens of messages by hand before they let it into the junk folder. Trust is formed from personal use; it cannot be built directly from statements and documents. A badge can start a trial. It cannot skip the trial.
Muir puts trust in automation on an update from experience. No experience, no input to the update.
Why it happens
A statement is someone else’s sample: the lab, someone else’s inbox, the previous version. One’s own inbox is a different distribution; the statement cannot replace those dozens of observations. People only weight right and wrong they have seen. So the same white paper is experience for the team that wrote it and advertising for the user. Lee and See’s attitude has to line up with capability; the lining-up happens in use, not in reading.
Documents can even get in the way: people who have read “96%” may shorten the trial to confirm the statement with one success, so the experience sample is narrower and overtrust arrives faster. A statement is not a neutral preface; it changes how experience is sampled.
Studying it
One group reads statements and badges then uses; one uses then gets statements; one only uses. Measure when handover matches the base rate, and whether “read it, so trial less” short sampling appears. Independent variables: timing of the statement, whether the statement carries this user’s slice, whether the trial is shortened. Dependent variables: number of own samples needed to reach calibration, rate of overtrust.
Own-sample count is this entry’s core. If statements could “build it directly,” that count should approach zero. It will not.
Where it stops holding
Experts in their own domain can treat someone else’s test report as an approximation to experience; documents do a bit more, and they still usually want one check of their own. In settings where a trial cannot be forced (a one-shot medical decision), the experience path does not exist; the trust question has to be rewritten as whether this system should be used at all, not how to build trust from documents. After a severe incident, new statements can still less build trust; only a new experience window can.
Applying it
- Design first use as an observable short trial, not as reading a trust centre first. The trial has to be able to show right and wrong.
- If a statement appears, it must carry a limit such as “not yet measured on mail like yours; look at twenty first,” so the statement does not replace sampling.
- Do not use a badge as the switch that turns auto-classifying on. The switch should appear after the user has seen a batch of their own results.
- Check: how many messages a new user who has read the badge actually looks at before turning auto on. If near zero, trust has been treated as a product of the statement; compare later misclassifications to see how fast overshoot arrived.
Related
- Same group: L5.09.1 A reasonable trust level varies with task and situation; there is no globally correct trust · L5.09.2 A streak of successes pushes trust above the system's actual reliability · L5.09.3 Undertrust shows up as repeated manual checking, whose cost is often ignored · L5.09.4 Showing typical failures can suppress overtrust, at the cost of short-term adoption
- Nearby: L5.03 Trust Calibration · L5.04 Collapse of Trust · L5.10 Asymmetric Effect of First Failures on Trust
- Search terms:
trust from experience·Muir·statements versus use
Cards in the same group
- L5.09.1A reasonable trust level varies with task and situation; there is no globally correct trust
- L5.09.2A streak of successes pushes trust above the system's actual reliability
- L5.09.3Undertrust shows up as repeated manual checking, whose cost is often ignored
- L5.09.4Showing typical failures can suppress overtrust, at the cost of short-term adoption