Confidence is worth showing only when the user can change the next action because of it
Aliases: actionable confidence · display threshold · confidence that flips an act
What it is
Every number on a screen spends attention. If a confidence cannot make the next step a different thing — check or not, send or not, pick A or pick B — it is noise wearing a measurement costume. Show it only if it changes the next act is the admission rule: write down the act it will flip; if you cannot, do not draw it.
This is display policy, not another encoding.
Why it happens
Transparency’s default pressure is “if we have it, give it.” Internals happen to hold a float, so it appears beside every message. The person’s decision is already fixed by something else (it must go out, there is only one candidate, they will edit it anyway). The score has no fulcrum and is still read as a guarantee or a threat. Information that cannot be acted on is not neutral: it spends the checking budget and manufactures “I already decided with a scientific number.”
The fulcrum has to be an act the UI allows. “When low, be careful” with no Careful act (no review queue, no hold, no handoff) is an empty fulcrum. A score on an empty fulcrum has only an emotional effect.
Studying it
List the next steps that actually exist in the product (send, save draft, assign, switch candidate). Then A/B with and without the score, and watch whether those acts’ distribution moves. If it does not, the score has no right to show in that scene. Independent variables: whether the act is clickable, whether the score crosses that act’s threshold. Dependent variables: act flips, gaze time, a later “what did this number help.”
Self-report of “it helped” with an unchanged act distribution is a typical false positive. The primary endpoint is the act.
Where it stops holding
Research, debug, and audit surfaces have “look at the distribution” as the next step, so a score has a fulcrum. Consumer one-shot generation with an empty or fully forced act set has none. The same score can qualify in preview and not at send-confirm, or the reverse: preview may only affect editing, confirm may affect outbound. Qualify per surface. This entry assumes the earlier misreading mechanisms hold. It is the gate on “so do we still show it.”
Applying it
- Write a table: each surface, each score, which clickable act it maps to, and the threshold. An empty row means that surface does not show a score.
- Low certainty must connect to a real act: hold, assign, a checklist that must be opened. If it will not connect, do not show low — it will only frighten.
- The act a high connects to must not be “skip every check,” unless that surface’s consequences allow it.
- Check: hide the score and watch the act distribution. If it matches the shown condition, take the score off that surface. If it differs, check that the difference is the act you wrote down — if people only press the same button more anxiously, it still does not qualify.
Related
- Same group: L1.08.1 Confidence is the model’s self-report; the user has no independent way to verify it · L1.08.2 Percents are read as frequency promises, and most model numbers are uncalibrated · L1.08.3 Bands are less over-read than continuous numbers, at the cost of hiding within-band differences · L1.08.4 Showing uniformly high confidence on a whole batch provides no discrimination
- Nearby: L1.03 Visualizing uncertainty · L5.05 The moderation principle of transparency · L1.04 Presenting confidence
- Search terms:
actionable confidence·display threshold·uncertainty that changes acts
Cards in the same group
- L1.08.1Confidence is the model’s self-report; the user has no independent way to verify it
- L1.08.2Percents are read as frequency promises, and most model numbers are uncalibrated
- L1.08.3Bands are less over-read than continuous numbers, at the cost of hiding within-band differences
- L1.08.4Showing uniformly high confidence on a whole batch provides no discrimination