L1.08.4uniform high confidence is no signaldesignresearch

Showing uniformly high confidence on a whole batch provides no discrimination

Aliases: collapsed confidence · all-high scores · no discrimination

What it is

Ten suggestions, each wearing 96%. The number cannot be used to decide which to check first, which to trust, which to drop. The score spends attention and yields no ranking. Uniform high confidence is no signal: showing it presupposes an actionable difference among items on the same screen. Without a difference, confidence is decoration — decoration that lifts trust in the whole batch.

Discrimination is relative. An absolute “all high” says nothing in a decision.

Why it happens

RLHF and instruction tuning shove mass toward high scores; the UI then displays the raw scores, and a batch pins to the ceiling. The work people do with scores is allocate scarce checking time. With no gradient on the ceiling, allocation falls back to reading order or position bias, and the column is wasted.

Worse, the level itself still speaks. Ten 96%s will not be read as “this column has no information.” They will be read as “this batch is solid.” Undiscriminating highs are an endorsement of the batch, more likely than showing nothing to switch checking off. A zero-entropy channel is still emitting a constant; here the constant is not silence, it is a steady “high.”

Studying it

Two packs: scores with a gradient (crossing an action threshold) versus all pinned in the high band. Measure: whether checking time follows the scores, trust in the batch, where planted errors are missed. Independent variables: scores shown or not, whether the highs are truly high after calibration or just ceiling-crushed. Dependent variables: use of discrimination, batch-level handover.

If people allocate checking by score when there is a gradient, and check the whole batch less when there is none, the harm of no discrimination is not “useless.” It is “used as a blurb.”

Where it stops holding

After honest calibration a whole batch is high, and the task allows light checking (low stakes, reversible), a constant high is truthful. At high stakes, even a true high should become “this batch did not yield a first-to-check,” not ten green lights. With a single output, discrimination is undefined and this entry does not apply; go back to whether an instance self-score is checkable and actionable. This entry does not discuss how to cut bands.

Applying it

  • Before showing, look at spread on this screen. With almost none, do not show instance scores; write “this batch did not pick out which line is surer.”
  • When checking time must be allocated, use an order or “look at these first,” not a row of equally full bars.
  • Calibrate or truncate a ceiling crush: cap how many items may sit in high, to force a relative difference; if none appears, turn the column off.
  • Check: screenshot ten full bars and ask “which do you check first.” If the answer is “whatever” or “top down,” the column is empty. Then ask “is this batch as a whole to be trusted” — “yes, very” means the empty column is still lifting trust.

Related

  • Same group: L1.08.1 Confidence is the model’s self-report; the user has no independent way to verify it · L1.08.2 Percents are read as frequency promises, and most model numbers are uncalibrated · L1.08.3 Bands are less over-read than continuous numbers, at the cost of hiding within-band differences · L1.08.5 Confidence is worth showing only when the user can change the next action because of it
  • Nearby: L1.04 Presenting confidence · L3.07 Multi-option generation and side-by-side comparison · L5.05 The moderation principle of transparency
  • Search terms: uniform high confidence · no discrimination · collapsed scores

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/L1.08.4