The higher the system’s accuracy, the worse actual human supervision — opposite the reason the gate exists
Aliases: reliability paradox at the gate · accuracy corrodes watch · sparse-negative dulling
What it is
A person sits in the loop because the system will err. The less it errs, the fewer errors the person sees, and both skill and motive to detect them fall. When the rare error arrives, the person on the gate is someone who can no longer catch it. Higher accuracy degrades supervision is the paradox of the intervention point: the closer to “almost no need for a person,” the less that person can work when needed.
Too-frequent gates stamp because of quantity. This is quality: negatives too sparse, detection dulls even when quantity is modest.
Why it happens
Signal detection ties misses and false alarms to the base rate. As the base rate nears 1 (almost always right), the optimal policy leans to release-all. People are not slacking; they are adapting to the true base rate. Bainbridge’s ironies land here on the gate: design leaves a person to handle the residual, and the residual is too thin to maintain the ability to handle it.
Products also sell accuracy. The louder the sell, the more the organisation cuts people, dwell, and training, and supervisory resource is withdrawn with the accuracy. Management amplifies the paradox.
Studying it
The same gate, system accuracy manipulated (say 70 / 90 / 99 percent), equally hard errors planted late. Dependent variables: detection, dwell, confidence. Independent variables: whether known practice negatives are interleaved, whether the base rate is told. Separate “tired” from “adapted to base rate” — still dull on a short shift at low frequency, and it is base rate.
A lab that densifies errors will hide the paradox. Field density is what generalises.
Where it stops holding
Accuracy high but errors clustered (one class still fails often) can keep vigilance on that class; the paradox is local. Accuracy high and errors random dulls most. If the person is not tasked to catch errors (only to sign), supervision quality does not exist to be corroded, whatever accuracy does. This entry does not treat whether liability was made explicit.
Applying it
- Do not use “already 99%” to cut people, time, or training on the gate. The residual is why the gate exists.
- Keep visible, harmless negatives or replays on purpose, to maintain detection calibration, like a safety drill, not only a post-accident review.
- Split supervision quality (catch of planted errors) and system accuracy as KPIs. Accuracy up and catch down is buying a blinder gate.
- Check: plot monthly accuracy against catch of planted errors. If the first rises and the second falls, the paradox is running. Stop using accuracy to argue “we can look less.”
Related
- Same group: L1.09.1 Where to put the intervention is decided by reversibility · L1.09.2 Intervenors need the context that reconstructs current state · L1.09.3 Too-frequent intervention turns review into a rubber stamp · L1.09.5 Putting a person in the loop also puts liability on them; that transfer must be said, not assumed
- Nearby: L4.02 Automation bias · L4.03 Automation complacency · L4.09 Skill degradation
- Search terms:
reliability paradox·base-rate at the gate·supervision quality
Cards in the same group
- L1.09.1Where to put the intervention is decided by reversibility; irreversible acts must have a gate before them
- L1.09.2Intervenors need the context that reconstructs current state, or they can only release blindly
- L1.09.3Too-frequent intervention turns review into a rubber stamp
- L1.09.5Putting a person in the loop also puts liability on them; that transfer must be said, not assumed