A10.09.1Human error rate quantification — scope and limitsresearch

Assigning a single error-probability number to a task only works under narrow, checkable conditions

Aliases: human error probability · HRA · THERP · performance shaping factors

What it is

Human Reliability Analysis (HRA) tries to attach a specific number to "how likely a person is to err at a given task" — a Human Error Probability (HEP), such as "reading this class of gauge, the error rate is about one in a hundred." Numbers like this were first developed in high-risk industries such as nuclear power and aviation, feeding into probabilistic risk assessment to decide whether a given operating step needs extra safeguarding. But the number isn't a universal constant: it's measured or estimated under a specific task structure, a specific stress level, a specific population, and specific environmental conditions. Lifted out of those conditions and applied to a different interface or a different user population, the number stops meaning anything.

Why it happens

The rate is so condition-dependent because a person's chance of erring is itself the joint product of task structure, time pressure, fatigue, interface clarity, and training level — not some intrinsic property that belongs to "the human" as a standalone factor. The same person might err once in a thousand tries on a well-structured task with prompt feedback, and an order of magnitude more often on a chaotic, time-pressured one. Early HRA methods such as THERP tried to model this dependency explicitly: break the task into elementary action units, look up or assign each one a baseline error probability, then adjust that baseline with a set of "performance shaping factors" — time pressure, ambient noise, procedure quality. That structure is itself evidence that a single number, detached from task structure and shaping factors, means nothing; reporting "the human error rate is X percent" without specifying the task structure and shaping factors it was measured under makes the number untrustworthy on its own.

Studying it

The typical HRA workflow combines task decomposition with expert judgment or historical data: break a complex operation into a sequence of elementary steps, look up or have domain experts estimate an error probability for each one, then multiply by performance shaping factors that reflect the specific situation at hand, arriving at an overall error-rate estimate. Methodological caveat: the great majority of these numbers come from historical incidents and lab data in specialized operating environments like nuclear power and aviation, where task structure differs enormously from a consumer software interface — professional operators go through standardized training on highly proceduralized steps, while ordinary users are doing exploratory, non-proceduralized operations. Applying an error-rate table from those industries directly to interface usability evaluation is, at bottom, an analogy across contexts rather than a measurement within the same population, and the resulting estimate's confidence interval is typically understated.

Where it stops holding

This kind of quantification holds up best under two conditions: the task is highly structured, its steps are enumerable, and a large body of historical data on the same class of operation already exists — a nuclear control room's standard operating procedure, say. Once a task is open-ended, lets users explore their own path, or the user population's experience level spans a wide range, describing "this interface's error rate" with a single number papers over individual and path differences. The number looks precise while its real confidence interval may span more than an order of magnitude. Treating such a number as a single acceptance gate — "the error rate must be under one percent to ship" — without stating the measurement conditions places more trust in the number's precision than it can bear.

Related

  • Same group: A10.09.2 treating "human error" as a conclusion masks systemic causes · A10.09.4 systemic causes should be examined before individual attribution · A10.09.5 skilled and novice users have different error profiles, so safeguards can't be one-size-fits-all
  • Nearby: A9.02 measuring workload · Y7.06 human factors verification and validation
  • Search terms: human error probability · HRA · THERP · performance shaping factors

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/A10.09.1