Y7.04.3Independent safety performance indicatorsdesignresearch

Safety indicators must be independent of production indicators

Aliases: safety KPI · leading safety indicator · process safety indicator

What it is

Independent safety performance indicators measure defense-barrier capability, risk exposure, or recovery readiness directly, rather than inferring safety from production-side results such as output, on-time rate, or injury-free days. "Independent" here refers to what the indicator measures and who governs it not being absorbed into production goals — it does not require safety data and operational data to be kept completely apart; the two can still be analyzed together.

Why it happens

Production results are frequent, positive, and immediate — output and on-time rate can be seen every day — while serious safety loss is rare and delayed. Treating "no accident" as proof of safety, in a system where accidents are already rare, keeps the metric sitting at the same value for long stretches: it neither rises when things get genuinely safer nor falls early when margin is quietly thinning, because it is insensitive to state change. It also rewards under-reporting and plain luck, since a missed report or a near escape both leave the metric looking just as good.

This is why lagging indicators (accident rate, injury rate) carry so little information in safety-critical systems: the denominator is the rare event of "something bad happened," which trends toward zero over the long run and shows no early warning. The information lies in leading indicators — how often the safety boundary is breached, the backlog of suppressed alarms, the accumulation of deferred maintenance, the quality with which critical tests are completed — because these measure capability states that exist before an accident and can still be changed, rather than the accident itself.

But choosing and weighting leading indicators has no standard answer, which is the flip side of their advantage over lagging indicators: a lagging indicator is objective (an accident is an accident), while a leading indicator requires expert judgment about which capability actually corresponds to a specific hazard. When safety indicators sit in the same appraisal system as production indicators, and the same decision-maker has discretion over how resources are split between them, that decision-maker has a built-in incentive, whenever resources compete, to let the definition and weighting of the safety indicator drift toward whatever does not hurt the production score — this is the same pressure that erodes safety boundaries, showing up at the level of indicator governance instead. Pick the wrong leading indicator and the organization learns to game the indicator itself — polishing the report, lowering the test standard until it always passes — rather than actually maintaining the underlying capability margin. This is exactly the phenomenon Goodhart's law describes.

Studying it

Test whether a candidate leading indicator actually explains or predicts degradation in barrier capability, rather than picking variables that merely sound relevant. After adopting a new indicator, keep watching for gaming (the reported number improves while the underlying capability does not), rising under-reporting, or resources being diverted to wherever numbers are easy to move rather than to the actual risk. Compare portfolios of indicators rather than searching for one universal number — in low-accident-rate settings, statistical correlation alone is underpowered, so validity needs to be backed by process evidence (checking on the ground whether the capability genuinely exists) and domain-expert judgment.

Where it stops holding

A leading indicator is not an accident predictor; it measures a capability state, not the probability that something will go wrong next. Too many indicators dilute attention, so that in practice none of them is really tracked. Easy-to-collect measures like report counts or training-completion rates only mean something when tied to content quality and actual capability — counting occurrences by itself carries no information. Governing safety indicators independently of production appraisal does not mean the safety function can set unworkable standards detached from production reality — standards detached from reality get learned around just as easily.

Applying it

  • For each major hazard, select a small number — not as many as possible — of barrier-capability indicators that map directly to it, and document the data source, the owner, and how the indicator could be gamed.
  • Present safety and production indicators side by side to the same decision-makers, but enforce a hard rule: production performance cannot offset a safety threshold that has not been met.
  • Periodically review whether each indicator, after being specifically optimized against, still represents the risk it was originally meant to measure, and retain the raw component data so gaming can be traced after the fact.

Related

  • Same group: Y7.04.1 Efficiency pressure erodes safety boundaries · Y7.04.2 Boundary erosion is invisible before an accident
  • Nearby: Y7.06 Human-factors verification and validation · Y2.08 Alarm-system performance metrics
  • Search terms: safety performance indicator · leading indicator · Goodhart's law

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Y7.04.3