Boundary erosion is invisible before an accident
Aliases: drift into failure · risk migration · safety-margin drift
What it is
Invisible safety-margin erosion is the gradual movement of a system's operating state toward a failure boundary that is unknown or not explicitly measured, while routine output and accident indicators still read normal. "Invisible" does not mean unobservable in principle — organizations typically measure output and whether an accident happened, not buffer margin, barrier health, or recovery capacity as objects tracked in their own right.
Why it happens
A safety boundary is rarely a single hard limit on one variable, the way a display can show "temperature must not exceed X." It is usually set jointly by several variables and the specific context at the time — an action that is safe under this load, this staffing level, and this equipment condition may not be safe under a different combination. That multivariate, context-dependent character is exactly what makes the boundary hard to compress into one number on a dashboard.
Rasmussen's drift model adds a second layer: each local adjustment — skipping a check that, this time, seems not to matter, trimming a maintenance window a little further — comes with a reason the person involved can fully justify for that specific situation. Taken one at a time, each is a reasonable adaptation to local conditions, not a lapse. The direction of the cumulative drift path only becomes visible in hindsight, once a long chain of such local adaptations is strung together; someone inside the process sees only "this one" adjustment, not the direction the whole series is heading.
More importantly, while it is happening, the system continuing to run normally is itself taken as evidence that "current practice is safe" — which in turn reinforces confidence to make the next adjustment in the same direction, on the reasoning that if this were really dangerous, something would already have gone wrong. This is a self-reinforcing, self-concealing process: nobody needs to hide the drift on purpose, because the drift's continuation depends precisely on this faulty inference — normal operation proves safety — and that inference is itself the mechanism that keeps the drift going, with no additional concealment required. Asymmetric feedback compounds this: crossing the boundary produces a strong, immediate signal, while the margin thinning out produces none at all until it is actually exhausted.
Studying it
Build time-varying margin proxies — standby-capacity utilization, deferred check and maintenance duration, frequency and duration of bypass activation, dwell time of open anomalies, and time to recover from a disturbance — and analyze them jointly with operating context (load, staffing, external pressure), watching whether the proxies move monotonically with contextual pressure.
Each proxy needs to be validated separately against a specific protective function; a variable should not be assumed to represent risk just because it is easy to collect — "easy to count" and "informative" are different properties. Before/after comparisons around an accident must guard against hindsight bias: it is easy, after the fact, to pick out a few indicators that "should have been obvious," but those same indicators may have shown up during long stretches of normal operation with nothing going wrong. Only after confirming that a proxy also discriminates within the no-accident sample can it be said to have provided advance warning.
Where it stops holding
Some boundaries can be defined cleanly by physical models and interlocks and displayed directly (variables like pressure or temperature with a clear physical limit); others can only be estimated with uncertain proxies, and there is always error between a proxy and the true boundary. Adding more instrumentation cannot remove unknown unknowns — some failure modes were never conceived of before being instrumented, so they fall outside any proxy's coverage. Combining multiple proxies into a single composite "safety score" can manufacture false precision, giving the impression that one number represents a multidimensional, context-dependent state.
Applying it
- Show trend lines for buffers and barriers together with uncertainty bands, not a single red/green light — a trend reveals the direction of drift far better than an instantaneous value.
- Set up combination triggers for review when several small degradations appear together (rising bypass use, deferred checks, and overtime occurring at once), and let reviewers drill into the individual components rather than seeing only a composite score.
- Back-test proxies against historical near misses and deliberate stress tests (raising pressure under controlled conditions and watching how the proxies respond): only proxies shown, on real pre-accident data, to have offered a usable window for action should be relied on; proxies that cannot pass this retrospective check should not be put into service.