Y7.01.1Layered defence failuredesignresearch

Accidents usually result from simultaneous failure of multiple defences

Aliases: defence in depth · Swiss cheese model · barrier failure

What it is

Layered defence failure is the standard explanation, in industrial and safety-critical systems, for why a major accident happens: the hazard is not released by one link acting alone, it passes through a sequence of prevention, detection, control, and mitigation barriers that should be independent, and it happens to find every one of them degraded or defeated at the same moment. The organizational accident model behind this (often shorthanded as the Swiss cheese model) already covers, in the human-factors literature, why stacked defences work the way they do and how a common-cause failure cancels that independence. What matters here is a different question: how a plant, a power station, or a hospital turns that model into an executable, day-to-day safety-management practice, rather than leaving it as a metaphor about slices of cheese.

Why it happens

On a real industrial site, a "layer" is not an abstraction — it maps onto a specific, heterogeneous set of mechanisms: operator training and qualification, written procedures, constraints built into the human-machine interface (an illegal operating path simply cannot be taken through the interface), physical interlocks and hardware voting logic, and independent supervisory checks and audits. These typically sit in different departments, different budget lines, different review cycles — which is exactly what is supposed to make them independent. But a safety management system (SMS) that manages each layer separately by department — training owns training, maintenance owns interlocks, quality owns audits — misses a real risk: a single upstream decision reaching into several layers at once. The common case is a budget cut whose reduction gets distributed across training hours, spare-parts inventory, maintenance frequency, and audit staffing. On paper these look like four independent adjustments made by four different departments; in fact they share one triggering cause and weaken together in the same window. That common-cause link never shows up in any single layer's routine records, because each department only sees "our layer is a bit tighter this year" — nobody assembles the full picture.

Studying it

Verifying that this framework is actually operating in an industrial setting is not a matter of re-arguing whether the Swiss cheese model holds; it means running an independence review: for a given hazard scenario, list every defence layer involved, tag each one's dependencies — personnel, funding source, supplier, information system — and check those dependency lists against each other for overlap. A common method pulls several years of budget adjustments, staffing changes, and supplier contract amendments and cross-checks each one against the defence list to see whether a single change touched more than one layer. Another method reviews near misses rather than only completed accidents: in a near miss, typically only one layer actually stopped the hazard, so it is worth checking whether that one surviving layer had already been weakened by the same pressure to the point of being close to failing, just not yet breached. Methodologically, a conclusion that "these two layers share a failure cause" tends to look obvious in hindsight; the review has to be run proactively against ledgers and contract records before an event, not reconstructed by an investigation team after the fact.

Where it stops holding

This practice fails in two situations. First, if the SMS treats the review as a one-time project gate — say, a dependency map drawn once before a new plant is commissioned — rather than something rerun continuously as budgets and staffing change, any common-cause decision made after that point is never captured, and the paper independence map drifts away from reality over time. Second, for scenarios that depend heavily on real-time coordination among on-site personnel, where the boundary between layers reorganizes dynamically during an abnormal event (an experienced crew informally covering for each other), a pre-drawn static dependency map cannot fully represent the defence structure that was actually in play at the time; the independence review only covers common-cause risk under routine configuration and cannot substitute for a separate analysis of dynamic adaptation.

Applying it

  • Merge the defence-layer ledgers that different departments keep separately into a single cross-departmental dependency table, tagging each layer's budget line, staffing, and external suppliers, so the independence review has real data to work from instead of being reconstructed after an event.
  • Add a step to the approval workflow for any decision touching budget, staffing, or outsourcing contracts that checks whether it affects more than one defence layer at once — owned by the SMS, not by the department proposing the decision.
  • Cross-check routine safety inspection, training, maintenance, and audit records on a regular schedule, looking specifically for multiple layers tightening or loosening in the same window, rather than checking each layer against its own standard in isolation.
  • How to check: pull one to two years of budget, staffing, and outsourcing-contract change records and cross-reference each entry against the defence list for a key hazard scenario; count how many changes actually touched more than one layer. If nearly every change is tagged against only one layer, the ledger is too coarse-grained and the review is not doing real work.

Related

  • Same group: Y7.01.2 Organizational decisions create latent conditions · Y7.01.3 Blaming individuals stops improvement
  • Nearby: A10.07 Swiss cheese model · Y4.02 Redundancy and voting mechanisms
  • Search terms: defense in depth · common-cause failure · safety management system

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Y7.01.1