Whether to stop safely or keep running degraded depends on weighing shutdown cost against exposure
Aliases: choosing fail-safe versus fail-operational · functional safety
What it is
Choosing between fail-safe and fail-operational weighs the cost of entering a safe state immediately against the exposure and available protection during continued operation — it is not a slogan applied by equipment category ("medical devices should be fail-operational," "industrial valves should be fail-safe"), because that shortcut ignores how much the right answer depends on the specific scenario. The actual question is which path produces lower overall risk after a fault.
Why it happens
Risk is roughly the probability that a fault goes undetected or unaddressed, multiplied by how long the system stays in that unaddressed state — the lower the fault coverage and the longer the exposure, the higher the risk. That is why fault probability alone is not enough; coverage and exposure duration together are what matter. Shutdown removes some energy input immediately, at the cost of potentially losing braking, cooling, treatment, or escape capability — and once that capability is gone, the exposure window stretches from "during the fault" to "from shutdown until capability is restored." Continued operation preserves function at the cost of prolonged exposure under degraded protection; if fault coverage in the degraded mode is markedly lower than in normal mode, risk keeps accumulating through that extended time. Mission phase shifts the balance precisely because it determines both how long a capability vacuum from shutdown would last and how much recovery resource remains available during continued operation — the same fault during climb-out and during level cruise produces very different exposure durations and available responses, so the right choice differs too.
Where it stops holding
Operating cost cannot be simply traded against life safety, and regulation or certification requirements often fix which path must be chosen in certain scenarios, leaving no room to weigh alternatives. Where the uncertainty in fault consequence is large, an average probability must not be allowed to hide a low-probability, high-consequence catastrophic scenario — an option that looks better in expectation can still be the wrong choice if its tail risk is severe enough.
Applying it
For every critical function, map the consequences of immediate shutdown, controlled degradation, and continued operation separately by mission phase, listing fault coverage, expected exposure duration, and exit conditions for each path.
- How to check: rerun this three-path review under the worst credible combination of faults plus a scenario where recovery resources happen to be unavailable (a needed spare part under repair, support staff already committed elsewhere), and verify whether the original weighting still holds under the least favourable conditions — not just accept a conclusion validated once against a typical scenario.
Related
Cards in the same group
- Y4.01.1Fail-safe means a fault drives the system into a safe state chosen in advance, not wherever it lands
- Y4.01.2Fail-operational means the system keeps working at reduced capability instead of stopping outright
- Y4.01.4Each subsystem's own sound failure choice can still add up to an incoherent whole without coordination