Y6.01.3Failure and degraded-mode trainingdesignresearch

Training scenarios must include failure and degradation

Aliases: degraded-mode training · failure-injection training · fallback training

What it is

Failure and degraded-mode training addresses how people detect lost capability, recover control, and operate under restrictions after automation, sensing, communications, or procedures fail. The central question is not merely which fault occurred, but which information remains trustworthy and which actions remain permitted. A less obvious design question sits alongside this: what ratio of normal-operation training to abnormal/degraded-operation training a curriculum should carry. If the bulk of training time goes to normal procedure and a single "failure drill" is dropped in as decoration, the operating strategy trainees build is fundamentally tuned to stable feedback, and there is correspondingly little experience to draw on when a real failure occurs.

Why it happens

Normal training builds strategies under stable feedback. In degradation, the same display may go stale, authority can shift, and backup channels carry different delays. Without experience of these boundaries, operators carry normal-mode assumptions forward, or switch blindly between fallback options. Progressive and combined failures further expose hidden dependence on a single information source.

A deeper mechanism concerns whether a scenario announces in advance that a fault will be injected. If trainees know this session will include a failure, they enter it in a troubleshooting mindset — doubting anomalous readings and cross-checking sooner than they would during ordinary work — which artificially inflates detection speed and response accuracy. A real degradation, by contrast, typically arrives while the operator is completely unguarded. This is the same class of bias as the heightened alertness produced by scripted rare events, except the variable here is whether the fault is pre-announced, not how rare the condition is. Whether fault injection is scattered unannounced through routine cycles or scheduled as a standalone assessment determines whether training builds genuine detection ability or a rehearsed response to a known test.

Studying it

Vary failure salience, propagation speed, channel combination, and recovery availability. Measure detection of lost capability, source selection, takeover, boundary violations, and coordination. Include latent failures with no explicit alarm, and let a safe stop count as a valid success rather than treating production recovery as the only acceptable outcome.

One key independent variable is whether the fault is announced: a control group is told in advance that a fault will be injected, an experimental group has the same scenario folded unannounced into routine sessions. Comparing detection latency and error-handling accuracy between the two quantifies how much ecological validity is lost to advance notice — a more informative measure than reporting detection rate alone.

Where it stops holding

  • Fault injection must be isolated from production systems, and training state must not leak into live operation.
  • More extreme combinations are not automatically better: implausible scenarios with no diagnostic basis train guessing, not diagnosis.
  • The share of degraded-mode scenarios has an upper bound too — if too much of a training cycle is spent in degraded conditions, trainees can start over-doubting normal scenarios, treating ordinary readings as precursors to failure. This is negative transfer running the other direction, detectable by monitoring the false-alarm rate on normal scenarios, not by guessing at the right ratio.
  • Training cannot substitute for the underlying design work of fault isolation, interlocks, and an intelligible degraded-mode display.

Applying it

  • Build a scenario matrix running from single-channel loss through partial capability loss to common-cause failure.
  • Set a baseline ratio of normal to degraded scenarios, tuned to the system's actual failure rate and consequence severity, rather than scheduling "one failure drill per term" by habit.
  • Fold some fault injections into unannounced windows mixed with other routine subjects, and use the detection time measured under that unannounced condition as a baseline closer to real performance.
  • Define residual capability, prohibited actions, the basis for judging a safe state, and recovery authority for every scenario.
  • Withhold some direct fault labels so trainees rely on timestamps, cross-checks, and system response to find the failure themselves.
  • Debrief the timeline from first anomalous cue to takeover, feed recurring source-trust errors back into interface design, and monitor the false-alarm rate on normal scenarios to catch the reverse negative-transfer effect of over-weighting degraded training.

Related

  • Same group: Y6.01.1 Rare operating conditions can only be trained through simulation · Y6.01.2 Simulation fidelity determines transfer
  • Nearby: Y4.01 Fail-safe and fail-operational behavior · Y3.09 Switching between manual and automatic control
  • Search terms: degraded-mode training · failure injection · automation surprise

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Y6.01.3