What makes speech intelligible is how far it rises above the background, not the noise level alone
Aliases: SNR · speech-to-noise ratio · intelligibility margin
What it is
Whether speech or an alert tone is intelligible depends on how many decibels it sits above the background noise — the signal-to-noise ratio (SNR). A 60 dB voice in a 50 dB office (SNR = +10 dB) is easy to follow; the same 60 dB voice on a 75 dB subway car (SNR = −15 dB) is almost entirely swamped. The absolute noise level — 50 dB or 75 dB — is not by itself the determining variable; how far the signal clears the noise floor is.
The intuitive mistake is to focus on how loud the environment itself is, as if quieter rooms are simply "easier" and louder rooms "harder." The variable that actually predicts intelligibility is the gap between signal and noise. Reporting a decibel figure without that comparison doesn't answer the question of whether something can be heard.
Why it happens
Detecting a sound means its energy in some frequency band must exceed the noise energy in that same band by enough of a margin. Below that margin, the continuous energy of the noise buries the fluctuations of the signal and the auditory system never registers that it occurred; above it, the signal's onsets, envelope, and spectral detail survive intact and intelligibility rises accordingly. Speech intelligibility is especially sensitive to SNR because the linguistic content is carried mainly by consonants, which are lower in energy and higher in frequency than vowels — exactly the components a noise floor swamps first. The earliest casualty of a shrinking SNR is not "can I tell someone is talking" but "can I tell 'she' from 'he.'"
Studying it
The standard paradigm presents speech material (word lists, sentences, digit strings) over controlled noise and systematically varies either the signal level or the noise level to manipulate SNR, then measures the SNR at which a target recognition rate (commonly 50%) is reached — the speech reception threshold (SRT). The independent variable is SNR (sometimes decomposed into signal level and noise level separately, since the two don't contribute symmetrically); the dependent variable is word/phoneme recognition accuracy or the SRT itself.
This paradigm is used to give a product's voice interaction or alert design a quantitative answer to "how noisy can it get before this stops working," rather than a vague claim about loudness.
Methodological caveat: lab noise is usually steady-state (pink noise or multi-talker babble) with a stable spectrum, while real-world noise is often intermittent and non-stationary — a siren, a door slamming. SNR in real settings fluctuates sharply over time, so a single lab-measured SRT is only an average reference point, not a guarantee.
Where it stops holding
- SNR only answers whether something can be heard. Even at an adequate SNR, fast speech rate, unfamiliar accents, or technical jargon independently reduce intelligibility — those are linguistic, not signal-level, limits.
- Spectral overlap matters as much as the overall ratio: the same aggregate SNR causes far greater intelligibility loss when the noise energy concentrates in the speech-critical band (roughly 500 Hz–4 kHz) than when it sits outside that band. A single SNR number hides this difference.
- Lab-measured SRTs come from cooperative listeners under steady-state noise; they are not a direct usability floor for a real product operating in non-stationary noise, only a starting reference for how much margin to design in.
Applying it
- When setting the output level for speech prompts, alerts, or voice interaction, don't target a fixed decibel number — target a margin above the typical noise level of the intended context. Measure the context's noise first, then set the level.
- Keep the alert's spectrum away from the frequency bands the target environment's noise concentrates in; separating spectra improves real-world intelligibility even when the raw SNR number is unchanged.
- How to verify it: record real noise from the target context (not a quiet test room), play it back under the candidate speech or alert, and ask listeners with no prior expectation whether and how much they can make out — testing against real-context SNR rather than a fixed lab noise floor.
Related
- Same group: A3.09.2 speech input and output degrade together under noise · A3.09.3 the social cost of audio output in quiet settings · A3.09.4 the SNR needed for intelligibility varies with content type and familiarity · A3.09.5 ambient noise floors swing sharply across contexts, defeating fixed volume settings · A3.09.6 active noise cancellation changes what reaches the ear, not the device's own output loudness setting
- Nearby: A3.04 auditory masking · A3.02 loudness perception and equal-loudness contours · A3.07 alarm discriminability
- Search terms:
signal-to-noise ratio·speech reception threshold·speech intelligibility·masking
Cards in the same group
- A3.09.2A noisy room garbles what a device says and what its microphone hears, both at once
- A3.09.3In a library or a sleeping room, the question isn't whether a sound can be heard but whether it should be
- A3.09.4A familiar sentence needs far less signal above the noise than an unfamiliar string of digits does
- A3.09.5One fixed volume setting can't survive a day that runs from a quiet bedroom to a subway car
- A3.09.6Noise cancellation quiets what reaches the ear; it doesn't touch the device's own volume setting