In noise, intelligibility comes first
Aliases: speech-in-noise TTS · clear-speech style · cabin intelligibility
What it is
Once the environment masks, synthetic speech’s first job is intelligibility, not naturalness. On a highway, engine, tire noise, and a radio left half-on sit on top of each other; a studio-warm, breathy prompt will swallow the next exit name at the driver’s seat. The object is whether the system’s mouth can still deliver words into the ear. It is not whether the microphone can still hear the user — that is the input side. People change how they talk in noise. TTS trained on booth read-speech often keeps talking as if the booth were still there.
Why it happens
Human talkers in noise go into a Lombard-like shift: more intensity, a brighter spectrum, a little slower, less breath. Booth recordings do not. Neural TTS further keeps breathy onsets, creak, and soft attacks — cues that add MOS in the quiet and sit exactly in the low-mid band tire and HVAC noise occupy, without carrying lexical identity. Intelligibility in noise is about glimpses: how many formant fragments still sit above the masker so a segment can be assembled. An intimate, breathy voice parks its energy where masking is cheapest. A flatter, more forward clear-speech style spends some humanness to buy more glimpses. Chasing naturalness as if the cabin were a living room buys pleasantness with the wrong acoustic budget.
Studying it
Speech-in-noise: overlay the target sentence on a controlled SNR. Use speech-shaped noise, a real cabin recording, or a station floor. The dependent measure is keyword identification or the speech reception threshold (SRT: SNR at 50% keywords), not MOS. Contrast a “studio-natural” style with a clear-speech style (slightly slower, more intense, less breath). The Speech Transmission Index (STI) can estimate the acoustics; it does not replace listeners.
Methodological catch: noise added on headphones is not a cabin — spatial layout, ego-noise, and speaker directivity differ. A radio as masker is under product control; tire noise is not. Do not pick a voice on quiet MOS and assume it will be “close enough” on a motorway.
Where it stops holding
In a quiet indoor IVR, naturalness can share the stage; not every prompt must become a PA voice. Hearing-aid users often need more high-frequency energy, not merely more level. Active-noise-cancelling headphones take the cabin masker off, and this priority loosens. Music or a podcast on the same speakers is a masker you can duck: ducking done, pressure on timbre drops a notch. When speech is fully gone people will look at a screen — a screen is a fallback, not a reason to give away output intelligibility.
Applying it
- For in-car, shop-floor, and station kiosks, pick voice and EQ on a keyword test with field noise injected, not on a quiet MOS bake-off.
- Ship a clear-speech setting: slightly slower, louder, less breath; switch to it when high noise is detected.
- When the product’s own music or podcast shares the speaker, duck at prompt onset. Do not wait for the user to mute media.
- How to check: play live prompts in a parked car with HVAC and a radio on. A passenger writes the exit number, street name, or gate. Wrong write-down means intelligibility has not been put first.
Related
- Same group: M3.01.1 Naturalness and intelligibility are two different measures · M3.01.3 Too much naturalness raises expectations of understanding · M3.01.4 Proper names and digit strings fail first on intelligibility · M3.01.5 Prosodic error hurts comprehension more than a dull voice · M3.01.6 Long listening accumulates auditory fatigue
- Nearby: C7.05 Noisy environments · C7.12 Recognition degradation in noise · M3.02 Speech rate and prosody
- Search terms:
intelligibility-first in noise·speech-in-noise·clear speech TTS
Cards in the same group
- M3.01.1Naturalness and intelligibility are two different measures
- M3.01.3Too much naturalness raises expectations of understanding
- M3.01.4Proper names and digit strings fail first on intelligibility
- M3.01.5Prosodic error hurts comprehension more than a dull voice
- M3.01.6Long listening accumulates auditory fatigue