Naturalness and intelligibility are two different measures
Aliases: MOS versus intelligibility · TTS quality · mean opinion score
What it is
Naturalness asks whether the voice sounds like a person. Intelligibility asks whether the words can be recovered. A bank IVR may use a warm neural voice to say “enter the last four digits of your card”; listeners award a high score while those four digits come out wrong. The first is typically scored with an ITU-T-style mean opinion score (MOS) from absolute category rating; the second needs transcription, keyword identification, or a diagnostic rhyme test. The two numbers describe different jobs on the same waveform. Neither stands in for the other.
Why it happens
MOS raters judge a global impression: timbre, breathiness, how smoothly units join. Intelligibility is phonemic and lexical recovery: place of articulation, vowel contrast, whether an unstressed syllable still holds a word. Neural TTS can match a speaker’s spectral envelope and F0 closely enough to sound human while smearing place cues on unstressed syllables — nicer to rate, harder to decode. Formant-era TTS can sound like a machine and still keep segmental contrasts that transcribe well. A quiet-room “sounds good” never passed through the identification pathway. Treating MOS as a proxy for “users understood” answers a recognition question with an aesthetic score.
Studying it
Run two tests on the same synthetic set. Naturalness: ITU-T P.800 absolute category rating (ACR); for TTS, the P.85 listening procedures. Intelligibility: transcription of semantically unpredictable sentences (SUS), the diagnostic rhyme test (DRT), or keyword identification. Independent variables: engine or vocoder, rate, telephone bandwidth. Dependent measures: MOS (1–5) and word or keyword accuracy, reported separately — do not blend them into one “quality” number.
Methodological catch: headphone MOS in a quiet lab inflates “pleasant”; transcription in the same quiet misses field failures. If raters see the transcript while they listen, the measure is not intelligibility. Do not assume the two axes correlate under neural TTS; the reason to split the tests is that they often fork.
Where it stops holding
On first contact, naturalness can be a brand claim of its own and need not yield. For amounts, accounts, irreversible confirms, intelligibility leads even if the voice is plainer. Call-centre QA listeners and naive users do not produce interchangeable MOS. Headphone MOS does not transfer to a lobby loudspeaker or a car hands-free speaker. Nonsense syllable lists will crush intelligibility without necessarily crushing “sounds nice”; the two corpora cannot be interpreted as one.
Applying it
- Accept a prompt set only with both numbers: a MOS panel, and write-down accuracy on last-four digits, amounts, and confirm verbs.
- A “nicer” voice that drops digit transcription has not passed. Do not let a MOS gain cover an intelligibility loss.
- Split the evaluation set: greetings and guides in one bin; account numbers, amounts, and names in another. Sentence-level accuracy buries the payload in high-frequency words.
- How to check: twenty live prompts; half the listeners give MOS, half write the payload immediately. MOS up and write-down down means the wrong axis was optimized.
Related
- Same group: M3.01.2 In noise, intelligibility comes first · M3.01.3 Too much naturalness raises expectations of understanding · M3.01.4 Proper names and digit strings fail first on intelligibility · M3.01.5 Prosodic error hurts comprehension more than a dull voice · M3.01.6 Long listening accumulates auditory fatigue
- Nearby: M3.02 Speech rate and prosody · M2.10 Persona and consistency
- Search terms:
naturalness versus intelligibility·MOS·TTS listening test