Proper names and digit strings fail first on intelligibility
Aliases: OOV TTS · name intelligibility · digit-string transcription
What it is
A sentence that transcribes cleanly can still drop its payload. A pharmacy kiosk says “give pickup code B7Q4, atorvastatin”: the carrier (“give,” “pickup code”) is repaired from context; the alphanumerics and the drug name go first. Proper names and digit strings have almost no contextual constraint. They are the first break on the intelligibility curve. Sentence-level word error can look acceptable because high-frequency function words prop up the average while the symbols that had to travel are already wrong.
Why it happens
Closed-class and high-frequency words have language-model and lexical support; a smeared vowel can still be fished out of the sentence. Hearing “please enter” as “please do enter” does not change the task. Names, drug names, and confirmation codes are nearly open vocabulary: B versus D, 15 versus 50, Li versus Lee — context does not help. The E-set of letters (B D E G P T V Z) is an old telephony problem; coarticulation on digits glues neighbouring places together. Names often take a weaker letter-to-sound path; a surname unseen in training is rendered as “something in that neighbourhood.” Failure is not uniform across words. It concentrates on the stretch that has no redundancy.
Studying it
Split the corpus into a carrier sentence + payload. The carrier is a fixed frame; the payload rotates through person names, drug names, amounts, alphanumeric codes. Score the payload alone; do not pass the item on sentence accuracy. Dependent measures: payload identification, and confusion matrices for letters and digits. Independent variables: channel bandwidth (wideband versus telephony), whether the listener already knows the name.
Listeners who know the name (the prescribing doctor, a friend) push the failure point later. Walk-up pickup users who do not know the drug name are the worst case. MOS is almost useless here: people can call the reading “natural” and still write the code wrong. Do not invent an engine’s “name accuracy percentage.”
Where it stops holding
If the payload is also on a screen, an ear error can be patched by the eye — that is another channel holding the load, not proof that the audio payload passed. In tone languages, surnames that differ only by tone fail earlier than in non-tone languages. Someone who just typed the code has a template in working memory to match against; a walk-up listener does not. Slowing the sentence can rescue some digits; it will not rescue a name whose letter-to-sound mapping is simply wrong.
Applying it
- Evaluation must isolate names, drugs, amounts, and codes. Do not sign off “intelligible” on sentence word error alone.
- Do not treat “read the code faster again” as the only recovery. Offer a second channel: show it, text it, or let the user speak it back.
- Put letter-to-sound failures for surnames into the lexicon; do not hope the listener will guess.
- How to check: twenty live codes and twenty drug names; listeners write only the payload. Carrier perfect and payload wrong is this failure point, not a globally bad engine.
Related
- Same group: M3.01.1 Naturalness and intelligibility are two different measures · M3.01.2 In noise, intelligibility comes first · M3.01.3 Too much naturalness raises expectations of understanding · M3.01.5 Prosodic error hurts comprehension more than a dull voice · M3.01.6 Long listening accumulates auditory fatigue
- Nearby: M3.07 Rate, pauses and prosody · M3.05 Voice and screen complementarity
- Search terms:
proper names and digit strings fail first·OOV intelligibility·digit transcription