Speech rate must be adjustable
Aliases: TTS speaking rate · rate control · words-per-minute setting
What it is
Synthetic speech rate is not a single site-wide default. The same engine: a commuter taking a news briefing will run it at about 1.4×; an older adult on a government IVR already finds factory speed rushed. In-car, the first reading of an unfamiliar street name is too fast at a rate that is fine on a known stretch. Rate has to be adjustable, and the adjustment has to persist, because there is no words-per-minute that fits every listener and every genre.
Why it happens
Comfortable listening rate for TTS is usually below conversational human speech: coarticulatory cues are weaker, so listeners need more time to assemble segments. At the other tail, familiar content, younger listeners, and ears-free users find a conservative default sluggish and will trade margin for time. Rate is not one MOS optimum; it is tangled with familiarity, hearing, and whether a transcript is glanceable. A locked wpm is a crowd mean that misses both tails. Time-scale modification (WSOLA, PSOLA) or a neural rate condition can move rate without huge pitch shift — that is the knob, not a one-time calibration.
Studying it
Method of adjustment: a navigation passage, a menu, a news brief; people set the speed to “still intelligible, one notch faster and I lose words.” Or a staircase: rate as independent variable, comprehension as dependent, find the notch where accuracy starts to drop. Other independents: age, L2 status, genre familiarity. Dependents: chosen rate, comprehension, preference.
A lab “please set your rate” overestimates how many people will find a buried control in the field. A one-time setting in a deep car menu is not a control beside every prompt. Do not use “first pass slow, listen-again fast” as the independent variable — that is a different design. The question here is whether a given listener can move the default at all.
Where it stops holding
Safety and legally required prompts may have a floor rate that cannot be rushed into unintelligibility. Extreme rate destroys stop-consonant closures. Some public announcements have a conventional official rate. Devices with no account (elevators, platforms) cannot do per-person settings; they pick a conservative default. Entertainment (audiobooks) and confirmations should not share one lock. People with hearing loss often need clearer, not merely slower; slow is only one lever.
Applying it
- Make rate a persistent, first-class setting, not three menus down. Also ship on-the-fly “slower / faster” as spoken commands.
- Do not lock navigation, messages, and entertainment to one rate.
- Pick the factory default for unfamiliar names and older ears, not for young listeners in a demo hall.
- How to check: of people who found the control, how many left the default. If most of them moved it, the default only served people who never found the knob. Spot-check comprehension at default versus the user’s setting.
Related
- Same group: M3.02.2 Pauses carry punctuation and structure · M3.02.3 Stress position changes the meaning of a sentence
- Nearby: M3.01 Naturalness and intelligibility of synthetic speech · M3.07 Rate, pauses and prosody · C7.06 Hands-free and eyes-free use
- Search terms:
adjustable speech rate·TTS speaking rate·time-scale modification