Long listening accumulates auditory fatigue
Aliases: TTS listening effort · listener fatigue · prolonged-playback wear
What it is
A few seconds of synthetic speech can pass; ten or twenty minutes is a different bill. Listening fatigue accumulates. A museum audio guide, or a commute of traffic plus calendar, still sounds clear in the first two minutes; later the listener drifts and tail items drop. The engine did not suddenly worsen; listening effort stacked up. Fatigue is not unintelligibility: people can still transcribe while finding the voice abrasive, and then they withdraw attention — and only then does intelligibility collapse with it.
Why it happens
Listening effort taxes attention even when words are recoverable; dual-task reaction time and pupil diameter pick it up. TTS usually has a narrower range of spectral and timing variation than a human talker, so the auditory system gets fewer of the micro-jitters that reset natural speech. Concatenative or vocoder nicks that are ignored in a five-second prompt become aversive over minutes. People slide from decoding to enduring to shutting the channel. That is not the same as a single answer being too long: answer length is an information budget; fatigue is an integral over a tour or a session. Short, exact replies help, but the same voice run for half an hour still integrates.
Studying it
Prolonged listening: fifteen to thirty minutes of TTS versus a human talker versus the same engine with more natural jitter inserted. Dual-task: listen while tracking or light driving. Dependent measures: a listening-effort scale or NASA-TLX, secondary-task RT, and comprehension on the first third versus the last third of the material. Pupillometry if available.
Tedium of the content and aversion to the voice entangle: the same script with a human control separates “dull copy” from “synthetic grind.” Twenty lab minutes are not a two-hour drive. Do not invent a minute-mark at which fatigue “starts.”
Where it stops holding
Self-chosen podcasts and interesting commentary turn aversive more slowly — agency and content matter. Listeners with hearing loss fatigue faster. Sub-minute IVR menus barely show the curve. Unpausable long playback amplifies fatigue, but that is a control problem stacked on top, not timbre alone. A highly characterful, affect-heavy voice used for long exposition often tires more than a plainer reading voice.
Applying it
- Anything longer than a one-shot confirm needs pause, skip, and an exit to screen or text.
- Keep the characterful, high-affect voice for openings; switch long informational stretches to a flatter reading style.
- Cut long tours or commute briefings into segments with an exit between them. Do not roll one waveform to the end.
- How to check: split comprehension into the first two minutes and the last two, plus a 1–7 “how tiring” item. Late accuracy down, early still fine, is fatigue, not an engine that cannot speak.
Related
- Same group: M3.01.1 Naturalness and intelligibility are two different measures · M3.01.2 In noise, intelligibility comes first · M3.01.3 Too much naturalness raises expectations of understanding · M3.01.4 Proper names and digit strings fail first on intelligibility · M3.01.5 Prosodic error hurts comprehension more than a dull voice
- Nearby: M3.04 Barge-in · M3.11 Trimming speech output · M3.02 Speech rate and prosody
- Search terms:
listening fatigue from prolonged TTS·listening effort·TTS wear-out
Cards in the same group
- M3.01.1Naturalness and intelligibility are two different measures
- M3.01.2In noise, intelligibility comes first
- M3.01.3Too much naturalness raises expectations of understanding
- M3.01.4Proper names and digit strings fail first on intelligibility
- M3.01.5Prosodic error hurts comprehension more than a dull voice