M3.01.6listening fatigue from prolonged TTSdesignresearch

Long listening accumulates auditory fatigue

Aliases: TTS listening effort · listener fatigue · prolonged-playback wear

What it is

A few seconds of synthetic speech can pass; ten or twenty minutes is a different bill. Listening fatigue accumulates. A museum audio guide, or a commute of traffic plus calendar, still sounds clear in the first two minutes; later the listener drifts and tail items drop. The engine did not suddenly worsen; listening effort stacked up. Fatigue is not unintelligibility: people can still transcribe while finding the voice abrasive, and then they withdraw attention — and only then does intelligibility collapse with it.

Why it happens

Listening effort taxes attention even when words are recoverable; dual-task reaction time and pupil diameter pick it up. TTS usually has a narrower range of spectral and timing variation than a human talker, so the auditory system gets fewer of the micro-jitters that reset natural speech. Concatenative or vocoder nicks that are ignored in a five-second prompt become aversive over minutes. People slide from decoding to enduring to shutting the channel. That is not the same as a single answer being too long: answer length is an information budget; fatigue is an integral over a tour or a session. Short, exact replies help, but the same voice run for half an hour still integrates.

Studying it

Prolonged listening: fifteen to thirty minutes of TTS versus a human talker versus the same engine with more natural jitter inserted. Dual-task: listen while tracking or light driving. Dependent measures: a listening-effort scale or NASA-TLX, secondary-task RT, and comprehension on the first third versus the last third of the material. Pupillometry if available.

Tedium of the content and aversion to the voice entangle: the same script with a human control separates “dull copy” from “synthetic grind.” Twenty lab minutes are not a two-hour drive. Do not invent a minute-mark at which fatigue “starts.”

Where it stops holding

Self-chosen podcasts and interesting commentary turn aversive more slowly — agency and content matter. Listeners with hearing loss fatigue faster. Sub-minute IVR menus barely show the curve. Unpausable long playback amplifies fatigue, but that is a control problem stacked on top, not timbre alone. A highly characterful, affect-heavy voice used for long exposition often tires more than a plainer reading voice.

Applying it

  • Anything longer than a one-shot confirm needs pause, skip, and an exit to screen or text.
  • Keep the characterful, high-affect voice for openings; switch long informational stretches to a flatter reading style.
  • Cut long tours or commute briefings into segments with an exit between them. Do not roll one waveform to the end.
  • How to check: split comprehension into the first two minutes and the last two, plus a 1–7 “how tiring” item. Late accuracy down, early still fine, is fatigue, not an engine that cannot speak.

Related

  • Same group: M3.01.1 Naturalness and intelligibility are two different measures · M3.01.2 In noise, intelligibility comes first · M3.01.3 Too much naturalness raises expectations of understanding · M3.01.4 Proper names and digit strings fail first on intelligibility · M3.01.5 Prosodic error hurts comprehension more than a dull voice
  • Nearby: M3.04 Barge-in · M3.11 Trimming speech output · M3.02 Speech rate and prosody
  • Search terms: listening fatigue from prolonged TTS · listening effort · TTS wear-out

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M3.01.6