Speech conveys meaning well but structure poorly
Aliases: speech output · auditory structure · TTS
What it is
Speech carries meaning well and structure poorly. Saying what a sentence means is easy; saying where that sentence sits in an information hierarchy is hard. Hierarchy, parallel items, and grouping are visible at a glance on screen, but once spoken they exist only if narrated item by item.
Why it happens
Speech is strictly serial. A listener cannot pause and re-read, nor compare the fifth item with the first, so earlier items must be held in working memory to establish relations. When the structure demands more held items than working memory affords, comprehension collapses. A screen externalizes structure in space and lifts that burden. Correspondingly, spoken structural markers—"first," "the following three"—matter because they temporarily encode spatial relations into a temporal stream.
Studying it
Recall and retrieval tasks can compare spoken and written presentation of the same content: whether users can restate the number of items, name which group an item belongs to, and how long it takes to locate a specific entry. Independent variables typically include list length, hierarchy depth, and the presence of structural markers. Report the speech rate, since it directly sets the working-memory load window.
Where it stops holding
For very short content, where order does not matter and the user needs a single meaning—"saved," "five new messages"—the structural disadvantage is negligible. When the content is itself a sequence (procedure steps, a timeline), speech and structure do not conflict, because temporal order is the semantics. The limitation targets hierarchy and juxtaposition, not all multi-item content.
Applying it
- Let speech carry conclusions, confirmations, and cues that must be understood; leave hierarchy, juxtaposition, and grouping to the screen.
- When a list must be spoken, use explicit numbering and grouping markers and keep the length within what can be held.
- Provide a text equivalent so structure can be re-examined spatially.
- Verification: after hearing a passage, have users restate item count and group membership; if structural recall errs far more than content recall, the structure load is too high.
Related
- Within the group: D2.06.2 Speech output occupies the auditory channel and cannot be browsed in parallel · D2.06.3 Speech output must support interruption and replay
- Adjacent: D2.07 The linear, unscannable nature of hearing · M1.05 Short-term memory limits in voice interfaces
- Search terms:
speech output·auditory structure·working memory load