Audition cannot scan, so spoken lists are expensive
Aliases: serial list cost · audition cannot scan · spoken-list access cost
What it is
The eyes can jump, return, and pin two rows. The ears cannot. An airline IVR reads eight Chengdu departure times in order; the caller cannot scan to “the afternoon ones,” only walk the whole timeline or speak to stop it. Audition cannot scan, so any spoken list’s cost grows linearly in time: each item owns the ear exclusively, skipping waits on a dialogue round-trip, going back means replay. The claim is the access cost of the reading channel. It is not “too many options to remember,” and not “comparison tasks should not use voice” — those hold on their own. The question here is: once you decide to read the inventory aloud, how expensive is the channel.
Why it happens
A visual list offers spatial random access. Audition is a tape that only runs forward. There is no pointer to a region of interest: you spend time until it arrives, or you approximate scanning with “stop / next,” coarse and slow. Information scent is delayed until the relevant item is spoken; until then the listener does not know whether to stay. Even when each item is immediately actionable (not a cross-comparison), a long read is still expensive because the region of interest cannot be pointed at. The cost is time and attention, not holding the whole set in working memory to choose — that is the screenless memory-load story. What is expensive here is the tape.
Studying it
Serial position in a spoken list: put the target at position k, measure time-to-find and misses. Independent variables: list length and target position. Dependent measures: search time, miss rate, abandonment, and at which item people speak to stop. Primacy and recency will show up; use them to diagnose how much time position k costs, not as an account of memory capacity.
A screen control gives scanning back, so you are no longer measuring a spoken list. Do not substitute satisfaction for time: people can call the reading “very clear” and quit at item three.
Where it stops holding
Two or three items, and the listener is waiting for a known label (“the eight o’clock”), the tape is still short and the cost is tolerable. When the user said “read them all,” scanning is not the goal; the reading was requested. Lean-back content — sports scores, radio time checks — is not meant to be jumped. Experts who already know the structure wait at a known slot; they brought their own index. Reading the list more slowly only lengthens the tape. It does not restore scanning.
Applying it
- Default to not reading a set aloud. Reading is expensive; it needs a reason each time.
- When you must read, stopping has to be as available as the items. Do not wait until the end.
- Do not treat “next item” as a scanning substitute for regional access (“I want the afternoon block”).
- How to check: time from list onset to user action, and which item the action landed on. If most actions are “stop, I’ll look it up,” the reading already cost more than it was worth.
Related
- Same group: M3.03.2 Give count and category before the items · M3.03.3 Beyond a few items, switch to filtering
- Nearby: M1.01 When voice-first is appropriate · M1.02 Memory load of screenless interaction · M3.08 Difficulty of reading long lists aloud
- Search terms:
spoken lists cannot be scanned·serial audio·list access cost