Comparison and browsing tasks are a poor fit for voice
Aliases: spoken browsing · audition cannot scan · serial audio comparison
What it is
Comparison and browsing need several candidates visible at once, with gaze jumping among them. Voice lines those candidates up on a one-way timeline: after the third restaurant, the first restaurant's price is no longer in the ear. If the task is essentially pick, compare, flip, voice as the main channel loses systematically. Setting a timer is not that task. Choosing among eight places by rating, distance, and queue length is.
Why it happens
Visual search is spatially parallel: the eyes can sweep a list, jump back, pin two rows against each other. Audition is temporally serial and almost without a jump-back — hearing item one again means asking the system to replay or remembering it yourself. Working memory holds few items at once; comparison wants several items' attributes co-present. People collapse to remembering only the last one heard, or only the first, with the middle squeezed out. That is not a failure of diction. The channel turned juxtaposition into a queue.
Browsing also has a control-of-pace problem. On a screen the user drives the scan: skip what is dull, enlarge what is interesting. Spoken browsing is paced by the reader; the user can only approximate scanning with “next” and “stop,” coarse and slow. The cross-check comparison needs (this wait versus that rating) requires the user to build a table that has no external store.
Studying it
Give the same candidate set under serial speech versus a visual list and score choice quality: whether the final pick is close to a pre-labeled better item, decision time, and replay requests. Independent variables: list length, number of attributes per item, whether barge-in replay is allowed. Also score position effects: are first and last items chosen abnormally often — memory voting instead of preference.
Eye tracking shows jump-back comparison in the visual condition; the speech analogue is how often people ask to hear “the second one” again. If those requests grow linearly with list length, the channel is using dialogue to patch missing space. Do not stop at satisfaction: people can call the reading “very clear” and still pick a worse item.
Where it stops holding
Two or three candidates with one or two attributes (“red or blue,” “leave now or in ten minutes”) keep serial comparison inside memory. When the user already has an external criterion (“the usual place”), so-called browsing is confirmation and voice is enough. If a screen is in peripheral vision and speech only reads a conclusion or the focused row, comparison is happening visually; this limit is about screenless main paths, not every interface that has a voice. Expert dispatchers compare inside an internalized model; they do not rebuild the list from the current read-out.
Applying it
- Detect whether the task is set comparison or scanning. If it is, put the set on a screen. Speech says “seven places, sorted by distance — hear the top three or look at the screen.”
- Screenless: collapse the set into an answerable question (“nearest or highest rated?”), filter N down to two or three, then read. Do not start at item one and read through eight.
- Do not treat “next track / next item” as a browsing substitute for decisions that need attribute cross-checks; that is only queue motion.
- How to check: same real candidates, half the people listen only, half see a list. If the listen-only group clusters on first or last items and cannot state key attributes of the ones not chosen, the task should not be voice-primary.
Related
- Same group: M1.01.1 Voice earns its keep when hands and eyes are already taken · M1.01.2 Voice prefers short commands over long procedures
- Nearby: M1.02 Memory load of screenless interaction · M3.03 Reading long lists aloud · M3.08 Difficulty of reading long lists aloud
- Search terms:
voice poorly supports comparison and browsing·serial audio·recency bias in spoken lists