M3.05.1screen for lists and detail, voice for conclusionsdesignresearch

Screen takes lists and detail; voice takes the conclusion

Aliases: voice gist screen detail · channel split · spoken headline

What it is

In a dialogue that has a display, the screen takes lists and detail; voice takes the conclusion. A shopping smart display says “cheapest is the twenty-four pack at eighty-nine”; the table, the comparison, the small print stay on screen. Speech is good at handing over one actionable judgement in time; vision is good at laying many fields down at once. That is channel complementarity, not one channel reciting the other. Speech’s information scent says whether looking is worth it; the screen is where looking pays off in detail.

Why it happens

Audition is good at a gist, a recommendation, a next act. Vision is good at juxtaposition, comparison, close reading. In a Wickens-style resource split the two can run together if they do not compete for the same code — if the mouth does not re-read the table the eyes are on. A conclusion is short in the ear; a list is wide on the screen. Hand the table to speech and juxtaposition becomes tape again; leave the conclusion only in a screen corner and the judgement is hidden in pixels that have to be searched. When the split fails, people either listen through the table or wait until they have read before they move, and neither channel’s strength is used.

Studying it

Three presentations: A, voice reads the table, screen blank; B, voice gives the conclusion, screen holds the table; C, both channels carry the whole table. Dependent measures: choice quality, decision time, whether gaze goes to the screen. Independent variable: number of rows. Do not make millisecond alignment of the two channels the independent variable; that is a timing question. Do not substitute “satisfied it sounded complete” for choice quality.

Use a real comparison or timetable with a pre-labeled better row. If in B the spoken conclusion contradicts the table, you measured conflict, not complementarity.

Where it stops holding

With no screen, this split does not apply; that needs a screenless design of its own. When the eyes are locked on a road, speech may have to carry more — a reallocation under occupancy, not a licence to read the table back into the ear. Legally spoken terms cannot live only on the display. In glare, or at a kiosk too far to read, people temporarily join the screenless case; the spoken line still has to stand as a conclusion. Experts who already know the table only need speech to report what changed.

Applying it

  • Write the spoken line as a headline: result + why + “details on screen.”
  • Put lists, prices, maps, and terms on the display. Do not use speech as a second table.
  • Cover the screen: the spoken line should still be a complete conclusion. Cover the speaker: the screen should still hold the list. Each keeps its job.
  • How to check: a cover-up walkthrough. Listen-only cannot reach a conclusion, or look-only cannot find the list, means the split was written backwards.

Related

  • Same group: M3.05.2 The two channels should not repeat each other verbatim · M3.05.3 Eyes-free needs its own design, not a degraded GUI
  • Nearby: M3.10 Multimodal complementarity of voice and screen · M3.03 Reading long lists aloud · M1.01 When voice-first is appropriate
  • Search terms: screen for lists and detail, voice for conclusions · multimodal complementarity · information scent

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M3.05.1