On-screen options need names voice can refer to
Aliases: speakable GUI · named cards · on-screen deixis
What it is
A delivery UI shows four logo-only cards. Someone looking at the screen says “the green one” or “the second.” If the voice layer has not bound visible attributes, ordinals, and titles to entities, looking and speaking do not meet. Speakable names for on-screen options require that every option the eyes are on has a name that can be said and heard — a visible title, a stable ordinal, a colour or position word that actually distinguishes. Complementarity is not “pictures on screen, conclusion in the mouth”; it is being able to make a deictic name for the object now displayed.
Why it happens
With a screen, referring strategy shifts to exophora: people say “this,” “the one on the left,” “the green shop,” assuming the system shares the current frame. Graphics that are only icons, thumbnails, or unlabelled colour blocks break that sharing — the person has a visual handle, speech has no lexical one. The recogniser hears words, not pixels; unless on-screen objects are compiled into speakable labels, the channels disconnect at reference.
Speakable names come in three layers, stable to fragile: the title printed on the card (read aloud), the ordinal on this screen (“the second”), visual attributes (“the green one”). Titles must be short and match the ASR vocabulary, not an internal SKU. Attributes are safe only when unique on the current screen; four greenish cards kill “the green one.” “This” still has to bind to focus or gaze, or four objects are tied.
Studying it
The same candidates in two UIs: icon plus internal name versus short visible title + ordinal. Elicit voice selection while looking. Code referring type (title, ordinal, colour, position, “this”) and selection success. Independent variables: title visible, title matching the ASR vocabulary, attribute unique on this screen. Eye tracking to confirm the corresponding card was actually looked at.
Split failure modes: looked but could not produce a legal name (missing speakable name); produced a legal name the system did not bind (linking). Do not fold them into “multimodal satisfaction.” Labs that set titles in huge subtitles overestimate read-aloud rates; real cards are often ten-point type.
Where it stops holding
If the option is not on screen (still loading, collapsed), there is no speakable name to offer — say “nothing selectable on screen now.” In listen-only, screen dark, referring should fall back to dialogue names; do not pretend “this” exists. When the user is far from the display and titles are unreadable, the visible-title layer dies; ordinals and names speech itself has spoken still work. Treating speakable names as “read the on-screen menu aloud” duplicates channels rather than giving a mouth-handle to objects the eyes already see.
Applying it
- Every voice-selectable on-screen object gets a short title people will actually say, an in-screen ordinal, and recogniser coverage for “the Nth,” “that + title,” and left/right/top/bottom.
- Colour, shape, and similar attributes count as legal handles only when unique on this screen; if they are not unique, do not visually invite people to say them.
- Titles in the user’s language, not internal names; when a title changes, the vocabulary and the screen change together.
- How to check: voice-select from real cards, no touch. Count “this / the green one / the second” and their success. Icon-only should collapse; adding short titles and ordinals should lift it. That lift is the speakable-name layer.
Related
- Same group: M3.10.1 Visual information is gone when gaze is off the screen · M3.10.3 The two channels have to line up in time
- Nearby: M3.05 Voice and Screen Complementarity · M1.09 Context Retention and Reference Resolution · M2.06 Discoverability and Help
- Search terms:
speakable names for on-screen options·deictic·speakable GUI