Referent candidates come from talk, screen, and the room
Aliases: exophora · screen deixis · situated reference
What it is
“Turn that off.” Two seconds ago the dialogue mentioned the kitchen light; the phone is highlighting an album; a floor lamp in the room is on. That points at three places at once. Multimodal referent candidates means the set for pronouns and demonstratives is not only prior nouns. It also includes the focused object on the current screen, and objects the device can perceive in the environment. Doing anaphora only against dialogue history binds a “that” that was aimed at the screen or the room to the wrong thing.
Why it happens
Spoken “it / that / the red one” lands on a salience list. In a voice system with a screen and a situation, three buffers feed that list. Dialogue: entities just mentioned or just confirmed. Screen: the focused row, playing artwork, a selected map pin. Environment: a lamp on in this room, a speaker that is sounding, the largest object in camera. Exophora is not a resolution failure. It is the speaker pointing out of the talk on purpose.
When the three sources compete, ranking needs a policy: recency in talk, visual focus, spatial nearness. “Prior mention wins” is not a default. Eyes on the screen saying “this” put salience in vision; walking past a lamp saying “turn it off” put it in the room; “order that restaurant again” put it in the dialogue. Merge the buffers into one unlabeled list and errors cannot be explained, nor can the system ask “the one on the screen, or the lamp.”
Studying it
Build a competing-referent arena: the same demonstrative has a plausible object in talk, on screen, and in the room. Elicit “make that louder,” “what is that.” Ground truth from gaze, pointing, or a later tap on “which one did you mean.” Report a confusion matrix: source the system bound × source the user wanted. Independent variables: whether the screen is in view, whether the environmental object is actually perceivable (is the lamp on this room’s account).
Do not rely on text coreference corpora. They have no screen focus and no room; they will score exophora as error. Eye tracking or first-person video checks whether a screen object was actually co-present, not to do recognition.
Where it stops holding
A speaker-only device has no screen buffer; the environment buffer still exists (“this room,” “the track that’s playing”). A phone in a pocket does not make screen objects visually co-present; they should drop out of the candidate set. In AR, “that” is almost always environment or the gaze cone; dialogue history should be down-weighted. When top scores are close and sources differ, asking beats guessing. Putting objects outside the currently perceivable set (a lamp upstairs, an album not on screen) into the environment buffer invents referents.
Applying it
- Keep three source-tagged candidate lists for “it / this / that / the red one.” Do not merge them into one anonymous list.
- If the top two candidates come from different sources and sit close in score, ask about source (“the track on screen, or the lamp”). Do not pick in silence.
- Clear the screen buffer when the screen is off or out of view. Down-weight a room’s environmental objects when the person has left that room.
- Sample “this / that” sessions. Label the source wanted and the source bound. Cross-source errors hurt more than near-misses inside one source; fix those first.
Related
- Same group: M1.09.2 Subject-dropped fragments need completing from the last turn · M1.09.3 New-turn slots need overwrite versus overlay
- Nearby: M1.04 Context Retention · M3.10 Multimodal Complementarity of Voice and Screen · C7.06 Hands-free and eyes-free use
- Search terms:
multimodal referent candidates·exophora·screen deixis