Speech supplies the action, pointing supplies the object
Aliases: deictic reference · speech plus gesture · multimodal reference
What it is
Speech says what to do, pointing says which object to do it to. The two are naturally complementary: in "delete this," the action comes from language and "this" must be filled in by pointing. Used separately each is incomplete; aligned, they produce an expression that needs no explicit naming.
Why it happens
The division is stable because the information forms differ. An action is an abstract predicate, well suited to language, where a limited vocabulary composes into many meanings; an object is a concrete spatial location, more precise and cheaper to indicate than to describe. Integration requires the two to align on the same referent—the system must know which object is being indicated and bind it to the spoken action. That also fixes the failure modes: when reference is unclear or timing is off, the system either picks the wrong object or falls back to a default.
Studying it
Measure reference-resolution accuracy: have users issue commands by speech plus pointing, record whether the system selects the right object and action, and analyze whether errors come from object or action resolution. Variables include pointing precision and stability, phrasing, candidate density, and whether users see their pointing feedback. Outcomes include resolution accuracy and the distribution of error types.
Where it stops holding
When few objects exist and language can name one uniquely—"open settings"—pointing is unnecessary. When objects are dense or spatial relations complex, pointing precision demands rise and visible feedback is needed so users can correct. If a user cannot use one modality, the other must accomplish the task alone, or the interaction is incomplete.
Applying it
- Let speech carry the action and pointing carry the object; do not require verbal description of spatial location.
- Show visible pointing feedback so users can confirm which object the system currently targets.
- In dense scenes, raise pointing precision requirements or offer a candidate list for confirmation.
- Verification: measure resolution accuracy by error source; errors clustered in object resolution mean pointing precision and feedback should be improved first.
Related
- Within the group: D5.02.2 Fusion requires alignment within a time window · D5.02.3 Failed fusion should fall back to one modality, not guess
- Adjacent: D5.01.2 Complementarity divides different kinds of information · C5.03 Reference resolution between gesture and speech
- Search terms:
deictic reference·speech and pointing·multimodal reference