How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human–Robot Interaction

Human-Robot Collaboration (HRC)Teleoperation & TelepresencePhysical-Digital Hybrid InteractionPhysicians, Nurses & CliniciansHCI Researchers

Paper Title

How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human–Robot Interaction

Publication Info

  • Topic area: Multimodal communication and intent recognition in human–robot interaction (HRI)
  • Keywords: Multimodal interaction, human–robot interaction, object reference, gaze tracking, pointing gestures, speech patterns, intent recognition, virtual reality, user behavior, interface design

Background and Problem

  • Problem / challenge: Current multimodal instruction-recognition algorithms are rigid and incomplete, failing to fully leverage natural human communicative cues such as gaze, gestures, and speech in unconstrained settings.
  • Significance: Robust interpretation of natural multimodal instructions is critical for enabling robots to perform complex tasks in dynamic environments, enhancing usability and interaction efficiency.
  • Motivation and related work: Prior research has focused on scripted multimodal schemes or voice-only systems, which limit adaptability and fail to capture spontaneous user behavior. This paper addresses the gap by studying natural multimodal behaviors and their implications for intent recognition.

Solution

  • Proposed approach: A VR-based study to analyze natural multimodal behaviors (speech, gaze, gestures, head movements) during object-referencing tasks in human–robot interaction.
  • Novelty:
    1. Systematic analysis of verbal and nonverbal behaviors in natural, unconstrained settings.
    2. Evaluation of spatial accuracy for various multimodal cues, revealing consistent and context-dependent patterns.
    3. Design implications for gaze-centered intent recognition algorithms and interactive systems with feedforward and feedback mechanisms.
  • Procedure and key techniques:
    1. Conducted a VR experiment with 30 participants performing household tasks (e.g., pick & place, cleaning).
    2. Manipulated spatial (distance, orientation) and contextual (referent complexity) variables.
    3. Collected multimodal data (speech, gaze, hand movements, head orientation) and analyzed patterns using statistical models.
    4. Compared localization accuracy of different multimodal vectors (e.g., gaze, hand-pointing) using various calculation methods.

Results

  • Concrete findings:
    • 21.7% of speech instructions lacked object attributes, relying on vague expressions like "this object."
    • Gaze vector achieved the highest localization accuracy (median error = 0.006 m with the minimum-distance method), outperforming all other vectors by at least an order of magnitude.
    • Participants dynamically adapted speech and multimodal behaviors based on distance and referent complexity (e.g., more descriptive features for distant or complex scenes).
    • Four multimodal coordination patterns were identified, with hand–eye coordination and target-locked attention being the most frequent.
  • Advantage over baselines:
    • Peak-interval-based method reduced gaze–target error to 0.08 m compared to 0.48 m for the speech-window average method.
    • Gaze-centered multimodal fusion outperformed traditional end-to-end approaches by leveraging behavior-informed heuristics.
  • Experiments / evaluation:
    • 2,430 instructions collected in a VR environment simulating household tasks.
    • Statistical analysis of multimodal data (e.g., linear mixed models, G-tests) to evaluate spatial and contextual effects on behavior and localization accuracy.
  • Limitations and future work:
    • Limited ecological validity due to VR-based simulation; future studies should validate findings in real-world settings.
    • Rule-based methods cannot fully capture user diversity; future work should develop AI models incorporating physiological and brain–computer interface signals.
    • Environmental noise, robot appearance, and interactivity were not included as variables.

Summary

This study investigates natural multimodal behaviors in human–robot interaction, focusing on how users refer to objects using speech, gaze, gestures, and head movements. Results highlight the dominance of vague speech expressions and the critical role of gaze for precise target localization, with gaze vectors achieving the highest accuracy. Temporal coordination between gaze and other modalities (e.g., hand-pointing) can further enhance intent recognition. The findings inform the design of gaze-centered multimodal algorithms and user-friendly interfaces with feedforward and feedback mechanisms, advancing both theoretical understanding and practical applications in HRI.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222361/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790946
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-Robot Collaboration (HRC), Teleoperation & Telepresence, Physical-Digital Hybrid Interaction
work
Professions
Physicians, Nurses & Clinicians, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers