AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XR

Identity & Avatars in XRAffective Human-Computer DialogueImmersion & Presence ResearchHCI ResearchersAI/ML Researchers & EngineersUI/UX Designers

Paper Title

AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XR

Publication Info

  • Topic area: Interactive hand gestures for XR conversational agents
  • Keywords: XR, hand gestures, conversational agents, spatial grounding, LLMs, co-speech gestures, user engagement, object interaction, embodied agents, gesture taxonomy

Background and Problem

  • Problem / challenge: Current conversational agents in XR struggle with spatial communication, relying heavily on text or speech, which creates cognitive burdens for users in spatial tasks. Existing systems lack versatile, interactive, and emotionally attuned behaviors that synchronize with open-ended dialogue.
  • Significance: Addressing this gap can improve user comprehension, engagement, and task performance in XR environments, enabling more intuitive human–AI interactions.
  • Motivation and related work: Prior efforts in XR assistants and gesture-based systems have been task-specific or limited to static overlays, failing to integrate dynamic, synchronized gestures. Human communication naturally integrates gestures for spatial reference and engagement, which inspired the development of AgentHands.

Solution

  • Proposed approach: AgentHands, an LLM-powered XR system that augments conversational agents with expressive, interactive hand gestures synchronized with speech, enabling spatially grounded and engaging interactions.
  • Novelty:
    1. A taxonomy of hand gestures and agent attributes for spatially grounded conversations in XR, derived from a formative study.
    2. An end-to-end system that generates and renders synchronized hand gestures aligned with verbal responses.
    3. Empirical validation through a user study demonstrating improved engagement and clarity over speech-only baselines.
  • Procedure and key techniques:
    • Pre-registration of objects in the XR environment using gaze and scene reconstruction.
    • Augmentation of LLM-generated responses with inline GestureEvents specifying gesture types, spatiality, and timing.
    • Real-time rendering of gestures synchronized with speech using a parametric animation engine.
    • A gesture library categorized into deictic, iconic, and expressive gestures based on a compositional taxonomy.

Results

  • Concrete findings:
    • AgentHands significantly improved spatial understanding (e.g., object and direction identification) with scores of 6.50 vs. 4.58 (baseline) on a 7-point scale.
    • Enhanced task comprehension and engagement, with users rating AgentHands 6.08 vs. 4.58 (baseline) for engagement.
    • Better memorability and clarity of instructions, with a score of 2.17 (AgentHands) vs. 3.50 (baseline) for difficulty remembering responses.
  • Advantage over baselines:
    • Statistically significant improvements in spatial clarity, task execution, and engagement compared to a speech-only agent.
    • Users reported that gestures made instructions easier to follow and warnings more noticeable.
  • Experiments / evaluation:
    • Within-subjects study with 12 participants performing two tasks: orchid care and 3D printer operation.
    • Metrics included spatial understanding, task clarity, engagement, and emotional expression, measured via Likert-scale surveys and SUS scores.
    • AgentHands achieved an average SUS score of 80.6 vs. 70.4 for the baseline.
  • Limitations and future work:
    • Reliance on pre-registration of objects limits adaptability to dynamic environments.
    • Current gesture library is manually curated; future work could explore LLM-based generation of novel gestures.
    • Limited evaluation scope; larger-scale and longitudinal studies are needed to assess long-term utility and generalizability.
    • Potential for deeper integration with XR operating systems and APIs for richer interaction capabilities.

Summary

AgentHands introduces a novel XR system that enhances conversational agents with synchronized, interactive hand gestures for spatially grounded tasks. By leveraging a taxonomy of hand gestures and an LLM-powered pipeline, the system improves user engagement, spatial understanding, and task clarity. A user study demonstrated significant advantages over a speech-only baseline, particularly in spatial reference clarity and user engagement. Future work includes enhancing scene understanding, expanding gesture generation, and conducting larger-scale evaluations. AgentHands represents a step toward more expressive and embodied XR interfaces.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222723/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790938
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Identity & Avatars in XR, Affective Human-Computer Dialogue, Immersion & Presence Research
work
Professions
HCI Researchers, AI/ML Researchers & Engineers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers