Towards Building Condition-Based Cross-Modality Intention-Aware Human-AI Cooperation under VR Environment

Social & Collaborative VRAR Navigation & Context AwarenessHuman-LLM CollaborationSoftware Engineers & DevelopersUI/UX DesignersHCI Researchers

Title of the Paper

Towards Building Condition-Based Cross-Modality Intention-Aware Human-AI Cooperation under VR Environment

Paper Information

  • Research Domain: Human-AI Cooperation, Virtual Reality (VR), Cross-Modality Intention Detection
  • Keywords: Virtual Reality, Human-AI Cooperation, Intention Detection, Action Generation, Multimodal Interaction, Condition-Driven, Large Language Models (LLM), Smart Shopping

Research Background and Problem Statement

  • Identified Problems and Challenges:

    1. Ambiguity in user expressions, such as "items suitable for a car trunk" or "the item I just saw."
    2. Insufficient accuracy of artificial intelligence (AI) in recognizing user intentions and generating recommendations.
    3. Difficulty in dynamically adjusting response content based on changes in user interaction.
  • Significance: Effectively capturing user intentions in a virtual reality environment and generating efficient information displays and recommendations based on these intentions can not only enhance user experience but also improve task completion efficiency.

  • Research Motivation and Related Work:

    • Progress has been made in multimodal user intention detection in virtual environments, particularly in the fields of eye-tracking, gesture recognition, and voice analysis.
    • Multimodal feedback (e.g., voice, visual tags) has been proven more effective than single-modal optimization in generating virtual responses based on user intentions. However, current methods often focus on single-modal optimization.
    • While large language models (LLMs) have been explored as conversational assistants in virtual scenarios, there are significant limitations in task efficiency and output control.

Proposed Solution

  • Proposed Method or Solution: The authors propose a condition-based cross-modality intention-aware human-AI cooperation framework, which includes the following core components:

    1. Intent Tuples: Used to construct an intention library, each tuple contains four core elements: intention, condition, intent prompt, and action prompt.
    2. 2-LLMs Framework: Two distinct large language models are used, one for intention recognition (Intent LLM) and the other for action generation (Action LLM), to reduce model load and improve output consistency.
  • Innovations:

    1. Introduction of a task description method centered on "conditions," explicitly representing the dynamic state of tasks.
    2. Separation of intention detection and action generation into Intent LLM and Action LLM, reducing the complexity and length of prompts through task decoupling.
    3. Dynamic matching of user intentions based on multimodal data (user eye movements, gestures, voice input), making it more aligned with everyday interaction scenarios compared to existing methods.
  • Implementation Steps and Key Techniques:

    1. Build an intention library and design intent tuples to cover common task requirements.
    2. Develop a smart furniture shopping system in a VR environment, including features like voice assistants, floating labels, red arrow indicators, and controller operations.
    3. Use Intent LLM to recognize user intentions, including processing multimodal inputs.
    4. Use Action LLM to generate dynamic multimodal responses based on intentions and present results to the user.

Research Outcomes

  • Specific Results:

    1. Successfully implemented a condition-driven VR furniture shopping system.
    2. Validated the efficiency of the proposed method through user experiments.
    3. Subjective user feedback indicated that the condition-driven multimodal cooperation framework provided higher satisfaction and reduced physical interaction burden.
  • Comparison with Existing Solutions and Advantages:

    1. By distributing intention detection and action generation across different LLMs, the prompt length for LLMs was reduced by 82.54%-92.87%, improving response accuracy.
    2. Compared to systems without condition-driven mechanisms, the proposed method improved task time efficiency by 30.15% and reduced error submissions by 29.06%.
    3. Supported more proactive user interaction methods, such as dynamically updated floating labels that help users quickly filter relevant objects.
  • Experiment or Evaluation Results:

    1. In terms of task completion time and error rates, the condition-driven system significantly outperformed traditional interaction systems.
    2. The system more efficiently recognized users' cross-modality expressions (64.7% of multimodal expressions were correctly identified, compared to 12.5% for the control system).
    3. Users were more inclined to engage in conversations with the AI, with interaction frequency increasing by approximately 48% compared to the control group.
  • Limitations and Future Directions:

    1. The current system uses a manually designed intention library; future work could explore automated generation of intent tuples or condition prompts.
    2. Modal support is not yet fully comprehensive; future research could incorporate physiological signals (e.g., EEG/EOG) to identify implicit intentions.
    3. The conditions in intent tuples currently include only limited product attributes; future work could add more metadata (e.g., user ratings, sales volume) to expand application scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147782/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642360
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Social & Collaborative VR, AR Navigation & Context Awareness, Human-LLM Collaboration
work
Professions
Software Engineers & Developers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers