VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guidance through Large Language Models

Social & Collaborative VRHuman-LLM CollaborationMuseum Curators & ArchivistsHCI Researchers

Document Title

VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guidance through Large Language Models

Document Information

  • Subject Area: Human-Computer Interaction, Virtual Reality (VR), Multi-modal Interaction, Cultural Heritage Technology
  • Keywords: Virtual Museum, Multi-modal Feedback, Large Language Models, User Personalization, Virtual Reality, Tour Guidance Experience, Human-Computer Interaction, Knowledge Enhancement, Museum Guidance, Interaction Design

Research Background and Issues

  • What problems or challenges did the authors identify?

    • Achieving effective tour guidance in virtual museums is challenging. Existing virtual tour systems often provide limited or semi-flexible interactions, such as preset routes or pre-written commentary, failing to fully utilize the flexibility and immersion of virtual environments.
    • Users have highly personalized tour guidance needs in different scenarios, requiring more natural interaction, multi-modal feedback, and services tailored to specific needs, which are inadequately supported by current research.
    • The potential of using Large Language Models (LLMs) for museum guidance has not been fully explored, particularly in providing contextualized multi-modal feedback to users.
  • Why is this issue important?

    • Virtual museums hold potential in education, cultural heritage preservation, and academic research, with their core focus on enhancing user experience in terms of immersion, engagement, and knowledge acquisition.
    • Designing tour guidance tailored to personalized needs can enhance user engagement and exploratory experiences in virtual museums, promoting the application of virtual technologies in real-world scenarios.
  • Research Motivation and Related Work

    • This study is inspired by the interaction demands of virtual museums, summarizing the limitations of existing interactive technologies and aligning with the research trend of multi-modal interaction to design a more flexible and adaptive virtual tour guidance system.

Solution

  • What methods or solutions did the authors propose?

    • The authors proposed VirtuWander, an innovative tour guidance system based on Large Language Models (LLMs) that supports multi-modal interaction in virtual environments.
    • The system employs a two-stage framework (context recognition and feedback generation) to transform users' natural language input into contextualized tour guidance and generate multi-modal feedback.
    • It provides five combinations of multi-modal feedback (e.g., voice + virtual avatar + text window) to meet different scenarios and needs.
  • What are the innovative aspects of this solution?

    • Combining multi-modal feedback with LLM capabilities enables more natural and intelligent user interaction.
    • By designing a framework for task categories (e.g., knowledge enhancement, navigation, personalized preferences) and feedback mechanisms (e.g., voice, text window, virtual screen), the system specifically addresses diverse user needs.
    • The system is designed and validated for the specific context of virtual museums, accommodating various tour scenarios (e.g., thematic tours, single artifact exploration, personalized customization).
  • What are the implementation steps? What key technologies were used?

    • Stage 1: Context Recognition

      • Convert users' natural language input into text via speech recognition.
      • Use LLMs to classify user tasks, such as knowledge enhancement or navigation.
      • Identify required contextual information (spatial information, semantic information, etc.).
    • Stage 2: Feedback Generation

      • Employ multiple LLM-based subtask bots to generate the multi-modal feedback users need through customized prompts.
      • Optimize feedback quality by integrating user behavior data and environmental context.
    • Multi-modal Feedback Design

      • Provide innovative feedback combinations, including voice, virtual avatars, text windows, navigation markers, maps, and highlighted areas.

Research Outcomes

  • What specific outcomes were achieved?

    • Developed the VirtuWander system and implemented five major multi-modal feedback combinations.
    • Demonstrated use cases in different scenarios (e.g., thematic tours, single artifact exploration, personalized tours), validating the system's flexibility and applicability.
    • User studies indicated that the VirtuWander system significantly enhanced the tour guidance experience, including knowledge acquisition, natural communication, and immersion.
  • What advantages does it have compared to existing solutions?

    • Compared to traditional virtual tour systems, VirtuWander offers real-time, personalized multi-modal interaction and intelligent knowledge supplementation.
    • By leveraging LLMs for task classification, context association, and feedback generation, the system supports more complex user needs.
  • What were the experimental or evaluation results?

    • Users rated the engagement and usability of multi-modal feedback (e.g., virtual screens, voice, highlighted areas) highly (Likert scale average score ≥5/7).
    • The system was perceived to enhance the tour guidance experience, particularly in natural interaction, task customization, and museum knowledge acquisition.
    • Users expressed strong interest in extending such systems to real-world applications in the future.
  • Limitations and Future Directions

    • Limitations:

      • The system has limited support for more complex physical interactions and multi-user collaboration.
      • The accuracy and standardization of LLM-generated responses need further improvement, especially in scenarios involving historical and cultural knowledge.
      • The system is primarily designed for flat art exhibits and does not yet address the guidance needs for 3D artifacts.
    • Future Directions:

      • Investigate new interaction input methods (e.g., gestures, haptics) to enhance the expressiveness of virtual tour guidance.
      • Explore a combination of active and passive feedback mechanisms to optimize feedback timing.
      • Expand the system's applicability (e.g., 3D artifacts, archaeological sites).
      • Consider using Augmented Reality (AR) technology to apply multi-modal feedback in real-world scenarios (e.g., airport guidance, mall navigation).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148243/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642235
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Social & Collaborative VR, Human-LLM Collaboration
work
Professions
Museum Curators & Archivists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers