VRCaptions: Design Captions for DHH Users in Multiplayer Communication in VR

Conversational ChatbotsSocial & Collaborative VRDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsSocial WorkersMuseum Curators & Archivists

Research Background and Problem

  • What problems or challenges did the authors identify?
    Deaf and Hard of Hearing (DHH) users face difficulties in accessing critical auditory information, such as speech content and non-speech cues (e.g., emotions, tone emphasis, and sound localization), in both real-world and virtual reality (VR) multi-user interactions. These types of information are crucial for effective collaboration. Existing remote conferencing tools and augmented reality (AR) devices have attempted to address these challenges but lack targeted designs for the needs of DHH users in VR multi-user environments. These include the development of more readable captions, efficient non-speech information transmission, and reduced cognitive load in interactive settings.

  • Why is this problem important?
    Auditory information is a vital component of collaboration and communication. DHH users often struggle to fully participate in complex interactive environments, which hinders their ability to engage in teamwork, gaming, and educational scenarios within VR. As VR and 360° video technologies become more widespread, ensuring equal participation for users with hearing difficulties is increasingly critical. Addressing this issue can enhance the inclusivity and user experience of VR systems.

  • Research Motivation and Related Work
    Although some studies have explored the communication needs of DHH users in real-world meetings and AR scenarios, research on VR multi-user interactions remains scarce. Previous work has focused on aspects such as caption timing, emotional information display, and localization in multiplayer games. However, these studies have not fully addressed how to implement comprehensive solutions for readable captions, speech, and non-speech information transmission in VR multi-user interactions. This gap motivated the research and design of the VRCaptions captioning system.


Solution

  • What methods or solutions did the authors propose?
    The authors proposed VRCaptions, a prototype captioning system designed for VR multi-user interaction scenarios. Using a user-centered design (UCD) approach, the system aims to meet DHH users' multidimensional needs for caption design, including readability, speech information transmission, and non-speech information transmission. The design focuses on seven key aspects:

    • Caption display position.
    • Sound localization.
    • Speaker turn-taking display.
    • Overlapping speech handling.
    • Caption delay feedback.
    • Chat history visualization.
    • Speaker identity representation.
  • What are the innovative aspects of this solution?

    • The design extends captioning solutions to VR multi-user interaction scenarios by incorporating the real needs and preferences of DHH users.
    • It integrates both speech and non-speech information while reducing cognitive load in complex interactive environments.
    • The solution includes user experience validation through evaluations in real gaming scenarios.
  • What are the implementation steps and key technologies used?

    • User Needs Assessment: Conducted a literature review and three rounds of co-design workshops to identify and confirm design requirements.
    • Design Refinement: Based on collected data, developed specific caption design options categorized into three dimensions (readability, speech, and non-speech information transmission).
    • User Preference Interviews: Engaged 13 DHH participants in semi-structured interviews to gather preference data on design options.
    • Caption System Development: Developed the VRCaptions prototype based on interview findings and integrated it into a multi-user cooperative VR escape room game.
    • User Validation: Evaluated the effectiveness of the caption design through gameplay experiences with mixed-hearing groups.

Research Findings

  • What specific outcomes were achieved?

    • Proposed and validated seven design directions, identifying the most preferred options among DHH users.
    • Developed and validated a prototype captioning system for VR multi-user environments (VRCaptions), achieving the following key features:
      • Displaying captions at the lower center of the view to enhance readability.
      • Showing captions in natural speaker order to help users follow conversations.
      • Providing a scrolling chat history for reviewing dialogue content.
      • Using avatars to identify speakers, offering a more intuitive recognition method.
      • Indicating sound localization through visual cues.
  • What advantages does it have over existing solutions?

    • It comprehensively addresses the needs for caption readability and non-speech information transmission, better supporting multi-user interactions.
    • The system is optimized to accommodate the unique characteristics of DHH users in complex VR interactive environments.
    • It offers flexible design features, such as the ability to hide or review chat history.
  • What were the experimental or evaluation results?
    During the experiments, DHH users endorsed the following design choices:

    • Placing captions at the lower center of the view.
    • Using natural order bubble captions to display speaker turns.
    • Implementing a hideable scrolling chat history.
    • Using avatars to identify speakers, which was more intuitive than using colors or nicknames.
    • Replacing mini-map sound localization with in-view visual cues.
      The experiments also demonstrated that VRCaptions significantly reduced information access barriers for DHH users in VR interaction scenarios.
  • Limitations and Future Directions

    • Technical Limitations: The current system relies on automatic speech recognition (ASR), which struggles with difficult-to-recognize speech. Future work could explore integrating lip-reading or sign language input technologies.
    • Scenario Limitations: The experiments were limited to multi-user gaming environments. Future research should expand to VR collaboration and education scenarios.
    • User Research: The study involved a small number of participants and validated only two sets of experimental data. Broader community-based large-scale user testing is needed.
    • Caption Delay Design: User feedback on delay-related designs was not collected, requiring further experimental validation.

Through this research, the authors not only expanded the understanding of DHH user needs but also established a preliminary user experience-centered VR captioning solution. This work provides a valuable reference for future developments in optimizing captioning systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188678/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714186
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Conversational Chatbots, Social & Collaborative VR, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Social Workers, Museum Curators & Archivists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers