Since U Been Gone: Augmenting Context-Aware Transcriptions for Re-Engaging in Immersive VR Meetings
Authors
Research Background and Issues
-
Identified Problems or Challenges:
The authors observed that while virtual reality (VR) and mixed reality (MR) meetings provide immersive experiences, users may experience interruptions or distractions due to internal or external factors. Such interruptions can lead to missed critical content and reduced engagement. Additionally, current VR meeting interface designs, such as transcription panels fixed to desktops or subtitles bound to avatars, present challenges: desktop-fixed panels disrupt social engagement, while avatar-bound subtitles fail to provide contextual information during interruptions. -
Significance:
With the growing demand for remote work and virtual collaboration, maintaining user immersion and social presence in VR meetings while enabling users to quickly re-engage after interruptions has become a critical issue. This is pivotal to determining whether virtual meeting platforms can effectively replace traditional face-to-face meetings. -
Research Motivation and Related Work:
The motivation lies in addressing engagement issues and facilitating re-engagement after interruptions in VR meeting environments. Previous studies have focused on transcription interface designs for video conferencing (e.g., meeting summaries, real-time interaction tools) and distraction management in VR, but have not explored how to integrate social presence with re-engagement needs in VR meetings.
Proposed Solution
-
Proposed Solution:
The authors designed and implemented a novel adaptive avatar-attached transcription interface—EngageSync. This system provides real-time information through transcription and contextually generated summaries during meetings, while offering interruption summaries to help users quickly re-engage, all while maintaining social presence with group members. -
Innovations:
- Automatically adjusts transcription modes (real-time transcription and concise summaries) based on user engagement status.
- Dynamically attaches the transcription panel to the virtual character (avatar) being gazed at by the user, rather than statically placing it on a desktop.
- Utilizes gaze- and gesture-based triggers to minimize distractions and support on-demand access to information.
- Generates key summaries using a language model (GPT-4 Turbo) to help users quickly grasp key points.
-
Implementation Techniques and Steps:
- Real-time Transcription and Summary Generation:
Audio transcription is performed using Google STT, and concise 15-word summaries are generated using GPT-4 Turbo. - User Attention Tracking and Interface Interaction Mechanisms:
Meta Movement SDK is employed to monitor users' gaze direction and detect whether users are gazing at avatars or triggering the transcription panel through gestures. - Two Mode Switching:
- Engagement Mode: Displays real-time transcription or recent summaries when the user focuses on an avatar.
- Re-engagement Mode: Generates summaries for all avatars who spoke during the interruption and deactivates them after being read.
- Multi-user Collaboration:
A real-time multi-user platform is built using Photon networking to ensure the synchronization of audio, avatar movements, and transcription and summary sharing.
- Real-time Transcription and Summary Generation:
Research Findings
-
Specific Results:
- Enhanced Social Presence: Compared to traditional desktop-fixed interfaces, EngageSync significantly increased the time spent gazing at avatars and improved the sense of social distribution. Users also exhibited higher focus on avatars.
- Faster Re-engagement: EngageSync significantly reduced re-engagement time compared to other interfaces, with particularly notable effects in medium-sized groups.
- Information Recall Performance: EngageSync performed best in helping users recall key points from speakers, especially in medium-sized groups, with stronger associative memory for speakers with fewer contributions.
-
Advantages:
- Outperformed desktop-fixed panels and traditional avatar-bound subtitles in terms of social presence (co-presence), attentional allocation, and re-engagement time.
- Provided automatically generated concise summaries, effectively addressing information overload issues in complex conversational environments.
-
Experimental or Evaluation Results:
- Two experimental scenarios compared three interfaces (EngageSync, Table-fixed, Avatar-fixed), recording metrics such as social presence, re-engagement time, and user memory performance.
- In medium-sized groups, EngageSync significantly outperformed other interfaces (e.g., re-engagement time effect p<0.001, statistically significant).
-
Limitations and Future Directions:
- Limitations:
- The fixed position of the text panel may not suit all users; future studies should test the effects of different positions.
- The experiment used pre-recorded meeting scripts rather than real-time conversations, which may affect the evaluation of user interaction and social presence.
- The accuracy of the transcription system is limited, particularly when speech recognition errors affect the quality of the summaries.
- Future Research:
- Extend experiments to real-time interactive scenarios to observe how users re-engage during face-to-face conversations.
- Explore the impact of different panel positions or dynamic interface layouts on user experience.
- Improve the accuracy of language model-generated summaries to enhance system reliability.
- Limitations:
The above analysis comprehensively summarizes the study's background, methodology, and findings while highlighting key insights and strategies for addressing the identified issues. This work holds significant research value in enhancing the effectiveness of immersive VR meetings.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- In VR meetings, how can users be supported to re-engage without interrupting social interaction?Category: XR Remote Collaboration and Telepresence CommunicationSimilar questionsarrow_forward
- Are transcription interfaces dynamically attached to virtual avatars more effective than fixed interfaces?Category: XR Remote Collaboration and Telepresence CommunicationSimilar questionsarrow_forward
- How can concise summaries help users quickly obtain key information and re-engage?Category: XR Remote Collaboration and Telepresence CommunicationSimilar questionsarrow_forward
Practical Problems
1- In VR meetings, distracted users miss content and struggle to quickly re-participate.Category: XR Remote Collaboration and Telepresence CommunicationSimilar questionsarrow_forward
- 80%
AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online Presentations
CHI '21· Social & Collaborative VR +1
- 67%
Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents
CUI '25· Social & Collaborative VR +1
- 60%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 60%
Beyond the Chat: Executable and Verifiable Text-Editing with LLMs
UIST '24· Human-LLM Collaboration
- 60%
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
UIST '24· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)