ChatDirector: Enhancing Video Conferencing with Space-Aware Scene Rendering and Speech-Driven Layout Transition
Authors
Social & Collaborative VRMixed Reality Workspaces
Document Title
ChatDirector: Enhancing Video Conferencing with Space-Aware Scene Rendering and Speech-Driven Layout Transition
Document Information
- Subject Area: Human-Computer Interaction (HCI), Remote Video Conferencing System Design
- Keywords: video conferencing, 3D avatar, remote presence, attention switching, depth map, depth estimation, machine learning, video communication, collaborative work, enhanced communication
Research Background and Problem
- Identified Issues or Challenges:
- Current remote video conferencing systems (RVCS) lack the sense of presence and visual cues inherent in face-to-face meetings.
- Traditional 2D grid layouts in video conferencing lack spatial awareness, requiring users to exert significant cognitive effort to interpret meeting information, leading to reduced attention and communication efficiency.
- Existing camera technologies and enhanced visual cues offer limited improvements, often confined to 2D representations or requiring expensive hardware.
- Significance: Video conferencing has become a ubiquitous tool in work, education, and social activities, yet its user experience struggles to match that of face-to-face meetings, thereby limiting communication efficiency and participant focus.
- Research Motivation and Related Work:
- Recent studies have focused on restoring visual cues (e.g., eye contact and head orientation) and spatial information in video conferencing but are often limited to technical prototypes or reliant on expensive hardware.
- This study aims to design a solution using conventional computing devices (e.g., laptops with standard cameras) to enhance spatial awareness and communication fluidity in video conferencing.
Solution
- Method or Solution:
- Propose a lightweight video conferencing system, ChatDirector, which uses standard 2D screens and RGB cameras to create a spatially aware shared virtual 3D meeting environment.
- Introduce a speech-driven decision tree layout algorithm (Speech-Driven Layout Transition) that automatically adjusts the virtual scene's layout and avatar actions to simulate attention shifts and natural conversational flow.
- Develop real-time RGB video-to-3D avatar reconstruction technology to generate 3D avatars and dynamically render the meeting scene.
- Innovations:
- Achieve spatial visualization without relying on specialized hardware (e.g., depth cameras or head-mounted devices).
- Design a dynamic layout transition algorithm integrated with speech detection to facilitate natural attention shifts.
- Propose a 3D virtual meeting layout distinct from traditional 2D grids, enhancing the sense of presence.
- Implementation Steps and Key Technologies:
- Spatial Scene Rendering Pipeline: Capture video streams from RGB cameras, use a lightweight depth estimation model to generate depth maps in real-time, and construct 3D avatars of users.
- Speech-Driven Layout Algorithm: Detect real-time speech activity and keywords to determine speaking states (e.g., silent, addressing the group, one-on-one conversation) and dynamically adjust layouts and avatar actions using a decision tree module.
- System Integration: Utilize WebRTC for real-time media streaming and Web Speech API for speech transcription.
Research Outcomes
- Specific Results:
- Developed a system prototype that enables a 3D virtual meeting experience without requiring specialized hardware.
- Quantitative experiments showed significantly improved communication fluidity and user engagement.
- Demonstrated diverse application scenarios, such as brainstorming sessions, online debates, and office hours for teaching.
- Advantages Over Existing Solutions:
- Achieves 3D spatial awareness on standard devices without expensive hardware and is compatible with XR systems like Meta Horizon Workrooms.
- High levels of automation and intelligence, dynamically capturing and adapting to speech activity.
- Experimental or Evaluation Results:
- User studies comparing ChatDirector with traditional 2D video conferencing systems (e.g., Google Meet) showed that ChatDirector more effectively supports attention switching and enhances meeting interactivity.
- The system's layout algorithm demonstrated excellent output accuracy (Layout State accuracy: 94.79%, Avatar State accuracy: 98.80%).
- Users gave higher ratings for engagement, presence, and responsiveness.
- Limitations and Future Directions:
- Current 3D avatar reconstruction is limited to the RGB camera's field of view and cannot present a full 3D representation of individuals.
- The decision tree layout algorithm has limited semantic understanding, which could be enhanced with NLP and LLM technologies in the future.
- Further exploration is needed for large-scale meeting scenarios, including group discussions, user anchoring, and visualization support for multitasking.
- Future work could investigate balancing user personalization and integrating additional meeting elements (e.g., screen sharing and real-time statistics).
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can spatial awareness be achieved in video conferencing using ordinary devices (such as RGB cameras and 2D screens) to enhance communication fluency?Category: Attention Sensing and State ClassificationSimilar questionsarrow_forward
- How can speech-driven layout algorithms dynamically adjust virtual meeting scenes to simulate attention shifts and natural conversation flow?Category: Attention Sensing and State ClassificationSimilar questionsarrow_forward
- How does 3D virtual meeting layout improve users' sense of engagement and presence compared to traditional 2D grid layouts?Category: Attention Sensing and State ClassificationSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Video conferencing lacks the spatial sense and communication fluency of face-to-face interaction, reducing efficiency.Category: Attention Sensing and State ClassificationSimilar questionsarrow_forward
- 100%
Poros: Configurable Proxies for Distant Interactions in VR
CHI '21· Social & Collaborative VR +1
- 100%
Evaluating the Extension of Wall Displays with AR for Collaborative Work
CHI '23· Social & Collaborative VR +1
- 100%
Being Social in VR Meetings: A Landscape Analysis of Current Tools
DIS '23· Social & Collaborative VR +1
- 100%
Sharing Play Spaces: Design Lessons from Reddit Posts Showing Virtual Reality in the Home
DIS '23· Social & Collaborative VR +1
- 100%
HyWay: Enabling Mingling in the Hybrid World
UbiComp '23· Social & Collaborative VR +1
- 100%
CoplayingVR: Understanding User Experience in Shared Control in Virtual Reality
UbiComp '24· Social & Collaborative VR +1
- 100%
ShARe: Enabling Co-Located Asymmetric Multi-User Interaction for Augmented Reality Head-Mounted Displays
UIST '20· Social & Collaborative VR +1
- 67%
VirtualSpace - Overloading Physical Space with Multiple Virtual Reality Users
CHI '18· Social & Collaborative VR +2
- 67%
On the Shoulder of the Giant: A Multi-Scale Mixed Reality Collaboration with 360 Video Sharing and Tangible Interaction
CHI '19· Social & Collaborative VR +2
- 67%
UndoPort: Exploring the Influence of Undo-Actions for Locomotion in Virtual Reality on the Efficiency, Spatial Understanding and User Experience
CHI '23· Social & Collaborative VR +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642110
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Social & Collaborative VR, Mixed Reality Workspaces
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers