ChatDirector: Enhancing Video Conferencing with Space-Aware Scene Rendering and Speech-Driven Layout Transition

Social & Collaborative VRMixed Reality Workspaces

Document Title

ChatDirector: Enhancing Video Conferencing with Space-Aware Scene Rendering and Speech-Driven Layout Transition

Document Information

  • Subject Area: Human-Computer Interaction (HCI), Remote Video Conferencing System Design
  • Keywords: video conferencing, 3D avatar, remote presence, attention switching, depth map, depth estimation, machine learning, video communication, collaborative work, enhanced communication

Research Background and Problem

  • Identified Issues or Challenges:
    1. Current remote video conferencing systems (RVCS) lack the sense of presence and visual cues inherent in face-to-face meetings.
    2. Traditional 2D grid layouts in video conferencing lack spatial awareness, requiring users to exert significant cognitive effort to interpret meeting information, leading to reduced attention and communication efficiency.
    3. Existing camera technologies and enhanced visual cues offer limited improvements, often confined to 2D representations or requiring expensive hardware.
  • Significance: Video conferencing has become a ubiquitous tool in work, education, and social activities, yet its user experience struggles to match that of face-to-face meetings, thereby limiting communication efficiency and participant focus.
  • Research Motivation and Related Work:
    1. Recent studies have focused on restoring visual cues (e.g., eye contact and head orientation) and spatial information in video conferencing but are often limited to technical prototypes or reliant on expensive hardware.
    2. This study aims to design a solution using conventional computing devices (e.g., laptops with standard cameras) to enhance spatial awareness and communication fluidity in video conferencing.

Solution

  • Method or Solution:
    1. Propose a lightweight video conferencing system, ChatDirector, which uses standard 2D screens and RGB cameras to create a spatially aware shared virtual 3D meeting environment.
    2. Introduce a speech-driven decision tree layout algorithm (Speech-Driven Layout Transition) that automatically adjusts the virtual scene's layout and avatar actions to simulate attention shifts and natural conversational flow.
    3. Develop real-time RGB video-to-3D avatar reconstruction technology to generate 3D avatars and dynamically render the meeting scene.
  • Innovations:
    • Achieve spatial visualization without relying on specialized hardware (e.g., depth cameras or head-mounted devices).
    • Design a dynamic layout transition algorithm integrated with speech detection to facilitate natural attention shifts.
    • Propose a 3D virtual meeting layout distinct from traditional 2D grids, enhancing the sense of presence.
  • Implementation Steps and Key Technologies:
    1. Spatial Scene Rendering Pipeline: Capture video streams from RGB cameras, use a lightweight depth estimation model to generate depth maps in real-time, and construct 3D avatars of users.
    2. Speech-Driven Layout Algorithm: Detect real-time speech activity and keywords to determine speaking states (e.g., silent, addressing the group, one-on-one conversation) and dynamically adjust layouts and avatar actions using a decision tree module.
    3. System Integration: Utilize WebRTC for real-time media streaming and Web Speech API for speech transcription.

Research Outcomes

  • Specific Results:
    1. Developed a system prototype that enables a 3D virtual meeting experience without requiring specialized hardware.
    2. Quantitative experiments showed significantly improved communication fluidity and user engagement.
    3. Demonstrated diverse application scenarios, such as brainstorming sessions, online debates, and office hours for teaching.
  • Advantages Over Existing Solutions:
    1. Achieves 3D spatial awareness on standard devices without expensive hardware and is compatible with XR systems like Meta Horizon Workrooms.
    2. High levels of automation and intelligence, dynamically capturing and adapting to speech activity.
  • Experimental or Evaluation Results:
    1. User studies comparing ChatDirector with traditional 2D video conferencing systems (e.g., Google Meet) showed that ChatDirector more effectively supports attention switching and enhances meeting interactivity.
    2. The system's layout algorithm demonstrated excellent output accuracy (Layout State accuracy: 94.79%, Avatar State accuracy: 98.80%).
    3. Users gave higher ratings for engagement, presence, and responsiveness.
  • Limitations and Future Directions:
    1. Current 3D avatar reconstruction is limited to the RGB camera's field of view and cannot present a full 3D representation of individuals.
    2. The decision tree layout algorithm has limited semantic understanding, which could be enhanced with NLP and LLM technologies in the future.
    3. Further exploration is needed for large-scale meeting scenarios, including group discussions, user anchoring, and visualization support for multitasking.
    4. Future work could investigate balancing user personalization and integrating additional meeting elements (e.g., screen sharing and real-time statistics).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148295/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642110
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Social & Collaborative VR, Mixed Reality Workspaces
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
10 related papers