GazeChat: Enhancing Virtual Conferences with Gaze Awareness and Interactive 3D Photos

Eye Tracking & Gaze InteractionSocial & Collaborative VRMixed Reality WorkspacesUI/UX Designers

Document Title

GazeChat: Enhancing Virtual Conferences with Gaze-aware 3D Photos

Document Information

  • Subject Area: Gaze awareness and enhancement technologies in virtual conferencing systems
  • Keywords: gaze awareness, eye contact interaction, video conferencing, networked collaboration, deep learning, privacy protection, user experience, low bandwidth, augmented reality technology

Research Background and Issues

  • Identified Problems or Challenges:
    • Users often turn off cameras during virtual conferences, leading to a lack of gaze information in communication.
    • Even with cameras on, traditional video conferencing fails to accurately convey "who is looking at whom."
    • Virtual conferences face challenges in balancing privacy protection and limited network bandwidth.
  • Significance:
    • Gaze awareness is a critical non-verbal communication cue that enhances interactivity and engagement.
    • Providing visual interaction solutions not only improves conference experiences but also balances privacy protection and bandwidth requirements.
  • Research Motivation and Related Work:
    • Existing technologies like GAZE-2 and TeleHuman require expensive multi-camera setups or specialized hardware environments, limiting widespread adoption.
    • Commercial software (e.g., Memoji) primarily focuses on avatar animation but neglects relative gaze awareness during conversations.
    • The authors explore a low-cost, accessible system that uses standard cameras to enable gaze awareness and optimize virtual conference experiences.

Solution

  • Method or Solution:
    • Propose "GazeChat," a virtual conferencing system that tracks user gaze using standard cameras and presents gaze awareness through 3D dynamic photos.
  • Innovations:
    • Utilize deep learning-based image synthesis methods to generate dynamic photos with varying gaze angles.
    • Employ a lightweight WebRTC framework to reduce network bandwidth usage while protecting user privacy.
    • Represent relative gaze information rather than absolute eye positions.
  • Implementation Steps and Key Technologies:
    1. Input Data: Users upload static profile photos and use cameras for real-time sessions.
    2. Depth Map and Image Synthesis: Use depth estimation models to generate visual images at different angles and animate eye movements through a pre-trained First Order Motion model.
    3. Real-time Gaze Tracking: Record users' gaze focus on the screen using WebGazer.js or Tobii eye-tracking devices.
    4. Rendering Module: Utilize 3D reconstruction and front-end rendering libraries like Three.js to display dynamic 3D avatars based on user gaze.
    5. Data Transmission Optimization: Servers only need to transmit minimal spatial position and audio data, reducing network load.

Research Outcomes

  • Specific Results:
    • GazeChat successfully visualizes "who is looking at whom," enhancing user engagement and communication experience.
    • Compared to traditional audio and video conferencing, GazeChat significantly improves social presence and conversation efficiency.
  • Advantages:
    • More bandwidth-efficient and privacy-protective than video conferencing.
    • More interactive and visually informative than audio conferencing.
    • Easy to operate and compatible with standard hardware environments (e.g., laptops, standard cameras).
  • Experimental Results:
    • User experiments show that GazeChat improves the perception of eye contact during conferences, with users rating the system highly for its novelty and entertainment value.
    • GazeChat outperforms traditional audio conferencing in social richness (e.g., interactivity, emotional feedback), user experience, and user engagement.
  • Limitations and Future Directions:
    • The current version only includes dynamic eye movement information, with limited support for other visual cues (e.g., facial expressions or body movements).
    • Gaze tracking algorithms rely on accurate calibration and are susceptible to user posture and ambient lighting conditions.
    • User studies involved a narrow age range and need to be expanded to other groups (e.g., students or elderly users).
    • Future systems could integrate more non-verbal cues (e.g., facial expressions, body posture) to enhance natural interaction.

This document provides a low-cost, lightweight, and multifunctional solution for virtual conferencing systems and offers significant insights for the development of future virtual collaboration scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/61357/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3472749.3474785
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Social & Collaborative VR, Mixed Reality Workspaces
work
Professions
UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers