Speech-Augmented Cone-of-Vision for Exploratory Data Analysis

Eye Tracking & Gaze InteractionSocial & Collaborative VRInteractive Data VisualizationUniversity Professors & ResearchersHCI Researchers

Document Title

Speech-Augmented Cone-of-Vision for Exploratory Data Analysis

Document Information

  • Domain: Human-Computer Interaction, Collaborative Virtual Reality, Data Visualization and Analysis
  • Keywords: Field of View (CoV), Multimodal Visual Attention, VR Collaborative Analysis, Speech Recognition, Eye Tracking

Research Background and Problem

  1. Identified Problems:

    • In collaborative environments, improving the understanding of others' visual attention is critical for successful collaboration.
    • Existing attention mechanisms based on field of view and head orientation cannot dynamically reflect different types of attention (e.g., focused vs. dispersed attention).
    • Eye-tracking-based visual cues may be limited by device configuration, calibration issues, and can be distracting to observers in certain scenarios.
  2. Significance:

    • With the increasing adoption of virtual reality devices and the demand for remote work during the COVID-19 pandemic, enhancing the accuracy of shared attention and visual cues in collaborative environments is particularly important, especially for low-cost VR devices without eye-tracking capabilities.
  3. Motivation and Related Work:

    • Combining multimodal cues (e.g., head orientation and speech content) to improve the accuracy of visual attention inference and enhance collaboration quality.
    • Previous research has shown that field-of-view visualization can enhance collaboration, but there is still room to explore how to dynamically adjust the field of view size to accommodate speech input.

Solution

  1. Proposed Method:

    • Introduced the "Speech-Augmented Cone-of-Vision (CoV+Speech)," which dynamically adjusts the field of view by combining users' head movements with collaborative speech input to improve attention inference.
    • Utilized speech recognition technology to capture keywords and narrow the field of view cone, focusing on specific areas of data visualization.
  2. Innovations:

    • Integrated speech recognition with multimodal visual attention, enabling the field of view cone to dynamically adapt to users' collaborative speech behavior.
    • Proposed a joint approach combining head orientation and speech semantics for visual attention inference, addressing the limitations of traditional single-cue methods based on head orientation or eye tracking.
  3. Implementation Steps and Key Techniques:

    • Field of View Cone Modeling: Used statistical models to generate the field of view cone based on head orientation.
    • Speech Processing: Captured keywords from users' speech in real-time using speech recognition, searched for keyword locations on an HTML page, and dynamically adjusted the field of view.
    • Dynamic Visual Cue Display: Plotted elliptical regions based on keyword distribution and interpolated to adjust the field of view cone display.

Research Outcomes

  1. Specific Results:

    • Proposed a visual attention cue that dynamically adjusts the field of view cone size, effectively supporting collaboration in multimodal environments.
    • Released a public dataset combining head orientation, eye-tracking data, and collaborative speech data.
  2. Comparison with Existing Solutions:

    • The field of view cone (CoV) demonstrated a 20% improvement in shared attention at the screen level compared to eye-gaze cursors, enhancing collaborative focus on screen-level regions.
    • The speech-augmented solution (CoV+Speech) significantly improved eye-tracking inference accuracy in non-real-time experiments, increasing precision by approximately 50 pixels and outperforming attention inference based solely on head orientation under certain conditions.
  3. Experimental or Evaluation Results:

    • Collected 179 interaction data samples under three experimental conditions (CoV, CoV+Speech, Eye-Gaze Cursor), validating the effectiveness of combining head orientation with speech input.
    • Surveys and interviews further highlighted the advantages of the field of view cone in initial visual alignment and the limitations of speech technology.
  4. Limitations and Future Directions:

    • Limitations: Real-time accuracy and latency issues in speech recognition affected the performance of CoV+Speech. Current technology cannot capture spatial semantic inputs (e.g., "top left corner").
    • Future Directions:
      • Develop hybrid approaches combining head orientation and eye-tracking technologies to better support dynamic attention representation.
      • Expand applications to 3D data, addressing occlusion and real-time environmental changes.
      • Optimize semantic processing technologies to support more complex spatial references in speech.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95852/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581283
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Social & Collaborative VR, Interactive Data Visualization
work
Professions
University Professors & Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers