Speech-Augmented Cone-of-Vision for Exploratory Data Analysis
Authors
Eye Tracking & Gaze InteractionSocial & Collaborative VRInteractive Data VisualizationUniversity Professors & ResearchersHCI Researchers
Document Title
Speech-Augmented Cone-of-Vision for Exploratory Data Analysis
Document Information
- Domain: Human-Computer Interaction, Collaborative Virtual Reality, Data Visualization and Analysis
- Keywords: Field of View (CoV), Multimodal Visual Attention, VR Collaborative Analysis, Speech Recognition, Eye Tracking
Research Background and Problem
-
Identified Problems:
- In collaborative environments, improving the understanding of others' visual attention is critical for successful collaboration.
- Existing attention mechanisms based on field of view and head orientation cannot dynamically reflect different types of attention (e.g., focused vs. dispersed attention).
- Eye-tracking-based visual cues may be limited by device configuration, calibration issues, and can be distracting to observers in certain scenarios.
-
Significance:
- With the increasing adoption of virtual reality devices and the demand for remote work during the COVID-19 pandemic, enhancing the accuracy of shared attention and visual cues in collaborative environments is particularly important, especially for low-cost VR devices without eye-tracking capabilities.
-
Motivation and Related Work:
- Combining multimodal cues (e.g., head orientation and speech content) to improve the accuracy of visual attention inference and enhance collaboration quality.
- Previous research has shown that field-of-view visualization can enhance collaboration, but there is still room to explore how to dynamically adjust the field of view size to accommodate speech input.
Solution
-
Proposed Method:
- Introduced the "Speech-Augmented Cone-of-Vision (CoV+Speech)," which dynamically adjusts the field of view by combining users' head movements with collaborative speech input to improve attention inference.
- Utilized speech recognition technology to capture keywords and narrow the field of view cone, focusing on specific areas of data visualization.
-
Innovations:
- Integrated speech recognition with multimodal visual attention, enabling the field of view cone to dynamically adapt to users' collaborative speech behavior.
- Proposed a joint approach combining head orientation and speech semantics for visual attention inference, addressing the limitations of traditional single-cue methods based on head orientation or eye tracking.
-
Implementation Steps and Key Techniques:
- Field of View Cone Modeling: Used statistical models to generate the field of view cone based on head orientation.
- Speech Processing: Captured keywords from users' speech in real-time using speech recognition, searched for keyword locations on an HTML page, and dynamically adjusted the field of view.
- Dynamic Visual Cue Display: Plotted elliptical regions based on keyword distribution and interpolated to adjust the field of view cone display.
Research Outcomes
-
Specific Results:
- Proposed a visual attention cue that dynamically adjusts the field of view cone size, effectively supporting collaboration in multimodal environments.
- Released a public dataset combining head orientation, eye-tracking data, and collaborative speech data.
-
Comparison with Existing Solutions:
- The field of view cone (CoV) demonstrated a 20% improvement in shared attention at the screen level compared to eye-gaze cursors, enhancing collaborative focus on screen-level regions.
- The speech-augmented solution (CoV+Speech) significantly improved eye-tracking inference accuracy in non-real-time experiments, increasing precision by approximately 50 pixels and outperforming attention inference based solely on head orientation under certain conditions.
-
Experimental or Evaluation Results:
- Collected 179 interaction data samples under three experimental conditions (CoV, CoV+Speech, Eye-Gaze Cursor), validating the effectiveness of combining head orientation with speech input.
- Surveys and interviews further highlighted the advantages of the field of view cone in initial visual alignment and the limitations of speech technology.
-
Limitations and Future Directions:
- Limitations: Real-time accuracy and latency issues in speech recognition affected the performance of CoV+Speech. Current technology cannot capture spatial semantic inputs (e.g., "top left corner").
- Future Directions:
- Develop hybrid approaches combining head orientation and eye-tracking technologies to better support dynamic attention representation.
- Expand applications to 3D data, addressing occlusion and real-time environmental changes.
- Optimize semantic processing technologies to support more complex spatial references in speech.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- In collaborative environments, how can combining head direction and voice input dynamically adjust field-of-view displays to improve attention inference accuracy?Category: XR Information Presentation and VisualizationSimilar questionsarrow_forward
- Can voice-enhanced field-of-view cone displays effectively support multimodal collaborative data visualization tasks?Category: XR Information Presentation and VisualizationSimilar questionsarrow_forward
- How do multimodal methods outperform single head-direction-based visual cues in improving collaborative focus alignment?Category: XR Information Presentation and VisualizationSimilar questionsarrow_forward
lightbulb
Practical Problems
1- On low-cost VR devices, users struggle to accurately perceive others' attention focus.Category: XR Information Presentation and VisualizationSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581283
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Social & Collaborative VR, Interactive Data Visualization
work
Professions
University Professors & Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers