SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research
Authors
Paper Title
SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research
Publication Info
- Topic area: Embodied cognition and spatial visual attention analysis in real-world environments.
- Keywords: Spatial visual attention, embodied cognition, 3D gaze tracking, HoloLens 2, multimodal analysis, Average Focus Weight, visualization, museum studies, ecological validity, open-source platform.
Background and Problem
- Problem / challenge: Current mobile eye-tracking platforms focus on 2D, first-person perspectives and lack integrated workflows for 3D spatial attention analysis. This limits the ability to analyze visual attention in embodied tasks within complex environments.
- Significance: Understanding spatial visual attention in 3D contexts is critical for embodied cognition research and practical applications such as museum design, education, and navigation.
- Motivation and related work: Previous research has explored 3D eye tracking and visualization but often focuses on isolated components rather than end-to-end solutions. Existing tools lack standardized frameworks for representing and analyzing spatial attention data, making workflows labor-intensive and difficult to scale.
Solution
- Proposed approach: SVATA (Spatial Visual Attention Tracking and Analysis), an open-source platform that provides an end-to-end workflow for collecting, analyzing, and visualizing 3D spatial attention data in embodied tasks.
- Novelty:
- Integration of multimodal signals (gaze, head pose, position) with world-referenced mapping onto reconstructed geometry.
- Introduction of the Average Focus Weight (AFW/m²) metric for quantifying spatial visual attention.
- Modular architecture supporting structured analysis and multidimensional visualization of spatial, behavioral, and temporal patterns.
- Deployment flexibility using HoloLens 2 and OpenXR for mobile sensing and desktop-based analysis.
- Procedure and key techniques:
- Spatial modeling and gaze-to-scene localization using ray-mesh intersections.
- Physiologically-informed Focus Weight model for representing attention distributions.
- Structured analysis modules for spatial, behavioral, and temporal dimensions.
- Interactive visualization tools for heatmaps, trajectories, and attention evolution.
Results
- Concrete findings:
- SVATA achieved a 97.5% success rate in capturing valid datasets (78 out of 80 sessions) during a museum deployment.
- Data completeness was 98.46%, and gaze ray validity was 97.07%.
- Temporal analysis revealed a 43% decline in attention density across sequential exhibition units.
- Spatial analysis differentiated engagement levels across media types, with interactive installations attracting higher attention (5928.95 AFW/m²) than static images (1953.49 AFW/m²).
- Advantage over baselines:
- Compared to existing systems, SVATA provides unified 3D spatial alignment, multimodal integration, and structured analysis, addressing gaps in scalability and interpretability.
- Experiments / evaluation:
- Study 1: In-the-wild deployment with 78 participants in two museum sites (guided-route and free exploration conditions).
- Study 2: Expert validation with 7 domain professionals, achieving a System Usability Scale (SUS) score of 77.1 (rated "Good").
- Limitations and future work:
- Hardware intrusiveness may affect naturalistic behavior; future iterations could use lighter sensors.
- Limited task and context coverage; future validation across diverse environments and tasks is needed.
- No independent validation of HoloLens 2 tracking precision; reliance on prior evaluations.
Summary
SVATA is an open-source platform designed to analyze spatial visual attention in embodied cognition tasks. It integrates multimodal signals with world-referenced 3D mapping, introduces the Average Focus Weight metric for quantifying attention distributions, and provides structured analysis and visualization tools. Evaluated through a museum deployment (N = 78) and expert validation (N = 7), SVATA demonstrated feasibility for capturing and analyzing attention data in real-world environments. Experts rated the system highly for its intuitive 3D mapping and actionable insights, though hardware intrusiveness and scalability remain areas for improvement. SVATA offers a robust toolkit for researchers and practitioners to study embodied visual behavior and inform design decisions in naturalistic settings.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)