SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research

Eye Tracking & Gaze InteractionHuman Pose & Activity RecognitionImmersion & Presence ResearchField StudiesMuseum Curators & ArchivistsHCI ResearchersSociologists & Anthropologists

Paper Title

SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research

Publication Info

  • Topic area: Embodied cognition and spatial visual attention analysis in real-world environments.
  • Keywords: Spatial visual attention, embodied cognition, 3D gaze tracking, HoloLens 2, multimodal analysis, Average Focus Weight, visualization, museum studies, ecological validity, open-source platform.

Background and Problem

  • Problem / challenge: Current mobile eye-tracking platforms focus on 2D, first-person perspectives and lack integrated workflows for 3D spatial attention analysis. This limits the ability to analyze visual attention in embodied tasks within complex environments.
  • Significance: Understanding spatial visual attention in 3D contexts is critical for embodied cognition research and practical applications such as museum design, education, and navigation.
  • Motivation and related work: Previous research has explored 3D eye tracking and visualization but often focuses on isolated components rather than end-to-end solutions. Existing tools lack standardized frameworks for representing and analyzing spatial attention data, making workflows labor-intensive and difficult to scale.

Solution

  • Proposed approach: SVATA (Spatial Visual Attention Tracking and Analysis), an open-source platform that provides an end-to-end workflow for collecting, analyzing, and visualizing 3D spatial attention data in embodied tasks.
  • Novelty:
    1. Integration of multimodal signals (gaze, head pose, position) with world-referenced mapping onto reconstructed geometry.
    2. Introduction of the Average Focus Weight (AFW/m²) metric for quantifying spatial visual attention.
    3. Modular architecture supporting structured analysis and multidimensional visualization of spatial, behavioral, and temporal patterns.
    4. Deployment flexibility using HoloLens 2 and OpenXR for mobile sensing and desktop-based analysis.
  • Procedure and key techniques:
    • Spatial modeling and gaze-to-scene localization using ray-mesh intersections.
    • Physiologically-informed Focus Weight model for representing attention distributions.
    • Structured analysis modules for spatial, behavioral, and temporal dimensions.
    • Interactive visualization tools for heatmaps, trajectories, and attention evolution.

Results

  • Concrete findings:
    • SVATA achieved a 97.5% success rate in capturing valid datasets (78 out of 80 sessions) during a museum deployment.
    • Data completeness was 98.46%, and gaze ray validity was 97.07%.
    • Temporal analysis revealed a 43% decline in attention density across sequential exhibition units.
    • Spatial analysis differentiated engagement levels across media types, with interactive installations attracting higher attention (5928.95 AFW/m²) than static images (1953.49 AFW/m²).
  • Advantage over baselines:
    • Compared to existing systems, SVATA provides unified 3D spatial alignment, multimodal integration, and structured analysis, addressing gaps in scalability and interpretability.
  • Experiments / evaluation:
    • Study 1: In-the-wild deployment with 78 participants in two museum sites (guided-route and free exploration conditions).
    • Study 2: Expert validation with 7 domain professionals, achieving a System Usability Scale (SUS) score of 77.1 (rated "Good").
  • Limitations and future work:
    • Hardware intrusiveness may affect naturalistic behavior; future iterations could use lighter sensors.
    • Limited task and context coverage; future validation across diverse environments and tasks is needed.
    • No independent validation of HoloLens 2 tracking precision; reliance on prior evaluations.

Summary

SVATA is an open-source platform designed to analyze spatial visual attention in embodied cognition tasks. It integrates multimodal signals with world-referenced 3D mapping, introduces the Average Focus Weight metric for quantifying attention distributions, and provides structured analysis and visualization tools. Evaluated through a museum deployment (N = 78) and expert validation (N = 7), SVATA demonstrated feasibility for capturing and analyzing attention data in real-world environments. Experts rated the system highly for its intuitive 3D mapping and actionable insights, though hardware intrusiveness and scalability remain areas for improvement. SVATA offers a robust toolkit for researchers and practitioners to study embodied visual behavior and inform design decisions in naturalistic settings.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222153/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791192
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Human Pose & Activity Recognition, Immersion & Presence Research, Field Studies
work
Professions
Museum Curators & Archivists, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers