Simulating Human Audiovisual Search Behavior

Eye Tracking & Gaze InteractionSonification & Auditory DisplayAffective Feedback & Emotion Regulation InterfacesUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Simulating Human Audiovisual Search Behavior

Publication Info

  • Topic area: Computational modeling of human audiovisual search behavior in immersive environments.
  • Keywords: Audiovisual search, resource rationality, POMDP, reinforcement learning, embodied cognition, human-computer interaction, virtual reality, multisensory integration, decision-making, simulation.

Background and Problem

  • Problem / challenge: Existing models of audiovisual search often treat perception and action in isolation, failing to capture how humans adaptively coordinate sensory and physical strategies under uncertainty.
  • Significance: Understanding and simulating human audiovisual search behavior can improve the design of interfaces and systems in immersive environments, such as VR, MR, and assistive technologies.
  • Motivation and related work: Prior work includes Bayesian models for sensory integration and active sensing frameworks, but these often neglect embodied actions or focus on single-modality tasks. This paper addresses the gap by integrating sensory inference with resource-rational decision-making in a unified framework.

Solution

  • Proposed approach: Sensonaut, a computational model of embodied audiovisual search, formalized as a Partially Observable Markov Decision Process (POMDP) and trained using reinforcement learning.
  • Novelty:
    1. A resource-rational model that unifies cue integration with decision-making under embodied action costs.
    2. Implementation of the model as a POMDP, simulating human-like search strategies and errors.
    3. A new dataset of human audiovisual search behavior in VR, systematically varying task complexity.
  • Procedure and key techniques:
    • Sensonaut integrates auditory and visual cues to update a belief distribution over target locations.
    • The model uses a resource-rational policy to balance information gain against physical costs (e.g., head turns, locomotion).
    • It employs reinforcement learning (PPO) to approximate optimal strategies under uncertainty.
    • Validation was conducted using a VR study where participants searched for sound-emitting targets under controlled conditions.

Results

  • Concrete findings:
    • Sensonaut reproduced key human behaviors, including high accuracy (94.7%), longer search times with distractors, and reliance on head turns over locomotion.
    • The model explained 35% of variance in accuracy, 26% in search time, 58% in head turns, and 6% in displacement.
    • It also replicated human error modes, such as occlusion-driven errors (35%) and confusion by distractors (33% combined).
  • Advantage over baselines:
    • Unlike prior models, Sensonaut captures the embodied dynamics of audiovisual search, including the trade-offs between accuracy, time, and effort.
    • It predicts not only success rates but also emergent search trajectories and characteristic human errors.
  • Experiments / evaluation:
    • A VR study with 12 participants (270 trials) systematically varied target angle, number of objects, and distractors.
    • Dependent variables included accuracy, search time, head turns, and displacement.
    • Sensonaut was evaluated on the same maps, with results compared to human data.
  • Limitations and future work:
    • The model lacks fine-grained embodied dynamics (e.g., continuous movement) and proximity bias observed in humans.
    • Visual likelihoods are treated as binary, leading to premature certainty in some cases.
    • Future work includes expanding the dataset, incorporating individual differences, and refining the action space.

Summary

This paper introduces Sensonaut, a computational model of embodied audiovisual search that integrates multisensory perception and resource-rational decision-making. Validated against human data from a VR study, the model reproduces key patterns of human behavior, including search strategies, accuracy, and error modes. Sensonaut offers a principled framework for simulating and understanding audiovisual search, with applications in XR, assistive systems, and interactive design. While the model aligns well with human behavior in most metrics, future improvements are needed to address discrepancies in locomotion and visual processing.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222686/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790614
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Sonification & Auditory Display, Affective Feedback & Emotion Regulation Interfaces
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers