MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR

Immersion & Presence ResearchSpatial Audio & 3D SoundGame Developers & DesignersMusicians, DJs & Sound Designers

Paper Title

MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR

Publication Info

  • Topic area: Real-time audio-visual sound separation and interaction in Extended Reality (XR).
  • Keywords: XR, audio-visual separation, sound interaction, real-time processing, cognitive load, speech intelligibility, cascaded architecture, user study, sound source localization.

Background and Problem

  • Problem / challenge: XR systems lack fine-grained, real-time control over auditory environments, leaving users overwhelmed by complex soundscapes. Existing sound separation methods are either computationally expensive, offline, or fail to provide interactive control.
  • Significance: Addressing this limitation can enhance scene awareness, social engagement, and user experience in XR, particularly in noisy or multi-source environments.
  • Motivation and related work: Prior work in sound separation has focused on offline processing or limited real-time capabilities, often without leveraging visual cues. HCI research has explored global sound filtering but lacks object-centric, interactive control. Machine learning approaches achieve high separation quality but are unsuitable for real-time XR due to latency and computational constraints.

Solution

  • Proposed approach: MoXaRt, a real-time XR system that integrates audio-visual cues for sound source separation and interactive soundscape control.
  • Novelty:
    1. Introduction of a cascaded audio-visual transformer model for real-time, object-guided sound separation.
    2. Development of a user interface for interactive volume control of individual sound sources.
    3. Creation of a new dataset featuring complex audio-visual scenarios for evaluation.
    4. Demonstration of significant improvements in speech intelligibility, cognitive load reduction, and user experience.
  • Procedure and key techniques:
    • Coarse audio-only separation into general categories (speech, music, noise).
    • Visual detection of faces and instruments to guide refinement networks.
    • Cascaded architecture for fine-grained separation of individual speakers and instruments.
    • Real-time processing pipeline with a latency of ~2 seconds, leveraging external PC computation and wireless data transmission.

Results

  • Concrete findings:
    • MoXaRt achieves a 36.2% improvement in listening comprehension (p = 0.0058) and significantly reduces cognitive load (p < 0.001).
    • Real-time model achieves a Word Error Rate (WER) of 0.4990, outperforming its architectural baseline (AudioScopeV2, WER = 0.5263).
    • DNSMOS scores indicate competitive perceptual audio quality.
  • Advantage over baselines:
    • Outperforms state-of-the-art models like AV-Mossformer2 and AudioScopeV2 in intelligibility and interactive capabilities.
    • Enables real-time, user-driven soundscape remixing, unlike passive or offline separation systems.
  • Experiments / evaluation:
    • Technical evaluation on a custom dataset of 30 recordings with up to five concurrent sources.
    • User study (N=22) across six XR scenarios, showing significant objective and subjective benefits.
  • Limitations and future work:
    • Visual dependency limits performance when sources are occluded or out of view.
    • Current implementation requires a tethered PC; optimization is needed for standalone XR devices.
    • Challenges in scaling to environments with more than four concurrent speakers or dynamic source tracking.

Summary

MoXaRt introduces a novel real-time system for audio-visual sound separation and interaction in XR, leveraging a cascaded architecture to isolate and remix individual sound sources. It significantly improves speech intelligibility and reduces cognitive load, validated through technical benchmarks and a 22-participant user study. While currently reliant on external computation and visual cues, MoXaRt demonstrates the potential for transforming auditory experiences in XR, with applications ranging from social interaction to AI-assisted tasks. Future work will focus on standalone deployment, scalability, and ethical considerations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222533/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791929
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Immersion & Presence Research, Spatial Audio & 3D Sound
work
Professions
Game Developers & Designers, Musicians, DJs & Sound Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers