MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Authors
Paper Title
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Publication Info
- Topic area: Real-time audio-visual sound separation and interaction in Extended Reality (XR).
- Keywords: XR, audio-visual separation, sound interaction, real-time processing, cognitive load, speech intelligibility, cascaded architecture, user study, sound source localization.
Background and Problem
- Problem / challenge: XR systems lack fine-grained, real-time control over auditory environments, leaving users overwhelmed by complex soundscapes. Existing sound separation methods are either computationally expensive, offline, or fail to provide interactive control.
- Significance: Addressing this limitation can enhance scene awareness, social engagement, and user experience in XR, particularly in noisy or multi-source environments.
- Motivation and related work: Prior work in sound separation has focused on offline processing or limited real-time capabilities, often without leveraging visual cues. HCI research has explored global sound filtering but lacks object-centric, interactive control. Machine learning approaches achieve high separation quality but are unsuitable for real-time XR due to latency and computational constraints.
Solution
- Proposed approach: MoXaRt, a real-time XR system that integrates audio-visual cues for sound source separation and interactive soundscape control.
- Novelty:
- Introduction of a cascaded audio-visual transformer model for real-time, object-guided sound separation.
- Development of a user interface for interactive volume control of individual sound sources.
- Creation of a new dataset featuring complex audio-visual scenarios for evaluation.
- Demonstration of significant improvements in speech intelligibility, cognitive load reduction, and user experience.
- Procedure and key techniques:
- Coarse audio-only separation into general categories (speech, music, noise).
- Visual detection of faces and instruments to guide refinement networks.
- Cascaded architecture for fine-grained separation of individual speakers and instruments.
- Real-time processing pipeline with a latency of ~2 seconds, leveraging external PC computation and wireless data transmission.
Results
- Concrete findings:
- MoXaRt achieves a 36.2% improvement in listening comprehension (p = 0.0058) and significantly reduces cognitive load (p < 0.001).
- Real-time model achieves a Word Error Rate (WER) of 0.4990, outperforming its architectural baseline (AudioScopeV2, WER = 0.5263).
- DNSMOS scores indicate competitive perceptual audio quality.
- Advantage over baselines:
- Outperforms state-of-the-art models like AV-Mossformer2 and AudioScopeV2 in intelligibility and interactive capabilities.
- Enables real-time, user-driven soundscape remixing, unlike passive or offline separation systems.
- Experiments / evaluation:
- Technical evaluation on a custom dataset of 30 recordings with up to five concurrent sources.
- User study (N=22) across six XR scenarios, showing significant objective and subjective benefits.
- Limitations and future work:
- Visual dependency limits performance when sources are occluded or out of view.
- Current implementation requires a tethered PC; optimization is needed for standalone XR devices.
- Challenges in scaling to environments with more than four concurrent speakers or dynamic source tracking.
Summary
MoXaRt introduces a novel real-time system for audio-visual sound separation and interaction in XR, leveraging a cascaded architecture to isolate and remix individual sound sources. It significantly improves speech intelligibility and reduces cognitive load, validated through technical benchmarks and a 22-participant user study. While currently reliant on external computation and visual cues, MoXaRt demonstrates the potential for transforming auditory experiences in XR, with applications ranging from social interaction to AI-assisted tasks. Future work will focus on standalone deployment, scalability, and ethical considerations.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)