MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
Authors
Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We present MIMOSA, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, MIMOSA automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix the errors in the locations of sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of MIMOSA exemplifies a human-AI collaboration approach that, instead of utilizing state-of-art end-to-end "black-box" ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user’s workflow. A lab user study with 15 participants demonstrates MIMOSA’s usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 86%
“It’s more of a vibe I’m going for”: Designing Text-to-Music Generation Interfaces for Video Creators
DIS '25· Generative AI (Text, Image, Music, Video) +3
- 83%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Reflection Across AI-based Music Composition
C&C '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Investigating Slowness as a Frame to Design Longer-Term Experiences with Personal Data: A Field Study of Olly
CHI '19· Generative AI (Text, Image, Music, Video) +1
- 67%
Music Creation by Example
CHI '20· Generative AI (Text, Image, Music, Video) +1
- 67%
The Sound Sketchpad: Expressively Combining Large and Diverse Audio Collections
IUI '21· Music Composition & Sound Design Tools +1
- 67%
SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and Exploration
IUI '24· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)