MoSound: An Interactive Tool for Generative Sound Design in Motion Graphics

Honorable Mention
Music Composition & Sound Design Tools3D Modeling & AnimationCreative Collaboration & Feedback SystemsMusicians, DJs & Sound DesignersUI/UX DesignersGame Developers & Designers

Paper Title

MoSound: An Interactive Tool for Generative Sound Design in Motion Graphics

Publication Info

  • Topic area: Generative sound design for motion graphics
  • Keywords: Motion graphics, sound design, generative AI, motion-to-sound mapping, interactive tools, sound synthesis, event detection, user study, motion tracking, creative workflows

Background and Problem

  • Problem / challenge: Existing tools for motion graphics focus on visual design, neglecting sound design, which is critical for enhancing motion graphics. Current generative sound techniques lack precise control and synchronization with visual events, making them unsuitable for abstract motion graphics.
  • Significance: Sound design enhances the emotional impact, memorability, and overall quality of motion graphics. However, creating synchronized and expressive sound effects is tedious and requires expertise, leaving novices at a disadvantage.
  • Motivation and related work: Prior works on motion graphics tools focus on visual aspects, while sound design tools like Foley and generative AI lack the temporal precision and flexibility needed for motion graphics. MoSound addresses this gap by integrating motion-driven sound synthesis with user control.

Solution

  • Proposed approach: MoSound, an interactive system that combines visual event detection, motion-to-sound mapping, and generative sound synthesis to create synchronized sound effects for motion graphics.
  • Novelty:
    1. A human-in-the-loop workflow for sound design in motion graphics, combining automation and user control.
    2. Motion-to-sound mapping that aligns visual events with generative sound effects.
    3. Insights from a user study showing MoSound’s accessibility for novices and utility as a prototyping tool for experts.
  • Procedure and key techniques:
    • Upload a motion graphics video.
    • Automatically detect visual events using a vision-language model (VLM).
    • Map motion features (e.g., position, velocity) to sound properties (e.g., volume, panning) to create a guide sound.
    • Use Sketch2Sound for generative sound synthesis based on guide sounds and textual descriptions.
    • Allow users to refine events, adjust mappings, and compose layered soundtracks.

Results

  • Concrete findings:
    • Event detection takes ~5 seconds per second of video; motion tracking takes ~9 seconds per second of video; sound synthesis takes ~10 seconds for four candidates per second of video.
    • Median onset deviation between guide and final sounds: 0.032 seconds; envelope correlation: 0.536 ± 0.332; panning correlation: 0.731 ± 0.191.
    • Semantic similarity of VLM-generated events: 0.505 ± 0.141; temporal IoU: 0.337 ± 0.128; center alignment error: 0.660 ± 0.247 seconds.
  • Advantage over baselines:
    • Outperforms generative sound techniques (e.g., FoleyCrafter, MMAudio) in producing coherent, synchronized sound effects for abstract motion graphics.
    • Provides automatic event detection and motion-driven sound synthesis, which competing methods lack.
  • Experiments / evaluation:
    • User study with 7 participants (3 experts, 4 novices): Average time to create sound effects was 9 minutes for a tutorial video and 21 minutes for a free-choice video.
    • Experts estimated MoSound reduced sound design time from 30 minutes to 5–10 minutes for short clips.
    • Quantitative Likert-scale feedback showed satisfaction with sound quality, motion tracking, and synchronization.
  • Limitations and future work:
    • Limited support for continuous textures and ambient layers.
    • Lacks fine-grained controls for layering, mixing, and timeline precision.
    • Event detection is semantically meaningful but temporally coarse.
    • Future work includes richer motion parameters, multimodal inputs, and better integration with professional workflows.

Summary

MoSound is an interactive tool that automates and enhances sound design for motion graphics by combining visual event detection, motion-to-sound mapping, and generative sound synthesis. It lowers the barrier for novices while providing experts with a prototyping tool. User studies and evaluations demonstrate its effectiveness in creating synchronized, high-quality sound effects, though improvements are needed for layering, continuous textures, and professional-grade controls. MoSound bridges the gap between AI-assisted automation and creative workflows, offering a foundation for future advancements in generative sound design.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222174/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791162
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Music Composition & Sound Design Tools, 3D Modeling & Animation, Creative Collaboration & Feedback Systems
work
Professions
Musicians, DJs & Sound Designers, UI/UX Designers, Game Developers & Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers