SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video

Generative AI (Text, Image, Music, Video)Music Composition & Sound Design ToolsCreative Collaboration & Feedback SystemsMusicians, DJs & Sound DesignersFilm & Animation ProducersUI/UX Designers

Paper Title

SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video

Publication Info

  • Topic area: AI-assisted sound design for video storytelling
  • Keywords: sound design, generative AI, video editing, soundscapes, text-to-SFX, multimodal AI, narrative analysis, creative tools, user control, iterative refinement

Background and Problem

  • Problem / challenge: Crafting effective soundscapes for video is complex and time-consuming, requiring expertise in sourcing, layering, and mixing sounds to align with narrative intent. Existing AI tools like text-to-SFX and video-to-audio systems often lack user control, produce pre-mixed outputs, and fail to account for narrative context.
  • Significance: Soundscapes are critical for storytelling, enhancing mood, emotion, and audience engagement. Making sound design accessible to non-experts can democratize high-quality video production and reduce reliance on stock audio.
  • Motivation and related work: Prior work includes text-to-audio and video-to-audio generation, but these tools often limit iterative refinement and user control. Research on sound design workflows and layering principles highlights the need for tools that integrate narrative analysis with flexible, editable soundscapes.

Solution

  • Proposed approach: SoundStager, an AI-assisted tool that generates story-driven soundscapes for video by analyzing narrative and visual features, offering editable sound layers and iterative refinement through conversational and manual controls.
  • Novelty:
    1. Combines narrative-driven video analysis with generative sound design.
    2. Introduces editable sound layers categorized by keynote, signal, soundmark, and archetypal sounds.
    3. Supports iterative refinement via chat-based and manual controls.
    4. Integrates sound generation, mixing, and visualization into a single workflow.
  • Procedure and key techniques:
    • Video Analysis: Hierarchical segmentation of video into cuts, scenes, and narrative themes, capturing metadata like emotional tone and character actions.
    • Sound Recommendation: Generates sound prompts for four soundscape categories, aligning with narrative and visual cues.
    • Mixing Recommendation: Provides initial mixing parameters (e.g., loudness, panning, fades) based on scene context and emotional tone.
    • Sound Generation: Uses a text-to-SFX model (Sketch2Sound) to produce multiple sound variations for user selection.
    • Interactive Interface: Combines timeline-based editing, chat-guided iteration, and visualization views (Stem, Mix, Focus) for flexible control.

Results

  • Concrete findings:
    • SoundStager significantly improved user ratings for control (M=4.58 vs. 2.92, p=0.006), narrative alignment (M=4.50 vs. 3.33, p=0.004), and efficiency (M=4.50 vs. 2.83, p=0.011) compared to baseline tools.
    • Users reported higher trust in AI outputs (M=3.92 vs. 3.08, p=0.023) and stronger intention to use SoundStager in future projects (M=4.67 vs. 3.50, p=0.006).
  • Advantage over baselines:
    • Streamlined workflows by integrating sound generation, mixing, and editing into one platform.
    • Greater creative control through editable layers, scene-based selection, and sound variations.
    • Context-aware sound recommendations aligned with narrative intent.
  • Experiments / evaluation:
    • Formative studies with 6 professional sound designers and 6 novice video creators informed the design.
    • A within-subjects study with 12 video creators compared SoundStager to a baseline toolset (Gemini, Adobe Firefly, Premiere Pro).
    • Metrics included Technology Acceptance Model (TAM) ratings, design guideline adherence, and qualitative feedback.
  • Limitations and future work:
    • Struggles with fine-grained temporal alignment and abstract video genres.
    • Overwhelming timelines for fast-cut videos.
    • Limited mixing capabilities compared to professional tools.
    • Future directions include editable sound prompts, sound-to-sound generation, palette-style previews, enhanced mixing support, and integration with existing editors.

Summary

SoundStager is an AI-assisted tool for creating story-driven soundscapes, combining narrative analysis, generative sound design, and iterative refinement through chat and manual controls. It significantly improves user control, efficiency, and narrative alignment compared to baseline tools. While effective for structured videos, challenges remain in handling abstract content, fine-grained alignment, and advanced mixing. Future work will focus on enhancing user customization, sound variation, and professional integration, making sound design more accessible and creatively empowering.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222160/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790870
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Music Composition & Sound Design Tools, Creative Collaboration & Feedback Systems
work
Professions
Musicians, DJs & Sound Designers, Film & Animation Producers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers