VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails
Authors
Paper Title
VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails
Publication Info
- Topic area: Generative AI for video soundtrack creation
- Keywords: generative music, video editing, text-to-music, music visualization, contextual thumbnails, accessibility, creative AI, music diversity, iterative refinement, user-centered design
Background and Problem
- Problem / challenge: Current text-to-music workflows struggle to meet the needs of video creators. Challenges include difficulty in crafting effective prompts, time-consuming review of generated tracks, and lack of accessible tools for creators with limited hearing.
- Significance: Soundtracks are essential for shaping video narratives and audience experiences. Efficient, accessible, and creative tools for soundtrack generation can significantly enhance video production workflows.
- Motivation and related work: Prior systems focus on generating music from text prompts but lack support for exploring diverse options, reviewing tracks in context, and iteratively refining outputs. Existing tools often rely on waveforms or generic thumbnails, which are insufficient for non-expert users or those with hearing impairments. This paper addresses these gaps by introducing VidTune.
Solution
- Proposed approach: VidTune, an interactive system that generates video soundtracks using text-to-music models, contextual thumbnails, and iterative refinement tools.
- Novelty:
- Introduction of contextual thumbnails that visually summarize music attributes in the context of the user’s video.
- A prompt expansion algorithm to generate diverse music options from user input.
- Iterative refinement tools, including natural language edits, variation, and blending of tracks.
- A music map for similarity-based exploration and organization of generated tracks.
- Procedure and key techniques:
- Prompt suggestions: VidTune analyzes video scenes to suggest keywords for music generation.
- Prompt expansion: Generates diverse variations of user prompts using a large multimodal model.
- Contextual thumbnails: Maps musical attributes (e.g., genre, tempo, mood) onto visual elements anchored in the video.
- Iterative refinement: Allows users to edit, vary, or blend tracks using natural language instructions.
- Music map: Projects tracks into a 2D space based on audio similarity for exploration and comparison.
Results
- Concrete findings:
- VidTune’s thumbnails scored higher in representing music (mean rating: 4.99 vs. 4.63) and aiding track selection (88.2% accuracy vs. 69.8% for baseline).
- Prompt expansion significantly increased music diversity for short prompts (e.g., cluster separation: 23.77 vs. 20.27, p < 0.05).
- VidTune reduced temporal demand and increased enjoyment in user studies (e.g., satisfaction: 6.42 vs. 5.42, p < 0.05).
- Advantage over baselines:
- VidTune outperformed a baseline text-to-music interface in helping users understand, compare, and remember music tracks.
- Contextual thumbnails provided more accurate and memorable representations of music compared to generic thumbnails.
- Experiments / evaluation:
- Technical evaluation: Assessed diversity of music from expanded prompts and music-thumbnail correspondence using 1,600 judgments.
- Controlled user study: Compared VidTune to a baseline with 12 participants, measuring cognitive load, creativity support, and thumbnail utility.
- Exploratory case study: Tested VidTune with 6 creators using their own videos, highlighting its accessibility and personalization benefits.
- Limitations and future work:
- Limited support for vocal tracks and fine-grained music editing.
- Occasional mismatches between thumbnails and music titles.
- Future work could explore beat-level synchronization, multimodal support for lyrics, and personalized recommendations.
Summary
VidTune is a generative AI system designed to streamline video soundtrack creation by combining text-to-music models with contextual thumbnails and iterative refinement tools. It enables users to explore diverse music options, visually compare tracks, and refine outputs efficiently. Technical evaluations and user studies demonstrate that VidTune improves music diversity, reduces review burden, and enhances user enjoyment compared to baseline tools. By making music more visual and accessible, VidTune supports a wide range of creators, including those with limited hearing, and fosters a more engaging and creative soundtrack generation process.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
“It’s more of a vibe I’m going for”: Designing Text-to-Music Generation Interfaces for Video Creators
DIS '25· Generative AI (Text, Image, Music, Video) +3
- 86%
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
C&C '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 71%
Automated Conversion of Music Videos into Lyric Videos
UIST '23· Music Composition & Sound Design Tools +1
- 63%
Vidmento: Scaffolded Expansion for Video Storytelling with Generative Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
Reflection Across AI-based Music Composition
C&C '24· Generative AI (Text, Image, Music, Video) +2
- 63%
Rewriting Video: Text-Driven Reauthoring of Video Footage
IUI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
Soundify: Matching Sound Effects to Video
UIST '23· Music Composition & Sound Design Tools +2
Based on Jaccard similarity of research subtopics & professions (≥60%)