VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails

Generative AI (Text, Image, Music, Video)Music Composition & Sound Design ToolsVideo Production & EditingCreative Collaboration & Feedback SystemsContent Creators (YouTubers, Podcasters)Musicians, DJs & Sound DesignersFilm & Animation Producers

Paper Title

VidTune: Creating Video Soundtracks with Generative Music and Video-Based Thumbnails

Publication Info

  • Topic area: Generative AI for video soundtrack creation
  • Keywords: generative music, video editing, text-to-music, music visualization, contextual thumbnails, accessibility, creative AI, music diversity, iterative refinement, user-centered design

Background and Problem

  • Problem / challenge: Current text-to-music workflows struggle to meet the needs of video creators. Challenges include difficulty in crafting effective prompts, time-consuming review of generated tracks, and lack of accessible tools for creators with limited hearing.
  • Significance: Soundtracks are essential for shaping video narratives and audience experiences. Efficient, accessible, and creative tools for soundtrack generation can significantly enhance video production workflows.
  • Motivation and related work: Prior systems focus on generating music from text prompts but lack support for exploring diverse options, reviewing tracks in context, and iteratively refining outputs. Existing tools often rely on waveforms or generic thumbnails, which are insufficient for non-expert users or those with hearing impairments. This paper addresses these gaps by introducing VidTune.

Solution

  • Proposed approach: VidTune, an interactive system that generates video soundtracks using text-to-music models, contextual thumbnails, and iterative refinement tools.
  • Novelty:
    1. Introduction of contextual thumbnails that visually summarize music attributes in the context of the user’s video.
    2. A prompt expansion algorithm to generate diverse music options from user input.
    3. Iterative refinement tools, including natural language edits, variation, and blending of tracks.
    4. A music map for similarity-based exploration and organization of generated tracks.
  • Procedure and key techniques:
    • Prompt suggestions: VidTune analyzes video scenes to suggest keywords for music generation.
    • Prompt expansion: Generates diverse variations of user prompts using a large multimodal model.
    • Contextual thumbnails: Maps musical attributes (e.g., genre, tempo, mood) onto visual elements anchored in the video.
    • Iterative refinement: Allows users to edit, vary, or blend tracks using natural language instructions.
    • Music map: Projects tracks into a 2D space based on audio similarity for exploration and comparison.

Results

  • Concrete findings:
    • VidTune’s thumbnails scored higher in representing music (mean rating: 4.99 vs. 4.63) and aiding track selection (88.2% accuracy vs. 69.8% for baseline).
    • Prompt expansion significantly increased music diversity for short prompts (e.g., cluster separation: 23.77 vs. 20.27, p < 0.05).
    • VidTune reduced temporal demand and increased enjoyment in user studies (e.g., satisfaction: 6.42 vs. 5.42, p < 0.05).
  • Advantage over baselines:
    • VidTune outperformed a baseline text-to-music interface in helping users understand, compare, and remember music tracks.
    • Contextual thumbnails provided more accurate and memorable representations of music compared to generic thumbnails.
  • Experiments / evaluation:
    • Technical evaluation: Assessed diversity of music from expanded prompts and music-thumbnail correspondence using 1,600 judgments.
    • Controlled user study: Compared VidTune to a baseline with 12 participants, measuring cognitive load, creativity support, and thumbnail utility.
    • Exploratory case study: Tested VidTune with 6 creators using their own videos, highlighting its accessibility and personalization benefits.
  • Limitations and future work:
    • Limited support for vocal tracks and fine-grained music editing.
    • Occasional mismatches between thumbnails and music titles.
    • Future work could explore beat-level synchronization, multimodal support for lyrics, and personalized recommendations.

Summary

VidTune is a generative AI system designed to streamline video soundtrack creation by combining text-to-music models with contextual thumbnails and iterative refinement tools. It enables users to explore diverse music options, visually compare tracks, and refine outputs efficiently. Technical evaluations and user studies demonstrate that VidTune improves music diversity, reduces review burden, and enhances user enjoyment compared to baseline tools. By making music more visual and accessible, VidTune supports a wide range of creators, including those with limited hearing, and fosters a more engaging and creative soundtrack generation process.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222873/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791572
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Music Composition & Sound Design Tools, Video Production & Editing, Creative Collaboration & Feedback Systems
work
Professions
Content Creators (YouTubers, Podcasters), Musicians, DJs & Sound Designers, Film & Animation Producers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers