AMUSE: Human-AI Collaborative Songwriting with Multimodal Inspirations
Best PaperAuthors
Generative AI (Text, Image, Music, Video)Music Composition & Sound Design Tools3D Modeling & AnimationMusicians, DJs & Sound DesignersFilm & Animation Producers
Research Background and Issues
What problems or challenges did the authors identify?
-
Lack of multimodal inspiration and creative support:
- The authors point out that songwriting is often driven by multimodal inspiration (e.g., images, narratives, or music), but existing music AI systems fail to adequately support creators in integrating these multimodal inputs into the creative process.
- Currently, generated music is typically non-editable audio files rather than symbolic music elements (e.g., chords or melodies) that can be reused in iterative creation.
- Most symbolic music generation systems only provide suggestions based on existing musical context, lacking support for multimodal inspiration.
-
Lack of training data:
- There is a shortage of training data that aligns multimodal inputs with symbolic music elements, hindering the development of generation models based on multimodal inputs.
Why is this issue important?
- Gap in inspiration transformation: Creators are inspired by multimodal inputs but are unable to efficiently translate them into concrete musical elements, potentially limiting creativity and efficiency in the creative process.
- Lack of creative control: Music creators need editable, modular tools to maintain ownership and flexibility over large-scale compositions, rather than relying on uncontrollable, monolithic generated outputs.
Research Motivation and Related Work
- Multimodal generation systems have been applied in artistic creation (e.g., illustration, writing), and the music domain is also exploring text-based music generation.
- However, most music generation tools either output non-editable audio files or lack integration of multimodal inspiration.
- This study addresses the gap in symbolic music generation tools by supporting multimodal inspiration while enabling iterative control by creators.
Solution
What methods or solutions did the authors propose?
-
Design of the Amuse Tool:
- Developed a collaborative songwriting assistant called Amuse, which integrates users' multimodal inputs (e.g., images, text, audio) to generate reusable chord progressions.
- Includes two main functional modules:
- Chord Generator: Extracts musical keywords from users' multimodal inputs and generates chord progressions that align with the keywords' style.
- Chord Transcriber: Transcribes chord progressions from audio.
-
Innovative generation methods:
- Multimodal language model generation: Utilizes large multimodal language models (e.g., GPT-4o) to generate chord progressions related to the extracted keywords.
- Rejection sampling filtering: Employs a single-modal model trained on real music data to filter the chord progressions generated by the LLM, enhancing musical coherence and diversity.
What are the innovative aspects of this solution?
- Generation method without aligned training data: Proposes a method that does not require aligned training data, combining the generalization capabilities of LLMs with the accuracy of single-modal models.
- Integration of multimodal and symbolic music generation: For the first time, explores chord generation based on multimodal inputs such as images, text, and audio, building upon existing symbolic music generation technologies.
- User guidance and transparency: Enhances the explainability and user control of the generation process through the use of keywords as an intermediary layer.
What are the implementation steps and key technologies used?
- User input: Users provide multimodal content (e.g., images, text, or audio).
- Keyword extraction: A large language model extracts musical keywords from the input.
- Chord generation:
- Based on the keywords, GPT-4o generates initial chord progressions.
- A note-based language model (trained on real music data) applies rejection sampling to retain chord progressions that align with musical distributions.
- Chord transcription: Audio files are transcribed into symbolic chords.
Research Outcomes
What specific outcomes were achieved?
-
Technical evaluation:
- Chord progressions generated by Amuse outperformed baseline methods in terms of diversity, relevance, and musical coherence.
- User studies demonstrated that Amuse effectively supports the transformation of multimodal inspiration into concrete musical elements.
-
User experience:
- Users felt a greater sense of control and creative agency when using Amuse.
- The tool's keyword feature was perceived as intuitive and transparent, significantly contributing to the explainability of the generated chords.
What advantages does it have compared to existing solutions?
- Multimodal support: Directly integrates multimodal inspiration rather than being limited to musical context.
- Editability and modularity: Provides symbolic music elements that users can control and flexibly apply.
- No need for aligned data: Solves the data scarcity issue through innovative methods without requiring large-scale multimodal-music alignment data.
What were the experimental or evaluation results?
-
Technical evaluation:
- Chord progressions generated by Amuse demonstrated superior musical coherence compared to traditional methods (validated using Jensen-Shannon Divergence).
- In user preference tests, Amuse-generated chords outperformed baselines in terms of keyword relevance.
-
User studies:
- Users experienced greater freedom and exploration in the creative process when using Amuse.
- The Chord Generator module was more popular, while the Chord Transcriber was used less frequently, indicating a preference for more flexible functionalities in actual creation.
Limitations and Future Directions
- Tool integration gaps: Users reported discontinuity when switching from Amuse to other tools (e.g., Aria), indicating a need for better integration of multimodal and contextual conditions.
- Limited task coverage: The current tool primarily supports chord generation and does not cover melody or lyric generation. Future work could expand to other musical elements.
- Scenario testing limitations:
- This study primarily tested in static environments and did not fully explore real-time creative scenarios.
- Long-term adoption has not yet been evaluated, and future research should explore more realistic and diverse use cases.
In summary, Amuse provides a novel approach and potential for human-computer collaborative music creation. Its support for multimodal inspiration significantly enhances the freedom and personalization of the creative process.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can users be supported in translating multimodal inspiration (e.g., images, text, or audio) into concrete chord progressions?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- Can diverse and musically coherent chords be generated without aligning large-scale multimodal music data?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- Do users gain greater creative control and transparency when generating chords guided by keywords?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Music creators struggle to efficiently translate multimodal inspiration into concrete musical elements with existing tools.Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- 67%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 60%
In a Silent Way: Communication Between AI and Improvising Musicians Beyond Sound
CHI '19· Generative AI (Text, Image, Music, Video) +1
- 60%
Understanding the Potentials and Limitations of Prompt-based Music Generative AI
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 60%
Expressive Communication: Evaluating Developments in Generative Models and Steering Interfaces for Music Creation
IUI '22· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713818
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Best Paper
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Music Composition & Sound Design Tools, 3D Modeling & Animation
work
Professions
Musicians, DJs & Sound Designers, Film & Animation Producers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers