Understanding the Potentials and Limitations of Prompt-based Music Generative AI
Authors
Research Background and Issues
-
Issues or Challenges:
The authors identify that current prompt-based music generation AI (GenAI) excels in interacting with users through natural language and efficiently generating music. However, it still faces limitations in conveying complex artistic intentions. The differences in the modes of expression between language and music make it challenging to generate music that meets composers' precise needs solely through language descriptions. -
Significance of the Issue:
As generative AI becomes increasingly prevalent in fields like visual arts and design, the potential of music generation AI is also being explored. Language-based music generation allows users to participate in music creation more easily, which is meaningful not only for professional composers but also provides opportunities for inexperienced users to engage in creative processes. However, to achieve true human-AI collaboration, these systems need to deeply understand creators' needs and offer customized support based on varying levels of expertise. -
Research Motivation and Related Work:
Although prompt-based GenAI has been widely applied in other domains, research on how users interact with music generation AI, especially prompt-based systems, is still in its early stages. Existing studies suggest that users' expertise significantly influences how they utilize AI-generated music. For instance, experts tend to use AI to validate ideas, while non-expert users require support to translate abstract concepts into concrete music. This study aims to further explore how to optimize these systems based on user needs.
Solution
-
Methods and Solutions:
The authors investigate three different types of interaction methods in music generation AI (language prompts, preset options, and melody input) to evaluate how these interaction types meet users' creative needs. The study involved 17 participants with varying musical backgrounds. -
Innovations:
- Proposed a design framework for optimizing music generation systems based on user expertise.
- Compared three primary interaction modes and revealed the potential and limitations of prompt-based systems.
- Analyzed users' unique application strategies in prompt-based systems and proposed specific design guidelines to address shortcomings.
-
Implementation Steps:
- Experiment Design: Conducted remote experiments with 17 participants, asking them to create music using commercial AI tools based on preset options (Boomy), melody input (DeepComposer), and language prompts (SUNO).
- Data Collection: Employed a step-by-step analysis method to record participants' screen interactions, issue logs, and music outputs, followed by questionnaires and semi-structured interviews.
- Data Analysis: Extracted participants' usage patterns, challenges, and preferences through quantitative and qualitative analysis to form comprehensive design recommendations.
Research Findings
-
Specific Findings:
- User Preferences: Professional composers preferred prompt-based systems for quickly validating concepts; beginners favored them for generating reference samples; non-expert users valued the system's ability to translate abstract concepts into actual music.
- Strengths and Limitations:
- Strengths: Prompt-based systems improved creative efficiency, enabling rapid realization of musical concepts, particularly helpful in the early ideation stage for experimental works and style exploration.
- Limitations: Systems struggled with expressing temporality and complex musical structures, and language prompts often failed to accurately convey artistic intentions.
- Experiment and Evaluation Results:
- Evaluations by experts and general listeners indicated that prompt-based music generation achieved high quality in rapid creation but imposed more constraints on experts during complex modifications.
- Music outputs from non-expert users received higher overall ratings, reflecting their preference for "listener-friendly" music over technical complexity.
-
Limitations and Future Directions:
- The small sample size and technical differences between specific AI tools may introduce bias.
- Functional limitations of different tools (e.g., SUNO's language descriptions may fail to capture complex musical features) highlight the need for a unified platform integrating the three interaction types.
- Future research should further explore how creators' identities shape their approaches to using AI for music creation.
Conclusion
This study provides comprehensive insights into the potential and limitations of prompt-based music generation AI. The findings demonstrate that users' expertise significantly influences their collaboration with AI. Based on these insights, the authors propose recommendations for multi-modal interaction design, conversational generation, and fine-grained temporal control interfaces to optimize user experience and support diverse creative needs. These recommendations lay a solid foundation for the future development of AI in creative domains, emphasizing the importance of user-centered design.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do users with different music creation backgrounds interact with prompt-based music generation AI?Category: Music and Audio Generative CreationSimilar questionsarrow_forward
- Which interaction mode—language prompts, preset options, or melody input—better meets users' creative needs?Category: Music and Audio Generative CreationSimilar questionsarrow_forward
- What are the limitations of prompt-based music generation AI in expressing complex artistic intent?Category: Music and Audio Generative CreationSimilar questionsarrow_forward
Practical Problems
1- Music generation AI struggles to meet creators' complex needs through language prompts.Category: Music and Audio Generative CreationSimilar questionsarrow_forward
- 100%
In a Silent Way: Communication Between AI and Improvising Musicians Beyond Sound
CHI '19· Generative AI (Text, Image, Music, Video) +1
- 100%
Expressive Communication: Evaluating Developments in Generative Models and Steering Interfaces for Music Creation
IUI '22· Generative AI (Text, Image, Music, Video) +1
- 75%
LoopMaker: Automatic Creation of Music Loops from Pre-recorded Music
CHI '18· Generative AI (Text, Image, Music, Video) +1
- 75%
Music Creation by Example
CHI '20· Generative AI (Text, Image, Music, Video) +1
- 75%
Entangling Entanglement: A Diffractive Dialogue on HCI and Musical Interactions
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 75%
Using Incongruous Genres to Explore Music Making with AI Generated Content
C&C '24· Generative AI (Text, Image, Music, Video) +2
- 75%
SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and Exploration
IUI '24· Generative AI (Text, Image, Music, Video) +2
- 60%
Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative Models
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 60%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 60%
AMUSE: Human-AI Collaborative Songwriting with Multimodal Inspirations
CHI '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)