Understanding the Potentials and Limitations of Prompt-based Music Generative AI

Generative AI (Text, Image, Music, Video)Music Composition & Sound Design ToolsMusicians, DJs & Sound Designers

Research Background and Issues

  • Issues or Challenges:
    The authors identify that current prompt-based music generation AI (GenAI) excels in interacting with users through natural language and efficiently generating music. However, it still faces limitations in conveying complex artistic intentions. The differences in the modes of expression between language and music make it challenging to generate music that meets composers' precise needs solely through language descriptions.

  • Significance of the Issue:
    As generative AI becomes increasingly prevalent in fields like visual arts and design, the potential of music generation AI is also being explored. Language-based music generation allows users to participate in music creation more easily, which is meaningful not only for professional composers but also provides opportunities for inexperienced users to engage in creative processes. However, to achieve true human-AI collaboration, these systems need to deeply understand creators' needs and offer customized support based on varying levels of expertise.

  • Research Motivation and Related Work:
    Although prompt-based GenAI has been widely applied in other domains, research on how users interact with music generation AI, especially prompt-based systems, is still in its early stages. Existing studies suggest that users' expertise significantly influences how they utilize AI-generated music. For instance, experts tend to use AI to validate ideas, while non-expert users require support to translate abstract concepts into concrete music. This study aims to further explore how to optimize these systems based on user needs.


Solution

  • Methods and Solutions:
    The authors investigate three different types of interaction methods in music generation AI (language prompts, preset options, and melody input) to evaluate how these interaction types meet users' creative needs. The study involved 17 participants with varying musical backgrounds.

  • Innovations:

    • Proposed a design framework for optimizing music generation systems based on user expertise.
    • Compared three primary interaction modes and revealed the potential and limitations of prompt-based systems.
    • Analyzed users' unique application strategies in prompt-based systems and proposed specific design guidelines to address shortcomings.
  • Implementation Steps:

    1. Experiment Design: Conducted remote experiments with 17 participants, asking them to create music using commercial AI tools based on preset options (Boomy), melody input (DeepComposer), and language prompts (SUNO).
    2. Data Collection: Employed a step-by-step analysis method to record participants' screen interactions, issue logs, and music outputs, followed by questionnaires and semi-structured interviews.
    3. Data Analysis: Extracted participants' usage patterns, challenges, and preferences through quantitative and qualitative analysis to form comprehensive design recommendations.

Research Findings

  • Specific Findings:

    • User Preferences: Professional composers preferred prompt-based systems for quickly validating concepts; beginners favored them for generating reference samples; non-expert users valued the system's ability to translate abstract concepts into actual music.
    • Strengths and Limitations:
      • Strengths: Prompt-based systems improved creative efficiency, enabling rapid realization of musical concepts, particularly helpful in the early ideation stage for experimental works and style exploration.
      • Limitations: Systems struggled with expressing temporality and complex musical structures, and language prompts often failed to accurately convey artistic intentions.
    • Experiment and Evaluation Results:
      • Evaluations by experts and general listeners indicated that prompt-based music generation achieved high quality in rapid creation but imposed more constraints on experts during complex modifications.
      • Music outputs from non-expert users received higher overall ratings, reflecting their preference for "listener-friendly" music over technical complexity.
  • Limitations and Future Directions:

    1. The small sample size and technical differences between specific AI tools may introduce bias.
    2. Functional limitations of different tools (e.g., SUNO's language descriptions may fail to capture complex musical features) highlight the need for a unified platform integrating the three interaction types.
    3. Future research should further explore how creators' identities shape their approaches to using AI for music creation.

Conclusion

This study provides comprehensive insights into the potential and limitations of prompt-based music generation AI. The findings demonstrate that users' expertise significantly influences their collaboration with AI. Based on these insights, the authors propose recommendations for multi-modal interaction design, conversational generation, and fine-grained temporal control interfaces to optimize user experience and support diverse creative needs. These recommendations lay a solid foundation for the future development of AI in creative domains, emphasizing the importance of user-centered design.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189087/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713762
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Music Composition & Sound Design Tools
work
Professions
Musicians, DJs & Sound Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers