MIDISpace: Finding Linear Directions in Latent Space for Music Generation

Honorable Mention
Generative AI (Text, Image, Music, Video)Human-LLM CollaborationMusic Composition & Sound Design ToolsMusicians, DJs & Sound DesignersAI/ML Researchers & Engineers

While recent works have shown that it is possible to find disentangled directions in the latent space of image generation networks, finding directions in the latent space of sequential models for music generation remains a largely unexplored topic. In this work, we propose a method for discovering linear directions in the latent space of a music generating Variational Auto-Encoder (VAE). We use PCA, a statistical method to transform the input data such that the variation along the new axes is maximized. We apply PCA on the latent space activations of our model and find largely disentangled directions that change the style and characteristics of the input music. Our experiments show that the found directions are often monotonic, global and encode fundamental musical characteristics such as colorfulness, speed and repetitiveness. Moreover, we propose a set of quantitative metrics to describe different musical styles and characteristics to evaluate our results. We show that the found directions decouple content and can be utilized for style transfer and conditional music generation tasks.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/cc/83023/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3527927.3532790
At a Glance

Paper Snapshot

fact_check
dataset
Source
C&C
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, Music Composition & Sound Design Tools
work
Professions
Musicians, DJs & Sound Designers, AI/ML Researchers & Engineers
article
Content Status
Abstract only
hub
Related Papers
3 related papers