SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and Exploration
Authors
Synthesizers are powerful tools that allow musicians to create dynamic and original sounds. Existing commercial interfaces for synthesizers typically require musicians to interact with complex low-level parameters or to manage large libraries of premade sounds. To address these challenges, we implement SynthScribe --- a fullstack system that uses multimodal deep learning to let users express their intentions at a much higher level. We implement features which address a number of difficulties, namely 1) searching through existing sounds, 2) creating completely new sounds, 3) making small but meaningful modifications to a given sound. This is achieved with three main features: a multimodal search engine for a large library of synthesizer sounds; a user centered genetic algorithm by which completely new sounds can be created and selected given the users preferences; a sound editing support feature which highlights and gives examples for key control parameters with respect to a text or audio based query. The results of our user studies show SynthScribe is capable of reliably retrieving and modifying sounds while also affording the ability to create completely new sounds that expand a musicians creative horizon.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can multimodal deep learning models enable natural synthesizer sound search and editing?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- How can genetic algorithms help users generate novel sounds beyond preset options?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- Which parameter groups are most critical for users' desired sound effects?Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
Practical Problems
1- Beginners struggle to intuitively operate complex synthesizer parameters or filter preset sounds.Category: Music Creation, Synthesis, and AI Generation ToolsSimilar questionsarrow_forward
- 80%
Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Agents in Concert: A Case-Study of Bringing AI to the Stage in Practice
IUI '26· Music Composition & Sound Design Tools +2
- 75%
In a Silent Way: Communication Between AI and Improvising Musicians Beyond Sound
CHI '19· Generative AI (Text, Image, Music, Video) +1
- 75%
Challenges of Music Score Writing and the Potentials of Interactive Surfaces
CHI '24· Music Composition & Sound Design Tools +1
- 75%
Understanding the Potentials and Limitations of Prompt-based Music Generative AI
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 75%
Expressive Communication: Evaluating Developments in Generative Models and Steering Interfaces for Music Creation
IUI '22· Generative AI (Text, Image, Music, Video) +1
- 67%
SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
A Design Space for Live Music Agents
CHI '26· Music Composition & Sound Design Tools +2
- 67%
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
C&C '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Reflection Across AI-based Music Composition
C&C '24· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)