Combining audio control and style transfer using latent diffusion
Nils Demerlé
Philippe Esling
Guillaume Doras
David Genova
demerle@ircam.fr
esling@ircam.fr
doras@ircam.fr
genova@ircam.fr
Control of generative models
2
Prompt-based control
MusicLM, MusicGen, StableAudio
Descriptor-based control
DDSP, MusicControlNet
Audio-based control
RAVE, SS-VAE
Transfers content to a different timbre
Existing methods
Goal
Audio based timbre extraction and transfer
Timbre
Structure
Nils Demerlé | ISMIR 2024
Model Overview
3
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Model Overview
4
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Model Overview
5
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Model Overview
6
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Structural bias
Model Overview
7
Structural bias
Two stage training
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Model Overview
8
Separating timbre and structure
Nils Demerlé | ISMIR 2024
Structural bias
Two stage training
Model Overview
9
Nils Demerlé | ISMIR 2024
Generator
Model Overview
10
Nils Demerlé | ISMIR 2024
Timbre
Model Overview
11
Nils Demerlé | ISMIR 2024
Structure
Model Overview
12
Nils Demerlé | ISMIR 2024
Structure
Model Overview
13
Nils Demerlé | ISMIR 2024
Separation
Experiments
14
Timbre
Nils Demerlé | ISMIR 2024
Structure
Tasks
Timbre transfer SS-VAE3, Music Style Transfer4
Transfer between structure and timbre audio targets
Experiments
15
Timbre
Structure
Nils Demerlé | ISMIR 2024
Tasks
Timbre transfer SS-VAE3, Music Style Transfer4
Transfer between structure and timbre audio targets
MIDI-to-audio Spectrogram Diffusion5
Synthesize audio from MIDI and timbre audio target
Experiments
16
Timbre
Structure
Nils Demerlé | ISMIR 2024
Similarity
Tasks
Timbre transfer SS-VAE3, Music Style Transfer4
Transfer between structure and timbre audio targets
MIDI-to-audio Spectrogram Diffusion5
Synthesize audio from MIDI and timbre audio target
Evaluation
Timbre
Timbre similarity metric between output and timbre target
Experiments
17
Timbre
Structure
Nils Demerlé | ISMIR 2024
Similarity
F1 Score
Basic
Pitch
Tasks
Timbre transfer SS-VAE3, Music Style Transfer4
Transfer between structure and timbre audio targets
MIDI-to-audio Spectrogram Diffusion5
Synthesize audio from MIDI and timbre audio target
Evaluation
Timbre
Timbre similarity metric between output and timbre target
Structure
Compare input ground-truth and output transcription
Evaluation using mir-eval onset F1 score
Experiments
18
Tasks
Timbre transfer SS-VAE3, Music Style Transfer4
Transfer between structure and timbre audio targets
MIDI-to-audio Spectrogram Diffusion5
Synthesize audio from MIDI and timbre audio target
Evaluation
Timbre
Timbre similarity metric between output and timbre target
Structure
Compare input ground-truth and output transcription
Evaluation using mir-eval onset F1 score
Datasets
Synthetic Data SLAKH 2100
Instrumental recordings synthesized from MIDI
Real Data Maestro, UMRP, GuitarSet
Real recordings from a variety of instruments
Timbre
Structure
Nils Demerlé | ISMIR 2024
Similarity
F1 Score
Basic
Pitch
Results
19
Nils Demerlé | ISMIR 2024
Timbre transfer
Results
20
Nils Demerlé | ISMIR 2024
Timbre transfer
Better quality and transfer than baselines
Results
21
Nils Demerlé | ISMIR 2024
Better quality and transfer than baselines
adversarial slightly degrades structure
Timbre transfer
Results
22
Nils Demerlé | ISMIR 2024
Better quality and transfer than baselines
adversarial slightly degrades structure
Better separation between timbre and structure
Timbre transfer
Applications
Demo and code
See you at the poster sessions !
References
1 : T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models�2 : K. Preechakul, N. Chatthee, S. Wizadwongsa, and S. Suwajanakorn, “Diffusion autoencoders: Toward a meaningful and decodable representation,” �3 : O. Cífka, A. Ozerov, U. S ̧ims ̧ekli, and G. Richard, “Self-supervised vq-vae for one-shot music style transfer,” in ICASSP 2021-2021�4 : S. Li, Y. Zhang, F. Tang, C. Ma, W. Dong, and C. Xu, “Music style transfer with time-varying inversion of diffusion models�5 : C. Hawthorne, I. Simon, A. Roberts, N. Zeghidour, J. Gardner, E. Manilow, and J. Engel, “Multi- instrument music synthesis with spectrogram diffusion,”
Nils Demerlé | ISMIR 2024