1 of 23

Combining audio control and style transfer using latent diffusion

Nils Demerlé

Philippe Esling

Guillaume Doras

David Genova

demerle@ircam.fr

esling@ircam.fr

doras@ircam.fr

genova@ircam.fr

2 of 23

Control of generative models

2

Prompt-based control

MusicLM, MusicGen, StableAudio

  • Impressive quality
  • Timbre and style are subjective
  • Limited to most common sounds

Descriptor-based control

DDSP, MusicControlNet

  • Expressive control
  • Limited to explicit descriptors
  • Require annotations

Audio-based control

RAVE, SS-VAE

Transfers content to a different timbre

Existing methods

  • Single or few timbre targets
  • Relatively low-quality

Goal

Audio based timbre extraction and transfer

Timbre

Structure

Nils Demerlé | ISMIR 2024

3 of 23

Model Overview

3

Separating timbre and structure

Nils Demerlé | ISMIR 2024

4 of 23

Model Overview

4

Separating timbre and structure

Nils Demerlé | ISMIR 2024

5 of 23

Model Overview

5

Separating timbre and structure

Nils Demerlé | ISMIR 2024

6 of 23

Model Overview

6

Separating timbre and structure

Nils Demerlé | ISMIR 2024

Structural bias

  • Temporal representation for structure
  • Global vector for timbre

7 of 23

Model Overview

7

Structural bias

  • Temporal representation for structure
  • Global vector for timbre

Two stage training

  • Timbre pre-training

Separating timbre and structure

Nils Demerlé | ISMIR 2024

8 of 23

Model Overview

8

Separating timbre and structure

Nils Demerlé | ISMIR 2024

Structural bias

  • Temporal representation for structure
  • Global vector for timbre

Two stage training

  • Timbre pre-training
  • Learn structure with an adversarial criteria that maximize the confusion of the timbre information

9 of 23

Model Overview

9

Nils Demerlé | ISMIR 2024

Generator

  • Diffusion model based on EDM1
  • Use a neural codec for quality and efficiency
  • Diffusion AutoEncoder2 with two encoders

10 of 23

Model Overview

10

Nils Demerlé | ISMIR 2024

Timbre

  • Timbre extracted from a different excerpt of the same recording to avoid learning time-varying features
  • Codec representation

11 of 23

Model Overview

11

Nils Demerlé | ISMIR 2024

Structure

  • Structure extracted from CQT for better pitch extraction

12 of 23

Model Overview

12

Nils Demerlé | ISMIR 2024

Structure

  • Structure extracted from CQT for better pitch extraction
  • Can alternatively use MIDI instead of audio

13 of 23

Model Overview

13

Nils Demerlé | ISMIR 2024

Separation

  • Encoders and Unet trained to minimize diffusion loss and maximize classification loss
  • Classifier tries to predict from

14 of 23

Experiments

14

Timbre

Nils Demerlé | ISMIR 2024

Structure

Tasks

Timbre transfer SS-VAE3, Music Style Transfer4

Transfer between structure and timbre audio targets

15 of 23

Experiments

15

Timbre

Structure

Nils Demerlé | ISMIR 2024

Tasks

Timbre transfer SS-VAE3, Music Style Transfer4

Transfer between structure and timbre audio targets

MIDI-to-audio Spectrogram Diffusion5

Synthesize audio from MIDI and timbre audio target

16 of 23

Experiments

16

Timbre

Structure

Nils Demerlé | ISMIR 2024

Similarity

Tasks

Timbre transfer SS-VAE3, Music Style Transfer4

Transfer between structure and timbre audio targets

MIDI-to-audio Spectrogram Diffusion5

Synthesize audio from MIDI and timbre audio target

Evaluation

Timbre

Timbre similarity metric between output and timbre target

17 of 23

Experiments

17

Timbre

Structure

Nils Demerlé | ISMIR 2024

Similarity

F1 Score

Basic

Pitch

Tasks

Timbre transfer SS-VAE3, Music Style Transfer4

Transfer between structure and timbre audio targets

MIDI-to-audio Spectrogram Diffusion5

Synthesize audio from MIDI and timbre audio target

Evaluation

Timbre

Timbre similarity metric between output and timbre target

Structure

Compare input ground-truth and output transcription

Evaluation using mir-eval onset F1 score

18 of 23

Experiments

18

Tasks

Timbre transfer SS-VAE3, Music Style Transfer4

Transfer between structure and timbre audio targets

MIDI-to-audio Spectrogram Diffusion5

Synthesize audio from MIDI and timbre audio target

Evaluation

Timbre

Timbre similarity metric between output and timbre target

Structure

Compare input ground-truth and output transcription

Evaluation using mir-eval onset F1 score

Datasets

Synthetic Data SLAKH 2100

Instrumental recordings synthesized from MIDI

Real Data Maestro, UMRP, GuitarSet

Real recordings from a variety of instruments

Timbre

Structure

Nils Demerlé | ISMIR 2024

Similarity

F1 Score

Basic

Pitch

19 of 23

Results

19

Nils Demerlé | ISMIR 2024

Timbre transfer

20 of 23

Results

20

Nils Demerlé | ISMIR 2024

Timbre transfer

Better quality and transfer than baselines

21 of 23

Results

21

Nils Demerlé | ISMIR 2024

Better quality and transfer than baselines

adversarial slightly degrades structure

Timbre transfer

22 of 23

Results

22

Nils Demerlé | ISMIR 2024

Better quality and transfer than baselines

adversarial slightly degrades structure

Better separation between timbre and structure

Timbre transfer

23 of 23

Applications

  • Style transfer between jazz, dub, rock, hip-hop
  • Improves upon MusicGen in cover detection and genre classification of the transfers
  • Streamable extension of the model

Demo and code

See you at the poster sessions !

References

1 : T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models2 : K. Preechakul, N. Chatthee, S. Wizadwongsa, and S. Suwajanakorn, “Diffusion autoencoders: Toward a meaningful and decodable representation,” 3 : O. Cífka, A. Ozerov, U. S ̧ims ̧ekli, and G. Richard, “Self-supervised vq-vae for one-shot music style transfer,” in ICASSP 2021-2021�4 : S. Li, Y. Zhang, F. Tang, C. Ma, W. Dong, and C. Xu, “Music style transfer with time-varying inversion of diffusion models�5 : C. Hawthorne, I. Simon, A. Roberts, N. Zeghidour, J. Gardner, E. Manilow, and J. Engel, “Multi- instrument music synthesis with spectrogram diffusion,”

Nils Demerlé | ISMIR 2024