1 of 29

Sonos Seminar Series • 1 September 2022

Style transfer of audio effects �with differentiable signal processing

Christian J. Steinmetz1,2

c.j.steinmetz@qmul.ac.uk

​

Nick J. Bryan2

Joshua D. Reiss1

1Queen Mary University of London

2Adobe Research

2 of 29

More people are creating audio content

Music

Podcasts

Short-form content

Sound for Video

🔊

3 of 29

Producing high quality audio requires expertise

Demand for high quality audio

4 of 29

Style transfer of audio effects

5 of 29

6 of 29

7 of 29

Audio production as a three stage process

1. Listen Perform an acoustic analysis of the input recording

2. Plan Establish an acoustic goal (style) considering the context

3. Execute Manipulate DSP controls to achieve this goal

8 of 29

Learning audio production by example

9 of 29

Differentiable signal processing

Backprop through DSP operations

  • Leveraging existing DSP tools and knowledge
  • High quality audio processing with few artifacts
  • Human understandable outputs that can be adjusted
  • Efficient and can easily run in real-time on CPU

10 of 29

1 Automatic differentiation

Explicitly define signal processing operations in autodiff framework

Engel, Jesse, et al. "DDSP: Differentiable digital signal processing." ICLR (2021).

11 of 29

2 Neural proxy

(1) Pretraining

Frozen DSP neural proxy

(2) Training

​

(3) Inference

Steinmetz, Christian J., et al. "Automatic multitrack mixing with a differentiable mixing console of neural audio effects." ICASSP, 2021.

12 of 29

3 Neural proxy hybrid

​

(3) Inference

(2) Training

Use original DSP during inference

13 of 29

4 Gradient approximation

Simultaneous perturbation stochastic approximation (SPSA)

Finite differences (FD)

Martínez Ramírez, Marco A., et al. "Differentiable signal processing with black-box audio effects." ICASSP, 2021.

14 of 29

Differentiable signal processing

​

  1. Automatic differentiation
  2. Neural proxy
  3. Neural proxy hybrid
  4. Gradient approximation

No existing comparison of these approaches in a unified setup.

15 of 29

Automatic differentiation audio effects

This can be approximated with �a FIR (frequency domain) filter

Estimate IIR filter response with DFT and apply as a frequency domain FIR filter

Nercessian, Shahan. "Neural parametric equalizer matching using differentiable biquads." Proc. Int. Conf. Digital Audio Effects (eDAFx-20). 2020.

16 of 29

Training details

RB-DSP Rule-based DSP

cTCN Conditional TCN

​

NP Neural Proxy

NP-HH Neural Proxy Half-hybrid

NP-FH Neural Proxy Full-hybrid

SPSA Gradient approximation

AD Automatic differentiation

Audio domain loss

Multi-resolution STFT

​

Training Datasets

Speech (LibriTTS)

Music (MTG-Jamendo)

​

Effects

6-band parametric EQ

Dynamic range compressor

​

Models

17 of 29

Experiments

  1. Synthetic production style transfer �(matching input and reference)
  2. Realistic production style transfer �(non-matching input and reference)
  3. Audio production representations �(audio production style classification)
  4. Computational complexity

18 of 29

Audio production style transfer

Synthetic

Realistic

Input

Reference

Input

Reference

High-level metrics

System

Prediction

System

Full Reference

Metric

Prediction

19 of 29

Evaluation metrics

PESQ Perceptual evaluation of speech quality

STFT Multi-resolution STFT error

General similarity

(full reference)

Spectral balance (EQ)�(high-level features)

Dynamics (Compression)�(high-level features)

MSD Large window log-mel spectrogram error

SCE Spectral centroid error

RMS Root mean square energy error

LUFS Perceptual loudness error

20 of 29

Synthetic audio production style transfer

out-of-domain datasets

21 of 29

Production style generation

For evaluating realistic style transfer

Styles are defined by distributions in the parameter space of the parametric EQ and dynamic range compressor.

Clean audio

Style dataset

EQ

DRC

22 of 29

Realistic audio production style transfer

23 of 29

Learning audio production representations

Frozen pretrained encoder

Linear classifier

24 of 29

Computational complexity

25 of 29

Differentiation approaches performance

  1. Rule-based DSP baseline outperformed by learned approaches
  2. Neural proxy hybrid approaches do not perform well
  3. Gradient approximation performs second best but struggles with instability
  4. Automatic differentiation performs best overall but is only an approximation of effects

​

26 of 29

Contributions

  1. The first audio effects style transfer method to integrate audio effects as differentiable operators, optimized end-to-end with an audio-domain loss

​

  • Self-supervised training that enables automatic audio production without labeled or paired training data

​

  • A benchmark of five differentiation strategies for audio effects, including compute cost, engineering difficulty, and performance

​

  • The development of novel neural proxy hybrid methods, and a differentiable dynamic range compressor.

27 of 29

Resources

28 of 29

Future directions

  1. Extend this approach with more differentiable effects (e.g. reverb, distortion, etc)
  2. Improved methods for training neural proxy (hybrids)
  3. Methods for handling dynamic construction of the processing chain
  4. Adapt this approach for multichannel use cases (e.g. multitrack mixing)
  5. Zero-shot adaptation to a new set of audio effects (can I use the plugins in my DAW?)

29 of 29

Sonos Seminar Series • 1 September 2022

Style transfer of audio effects �with differentiable signal processing

Christian J. Steinmetz1,2

c.j.steinmetz@qmul.ac.uk

​

Nick J. Bryan2

Joshua D. Reiss1

1Queen Mary University of London

2Adobe Research