1 of 110

Slides by Xin Wang

National Institute of Informatics

© 2019, Xin Wang. All rights reserved.

This work is licensed under the Creative Commons Attribution 3.0 license.

See http://creativecommons.org/ for details.

Note: Natural Japanese speech data belonging to ATR Ximera corpus are deleted in this public available version

2 of 110

Neural Waveform Modeling�from our experiences in text-to-speech application

2

contact: wangxin@nii.ac.jp

we welcome critical comments, suggestions, and discussion

Xin WANG with Shinji Takaki and Junichi Yamagishi

National Institute of Informatics, Japan

NLP lecture series, IIS

Erlangen Germany, 2019

3 of 110

3

  • Postdoc, Yamagishi-lab, NII
  • Research keywords:
    • Text-to-speech synthesis (TTS)
      1. Neural network
      2. Hidden Markov model
    • Speech anti-spoofing

Self-introduction

WANG Xin

Pronunciation

one

shin

4 of 110

4

Contents

Introduction

Theory

Practice

Summary

  • AR & flow-based models
  • No AR nor flow
  • WaveNet
  • Neural source-filter model
  • Beyond speech
  • Future direction

5 of 110

5

Text-to-speech synthesis

http://www.hawking.org.uk/the-computer.html

https://hackaday.com/2018/05/10/googles-duplex-ai-has-conversation-indistinguishable-from-humans/

Text

TTS

Speech waveform

Introduction

6 of 110

6

Text

TTS

Speech waveform

Text-to-speech synthesis

    • Statistical parametric speech synthesis 1

Marianna made the marmalade

Linguistic features

Acoustic features

Front-end

(Text-analyzer)

Back-end

Waveform

generator

Acoustic models

Text

/m/ /ɛ/ /r/ …

H* on Marianna …

(S (NP (N Marianna))

(VP (V made)

(NP (ART the))

(N marmalade))))

Mel-spectrum, F0,

Band-aperiodicity, etc.

1. H. Zen, K. Tokuda, and A. W. Black. Statistical parametric speech synthesis. Speech Communication, 51:1039–1064, 2009.

Introduction

7 of 110

7

Text-to-speech synthesis

    • Recent TTS frameworks

Front-end

(Text-analyzer)

Back-end

Waveform

generator

Acoustic models

Text

Trimmed front-end

‘end-to-end’ TTS system

Text

Waveform

module

Attention-based

acoustic model

Front-end

(Text-analyzer)

Unified back-end

Text

Waveform

module

Pre-processing

A. van den Oord, et al. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.

Y. Wang, et al. Tacotron: Towards end-to-end speech synthesis. In Proc. Interspeech, pages 4006–4010, 2017.

J. Shen, et al. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions. In Proc. ICASSP, pages 4779–4783, 2017

Introduction

8 of 110

8

Spectral features, F0, etc.

Neural waveform modeling

Introduction

9 of 110

9

Neural waveform modeling

1

2

3

4

T

Waveform

values

Neural waveform models

Introduction

10 of 110

10

Naïve neural waveform model

1

2

3

4

T

Convolution network (CNN) / Recurrent network (RNN)

Mean-square-error (MSE) / Cross-entropy (CE)

Waveform

values

Introduction

11 of 110

11

Naïve neural waveform model

1

2

3

4

T

Introduction

12 of 110

12

Introduction

  • Reference in appendix
  • Tutorial slides: https://www.slideshare.net/jyamagis/

Autoregressive (AR) neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

Multi-head

CNN

No AR, nor flow

Neural source-filter

Model (NSF)

  • No explicit

Improve naïve model

Naïve model

Flow-based model

FloWaveNet

WaveGlow

ClariNet

Parallel

WaveNet

GELP

Theoretical interpretation

Practical issues

13 of 110

13

Contents

Introduction

Theory

Practice

Summary

  • AR & flow-based models
  • Neural source filter model
  • WaveNet
  • Neural source-filter model
  • Beyond speech
  • Future work

14 of 110

14

MCNN

GELP

Theory: AR neural waveform model

Flow-based model

FloWaveNet

WaveGlow

No AR, nor flow

Neural source-filter

Model (NSF)

  • No explicit
  • Spectral-domain training criterion
  • Source-filter architecture

ClariNet

Parallel

WaveNet

Naïve model

AR neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

Jordan network

Michael I. Jordan. Serial order: A parallel distributed processing approach. Technical Report 8604, Institute for Cognitive Science, 1986.

Overview

15 of 110

15

Theory: AR neural waveform model

General idea

    • Training: teacher forcing 1

1

2

3

4

T

1

2

3

Natural

waveform

1 R. J. Williams and D. Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.

16 of 110

16

General idea

    • Training: teacher forcing 1

1

2

3

4

T

1

2

3

1 R. J. Williams and D. Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.

Theory: AR neural waveform model

17 of 110

17

General idea

    • Sequential generation

1

2

1

2

3

4

3

T

Generated

waveform

Theory: AR neural waveform model

18 of 110

18

MCNN

GELP

Flow-based model

FloWaveNet

WaveGlow

No AR, nor flow

Neural source-filter

Model (NSF)

  • No explicit
  • Spectral-domain training criterion
  • Source-filter architecture

ClariNet

Parallel

WaveNet

Naïve model

    • WaveNet
      • Tractable probability & powerful AR dependency
      • Slow sequential generation & only left-to-right dependency
    • WaveRNN 1
      • Batch-sampling: faster generation
      • Subscale-dependency: more than left-to-right dependency
    • LPCNet & GlotNet 2,3
      • Classical AR + neural AR

AR neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

1. N. Kalchbrenner, et al Efficient neural audio synthesis. In Proc. ICML, volume 80, pages 2410–2419, 10–15 Jul 2018.

2. J.-M. Valin and J. Skoglund. LPCNet: Improving neural speech synthesis through linear prediction. In Proc. ICASSP, pages 5891–5895, 2019.

3. L. Juvela, et al . Speaker-independent raw waveform model for glottal excitation. In Proc. Interspeech 2018, pages 2012–2016, 2018.

Theory: AR neural waveform model

19 of 110

19

MCNN

GELP

Flow-based model

FloWaveNet

WaveGlow

Theory: flow-based models

No AR, nor flow

Neural source-filter

Model (NSF)

  • No explicit
  • Spectral-domain training criterion
  • Source-filter architecture

ClariNet

Parallel

WaveNet

Naïve model

AR neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

Flow-based model

FloWaveNet

WaveGlow

    • Fast generation?

20 of 110

20

Revisit AR model

    • Consider an AR model using a Gaussian distribution

1

2

3

T

1

2

3

NN

1

2

3

T

1

2

3

NN

Training

Generation

G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.

Theory: flow-based models

Or equivalently

21 of 110

21

Revisit AR model

    • Consider an AR model using a Gaussian distribution

1

2

3

T

1

2

3

NN

1

2

3

T

1

2

3

NN

Training

Generation

  • z-1 denotes time delay
  • See proof of in appendix

NN

z-1

H(.)

NN

z-1

H-1(.)

G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.

Theory: flow-based models

22 of 110

22

Revisit AR model

    • Consider an AR model using a Gaussian distribution

1

2

3

T

1

2

3

NN

1

2

3

T

1

2

3

NN

Training

Generation

NN

z-1

H(.)

NN

z-1

H-1(.)

  • z-1 denotes time delay
  • See proof of in appendix

G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.

Such an AR model is a flow-based model

Training:

  1. Transform o1:T to n1:T
  2. Maximizing n1:T likelihood over N(nt 0, 1)

Generation:

  1. Sample nt from N(nt 0, 1)
  2. Transform nt to ot
  3. Repeat from t=1 to t=T

Theory: flow-based models

23 of 110

23

Theory: flow-based models

From AR to Inverse AR flow-based model

  • z-1 denotes time delay

NN

z-1

H(.)

Training

Generation

NN

z-1

H(.)

NN

z-1

H-1(.)

NN

z-1

H-1(.)

AR flow

Inverse-AR flow

24 of 110

24

Theory: flow-based models

From AR to Inverse AR flow-based model

  • z-1 denotes time delay

NN

z-1

H(.)

Training

Generation

NN

z-1

H-1(.)

AR flow

NN

z-1

H(.)

NN

z-1

H-1(.)

Inverse-AR flow

      • O(1)
      • O(T)
      • O(1)
      • O(T)

Knowledge distilling

Parallel WaveNet & ClariNet

25 of 110

25

MCNN

No AR, nor flow

Neural source-filter

Model (NSF)

Naïve model

AR neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

    • WaveGlow1 & FloWaveNet2
  • Fast generation & slow training
    • Parallel WaveNet3 & ClariNet4
  • Knowledge-distilling is complicated

1. R. Prenger, R. Valle, and B. Catanzaro. Waveglow: A flow-based generative network for speech synthesis. ICASSP 2019, 2018.

2. S. Kim, S.-g. Lee, J. Song, and S. Yoon. Flowavenet: A generative flow for raw audio. ICML 2019.

3. A. van den Oord, Y. Li, I. Babuschkin, et. al.. Parallel WaveNet: Fast high-fidelity speech synthesis. Proc. ICML, pages 3918–3926, 2018.

4. W. Ping, K. Peng, and J. Chen. Clarinet: Parallel wave generation in end-to-end text-to-speech. ICLR, 2018.

Inverse AR flow

FloWaveNet

WaveGlow

ClariNet

Parallel

WaveNet

Theory: flow-based models

26 of 110

26

  • Faster training & generation
  • Easy to implement

Inverse AR flow

FloWaveNet

WaveGlow

Theory: neural source-filter model

No AR, no flow

Neural source-filter

Model (NSF) 1

  • Source-filter architecture
  • Spectral-domain training criterion

Naïve model

AR neural model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

ClariNet

Parallel

WaveNet

1. X. Wang, et al. Neural source-filter-based waveform model for statistical para- metric speech synthesis. In Proc. ICASSP, pages 5916–5920, 2019.

2. S. O ̈. Arık, et. al. Fast spectrogram inversion using multi-head convolutional neural networks. IEEE Signal Processing Letters, 26(1):94–98, 2018.

3. J. Lauri, et. al. GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram, Proc. Interspeech, 2019

MCNN2

GELP3

27 of 110

27

Theory: neural source-filter model

1

2

3

4

T

General idea

  • No AR or inverse AR flow

1

2

3

4

T

‘Filter’

Natural

waveform

Generated

waveform

1

2

3

4

T

F0/pitch

‘Source’

28 of 110

28

Theory: neural source-filter model

1

2

3

4

T

General idea

  • Based on short-time Fourier transform (STFT)

Generated

waveform

Natural

waveform

1

2

3

4

T

Spectral distance

1

2

3

4

T

F0/pitch

29 of 110

29

1

2

3

4

T

F0/pitch

Theory: neural source-filter model

1

2

3

4

T

Probabilistic interpretation?

Generated

waveform

Natural

waveform

1

2

3

4

T

Spectral distance

What is the ?

30 of 110

30

Theory: neural source-filter model

1

2

3

4

T

Probabilistic interpretation?

  • Spectral distance

1

2

3

4

T

Framing

Framing

Spectral

distance

FFT

FFT

  • , where D is frame length. where K is FFT points.

31 of 110

31

Theory: neural source-filter model

1

2

3

4

T

Probabilistic interpretation?

1

2

3

4

T

Framing

Framing

FFT

FFT

Likelihood over Gaussian

  • For explanation, denotes spectral power vector
  • , where D is frame length. where K is FFT points.

32 of 110

32

Theory: neural source-filter model

Probabilistic interpretation?

1

2

3

4

T

Framing

FFT

1

2

3

4

T

Framing

FFT

Likelihood over Gaussian

  • , where K is FFT points
  • For explanation, denotes spectral power vector

33 of 110

33

Naïve model

AR model

Inverse-AR flow

NSF

Theory in summary

34 of 110

34

Contents

Introduction

Theory

Practice

Summary

  • AR & flow-based models
  • Neural source filter model
  • WaveNet
  • Neural source-filter model
  • Beyond speech
  • Future work

35 of 110

35

Practice: WaveNet

WaveNet variants

    • Discretized or continuous-valued waveforms

      • Two practical issues:
        1. How to generate waveform samples?
        2. How to train WaveNet Gaussian?

1

2

1

2

3

4

3

1

2

1

2

3

4

3

GMM/Gaussian

Softmax

36 of 110

36

Practice: WaveNet

WaveNet variants

    • Discretized or continuous-valued waveforms

    • Other variants
      • WaveNet using mixture of logistic distribution 1
      • WaveNet + Spline 2
      • Quantization noise shaping 3, related noise shaping method 4

1. T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma. Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications. arXiv preprint arXiv:1701.05517, 2017.

2. Y. Agiomyrgiannakis. B-spline PDF: A generalization of histograms to continuous density models for generative audio networks. In Proc. ICASSP, pages 5649–5653. IEEE, 2018.

3. T. Yoshimura, et al. Mel-cepstrum-based quantization noise shaping applied to neural-network-based speech waveform synthesis. IEEE TASLP, 26(7):1173–1180, 2018.

4. K. Tachibana, T. Toda, Y. Shiga, and H. Kawai. An investigation of noise shaping with perceptual weighting for WaveNet-based speech generation. In Proc. ICASSP, pages 5664–5668. IEEE, 2018.

1

2

1

2

3

4

3

Softmax

1

2

1

2

3

4

3

GMM/Gaussian

37 of 110

37

Generation strategy

    • WaveNet-softmax

      • Generation as a search problem

      • Search space: 256T for 8-bits waveform of length T

1

2

1

2

3

4

3

Practice: WaveNet

1

2

3

4

38 of 110

38

Generation strategy

    • WaveNet-softmax

      • Sub-optimal search by
        • Exploitation
        • Exploration
        • Or mix of both

Practice: WaveNet

1

2

3

4

Random sampling

Greedy search

39 of 110

39

Generation strategy

    • WaveNet-softmax

Waveform levels (0-1024)

Waveform levels (0-1024)

Practice: WaveNet

40 of 110

59

Practice: WaveNet

Generation method

    • Experiments on WaveNet vocoder

How about

  1. Exploration in unvoiced steps
  2. Exploitation in randomly selected voiced steps

41 of 110

41

Practice: WaveNet

Natural

Greedy

search

Random

sampling

Mixed

approach

42 of 110

42

Practice: WaveNet

Natural

Greedy

search

Random

sampling

Mixed

approach

43 of 110

43

Generation strategy

    • WaveNet-softmax
      • Exploitation & exploration
      • Other strategy: temperature of softmax 1

    • WaveNet-Gaussian
      • Infinite search space: the best is impossible

      • Same strategy as WaveNet-softmax

Practice: WaveNet

Greedy best?

Sampling?

1. Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu. FFTNet: A real-time speaker-dependent neural vocoder. In Proc. ICASSP, pages 2251–2255, 2018.

44 of 110

44

Practice: WaveNet

Training stability

    • WaveNet-Gaussian

  • Maximum likelihood training is risky: very large gradients

NN

1

2

1

2

3

4

3

45 of 110

45

Practice: WaveNet

Training stability

    • WaveNet-Gaussian

46 of 110

46

Negative log-likelihood

Practice: WaveNet

Training stability

    • WaveNet-Gaussian
      • Toy experiment
        • 1 utterance & network well-initialized
        • Variance floor applied

Epoch

47 of 110

47

Network training

 

Epoch

48 of 110

48

Part 1: implicit source-filter model

WaveNet

Illness of fitting Gaussian

    • Why joint learning is unstable? Toy experiment
      • Use the MSE network
      • Fit only one utterance

Epoch

 

49 of 110

49

Practice: WaveNet

Training stability

    • WaveNet-Gaussian

      • Our two-steps strategy
        1. Train blue part with
        2. Train red part only

      • Gradient will be mild
        • Minimizes while keep gradient mild
        • Gradient not explode when

NN

1

2

1

2

3

4

3

50 of 110

50

Training stability

    • WaveNet-Gaussian
      • Experiment: 5 hours data

Practice: WaveNet

Negative log-likelihood

Naïve strategy

Our

strategy

Epoch

On training set

On validation set

Step 1

Step 2

51 of 110

51

Training stability

    • WaveNet-Gaussian
      • Experiment: 5 hours data

Practice: WaveNet

Negative log-likelihood

Naïve strategy

Our

strategy

Epoch

On training set

On validation set

Step 1

Step 2

52 of 110

52

Generation strategy

Training WaveNet-Gaussian

Practice: WaveNet

Greedy best?

Sampling?

Exploitation + exploration

Keep gradients mild

NN

1

2

1

2

3

4

3

53 of 110

53

Contents

Introduction

Theory

Practice

Summary

  • AR & flow-based models
  • Neural source filter model
  • WaveNet
  • Neural source-filter model
  • Beyond speech
  • Future work

54 of 110

54

Practice: NSF

1

2

3

4

T

General idea

  • Spectral-domain training criterion
  • Source-filter structure

Generated

waveform

Natural

waveform

1

2

3

4

T

Spectral distance

1

2

3

4

T

F0/pitch

55 of 110

55

Practice: NSF

Common structure

  • No AR or inverse AR
  • No knowledge distilling

Spectral features & F0

Condition module

Source module

Filter module

Frequency-domain distance

Natural

waveform

Generated

waveform

F0 infor.

Spectral infor.

Generated waveform

Gradients

56 of 110

56

Practice: NSF

Common structure

  • Condition module: input feature pre-process

Spectral features & F0

Source module

Filter module

Frequency-domain distance

Natural

waveform

Generated

waveform

Up sampling

Up sampling

Bi-LSTM

CONV

F0

Generated waveform

Gradients

Up sampling

Dimension change

Temporal smoothing

Cat.

57 of 110

57

Practice: NSF

Common structure

  • Source module: generate a sine waveform given F0

  • FF: feedforward layer with Tanh

Spectral features & F0

Filter module

Frequency-domain distance

Natural

waveform

Generated

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

Generated waveform

Gradients

Up sampling

Bi-LSTM

CONV

Cat.

F0

58 of 110

58

Spectral features & F0

Filter module

Frequency-domain distance

Natural

waveform

Generated

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

Generated waveform

Gradients

Up sampling

Bi-LSTM

CONV

Cat.

F0

Practice: NSF

Common structure

Random initial phase

Sampling rate

Noise

FF

Sine

generator

Fundamental component

Voiced:

Unvoiced: noise

59 of 110

59

Practice: NSF

Common structure

  • Error metric

Spectral features & F0

Frequency-domain distance

Natural

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

Filter module

Generated waveform

Gradients

Compute frequency-domain distance

Compute gradients for SGD

Up sampling

Bi-LSTM

CONV

Cat.

F0

60 of 110

60

Practice: NSF

Common structure

  • Based on short-time Fourier transform

Spectral features & F0

Natural

waveform

Generated

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

FFT

Framing

FFT

Framing

iFFT

De-framing

Filter module

Up sampling

Bi-LSTM

CONV

Cat.

F0

61 of 110

61

Practice: NSF

Common structure

  • Different frame shifts / window lengths / FFT points
  • Homogenous distances

FFT

Framing

FFT

Framing

iFFT

De-framing

FFT

Framing

FFT

Framing

iFFT

De-framing

FFT

Framing

FFT

Framing

iFFT

De-framing

+

62 of 110

62

Practice: NSF

Common structure

  • Different NSF models, different neural filter modules

Spectral features & F0

Natural

waveform

Generated

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

FFT

Framing

FFT

Framing

iFFT

De-framing

Filter module

Up sampling

Bi-LSTM

CONV

Cat.

F0

Filter module

63 of 110

63

Practice: NSF

Common structure

Spectral features & F0

Natural

waveform

Generated

waveform

Up sampling

Noise

FF

Sine

generator

harmonics

FFT

Framing

FFT

Framing

iFFT

De-framing

Filter module

Up sampling

Bi-LSTM

CONV

Cat.

F0

Filter module

NSF models

Baseline NSF

(b-NSF)

Simplified NSF

(s-NSF)

Harmonic-plus-noise NSF

(hn-NSF)

hn-NSF

Ver.1

hn-NSF with ver.2

ICASSP 2019

Journal paper submitted

SSW 2019

64 of 110

64

Practice: NSF

Baseline and simplified NSF

  • Baseline filter block follows WaveNet / ClariNet

  • Baseline filter block can be simplified

Baseline

filter block 1

Baseline

filter block 2

Baseline

filter block 5

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

b-NSF

s-NSF

simplify

65 of 110

65

simplify

Practice: NSF

Baseline and simplified NSF

Baseline

filter block 2

Baseline

filter block 5

Simplified

filter block 2

Simplified

filter block 5

b-NSF

s-NSF

  • Element-wise multiplication

Baseline

filter block 1

Simplified

filter block 1

Simplified filter block

Dilated CONV

+

FF

FF

Dilated CONV

+

Baseline filter block

Dilated CONV

+

Tanh

Sigmoid

FF

FF

FF

+

Dilated CONV

+

Tanh

Sigmoid

FF

FF

+

+

FF

66 of 110

66

Practice: NSF

Baseline and simplified NSF

  • Both models:
    1. Strong harmonics in high-frequency bands
    2. Awful unvoiced (fricative) sounds
  • Model ‘overfitted’ to voiced sounds?

Baseline

filter block 1

Baseline

filter block 2

Baseline

filter block 5

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

b-NSF

s-NSF

simplify

67 of 110

67

Practice: NSF

Harmonic-plus-noise NSF

  • HP, LP: high- and low-pass finite-impulse-response (FIR) filter

Baseline

filter block 1

Baseline

filter block 2

Baseline

filter block 5

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

b-NSF

s-NSF

hn-NSF

simplify

upgrade

68 of 110

68

Baseline

filter block 2

Baseline

filter block 5

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

b-NSF

s-NSF

hn-NSF

simplification

improvement

Baseline

filter block 1

Simplified

filter block 1

Simplified

filter block 1

Practice: NSF

Harmonic-plus-noise NSF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Maximum voicing frequency

(MVF)

hn-NSF

69 of 110

69

Practice: NSF

Harmonic-plus-noise NSF

    • Version I: choose MVF based on u/v

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

u/v flag

For voiced sounds

For unvoiced sounds

Condition module for hn-NSF

Fixed MVFs

70 of 110

70

Practice: NSF

Harmonic-plus-noise NSF

    • Version II: predict MVF from input features

  • Predict MVF from condition module (SSW paper)
  • From MVF to FIR filter coefficients (SSW paper)

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Condition module for hn-NSF

sinc

Hamming window

Gain norm.

HP

LP

71 of 110

71

Practice: NSF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

u/v flag

Condition module for hn-NSF

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Condition module for hn-NSF

72 of 110

72

Practice: NSF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

u/v flag

Condition module for hn-NSF

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Condition module for hn-NSF

73 of 110

73

Practice: NSF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

u/v flag

Condition module for hn-NSF

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Condition module for hn-NSF

74 of 110

74

Practice: NSF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

u/v flag

Condition module for hn-NSF

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

Simplified

filter block 5

noise

+

HP

LP

Condition module for hn-NSF

75 of 110

75

Practice: NSF

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

NSF is a deep-residual network

76 of 110

76

Practice: NSF

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

NSF is a deep-residual network

77 of 110

77

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

78 of 110

78

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

79 of 110

79

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

80 of 110

80

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

81 of 110

81

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

82 of 110

82

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

83 of 110

83

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

84 of 110

84

Spectral features & F0

Up sampling

Noise

FF

Sine

generator

harmonics

Up sampling

Bi-LSTM

CONV

Cat.

F0

MVF

Simplified

filter block 1

Simplified

filter block 2

Simplified

filter block 5

noise

+

HP

LP

Simplified

filter block 5

Condition module for proposed hn-NSF

Source module

Practice: NSF

NSF is a deep-residual network

85 of 110

85

Configuration

    • Data and features

    • Models

Corpus

Size

Note

ATR Ximera F009 [1]

15 hours

16kHz, Japanese, neutral style

Feature

Dimension

Acoustic

Mel-generalized cepstrum coefficients (MGC)

or

Mel-spectra

60

80

F0

1

Practice: comparison

WaveNet

softmax

b-NSF

hn-NSF

trainable MVF

WaveNet

Gaussian

s-NSF

hn-NSF

fixed MVF

WORLD

vocoder

86 of 110

86

Speech quality (ICASSP)

  • 245 paid evaluators, 1450 evaluation sets

Practice: comparison

Copy-synthesis

Pipeline TTS

WaveNet

softmax

b-NSF

hn-NSF

trainable MVF

WaveNet

Gaussian

s-NSF

hn-NSF

fixed MVF

WORLD

vocoder

WORLD

vocoder

WaveNet

softmax

WaveNet

Gaussian

b-NSF

87 of 110

87

Speech quality (Journal paper submitted)

  • >150 paid evaluators
  • s-NSF did badly on unvoiced sounds

Practice: comparison

WaveNet

softmax

b-NSF

s-NSF

hn-NSF

fixed MVF

WaveNet

softmax

b-NSF

hn-NSF

trainable MVF

WaveNet

Gaussian

s-NSF

hn-NSF

fixed MVF

WORLD

vocoder

88 of 110

88

Speech quality (SSW 2019)

  • >150 paid evaluators

Practice: comparison

  • Copy-synthesis
  • Pipeline TTS

WaveNet

softmax

b-NSF

hn-NSF

trainable MVF

WaveNet

Gaussian

s-NSF

hn-NSF

fixed MVF

WORLD

vocoder

WaveNet

softmax

hn-NSF

trainable MVF

hn-NSF

fixed MVF

Natural

89 of 110

89

Generation speed

  • Mem-save mode: allocate and release GPU memory layer by layer (limited by our CUDA implemetation)
  • Normal mode: allocate GPU memory once

How many waveform points can be generated in 1s (Tesla p100)?

Practice: comparison

WaveNet

softmax

b-NSF

hn-NSF

trainable MVF

WaveNet

Gaussian

s-NSF

hn-NSF

fixed MVF

WORLD

vocoder

90 of 110

90

Contents

Introduction

Theory

Practice

Summary

  • AR & flow-based models
  • Neural source filter model
  • WaveNet
  • Neural source-filter model
  • Beyond speech
  • Future work

91 of 110

91

Summary

AR model

WaveRNN

SampleRNN

FFTNet

WaveNet

LPCNet

ExcitNet

GlotNet

Multi-head

CNN

No AR, no flow

Neural source-filter

Model (NSF)

  • No explicit

Naïve model

Inverse AR flow

FloWaveNet

WaveGlow

ClariNet

Parallel

WaveNet

GELP

92 of 110

92

Beyond speech

(c.f. HTS Slides, by HTS Working Group)

Source module

Filter module

93 of 110

93

Beyond speech

Music performance

    • Training
  • URPM dataset1
    • ground-truth F0
    • 13 instruments
    • solo recording

  • One model for all instruments

1 University of Rochester Multi-Modal Music Performance (URMP) Dataset http://www2.ece.rochester.edu/projects/air/projects/URMP.html

Neural

waveform

model

F0

Mel-spectra

94 of 110

Natural

b-NSF

S-NSF

hn-NSF

trainable MVF

Violin

Viola

Oboe

Trumpet

Saxophone

Beyond speech

Music performance

    • Testing with natural Mel-spectra and F0 as input

WaveNet

95 of 110

Natural

b-NSF

S-NSF

hn-NSF

trainable MVF

Horn

Trombone

Tuba

Clarinet

Flute

Beyond speech

Music performance

    • Testing with natural Mel-spectra and F0 as input

96 of 110

96

Future direction

(c.f. HTS Slides, by HTS Working Group)

97 of 110

Questions & Comments �are always Welcome!�

97

98 of 110

98

Reference

WaveNet: A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.

SampleRNN: S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio. Samplernn: An unconditional end-to-end neural audio generation model. arXiv preprint arXiv:1612.07837, 2016.

WaveRNN: N. Kalchbrenner, E. Elsen, K. Simonyan, et.al. Efficient neural audio synthesis. In J. Dy and A. Krause, editors, Proc. ICML, volume 80 of Proceedings of Machine Learning Research, pages 2410–2419, 10–15 Jul 2018.

FFTNet: Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu. FFTNet: A real-time speaker-dependent neural vocoder. In Proc. ICASSP, pages 2251–2255. IEEE, 2018.

Universal vocoder: J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, and R. Barra-Chicote. Robust universal neural vocoding. arXiv preprint arXiv:1811.06292, 2018.

Subband WaveNet: T. Okamoto, K. Tachibana, T. Toda, Y. Shiga, and H. Kawai. An investigation of subband wavenet vocoder covering entire audible frequency range with limited acoustic features. In Proc. ICASSP, pages 5654–5658. 2018.

Parallel WaveNet: A. van den Oord, Y. Li, I. Babuschkin, et. al.. Parallel WaveNet: Fast high-fidelity speech synthesis. In Proc. ICML, pages 3918–3926, 2018.

ClariNet: W. Ping, K. Peng, and J. Chen. Clarinet: Parallel wave generation in end-to-end text-to-speech. arXiv preprint arXiv:1807.07281, 2018.

FlowWaveNet: S. Kim, S.-g. Lee, J. Song, and S. Yoon. Flowavenet: A generative flow for raw audio. arXiv preprint arXiv:1811.02155, 2018.

WaveGlow: R. Prenger, R. Valle, and B. Catanzaro. Waveglow: A flow-based generative network for speech synthesis. arXiv preprint arXiv:1811.00002, 2018.

RNN+STFT: S. Takaki, T. Nakashika, X. Wang, and J. Yamagishi. STFT spectral loss for training a neural speech waveform model. In Proc. ICASSP (submitted), 2018.

NSF: X. Wang, S. Takaki, and J. Yamagishi. Neural source-filter-based waveform model for statistical para- metric speech synthesis. arXiv preprint arXiv:1810.11946, 2018.

LP-WavNet: M.-J. Hwang, F. Soong, F. Xie, X. Wang, and H.-G. Kang. Lp-wavenet: Linear prediction-based wavenet speech synthesis. arXiv preprint arXiv:1811.11913, 2018.

GlotNet: L. Juvela, V. Tsiaras, B. Bollepalli, M. Airaksinen, J. Yamagishi, and P. Alku. Speaker-independent raw waveform model for glottal excitation. arXiv preprint arXiv:1804.09593, 2018.

ExcitNet: E. Song, K. Byun, and H.-G. Kang. Excitnet vocoder: A neural excitation model for parametric speech synthesis systems. arXiv preprint arXiv:1811.04769, 2018.

LPCNet: J.-M. Valin and J. Skoglund. Lpcnet: Improving neural speech synthesis through linear prediction. arXiv preprint arXiv:1810.11846, 2018.

MCNN: S. O ̈. Arık, H. Jun, and G. Diamos. Fast spectrogram inversion using multi-head convolutional neural networks. IEEE Signal Processing Letters, 26(1):94–98, 2018.

GELP: J. Lauri, et. al. GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram, Proc. Interspeech, 2019

99 of 110

99

Reference

By Lauri Juvela, Aalto University

100 of 110

DFT

Framing/

windowing

DFT

Framing/

windowing

Generated waveform

Natural waveform

N frames

K DFT bins

K-points

DFT

Frame Length M

Padding

K-M

0

0

0

0

0

0

0

0

0

0

Framing/

windowing

Complex-value domain

Real-value domain

Appendix

Training criterion

101 of 110

X

=

1st Frame

2nd Frame

Nth Frame

T rows

M (frame length)

Frame

shift

NM

columns

0

0

0

0

0

0

0

0

0

0

0

0

Appendix

Training criterion

102 of 110

Training criterion

DFT

Framing/

windowing

DFT

Framing/

windowing

Generated waveform

Natural waveform

N frames

K DFT bins

K-points

iDFT

Frame Length M

De-framing/

windowing

inverseDFT

De-framing

/windowing

Gradients

Gradients w.r.t. zero-padded part

Not used in de-framing/windowing

Padding

K-M

Complex-value domain

Real-value domain

Appendix

103 of 110

103

flow-based models

Recap AR model

    • Consider a WaveNet using a Gaussian distribution

      • Because , we have

1

2

3

T

1

2

3

NN

  • z-1 denotes time delay

NN

z-1

H-1(.)

104 of 110

104

Flow-based models

105 of 110

105

flow-based models

Recap AR model

    • Consider a WaveNet using a Gaussian distribution

      • Because , we have

      • Therefore

  • z-1 denotes time delay

Triangle-matrix,

as nt depends on o<t

106 of 110

106

flow-based models

Recap AR model

    • Consider a WaveNet using a Gaussian distribution

      • So:

  • z-1 denotes time delay

107 of 110

107

flow-based models

Inverse-AR flow

      • Because , we have

  • z-1 denotes time delay

NN

z-1

H-1(.)

Triangle-matrix,

as nt depends on ot

108 of 110

108

flow-based models

Inverse-AR flow

      • Therefore

  • z-1 denotes time delay

NN

z-1

H-1(.)

109 of 110

109

flow-based models

AR flow vs inverse-AR

  • z-1 denotes time delay

NN

z-1

H-1(.)

NN

z-1

H-1(.)

110 of 110

110

flow-based models

  • z-1 denotes time delay

NN

z-1

H-1(.)

NN

z-1

H-1(.)

AR flow

AR flow vs inverse-AR

Inverse-AR flow