Slides by Xin Wang
National Institute of Informatics
© 2019, Xin Wang. All rights reserved.
This work is licensed under the Creative Commons Attribution 3.0 license.
See http://creativecommons.org/ for details.
Note: Natural Japanese speech data belonging to ATR Ximera corpus are deleted in this public available version
Neural Waveform Modeling�from our experiences in text-to-speech application
2
contact: wangxin@nii.ac.jp
we welcome critical comments, suggestions, and discussion
Xin WANG with Shinji Takaki and Junichi Yamagishi
National Institute of Informatics, Japan
NLP lecture series, IIS
Erlangen Germany, 2019
3
Self-introduction
WANG Xin
Pronunciation | one | shin |
鑫
王
4
Contents
Introduction
Theory
Practice
Summary
5
Text-to-speech synthesis
http://www.hawking.org.uk/the-computer.html
https://hackaday.com/2018/05/10/googles-duplex-ai-has-conversation-indistinguishable-from-humans/
Text
TTS
Speech waveform
Introduction
6
Text
TTS
Speech waveform
Text-to-speech synthesis
Marianna made the marmalade
Linguistic features
Acoustic features
Front-end
(Text-analyzer)
Back-end
Waveform
generator
Acoustic models
Text
/m/ /ɛ/ /r/ …
H* on Marianna …
(S (NP (N Marianna))
(VP (V made)
(NP (ART the))
(N marmalade))))
Mel-spectrum, F0,
Band-aperiodicity, etc.
1. H. Zen, K. Tokuda, and A. W. Black. Statistical parametric speech synthesis. Speech Communication, 51:1039–1064, 2009.
Introduction
7
Text-to-speech synthesis
Front-end
(Text-analyzer)
Back-end
Waveform
generator
Acoustic models
Text
Trimmed front-end
‘end-to-end’ TTS system
Text
Waveform
module
Attention-based
acoustic model
Front-end
(Text-analyzer)
Unified back-end
Text
Waveform
module
Pre-processing
A. van den Oord, et al. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
Y. Wang, et al. Tacotron: Towards end-to-end speech synthesis. In Proc. Interspeech, pages 4006–4010, 2017.
J. Shen, et al. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions. In Proc. ICASSP, pages 4779–4783, 2017
Introduction
8
Spectral features, F0, etc.
Neural waveform modeling
Introduction
9
Neural waveform modeling
1
2
3
4
T
…
Waveform
values
Neural waveform models
…
Introduction
10
Naïve neural waveform model
…
1
2
3
4
T
…
Convolution network (CNN) / Recurrent network (RNN)
Mean-square-error (MSE) / Cross-entropy (CE)
Waveform
values
Introduction
11
Naïve neural waveform model
…
…
1
2
3
4
T
…
Introduction
12
Introduction
Autoregressive (AR) neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
Multi-head
CNN
No AR, nor flow
Neural source-filter
Model (NSF)
Improve naïve model
Naïve model
Flow-based model
FloWaveNet
WaveGlow
ClariNet
Parallel
WaveNet
GELP
Theoretical interpretation
Practical issues
13
Contents
Introduction
Theory
Practice
Summary
14
MCNN
GELP
Theory: AR neural waveform model
Flow-based model
FloWaveNet
WaveGlow
No AR, nor flow
Neural source-filter
Model (NSF)
ClariNet
Parallel
WaveNet
Naïve model
AR neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
Jordan network
Michael I. Jordan. Serial order: A parallel distributed processing approach. Technical Report 8604, Institute for Cognitive Science, 1986.
Overview
15
Theory: AR neural waveform model
General idea
1
2
3
4
T
…
…
…
1
2
3
Natural
waveform
1 R. J. Williams and D. Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.
16
General idea
1
2
3
4
T
…
…
…
1
2
3
1 R. J. Williams and D. Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.
Theory: AR neural waveform model
17
General idea
…
…
1
2
1
2
3
4
3
T
…
Generated
waveform
Theory: AR neural waveform model
18
MCNN
GELP
Flow-based model
FloWaveNet
WaveGlow
No AR, nor flow
Neural source-filter
Model (NSF)
ClariNet
Parallel
WaveNet
Naïve model
AR neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
1. N. Kalchbrenner, et al Efficient neural audio synthesis. In Proc. ICML, volume 80, pages 2410–2419, 10–15 Jul 2018.
2. J.-M. Valin and J. Skoglund. LPCNet: Improving neural speech synthesis through linear prediction. In Proc. ICASSP, pages 5891–5895, 2019.
3. L. Juvela, et al . Speaker-independent raw waveform model for glottal excitation. In Proc. Interspeech 2018, pages 2012–2016, 2018.
Theory: AR neural waveform model
19
MCNN
GELP
Flow-based model
FloWaveNet
WaveGlow
Theory: flow-based models
No AR, nor flow
Neural source-filter
Model (NSF)
ClariNet
Parallel
WaveNet
Naïve model
AR neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
Flow-based model
FloWaveNet
WaveGlow
20
Revisit AR model
1
2
3
T
1
2
3
NN
1
2
3
T
1
2
3
NN
Training
Generation
G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.
Theory: flow-based models
Or equivalently
21
Revisit AR model
1
2
3
T
1
2
3
NN
1
2
3
T
1
2
3
NN
Training
Generation
NN
z-1
H(.)
NN
z-1
H-1(.)
G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.
Theory: flow-based models
22
Revisit AR model
1
2
3
T
1
2
3
NN
1
2
3
T
1
2
3
NN
Training
Generation
NN
z-1
H(.)
NN
z-1
H-1(.)
G. Papamakarios, T. Pavlakou, and I. Murray. Masked autoregressive flow for density estimation. Proc NIPS, pages 2338–2347, 2017.
Such an AR model is a flow-based model
Training:
Generation:
Theory: flow-based models
23
Theory: flow-based models
From AR to Inverse AR flow-based model
NN
z-1
H(.)
Training
Generation
NN
z-1
H(.)
NN
z-1
H-1(.)
NN
z-1
H-1(.)
AR flow
Inverse-AR flow
24
Theory: flow-based models
From AR to Inverse AR flow-based model
NN
z-1
H(.)
Training
Generation
NN
z-1
H-1(.)
AR flow
NN
z-1
H(.)
NN
z-1
H-1(.)
Inverse-AR flow
Knowledge distilling
Parallel WaveNet & ClariNet
25
MCNN
No AR, nor flow
Neural source-filter
Model (NSF)
Naïve model
AR neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
1. R. Prenger, R. Valle, and B. Catanzaro. Waveglow: A flow-based generative network for speech synthesis. ICASSP 2019, 2018.
2. S. Kim, S.-g. Lee, J. Song, and S. Yoon. Flowavenet: A generative flow for raw audio. ICML 2019.
3. A. van den Oord, Y. Li, I. Babuschkin, et. al.. Parallel WaveNet: Fast high-fidelity speech synthesis. Proc. ICML, pages 3918–3926, 2018.
4. W. Ping, K. Peng, and J. Chen. Clarinet: Parallel wave generation in end-to-end text-to-speech. ICLR, 2018.
Inverse AR flow
FloWaveNet
WaveGlow
ClariNet
Parallel
WaveNet
Theory: flow-based models
26
Inverse AR flow
FloWaveNet
WaveGlow
Theory: neural source-filter model
No AR, no flow
Neural source-filter
Model (NSF) 1
Naïve model
AR neural model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
ClariNet
Parallel
WaveNet
1. X. Wang, et al. Neural source-filter-based waveform model for statistical para- metric speech synthesis. In Proc. ICASSP, pages 5916–5920, 2019.
2. S. O ̈. Arık, et. al. Fast spectrogram inversion using multi-head convolutional neural networks. IEEE Signal Processing Letters, 26(1):94–98, 2018.
3. J. Lauri, et. al. GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram, Proc. Interspeech, 2019
MCNN2
GELP3
27
Theory: neural source-filter model
1
2
3
4
T
…
…
General idea
…
1
2
3
4
T
‘Filter’
Natural
waveform
Generated
waveform
1
2
3
4
T
F0/pitch
‘Source’
28
Theory: neural source-filter model
1
2
3
4
T
…
…
General idea
…
Generated
waveform
Natural
waveform
1
2
3
4
T
Spectral distance
…
…
1
2
3
4
T
F0/pitch
29
…
…
1
2
3
4
T
F0/pitch
Theory: neural source-filter model
1
2
3
4
T
…
Probabilistic interpretation?
Generated
waveform
Natural
waveform
1
2
3
4
T
Spectral distance
…
…
What is the ?
30
Theory: neural source-filter model
1
2
3
4
T
…
Probabilistic interpretation?
1
2
3
4
T
…
Framing
Framing
Spectral
distance
FFT
FFT
31
Theory: neural source-filter model
1
2
3
4
T
…
Probabilistic interpretation?
1
2
3
4
T
…
Framing
Framing
FFT
FFT
Likelihood over Gaussian
32
Theory: neural source-filter model
Probabilistic interpretation?
1
2
3
4
T
…
Framing
FFT
1
2
3
4
T
…
Framing
FFT
Likelihood over Gaussian
33
Naïve model
AR model
Inverse-AR flow
NSF
Theory in summary
34
Contents
Introduction
Theory
Practice
Summary
35
Practice: WaveNet
WaveNet variants
1
2
1
2
3
4
3
1
2
1
2
3
4
3
GMM/Gaussian
Softmax
36
Practice: WaveNet
WaveNet variants
1. T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma. Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications. arXiv preprint arXiv:1701.05517, 2017.
2. Y. Agiomyrgiannakis. B-spline PDF: A generalization of histograms to continuous density models for generative audio networks. In Proc. ICASSP, pages 5649–5653. IEEE, 2018.
3. T. Yoshimura, et al. Mel-cepstrum-based quantization noise shaping applied to neural-network-based speech waveform synthesis. IEEE TASLP, 26(7):1173–1180, 2018.
4. K. Tachibana, T. Toda, Y. Shiga, and H. Kawai. An investigation of noise shaping with perceptual weighting for WaveNet-based speech generation. In Proc. ICASSP, pages 5664–5668. IEEE, 2018.
1
2
1
2
3
4
3
Softmax
1
2
1
2
3
4
3
GMM/Gaussian
37
Generation strategy
1
2
1
2
3
4
3
Practice: WaveNet
1
2
3
4
…
…
…
…
38
Generation strategy
Practice: WaveNet
1
2
3
4
…
…
…
…
Random sampling
Greedy search
39
Generation strategy
Waveform levels (0-1024)
Waveform levels (0-1024)
Practice: WaveNet
59
Practice: WaveNet
Generation method
How about
41
Practice: WaveNet
Natural
Greedy
search
Random
sampling
Mixed
approach
42
Practice: WaveNet
Natural
Greedy
search
Random
sampling
Mixed
approach
43
Generation strategy
Practice: WaveNet
Greedy best?
Sampling?
1. Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu. FFTNet: A real-time speaker-dependent neural vocoder. In Proc. ICASSP, pages 2251–2255, 2018.
44
Practice: WaveNet
Training stability
NN
1
2
1
2
3
4
3
45
Practice: WaveNet
Training stability
46
Negative log-likelihood
Practice: WaveNet
Training stability
Epoch
47
Network training
Epoch
48
Part 1: implicit source-filter model
WaveNet
Illness of fitting Gaussian
Epoch
49
Practice: WaveNet
Training stability
NN
1
2
1
2
3
4
3
50
Training stability
Practice: WaveNet
Negative log-likelihood
Naïve strategy
Our
strategy
Epoch
On training set
On validation set
Step 1
Step 2
51
Training stability
Practice: WaveNet
Negative log-likelihood
Naïve strategy
Our
strategy
Epoch
On training set
On validation set
Step 1
Step 2
52
Generation strategy
Training WaveNet-Gaussian
Practice: WaveNet
Greedy best?
Sampling?
Exploitation + exploration
Keep gradients mild
NN
1
2
1
2
3
4
3
53
Contents
Introduction
Theory
Practice
Summary
54
Practice: NSF
1
2
3
4
T
…
…
General idea
…
Generated
waveform
Natural
waveform
1
2
3
4
T
Spectral distance
…
…
1
2
3
4
T
F0/pitch
55
Practice: NSF
Common structure
Spectral features & F0
Condition module
Source module
Filter module
Frequency-domain distance
Natural
waveform
Generated
waveform
F0 infor.
Spectral infor.
Generated waveform
Gradients
56
Practice: NSF
Common structure
Spectral features & F0
Source module
Filter module
Frequency-domain distance
Natural
waveform
Generated
waveform
Up sampling
Up sampling
Bi-LSTM
CONV
F0
Generated waveform
Gradients
Up sampling
Dimension change
Temporal smoothing
Cat.
57
Practice: NSF
Common structure
Spectral features & F0
Filter module
Frequency-domain distance
Natural
waveform
Generated
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
Generated waveform
Gradients
Up sampling
Bi-LSTM
CONV
Cat.
F0
58
Spectral features & F0
Filter module
Frequency-domain distance
Natural
waveform
Generated
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
Generated waveform
Gradients
Up sampling
Bi-LSTM
CONV
Cat.
F0
Practice: NSF
Common structure
…
Random initial phase
Sampling rate
Noise
FF
Sine
generator
Fundamental component
Voiced:
Unvoiced: noise
59
Practice: NSF
Common structure
Spectral features & F0
Frequency-domain distance
Natural
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
Filter module
Generated waveform
Gradients
Compute frequency-domain distance
Compute gradients for SGD
Up sampling
Bi-LSTM
CONV
Cat.
F0
60
Practice: NSF
Common structure
Spectral features & F0
Natural
waveform
Generated
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
FFT
Framing
FFT
Framing
iFFT
De-framing
Filter module
Up sampling
Bi-LSTM
CONV
Cat.
F0
61
Practice: NSF
Common structure
FFT
Framing
FFT
Framing
iFFT
De-framing
FFT
Framing
FFT
Framing
iFFT
De-framing
FFT
Framing
FFT
Framing
iFFT
De-framing
+
62
Practice: NSF
Common structure
Spectral features & F0
Natural
waveform
Generated
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
FFT
Framing
FFT
Framing
iFFT
De-framing
Filter module
Up sampling
Bi-LSTM
CONV
Cat.
F0
Filter module
63
Practice: NSF
Common structure
Spectral features & F0
Natural
waveform
Generated
waveform
Up sampling
Noise
FF
Sine
generator
harmonics
FFT
Framing
FFT
Framing
iFFT
De-framing
Filter module
Up sampling
Bi-LSTM
CONV
Cat.
F0
Filter module
NSF models
Baseline NSF
(b-NSF)
Simplified NSF
(s-NSF)
Harmonic-plus-noise NSF
(hn-NSF)
hn-NSF
Ver.1
hn-NSF with ver.2
ICASSP 2019
Journal paper submitted
SSW 2019
64
Practice: NSF
Baseline and simplified NSF
Baseline
filter block 1
Baseline
filter block 2
Baseline
filter block 5
…
Simplified
filter block 1
Simplified
filter block 2
Simplified
filter block 5
…
b-NSF
s-NSF
simplify
65
simplify
Practice: NSF
Baseline and simplified NSF
Baseline
filter block 2
Baseline
filter block 5
…
Simplified
filter block 2
Simplified
filter block 5
…
b-NSF
s-NSF
Baseline
filter block 1
Simplified
filter block 1
Simplified filter block
Dilated CONV
+
FF
…
FF
Dilated CONV
+
Baseline filter block
Dilated CONV
+
Tanh
Sigmoid
•
FF
FF
FF
+
Dilated CONV
+
Tanh
Sigmoid
•
FF
FF
+
…
+
FF
66
Practice: NSF
Baseline and simplified NSF
Baseline
filter block 1
Baseline
filter block 2
Baseline
filter block 5
…
Simplified
filter block 1
Simplified
filter block 2
Simplified
filter block 5
…
b-NSF
s-NSF
simplify
67
Practice: NSF
Harmonic-plus-noise NSF
Baseline
filter block 1
Baseline
filter block 2
Baseline
filter block 5
…
Simplified
filter block 1
Simplified
filter block 2
Simplified
filter block 5
…
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
b-NSF
s-NSF
hn-NSF
simplify
upgrade
68
Baseline
filter block 2
Baseline
filter block 5
…
Simplified
filter block 2
Simplified
filter block 5
…
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
b-NSF
s-NSF
hn-NSF
simplification
improvement
Baseline
filter block 1
Simplified
filter block 1
Simplified
filter block 1
Practice: NSF
Harmonic-plus-noise NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Maximum voicing frequency
(MVF)
hn-NSF
69
Practice: NSF
Harmonic-plus-noise NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
u/v flag
For voiced sounds
For unvoiced sounds
Condition module for hn-NSF
Fixed MVFs
70
Practice: NSF
Harmonic-plus-noise NSF
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Condition module for hn-NSF
sinc
Hamming window
Gain norm.
HP
LP
71
Practice: NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
u/v flag
Condition module for hn-NSF
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Condition module for hn-NSF
72
Practice: NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
u/v flag
Condition module for hn-NSF
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Condition module for hn-NSF
73
Practice: NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
u/v flag
Condition module for hn-NSF
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Condition module for hn-NSF
74
Practice: NSF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
u/v flag
Condition module for hn-NSF
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
Simplified
filter block 5
noise
+
HP
LP
Condition module for hn-NSF
75
Practice: NSF
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
NSF is a deep-residual network
76
Practice: NSF
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
NSF is a deep-residual network
77
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
78
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
79
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
80
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
81
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
82
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
83
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
84
Spectral features & F0
Up sampling
Noise
FF
Sine
generator
harmonics
Up sampling
Bi-LSTM
CONV
Cat.
F0
MVF
Simplified
filter block 1
Simplified
filter block 2
…
Simplified
filter block 5
noise
+
HP
LP
Simplified
filter block 5
Condition module for proposed hn-NSF
Source module
Practice: NSF
NSF is a deep-residual network
85
Configuration
Corpus | Size | Note |
ATR Ximera F009 [1] | 15 hours | 16kHz, Japanese, neutral style |
| Feature | Dimension |
Acoustic | Mel-generalized cepstrum coefficients (MGC) or Mel-spectra | 60 80 |
F0 | 1 |
Practice: comparison
WaveNet
softmax
b-NSF
hn-NSF
trainable MVF
WaveNet
Gaussian
s-NSF
hn-NSF
fixed MVF
WORLD
vocoder
86
Speech quality (ICASSP)
Practice: comparison
Copy-synthesis
Pipeline TTS
WaveNet
softmax
b-NSF
hn-NSF
trainable MVF
WaveNet
Gaussian
s-NSF
hn-NSF
fixed MVF
WORLD
vocoder
WORLD
vocoder
WaveNet
softmax
WaveNet
Gaussian
b-NSF
87
Speech quality (Journal paper submitted)
Practice: comparison
WaveNet
softmax
b-NSF
s-NSF
hn-NSF
fixed MVF
WaveNet
softmax
b-NSF
hn-NSF
trainable MVF
WaveNet
Gaussian
s-NSF
hn-NSF
fixed MVF
WORLD
vocoder
88
Speech quality (SSW 2019)
Practice: comparison
WaveNet
softmax
b-NSF
hn-NSF
trainable MVF
WaveNet
Gaussian
s-NSF
hn-NSF
fixed MVF
WORLD
vocoder
WaveNet
softmax
hn-NSF
trainable MVF
hn-NSF
fixed MVF
Natural
89
Generation speed
How many waveform points can be generated in 1s (Tesla p100)?
Practice: comparison
WaveNet
softmax
b-NSF
hn-NSF
trainable MVF
WaveNet
Gaussian
s-NSF
hn-NSF
fixed MVF
WORLD
vocoder
90
Contents
Introduction
Theory
Practice
Summary
91
Summary
AR model
WaveRNN
SampleRNN
FFTNet
WaveNet
LPCNet
ExcitNet
GlotNet
Multi-head
CNN
No AR, no flow
Neural source-filter
Model (NSF)
Naïve model
Inverse AR flow
FloWaveNet
WaveGlow
ClariNet
Parallel
WaveNet
GELP
92
Beyond speech
(c.f. HTS Slides, by HTS Working Group)
Source module
Filter module
93
Beyond speech
Music performance
1 University of Rochester Multi-Modal Music Performance (URMP) Dataset http://www2.ece.rochester.edu/projects/air/projects/URMP.html
Neural
waveform
model
F0
Mel-spectra
| Natural | b-NSF | S-NSF | hn-NSF trainable MVF |
Violin | | | | |
Viola | | | | |
Oboe | | | | |
Trumpet | | | | |
Saxophone | | | | |
Beyond speech
Music performance
WaveNet
| Natural | b-NSF | S-NSF | hn-NSF trainable MVF |
Horn | | | | |
Trombone | | | | |
Tuba | | | | |
Clarinet | | | | |
Flute | | | | |
Beyond speech
Music performance
96
Future direction
(c.f. HTS Slides, by HTS Working Group)
Questions & Comments �are always Welcome!�
97
98
Reference
WaveNet: A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
SampleRNN: S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio. Samplernn: An unconditional end-to-end neural audio generation model. arXiv preprint arXiv:1612.07837, 2016.
WaveRNN: N. Kalchbrenner, E. Elsen, K. Simonyan, et.al. Efficient neural audio synthesis. In J. Dy and A. Krause, editors, Proc. ICML, volume 80 of Proceedings of Machine Learning Research, pages 2410–2419, 10–15 Jul 2018.
FFTNet: Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu. FFTNet: A real-time speaker-dependent neural vocoder. In Proc. ICASSP, pages 2251–2255. IEEE, 2018.
Universal vocoder: J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, and R. Barra-Chicote. Robust universal neural vocoding. arXiv preprint arXiv:1811.06292, 2018.
Subband WaveNet: T. Okamoto, K. Tachibana, T. Toda, Y. Shiga, and H. Kawai. An investigation of subband wavenet vocoder covering entire audible frequency range with limited acoustic features. In Proc. ICASSP, pages 5654–5658. 2018.
Parallel WaveNet: A. van den Oord, Y. Li, I. Babuschkin, et. al.. Parallel WaveNet: Fast high-fidelity speech synthesis. In Proc. ICML, pages 3918–3926, 2018.
ClariNet: W. Ping, K. Peng, and J. Chen. Clarinet: Parallel wave generation in end-to-end text-to-speech. arXiv preprint arXiv:1807.07281, 2018.
FlowWaveNet: S. Kim, S.-g. Lee, J. Song, and S. Yoon. Flowavenet: A generative flow for raw audio. arXiv preprint arXiv:1811.02155, 2018.
WaveGlow: R. Prenger, R. Valle, and B. Catanzaro. Waveglow: A flow-based generative network for speech synthesis. arXiv preprint arXiv:1811.00002, 2018.
RNN+STFT: S. Takaki, T. Nakashika, X. Wang, and J. Yamagishi. STFT spectral loss for training a neural speech waveform model. In Proc. ICASSP (submitted), 2018.
NSF: X. Wang, S. Takaki, and J. Yamagishi. Neural source-filter-based waveform model for statistical para- metric speech synthesis. arXiv preprint arXiv:1810.11946, 2018.
LP-WavNet: M.-J. Hwang, F. Soong, F. Xie, X. Wang, and H.-G. Kang. Lp-wavenet: Linear prediction-based wavenet speech synthesis. arXiv preprint arXiv:1811.11913, 2018.
GlotNet: L. Juvela, V. Tsiaras, B. Bollepalli, M. Airaksinen, J. Yamagishi, and P. Alku. Speaker-independent raw waveform model for glottal excitation. arXiv preprint arXiv:1804.09593, 2018.
ExcitNet: E. Song, K. Byun, and H.-G. Kang. Excitnet vocoder: A neural excitation model for parametric speech synthesis systems. arXiv preprint arXiv:1811.04769, 2018.
LPCNet: J.-M. Valin and J. Skoglund. Lpcnet: Improving neural speech synthesis through linear prediction. arXiv preprint arXiv:1810.11846, 2018.
MCNN: S. O ̈. Arık, H. Jun, and G. Diamos. Fast spectrogram inversion using multi-head convolutional neural networks. IEEE Signal Processing Letters, 26(1):94–98, 2018.
GELP: J. Lauri, et. al. GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-spectrogram, Proc. Interspeech, 2019
99
Reference
By Lauri Juvela, Aalto University
DFT
Framing/
windowing
DFT
Framing/
windowing
Generated waveform
Natural waveform
…
N frames
…
K DFT bins
K-points
DFT
Frame Length M
Padding
K-M
0
0
0
0
0
0
0
0
0
0
Framing/
windowing
Complex-value domain
Real-value domain
Appendix
Training criterion
X
=
…
1st Frame
2nd Frame
Nth Frame
T rows
M (frame length)
…
…
Frame
shift
NM
columns
…
…
0
0
0
0
0
0
0
0
0
0
0
0
Appendix
Training criterion
Training criterion
DFT
Framing/
windowing
DFT
Framing/
windowing
Generated waveform
Natural waveform
…
N frames
…
K DFT bins
K-points
iDFT
Frame Length M
De-framing/
windowing
inverseDFT
De-framing
/windowing
Gradients
Gradients w.r.t. zero-padded part
Not used in de-framing/windowing
Padding
K-M
Complex-value domain
Real-value domain
Appendix
103
flow-based models
Recap AR model
1
2
3
T
1
2
3
NN
NN
z-1
H-1(.)
104
Flow-based models
105
flow-based models
Recap AR model
Triangle-matrix,
as nt depends on o<t
106
flow-based models
Recap AR model
107
flow-based models
Inverse-AR flow
NN
z-1
H-1(.)
Triangle-matrix,
as nt depends on ot
108
flow-based models
Inverse-AR flow
NN
z-1
H-1(.)
109
flow-based models
AR flow vs inverse-AR
NN
z-1
H-1(.)
NN
z-1
H-1(.)
110
flow-based models
NN
z-1
H-1(.)
NN
z-1
H-1(.)
AR flow
AR flow vs inverse-AR
Inverse-AR flow