2022.12.16
제네시스랩 AI연구팀
신재영
NeurIPS 2020
Text-To-Speech (TTS)
Introduction
Ma, Mingbo, et al. “Incremental Text-to-Speech Synthesis with Prefix-to-Prefix Framework”, Findings of EMNLP 2020
Autoregressive Models
Introduction
Autoregressive TTS models have a few difficulties in deploying them directly in real-time services.
Most of the autoregressive models show a lack of robustness in some cases.
When an input text includes repeated words, autoregressive TTS models produce serious attention errors.
Shen, Jonathan, et al. “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”, ICASSP 2018
Tacotron2
Non-autoregressive Models
Introduction
FastSpeech
Pre-defined
Alignment
from
Autoregressive model
Ren, Yi, et al. “FastSpeech: Fast, Robust and Controllable Text to Speech”, NeurIPS 2019
Motivation
Introduction
Glow-TTS jointly learns to align text and speech, and generate speech in parallel
Invertible Decoder
Most Likely Alignment
Normalizing Flows
Background
Dinh, Laurent, et al., “DENSITY ESTIMATION USING REAL NVP”, ICLR 2017
Glow
Invertible Transformation
Flow-based
Flow-based TTS Modeling
x and c denote the input mel spectrogram and text sequence, respectively
Prior Distribution
Flow-based Decoder
Goal
Maximum Likelihood Estimation
Exact log likelihood by Normalizing Flows
Objective
Flow-based
Monotonic Alignment Search
Monotonic Alignment Search
Glow-TTS is designed to generate a mel-spectrogram conditioned on
a monotonic and non-skipping alignment between text and speech representations.
→ Find the best one among all possible alignments
Monotonic Alignment Search
Viterbi Algorithm
A dynamic programming method to search for the most probable alignment between text and the latent representation of speech
It runs efficiently on CPU and amounts to less than 2% of the total training time
Monotonic Alignment Search
Alignment Generation via Duration Predictor
Estimate the alignment by only using text information
Alignment Generation
Procedure
Glow-TTS
Glow-TTS jointly learns to align text and speech, and generate speech in parallel
Audio Quality
Results
Glow-TTS shows comparable performance to Tacotron 2
Inference Time and Length Robustness
Results
Nearly constant inference time
Robust to extremely long utterance generation
THANK YOU FOR LISTENING