1 of 15

2022.12.16

제네시스랩 AI연구팀

신재영

NeurIPS 2020

2 of 15

Text-To-Speech (TTS)

Introduction

Ma, Mingbo, et al. “Incremental Text-to-Speech Synthesis with Prefix-to-Prefix Framework”, Findings of EMNLP 2020

3 of 15

Autoregressive Models

Introduction

Autoregressive TTS models have a few difficulties in deploying them directly in real-time services.

Most of the autoregressive models show a lack of robustness in some cases.

When an input text includes repeated words, autoregressive TTS models produce serious attention errors.

Shen, Jonathan, et al. “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”, ICASSP 2018

Tacotron2

4 of 15

Non-autoregressive Models

Introduction

FastSpeech

Pre-defined

Alignment

from

Autoregressive model

Ren, Yi, et al. “FastSpeech: Fast, Robust and Controllable Text to Speech”, NeurIPS 2019

5 of 15

Motivation

Introduction

Glow-TTS jointly learns to align text and speech, and generate speech in parallel

Invertible Decoder

Most Likely Alignment

6 of 15

Normalizing Flows

Background

Dinh, Laurent, et al., “DENSITY ESTIMATION USING REAL NVP”, ICLR 2017

Glow

Invertible Transformation

7 of 15

Flow-based

Flow-based TTS Modeling

x and c denote the input mel spectrogram and text sequence, respectively

Prior Distribution

Flow-based Decoder

Goal

 

8 of 15

Maximum Likelihood Estimation

Exact log likelihood by Normalizing Flows

Objective

Flow-based

9 of 15

Monotonic Alignment Search

Monotonic Alignment Search

Glow-TTS is designed to generate a mel-spectrogram conditioned on

a monotonic and non-skipping alignment between text and speech representations.

Find the best one among all possible alignments

Monotonic Alignment Search

10 of 15

Viterbi Algorithm

A dynamic programming method to search for the most probable alignment between text and the latent representation of speech

It runs efficiently on CPU and amounts to less than 2% of the total training time

Monotonic Alignment Search

11 of 15

Alignment Generation via Duration Predictor

Estimate the alignment by only using text information

Alignment Generation

12 of 15

Procedure

Glow-TTS

Glow-TTS jointly learns to align text and speech, and generate speech in parallel

13 of 15

Audio Quality

Results

 

Glow-TTS shows comparable performance to Tacotron 2

  • Single Speaker Dataset : LJSpeech (13,100 short audio clips, total 24h)
  • Multi Speaker Dataset : train-clean-100 subset of the LibriTTS corpus(247 speakers, total 54 h)
  • Text Dataset : 227 utterances from the book Harry Potter and the Philosopher’s Stone (maximum length: 800 ↑)

14 of 15

Inference Time and Length Robustness

Results

Nearly constant inference time

Robust to extremely long utterance generation

  • Single Speaker Dataset : LJSpeech (13,100 short audio clips, total 24h)
  • Multi Speaker Dataset : train-clean-100 subset of the LibriTTS corpus(247 speakers, total 54 h)
  • Text Dataset : 227 utterances from the book Harry Potter and the Philosopher’s Stone (maximum length: 800 ↑)

15 of 15

THANK YOU FOR LISTENING