2023.07.07
제네시스랩 AI연구팀
신재영
PMLR 2022
Motivation
Introduction
Background
Introduction
VITS
JaeHyeon Kim, et al. “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”, PMLR 2021
Background
Introduction
Speaker Consistency Loss
Detai Xin, et al. “Cross-lingual Speaker Adaptation using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis”, INTERSPEECH 2021
Architecture
Overall Structure
Architecture
Speaker Encoder
Hee Soo Heo, et al. “Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020”, Arxiv 2020
H/ASP model configuration
Method
Objective Function
�
Experiment
Audio datasets
�
Experiment
Setup
�
Experiment
Setup
Detai Xin, et al. “Cross-lingual Speaker Adaptation using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis”, INTERSPEECH 2021
Zero-shot Multi-Speaker TTS
Results
between the speaker embeddings of two audios extracted from the speaker encoder. �It ranges from -1 to 1, and a larger value indicates a stronger similarity
Contribution
Results
Limitation
Results
THANK YOU FOR LISTENING