2023.06.09
제네시스랩 AI연구팀
신재영
ICASSP 2022
What Is Grapheme-to-Phoneme?
Background
Limitations of Previous Studies
Introduction
Motivation
Introduction
GBERT
Introduction
Introduction
GBERT vs BERT
Method
Pre-training GBERT for G2P Conversion
Method
Fine-tuning GBERT for G2P Conversion
Method
BERT-fused model
Jinhua Zhu, et al. “Incorporating BERT into Neural Machine Translation”, ICLR 2020
Method
BERT-fused model
Jinhua Zhu, et al. “Incorporating BERT into Neural Machine Translation”, ICLR 2020
Method
Fusing GBERT into Transformer-based G2P Model
Datasets
Experiments
WER and PER results
Results
※Low-resource : randomly sampled 1000 records from the original training set
WER and PER results
Results
Conclusion
Results
THANK YOU FOR LISTENING
Limitations of Previous Studies
Introduction
The hierarchical prosody annotation adopted in this work categorizes the prosodic boundaries of Mandarin speech into five levels, which from low to high are Character (CC), Lexicon Word (LW), Prosodic Word (PW), Prosodic Phrase (PPH) and Intonational Phrase (IPH). Prosodic Word (PW), Prosodic Phrase (PPH) and Intonational Phrase (IPH) correspond to three different lengths of pause in speech from short to long. Lexicon Word (LW) indicates syntactic boundary between words, and Chinese Character (CC) is the smallest unit of Chinese.