AALBERT
Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
Recap: ALBERT
ALBERT
1.Factorize Embedding Matrix
Original BERT:
30000 x 768 = 23.04M
ALBERT:
30000 x128 = 3.8M
128 x 768 = 0.098M
Total: 3.898M
Reduce Parameters !
ALBERT
2.Shared Same Parameters across Layer
1/ 12 BERT Parameters on Layer
Reduce Parameters !!!
Model Configuration
AALBERT
v.s ALBERT
v.s Mockingjay
Configuration
Pre-Training Stage
LibriSpeech 360 hours dataset, 500k step, batch size 48.
Phoneme Classification
Phoneme Classification task
Weighted-sum and Fine-tune version
Different Proportion of training data
(Weighted-sum) (Fine-tune)
Speaker Identification
Utterance-level
T-sne visualization
Frame-level
Overall Performance
Probing Tasks