1 of 20

AALBERT

Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation

2 of 20

Recap: ALBERT

  • Factorize Embedding Matrix
  • Share Parameters across layer
  • Model Configuration between BERT

3 of 20

ALBERT

1.Factorize Embedding Matrix

Original BERT:

30000 x 768 = 23.04M

ALBERT:

30000 x128 = 3.8M

128 x 768 = 0.098M

Total: 3.898M

Reduce Parameters !

4 of 20

ALBERT

2.Shared Same Parameters across Layer

1/ 12 BERT Parameters on Layer

Reduce Parameters !!!

5 of 20

Model Configuration

6 of 20

AALBERT

7 of 20

v.s ALBERT

8 of 20

v.s Mockingjay

9 of 20

Configuration

10 of 20

Pre-Training Stage

LibriSpeech 360 hours dataset, 500k step, batch size 48.

11 of 20

Phoneme Classification

  • Weighted-sum and fine-tune feature extraction
  • Different Proportion of training data

12 of 20

Phoneme Classification task

  • Utilizing MLP classifier behind representation to train phoneme classification task.
  • Weighted-sum, Fine-tune.

13 of 20

Weighted-sum and Fine-tune version

14 of 20

Different Proportion of training data

(Weighted-sum) (Fine-tune)

15 of 20

Speaker Identification

  • Utterance-level
  • Frame-level
  • Overall Performance

16 of 20

Utterance-level

  1. Utilizing mean pooling over an utterance to generate utterance-level representation.
  2. Simple linear classifier need to train in the Utterance-level speaker identification

17 of 20

T-sne visualization

18 of 20

Frame-level

  1. Classify Each frame-level representation to corresponding speaker.
  2. Simple linear Classifier need to train in the frame-level speaker identification

19 of 20

Overall Performance

20 of 20

Probing Tasks