1 of 16

Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results

Antti Tarvainen, Harri Valpola

NIPS 2017

Presenter: Tolunay Durmuş

METU CENG 501 Deep Learning

In-class paper presentation

2 of 16

Outline

  • Introduction
  • Background
  • Mean-Teacher Model and Its Method
  • Experiments
  • Conclusion

2

METU CENG 501 Deep Learning - Paper presentation

3 of 16

Introduction

  • The problem that the paper aims to solve is to increase the classification accuracy of semi-supervised learning with less labeled data.

  • Supervised learning: Every input has a label.
    • Adding noise, regularization, drop-out

  • Semi-supervised learning : A few of the inputs are labeled, and most of them are unlabeled.

3

METU CENG 501 Deep Learning - Paper presentation

4 of 16

Backgrounds

  • Γ model: Rasmus et al. (2015)
    • the ladder network
    • teacher and student
    • consistency cost

4

METU CENG 501 Deep Learning - Paper presentation

5 of 16

      • Π model
      • Temporal Ensembling
      • Virtual Adversarial Training
      • The Mean Teacher

5

: Laine & Aila (2016)

: Miyato et al. (2017)

METU CENG 501 Deep Learning - Paper presentation

6 of 16

6

METU CENG 501 Deep Learning - Paper presentation

7 of 16

7

METU CENG 501 Deep Learning - Paper presentation

8 of 16

Mean-Teacher

8

METU CENG 501 Deep Learning - Paper presentation

9 of 16

Experiments

9

Error rate percentage on SVHN over 10 runs (4 runs when using all labels).

METU CENG 501 Deep Learning - Paper presentation

10 of 16

10

Error rate percentage on CIFAR-10 over 10 runs (4 runs when using all labels).

METU CENG 501 Deep Learning - Paper presentation

11 of 16

11

Smoothened classification cost (top) and classification error (bottom) of Mean Teacher and our baseline Π model on SVHN over the first 100000 training steps.

METU CENG 501 Deep Learning - Paper presentation

12 of 16

12

Validation error on 250-label SVHN over four runs per hyper-parameter setting and their means.

METU CENG 501 Deep Learning - Paper presentation

13 of 16

13

Error rate percentage of ResNet Mean Teacher compared to the state of the art.

(over 10 runs on CIFAR-10 and validation over 2 runs on ImageNet)

METU CENG 501 Deep Learning - Paper presentation

14 of 16

Conclusion

  • Large Dataset and on-line learning
  • Learning speed
  • Classification accuracy
  • Scalability

14

METU CENG 501 Deep Learning - Paper presentation

15 of 16

REFERENCES

  • Rasmus, A., Valpola, H., Honkala, M., Berglund, M., & Raiko, T. (2015). Semi-supervised learning with ladder networks. arXiv preprint arXiv:1507.02672.
  • Tarvainen, A., & Valpola, H. (2017). Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780.
  • Laine, S., & Aila, T. (2016). Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242.
  • Semi-supervised learning. 2021. Wikipedia. https://en.wikipedia.org/wiki/Semi-supervised_learning.

15

METU CENG 501 Deep Learning - Paper presentation

16 of 16

Thank you for listening.

METU CENG 501 Deep Learning - Paper presentation