1 of 10

Weighted Ensemble Self-Supervised Learning

Yangjun Ruan*, Saurabh Singh, Warren Morningstar, Alexander A. Alemi�Sergey Ioffe, Ian Fischer, Joshua V. Dillon

ICLR 2023

2 of 10

Overview

Ensembling has proven a simple yet effective method for…

Self-supervised learning (SSL) has emerged as the dominating ML paradigm

2

  • Improving model accuracy (left)
  • Capturing predictive uncertainty (right)
  • … yet mostly in supervised learning!
  • Ensembling for SSL is underexplored
  • It is not obvious where and how to ensemble

3 of 10

Overview

We develop an efficient ensemble method tailored for SSL

3

  • Significantly improve downstream performance, particularly for few-shot learning
  • No inference cost during downstream evaluation
  • Potentially applicable to a wide range of SSL methods

4 of 10

Method - Motivation

Many SSL methods utilize a projection head for learning better representations

4

  • Examples: SimCLR, BYOL, DINO, …
  • Crucial for downstream performance
  • Thrown away during downstream evaluation ⇒ No inference cost!

5 of 10

Method - Where to Ensemble?

Only ensemble the projection heads (and optionally other non-encoder parts)

5

  • Ensemble heads facilitate the learning of a single encoder
  • No additional inference overhead

6 of 10

Method - How to Ensemble?

A data-dependant weighted objective for learning the ensemble

6

  • Weight each teacher/student pair to enable preferential treatment
  • Different weightings are used based on the inputs

7 of 10

Method - How to Weight?

We explore numerous weighting schemes

7

  • Uniform weighting
    • Treat all teacher/student pairs equally
  • Probability weighting
    • Favor student head with maximum prediction confidence
  • Entropy weighting
    • Favor teacher head with minimum prediction entropy

8 of 10

Experiment - Setup

We demonstrate the effectiveness of our methods with an extensive study

8

  • Two state-of-the-art SSL methods
    • DINO (Caron et al., 2021)
    • MSN (Assran et al., 2022)
  • Multiple metrics on ImageNet-1K
    • Decodability: Linear or k-NN evaluation
    • Label efficiency: Few-shot evaluation (1/2/5-shot, 1% data)
  • Improved baselines with careful tuning for fair comparison
    • A smaller projection head significantly improves few-shot performance!

9 of 10

Experiment - Controlled Study

Better weighting schemes encourage more diverse ensembles

Comparison of different weighting schemes with DINO ViT-S/16

Visualization of ensemble diversity

9

10 of 10

Experiment - Improving the SOTA

Our method consistently improves the SOTA

DINO ViT-S/16

DINO ViT-B/8

MSN ViT-S/16

Our improvements include both baseline improvements (dark) and ensembling (light)

10

  • Across all setups and metrics
  • Substantial gains in few-shot evaluation