1 of 22

Author: Ekaterina Lobacheva, Nadezhda Chirkova, Maxim Kodryan, Dmitry Vetrov

Presented by: Tianyu Zhang & Yusong Wu

1

2 of 22

Introduction

  • Limitation of neural networks
    • Overconfidence
    • Vulnerability to adversarial attacks
    • Overfitting
  • Ensemble works but what do we do when resources are limited (fixed)
  • Metric of the quality of uncertainty estimation
    • calibrated negative log-likelihood

2

3 of 22

Introduction

  • Fix the network size s + increase the ensemble size n
  • Fix the ensemble size n + increase the network size s
  • Fix the ratio between the network size s and the ensemble size n + increase total number of parameters

3

4 of 22

Contributions

  • Derive the conditions under which CNLL follows a power law
  • Present the power laws on the whole considered range of their arguments
    • CNLL(n), n is the ensemble size
    • CNLL(s), s is the network size
    • CNLL(m), m is the total parameter count
  • Make important conclusions: given fixed resources 1 large NN < more small NN
  • CNLL power laws and the optimal memory split can be predicted

4

5 of 22

Notations

  • [def] Power Law
    • X�
    • X�
    • X�
    • X�
  • [def] the Optimal Memory Split: given fixed budget, the best way to balance (split) ensemble size n and the network size s

5

6 of 22

Theoretical view

  • The model-average NLL of an ensemble of size n for the given object����
  • ����

6

7 of 22

Theoretical view

  • X�
  • The comparison of the NLLs of different models with suboptimal softmax temperature may lead to an arbitrary ranking of the models�
  • The comparison should only be performed after calibration��
  • ����

7

8 of 22

Theoretical view

  • ���

8

9 of 22

Theoretical view

  • �������������
  • the difference between the values of LE-NLLn and CNLLn is negligible in practice.

9

10 of 22

Experiments Setup

  • Networks: WideResNet, VGG16
  • Datasets: CIFAR-10, CIFAR-100
  • Network size change: filter width and fully-connect neurons
  • Standard configuration: 15.3M / 36.8M parameters

  • Training details:
    • Grid search weight decay and dropout for every network
    • Batch size of 128, 200 epochs, annealing learning schedule
    • Train multiple groups of ensembles and average performance

10

11 of 22

Approximating Power Law

  • Given sequence of data points:
  • Fit in power law:

  • Finding a,b,c: optimize following objective using BFGS:

  • Equivalent to fitting the linear regression model in log space

11

12 of 22

Experiment Questions

  • Factors:
    • Ensemble size
    • Network size
    • Number of parameters
  • Performance metrics:
    • NLL
    • CNLL
  • Goal:
    • Is there power law between factors and performance metrics?
    • If so, how effective it is for each pair?
    • Can we predict with power law to find optimal ensemble?

12

13 of 22

NLL - Ensemble Size

  • NLL &CNLL - ensemble size has power law in all cases
  • CNLL is very close to lower envelope of NLL
  • (VGG on CIFAR-100)

13

14 of 22

NLL - Ensemble Size

  • NLL &CNLL - ensemble size has power law in all cases

14

15 of 22

NLL - Network Size, single network

  • CNLL - network size has power law
  • NLL - network size does not have power law (double descent)

15

16 of 22

NLL - Network Size, multiple networks

  • NLL&CNLL - network size does not has power law
  • Larger networks starts have increasing CNLL
  • Probably due to lack of regularization

16

17 of 22

NLL - Parameter Count

  • CNLL - parameter count has power law

17

18 of 22

NLL - Parameter Count

  • Fior a fixed memory budget, there is an optimal point
  • Train one large network < train ensemble of small networks
  • Also apply for accuracy

18

19 of 22

Prediction Using Power Laws

  • CNLL can be accurately predicted

19

20 of 22

Prediction Using Power Laws

  • Optimal ensemble can be predicted accurately

20

21 of 22

Conclusion

  • The CNLL as a function of n follows a power law on the wide finite range of n, starting from n = 1, but with the power parameter slightly higher than the one derived theoretically�
  • The CNLL of a single network follows a power law as a function of the network size s on the whole reasonable range of network sizes, with the power parameter approximately the same as derived�
  • The CNLL also follows a power law as a function of the total parameter count (memory budget)�
  • For a given memory budget, the number of networks in the optimal memory split is usually much higher than one, and can be predicted using the discovered power laws.

21

22 of 22

Questions?