1 of 48

Sayak Paul | Deep Learning Associate at PyImageSearch

DevFest Goa 2019, September 29

Training Neural Nets: a Hacker’s Perspective

2 of 48

Agenda

  • The motivation
  • Being thorough about training neural nets
    • Training a neural network
    • Maintaining a healthy prototyping process
    • Gradually increasing model complexity
    • ...
  • Guiding lights

3 of 48

The motivation

4 of 48

The motivation

5 of 48

The motivation

“Deep learning neural networks have become easy to define and fit, but are still hard to configure.”

6 of 48

Training a neural network

  • Implementation bugs

7 of 48

Training a neural network

  • Implementation bugs

Wrong shuffling!

8 of 48

Training a neural network

  • Implementation bugs

9 of 48

Training a neural network

  • Implementation bugs

Feature interpretation

10 of 48

Training a neural network

  • Implementation bugs

11 of 48

Training a neural network

  • Implementation bugs
    • Wrong activation function & initialization
      • tanh + He -> tanh + Xavier
      • ReLU + Xavier -> ReLU + He

12 of 48

Training a neural network

  • Implementation bugs
    • Wrong activation function & initialization
    • Forgetting to zero out gradients (PyTorch)

No zero outs!

13 of 48

Training a neural network

  • Implementation bugs
    • Wrong activation function & initialization
    • Forgetting to zero out gradients (PyTorch)

14 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
    • Too high learning rate leads to numerical instability
      • Loss quantity coming out as NaNs

15 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
    • Too high learning rate leads to numerical instability
    • Too few epochs
      • Network not trained enough

16 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
    • Too high learning rate leads to numerical instability
    • Too few epochs
    • Too large batch size for a small dataset

17 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
  • Dataset construction and others
    • Dataset distribution

18 of 48

Training a neural network

19 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
  • Dataset construction and others
    • Dataset distribution
    • Effects of random data augmentation

20 of 48

Training a neural network

21 of 48

Training a neural network

  • Implementation bugs
  • Model’s sensitivity towards hyperparameter choices
  • Dataset construction and others
    • Dataset distribution
    • Effects of random data augmentation
    • Wrong normalization statistics
    • Label noise
    • Class imbalance

22 of 48

Maintaining a healthy prototyping process

  • Write code quickly
    • Reuse existing codebases for quick baselines
    • But be very careful

23 of 48

Maintaining a healthy prototyping process

  • Write code quickly
  • Run experiments and keep track of what you tried

24 of 48

Maintaining a healthy prototyping process

  • Write code quickly
  • Run experiments and keep track of what you tried

25 of 48

Maintaining a healthy prototyping process

  • Write code quickly
  • Run experiments and keep track of what you tried
  • Analyze model behavior. Did it do what you wanted?

26 of 48

Overfitting a single batch of data

  • The loss could go up instead of down

27 of 48

Overfitting a single batch of data

  • The loss could go up instead of down
  • The loss could go down for a while, then explode

28 of 48

Overfitting a single batch of data

  • The loss could go up instead of down
  • The loss could go down for a while, then explode
  • The loss could oscillate across a region

29 of 48

Overfitting a single batch of data

  • The loss could go up instead of down
  • The loss could go down for a while, then explode
  • The loss could oscillate across a region
  • The loss could get down to a scalar quantity (0.01, for example) and not get any better than that

30 of 48

Model complexity as a f(order)

Deciding on a model architecture:

31 of 48

Model complexity as a f(order)

Deciding on a model architecture:

  • Time to train the network
  • Size of the final network
  • Inference speed
  • Accuracy

32 of 48

Model complexity as a f(order)

Some points to remember here:

  • Ramping up the model complexity gradually

33 of 48

Model complexity as a f(order)

Visualize performance

Visualize performance

Deep learning is a shout in the void and the oblivion is inevitable!

34 of 48

Model complexity as a f(order)

Some points to remember here:

  • Ramping up the model complexity gradually

35 of 48

Model complexity as a f(order)

Some points to remember here:

  • Ramping up the model complexity gradually
  • Experimentation with several random subsets
  • Human evaluation
  • Visualizing the intermediate activations of the model

36 of 48

Model complexity as a f(order)

37 of 48

Chasing the hyperparameters

  • Declarative configuration

38 of 48

Chasing the hyperparameters

  • Declarative configuration

39 of 48

Chasing the hyperparameters

  • Declarative configuration
  • Organizing the hyperparameter tuning process

40 of 48

Chasing the hyperparameters

  • Declarative configuration
  • Organizing the hyperparameters’ search process

41 of 48

Going beyond ...

  • Model ensembling

42 of 48

Going beyond ...

  • Model ensembling
  • Knowledge distillation

43 of 48

Going beyond ...

  • Model ensembling
  • Knowledge distillation
  • Lottery ticket hypothesis

44 of 48

Going beyond ...

  • Model ensembling
  • Knowledge distillation
  • Lottery ticket hypothesis
  • Model quantization

45 of 48

Summary

  • Training (and debugging) neural nets
  • Healthy model prototyping
  • The idea of overfitting a single batch of data
  • Practical hyperparameter tuning
  • Squeezing the best out of a network

46 of 48

References

47 of 48

References

Slides are available here: http://bit.ly/DFGoa19

48 of 48

See you next time

Find me here:

sayak.dev

Thank you very much :)