1 of 90

Training CNNs

Adapted from Deepak Pathak

2 of 90

Slide Credit: Kris Kitani

3 of 90

Slide Credit: Kris Kitani

4 of 90

Slide Credit: Kris Kitani

5 of 90

Slide Credit: Kris Kitani

6 of 90

Slide Credit: Kris Kitani

7 of 90

Slide Credit: Kris Kitani

8 of 90

Slide Credit: Kris Kitani

9 of 90

Slide Credit: Kris Kitani

10 of 90

Last Lecture: CNN Architectures

CNNs: Convolutional Neural Networks

  • Layers in Standard CNN
    • Convolution
    • Pooling
    • FC
    • Dropout
  • Popular Architectures
    • LeNet
    • AlexNet
    • VGG
    • (New) GoogleNet
    • (New) ResNet

11 of 90

AlexNet: How Many Parameters?

Slide Credit: David Fouhey

12 of 90

AlexNet: How Many Parameters?

Slide Credit: David Fouhey

  • 62.4 million parameters
  • 6 days to train (2012)
  • 8 convolutional layers
  • Vast majority are in fully-connected layers
  • The paper notes that removing convolutions is disastrous for performance.

13 of 90

14 of 90

AlexNet: Training Trick #1 DropOut

15 of 90

AlexNet: Training Trick #2 Data Augmentation

Slide Credit: Kris Kitani

16 of 90

VGG-16: Going Deeper!

Slide Credit: David Fouhey

nn.Sequential(nn.Conv2d, nn.BatchNorm2d, nn.ReLU, nn.MaxPool2d)

All 3x3 convolutions

17 of 90

VGG-16: What are Convolutions Learning?

What features can you see?

18 of 90

VGG-16: Emerging ‘Rule of Thumb’

  • Convolutions all standardized to 3x3
  • Convolutions always followed by ReLU
  • Stack several convolutions at same resolution
  • Downsample by 2, increase channels by 2 to create a bottleneck

19 of 90

Receptive Field of 3 Stacks of 3x3 Convolutions?

Slide Credit: Kris Kitani

20 of 90

Receptive Field of 3 Stacks of 3x3 Convolutions?

Slide Credit: Kris Kitani

21 of 90

What do you Gain by Stacking Convolutions?

Slide Credit: Kris Kitani

22 of 90

VGGNet: Training Trick #3 Pre-Training

  • Training VGGNet-{13, 16, 19} was incredibly challenging. Why?

23 of 90

Training Deeper Networks

Every backpropagation step multiplies the gradient by the local gradient

1 * d * d * d … * d = dn-1

If d << 1, n is big → Vanishing Gradients

24 of 90

Training Deeper Networks

Every backpropagation step multiplies the gradient by the local gradient

1 * d * d * d … * d = dn-1

If d >> 1, n is big → Exploding Gradients

25 of 90

VGGNet: Training Trick #3 Pre-Training

  • Training VGGNet-{13, 16, 19} was incredibly challenging. Why?
    • Back-propagation in large networks is unstable

  • Soln: Train VGG-11

Use weights to initialize VGG-13

Train VGG-13

Train VGG-19

26 of 90

VGGNet: Training Trick #4 Batch Normalization

Slide Credit: Justin Johnson

27 of 90

VGGNet: Training Trick #4 Batch Normalization

Slide Credit: Justin Johnson

Problem #1: What if zero-mean and unit variance is too hard of a constrain?

28 of 90

Slide Credit: Justin Johnson

29 of 90

Slide Credit: Justin Johnson

Problem #2: Mean and variance estimates depend on the mini-batch. We only see one example at a time during test-time.

30 of 90

Batch Normalization at Test Time

Slide Credit: Justin Johnson

31 of 90

Batch Norm at Training V. Test Time

Slide Credit: Justin Johnson

Training

  • Take a batch, compute mean and variance
  • Normalize input with computed mean and variance
  • Feed normalized input to activation layer

Testing

  • Compute mean and variance

  • Normalize input with computed mean and variance
  • Feed normalized input to activation layer

32 of 90

Batch Normalization Improves Training

Slide Credit: Justin Johnson

  1. Deep networks are much easier to train
  2. Allows for higher learning rates, faster convergence
  3. Acts as regularization during training
  4. Not well-understood theoretically (yet)
  5. Behaves differently during training and testing (common bugs)

33 of 90

Alternate Normalization Techniques

34 of 90

GoogLeNet: Why Only 3x3 Convolution?

35 of 90

GoogLeNet: Everything But the Kitchen Sink

36 of 90

Intermediate Loss Layer Prevents Vanishing Gradients

Stagewise training in 3 parts

37 of 90

ImageNet 14’: VGGNet: 7.30% Error, GoogLeNet: 6.67% Error

38 of 90

ResNet: Scaling to 100+ Layers

  • Simply stacking 3x3 convolutions (like VGG-Net) does not provide improvement
  • 56-layer net has higher training error and test error than 20-layer net

39 of 90

Deep Residual Learning

If x is optimal, we can set weight layers to 0. It's easy to “ignore” extra model capacity.

If x is close to optimal, it's easier to find small weights.

Ensembling Effect: This could be seen as an ensemble of networks of varying depth

Identity serves as a “gradient highway” to preserve strong signal from output to earlier layers during back propagation.

40 of 90

41 of 90

42 of 90

43 of 90

DenseNet: Connecting All Layers with Skip Connections

44 of 90

Training Neural Networks

  • Data Normalization
  • Weight Initialization
  • Activation Functions
  • Interpreting Loss Curves
  • Adaptive Optimizers
  • Regularization
  • Feature Normalization
  • Model Selection
  • Hyperparameter Tuning

45 of 90

Input Data Preprocessing: Subtract Mean, Divide by STD

46 of 90

Weight Initialization: Constant Values

Q: What if we initialize all weights/biases to a constant value?

47 of 90

Weight Initialization: Constant Values

Q: What if we initialize all weights/biases to a constant value?

A: We cannot break symmetry of weight values. Gradients are the same at each location.

48 of 90

Weight Initialization: Random Values

Q: What if we initialize all weights/biases to a (small) random value?

49 of 90

Weight Initialization: Random Values

Q: What if we initialize all weights/biases to a (small) random value?

A: This is usually fine for shallow networks, but becomes problematic for deep networks.

50 of 90

Weight Initialization for Deeper Models

Slide Credit: Kris Kitani

51 of 90

52 of 90

Slide Credit: Rawal Khirodkar

53 of 90

Understanding Loss Curves

54 of 90

Choosing the Correct Learning Rate

Slide Credit: Rawal Khirodkar

55 of 90

Learning Rate Scheduling

  • Dropping learning rate helps settle into a local minima
  • Increasing learning rate helps get out of a local minima

56 of 90

Learning Rate Scheduling Varieties

  • Personal Favorite: Decay on (Validation Accuracy) Plateau

57 of 90

Diagnosing Loss Curves

58 of 90

Diagnosing Loss Curves

59 of 90

Diagnosing Loss Curves

60 of 90

Diagnosing Loss Curves

61 of 90

Diagnosing Loss Curves

62 of 90

Diagnosing Loss Curves

63 of 90

Diagnosing Loss Curves

64 of 90

Diagnosing Loss Curves

65 of 90

Diagnosing Loss Curves

66 of 90

Diagnosing Loss Curves

67 of 90

Diagnosing Loss Curves

68 of 90

Diagnosing Loss Curves

69 of 90

Choosing Hyperparameters

Learning Rate: Look at Loss Curves (1e-3 Safe Bet)

Batch Size: As Large as Possible

# of Epoch: Until Validation Accuracy Saturates

Network Architecture: ResNet (Safe Bet)

Optimizer: Adam (Safe Bet)

For Everything Else: Experimental Validation

  • Grid Search
  • Random Search
  • Bayesian Optimization

Which is better? Why?

70 of 90

Choosing Hyperparameters

Learning Rate: Look at Loss Curves (1e-3 Safe Bet)

Batch Size: As Large as Possible

# of Epoch: Until Validation Accuracy Saturates

Network Architecture: ResNet (Safe Bet)

Optimizer: Adam (Safe Bet)

For Everything Else: Experimental Validation

  • Grid Search
  • Random Search
  • Bayesian Optimization

For low dimensional search spaces, random search works better in practice.

  • RS: Search 9 distinct values of the important parameter
  • GS: Search 3 distinct values of the important parameter

71 of 90

Optimizers: Navigating the Loss Ravines

72 of 90

Drawbacks of Gradient Descent

Slide Credit: Rawal Khirodkar

  • Long, narrow ravines

  • Lots of sloshing around the walls

  • Only a small derivative in the direction of steepest descent

73 of 90

74 of 90

75 of 90

76 of 90

77 of 90

78 of 90

79 of 90

80 of 90

81 of 90

82 of 90

83 of 90

84 of 90

85 of 90

Regularization: Limit the Complexity of your Model

86 of 90

Regularization: Toy Example

87 of 90

Regularization: Toy Example

88 of 90

Regularization: Not Too Little, Not Too Much

89 of 90

Training Neural Networks Summary

  • Data Normalization
  • Weight Initialization
  • Activation Functions
  • Interpreting Loss Curves
  • Adaptive Optimizers
  • Regularization
  • Feature Normalization
  • Model Selection
  • Hyperparameter Tuning

Training Neural Networks is an Art and a Science

90 of 90

Personal Experience Debugging Neural Networks

  1. Make small delta’s between experiments to localize errors
  2. Establish the invariants. What must be true? (Assert Statement)
  3. Know what the output should look like. Is the result reasonable?
  4. Most bugs stem from the data loader and pre-processing
  5. Visualize the output! Get friendly with the pixels
  6. Don’t try to reimplement papers. Use off the shelf code whenever possible.
  7. Fix the (PyTorch, NumPy, OS) random seed when debugging
  8. Log everything! Tensorboard and W&B are great
  9. Make sure you are running inference in eval mode
  10. Learn to read and google error messages

Complex Error Logs = Simple Bugs