Training CNNs
Adapted from Deepak Pathak
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Slide Credit: Kris Kitani
Last Lecture: CNN Architectures
CNNs: Convolutional Neural Networks
AlexNet: How Many Parameters?
Slide Credit: David Fouhey
AlexNet: How Many Parameters?
Slide Credit: David Fouhey
AlexNet: Training Trick #1 DropOut
AlexNet: Training Trick #2 Data Augmentation
Slide Credit: Kris Kitani
VGG-16: Going Deeper!
Slide Credit: David Fouhey
nn.Sequential(nn.Conv2d, nn.BatchNorm2d, nn.ReLU, nn.MaxPool2d)
All 3x3 convolutions
VGG-16: What are Convolutions Learning?
What features can you see?
VGG-16: Emerging ‘Rule of Thumb’
Receptive Field of 3 Stacks of 3x3 Convolutions?
Slide Credit: Kris Kitani
Receptive Field of 3 Stacks of 3x3 Convolutions?
Slide Credit: Kris Kitani
What do you Gain by Stacking Convolutions?
Slide Credit: Kris Kitani
VGGNet: Training Trick #3 Pre-Training
Training Deeper Networks
Every backpropagation step multiplies the gradient by the local gradient
1 * d * d * d … * d = dn-1
If d << 1, n is big → Vanishing Gradients
Training Deeper Networks
Every backpropagation step multiplies the gradient by the local gradient
1 * d * d * d … * d = dn-1
If d >> 1, n is big → Exploding Gradients
VGGNet: Training Trick #3 Pre-Training
Use weights to initialize VGG-13
Train VGG-13
…
Train VGG-19
VGGNet: Training Trick #4 Batch Normalization
Slide Credit: Justin Johnson
VGGNet: Training Trick #4 Batch Normalization
Slide Credit: Justin Johnson
Problem #1: What if zero-mean and unit variance is too hard of a constrain?
Slide Credit: Justin Johnson
Slide Credit: Justin Johnson
Problem #2: Mean and variance estimates depend on the mini-batch. We only see one example at a time during test-time.
Batch Normalization at Test Time
Slide Credit: Justin Johnson
Batch Norm at Training V. Test Time
Slide Credit: Justin Johnson
Training
Testing
Batch Normalization Improves Training
Slide Credit: Justin Johnson
Alternate Normalization Techniques
GoogLeNet: Why Only 3x3 Convolution?
GoogLeNet: Everything But the Kitchen Sink
Intermediate Loss Layer Prevents Vanishing Gradients
Stagewise training in 3 parts
ImageNet 14’: VGGNet: 7.30% Error, GoogLeNet: 6.67% Error
ResNet: Scaling to 100+ Layers
Deep Residual Learning
If x is optimal, we can set weight layers to 0. It's easy to “ignore” extra model capacity.
If x is close to optimal, it's easier to find small weights.
Ensembling Effect: This could be seen as an ensemble of networks of varying depth
Identity serves as a “gradient highway” to preserve strong signal from output to earlier layers during back propagation.
DenseNet: Connecting All Layers with Skip Connections
Training Neural Networks
Input Data Preprocessing: Subtract Mean, Divide by STD
Weight Initialization: Constant Values
Q: What if we initialize all weights/biases to a constant value?
Weight Initialization: Constant Values
Q: What if we initialize all weights/biases to a constant value?
A: We cannot break symmetry of weight values. Gradients are the same at each location.
Weight Initialization: Random Values
Q: What if we initialize all weights/biases to a (small) random value?
Weight Initialization: Random Values
Q: What if we initialize all weights/biases to a (small) random value?
A: This is usually fine for shallow networks, but becomes problematic for deep networks.
Weight Initialization for Deeper Models
Slide Credit: Kris Kitani
Slide Credit: Rawal Khirodkar
Understanding Loss Curves
Choosing the Correct Learning Rate
Slide Credit: Rawal Khirodkar
Learning Rate Scheduling
Learning Rate Scheduling Varieties
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Diagnosing Loss Curves
Choosing Hyperparameters
Learning Rate: Look at Loss Curves (1e-3 Safe Bet)
Batch Size: As Large as Possible
# of Epoch: Until Validation Accuracy Saturates
Network Architecture: ResNet (Safe Bet)
Optimizer: Adam (Safe Bet)
For Everything Else: Experimental Validation
Which is better? Why?
Choosing Hyperparameters
Learning Rate: Look at Loss Curves (1e-3 Safe Bet)
Batch Size: As Large as Possible
# of Epoch: Until Validation Accuracy Saturates
Network Architecture: ResNet (Safe Bet)
Optimizer: Adam (Safe Bet)
For Everything Else: Experimental Validation
For low dimensional search spaces, random search works better in practice.
Optimizers: Navigating the Loss Ravines
Drawbacks of Gradient Descent
Slide Credit: Rawal Khirodkar
Regularization: Limit the Complexity of your Model
Regularization: Toy Example
Regularization: Toy Example
Regularization: Not Too Little, Not Too Much
Training Neural Networks Summary
Training Neural Networks is an Art and a Science
Personal Experience Debugging Neural Networks
Complex Error Logs = Simple Bugs