Workshop on Hands-on Deep Learning Coding and Code Management
Organized by
Center for Computational & Data Sciences, IUB
Day 1
Who we are?
Dr. AKM Mahbubur Rahman
Associate Professor, Director Data Science wing
Research Assistants
Md Fahim
Mir Sazzat Hossain
Kishor Kumar Bhaumik
Minhajul Islam
Rafat Hasan Khan
Shadman Rohan
Tahmid Hasan Fuad
Jahir Sadik Monon
Why we are here?
Preferred skills
Disclaimer: Some slides are modified and adopted from CSE231n (CS231n: Deep Learning for Computer Vision) , Stanford University
Day 1, Session 1
Visual Recognition
Image Classification
Visual Recognition
Image Classification
Visual Recognition
Image Classification
Visual Recognition
Image Classification
This image is licensed under CC BY-NC-SA 2.0; changes made
Person
Hammer
This image is licensed under CC BY-SA 2.0; changes made
Person
Bike
Person on Bike
This image is licensed under CC BY-SA 3.0; changes made
Image Classification pipeline
The Problem: Semantic Gap
This image by Nikita is licensed under CC-BY 2.0
What the computer sees
An image is just a big grid of numbers between [0, 255]:
e.g. 800 x 600 x 3 (3 channels RGB)
An image classifier
Machine Learning: Data-Driven Approach
14
Example training set
First classifier: Nearest Neighbor
Memorize all data and labels
Predict the label of the most similar training image
Example Dataset: CIFAR10
10 classes
50,000 training images
10,000 testing images
Test images and nearest neighbors
Alex Krizhevsky, “Learning Multiple Layers of Features from Tiny Images”, Technical Report, 2009.
Nearest Neighbor classifier
for two dimensional data
Nearest Neighbor classifier
for two dimensional data
Distance Metric to compare images
19
Fei-Fei Li & Justin Johnson & Serena Yeung
L1 distance:
Lecture 2 -
add
We generally use L2 distance as distance metric
Build nearest neighbor classifier
in numpy
TO DO (10 minutes)
K - Nearest Neighbors classifier
for two dimensional data
TO DO (20 minutes)
Evaluation in Test images
Evaluation in Test images
Finally ..
Putting everything together in OOP structure (class, methods)
April 5, 2018
27
Lecture 2 -
KNN is not good at all….
Linear Classification
f(x,W) = Wx + b
x is an Array of 32x32x3 numbers (3072 numbers total)
W is 10 x 3072
b is 10 x 1
Linear Classifier: Three Viewpoints
29
Fei-Fei Li & Justin Johnson & Serena Yeung
f(x,W) = Wx + b
Algebraic Viewpoint
Visual Viewpoint
Geometric Viewpoint
One template per class
Hyperplanes cutting up space
Lecture 2 -
Softmax Classifier
(Multinomial Logistic Regression)
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Classifier
(Multinomial Logistic Regression)
Want to interpret raw classifier scores as probabilities
Softmax Function
cat car frog
3.2
5.1
-1.7
Softmax Classifier
(Multinomial Logistic Regression)
32
Fei-Fei Li & Justin Johnson & Serena Yeung
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Function
Probabilities must be >= 0
24.5
164.0
0.18
exp
April 10, 2018
unnormalized probabilities
Softmax Classifier
(Multinomial Logistic Regression)
33
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Function
24.5
164.0
0.18
0.13
0.87
0.00
exp
normalize
unnormalized
probabilities
Fei-Fei Li & Justin Johnson & Serena Yeung
probabilities
April 10, 2018
Probabilities must be >= 0
Probabilities must sum to 1
Softmax Classifier
(Multinomial Logistic Regression)
34
Fei-Fei Li & Justin Johnson & Serena Yeung
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Function
24.5
164.0
0.18
0.13
0.87
0.00
exp
normalize
log-probabilities / logits
probabilities
April 10, 2018
unnormalized
Probabilities must be >= 0
Probabilities must sum to 1
probabilities
Unnormalized
Softmax Classifier
(Multinomial Logistic Regression)
35
Fei-Fei Li & Justin Johnson & Serena Yeung
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Function
24.5
164.0
0.18
0.13
0.87
0.00
exp
normalize
Probabilities must be >= 0
Probabilities must sum to 1
1.00
0.00
0.00
compare
Unnormalized
log-probabilities / logits
unnormalized
probabilities
probabilities
Correct
probs
April 10, 2018
Softmax Classifier
(Multinomial Logistic Regression)
36
cat car frog
3.2
5.1
-1.7
Want to interpret raw classifier scores as probabilities
Softmax Function
Maximize probability of correct class
Putting it all together:
Whole Picture & CE Loss Calculation
Total loss:
Optimization
Lecture 3 -
38
April 10, 2018
Find the W and b that will give the total loss minimum.
How?
Follow the slope
Lecture 3 -
39
April 10, 2018
Gradient Descent
Gradient Calculation
Parameter updates
Training Procedure
(Learning the parameters)
Mini-batch Stochastic Gradient Descent (SGD) Loop:
Visualize the weights
TODO
Convolutional Neural Network
Day 1, Session 2
Big Picture
Illustration of LeCun et al. 1998 from CS231n 2017 Lecture 1
Convolution Layer
32
3
32
depth
32x32x3 image -> preserve spatial structure
width
height
Convolution Layer
32x32x3 image
5x5x3 filter
32
Convolve the filter with the image
i.e. “slide over the image spatially, computing dot products”
32
3
Convolution Layer
32x32x3 image
5x5x3 filter
32
Convolve the filter with the image
i.e. “slide over the image spatially, computing dot products”
Filters always extend the full depth of the input volume
32
3
Convolution Layer
32
32x32x3 image 5x5x3 filter
32
1 number:
the result of taking a dot product between the filter and a small 5x5x3 chunk of the image
(i.e. 5*5*3 = 75-dimensional dot product + bias)
3
consider a second, green filter
Convolution Layer
32
32
3
32x32x3 image 5x5x3 filter
convolve (slide) over all spatial locations
activation maps
1
28
28
Convolution Demo
32
32
3
Convolution Layer
activation maps
6
28
28
We stack these up to get a “new image” of size 28x28x6!
We have six filters
Summary of convolutional layer
Sub-sampling/Pooling Layer
Max Pooling
1 | 1 | 2 | 4 |
5 | 6 | 7 | 8 |
3 | 2 | 1 | 0 |
1 | 2 | 3 | 4 |
Single depth slice
x
y
max pool with 2x2 filters and stride 2
6 | 8 |
3 | 4 |
Summary of maxpool layer
Common settings:
F = 2, S = 2
F = 3, S = 2
Convolutional Net
ConvNet is a sequence of Convolution Layers, interspersed with activation functions
32
32
3
28
28
6
CONV, ReLU
e.g. 6 5x5x3 filters
Convolutional Net
32
32
3
CONV, ReLU
e.g. 6 5x5x3 filters
28
28
6
CONV, ReLU
e.g. 10 5x5x6 filters
CONV, ReLU
….
10
24
24
ReLu Activation Function
VGG 16
Fully Connected Layer (FC layer)
- Contains neurons that connect to the entire input volume, as in ordinary Neural Networks
Demo: http://cs231n.stanford.edu/
Developing a CNN
for image classification
Experiments
Loss curve investigation
Loss curve investigation
TODO
CIFAR10 Image classification
using pretrained VGG16