1 of 68

Workshop on Hands-on Deep Learning Coding and Code Management

Organized by

Center for Computational & Data Sciences, IUB

Day 1

2 of 68

Who we are?

Dr. AKM Mahbubur Rahman

Associate Professor, Director Data Science wing

Research Assistants

Md Fahim

Mir Sazzat Hossain

Kishor Kumar Bhaumik

Minhajul Islam

Rafat Hasan Khan

Shadman Rohan

Tahmid Hasan Fuad

Jahir Sadik Monon

3 of 68

Why we are here?

  • Get introduced to deep learning programming
  • Practice programming
  • Develop deep learning models
  • Training and fine tuning deep learning models
  • Hands on experiment design and result analysis
  • Guidelines for standard coding practice for deep learning

4 of 68

Preferred skills

  • Python basics with numpy
  • Finished Numerical Methods course
  • Linear Algebra with vector notations
  • Matplotlib, pyplot for visualization
  • AI, ML, Data Mining courses

Disclaimer: Some slides are modified and adopted from CSE231n (CS231n: Deep Learning for Computer Vision) , Stanford University

5 of 68

Day 1, Session 1

  • Numpy Recap
  • Image Classification pipeline, CIFAR10
  • KNN Classifier
  • Linear Classifier from scratch
  • Linear Classifier with pytorch

6 of 68

Visual Recognition

Image Classification

7 of 68

Visual Recognition

Image Classification

8 of 68

Visual Recognition

Image Classification

9 of 68

Visual Recognition

Image Classification

10 of 68

  • Object detection
  • Action classification
  • Image captioning

This image is licensed under CC BY-NC-SA 2.0; changes made

Person

Hammer

This image is licensed under CC BY-SA 2.0; changes made

Person

Bike

Person on Bike

This image is licensed under CC BY-SA 3.0; changes made

11 of 68

Image Classification pipeline

12 of 68

The Problem: Semantic Gap

This image by Nikita is licensed under CC-BY 2.0

What the computer sees

An image is just a big grid of numbers between [0, 255]:

e.g. 800 x 600 x 3 (3 channels RGB)

13 of 68

An image classifier

  • Unlike e.g. sorting a list of numbers
  • classify_image function should have some magic

14 of 68

Machine Learning: Data-Driven Approach

14

  1. Collect a dataset of images and labels
  2. Use Machine Learning to train a classifier
  3. Evaluate the classifier on new images

Example training set

15 of 68

First classifier: Nearest Neighbor

Memorize all data and labels

Predict the label of the most similar training image

16 of 68

Example Dataset: CIFAR10

10 classes

50,000 training images

10,000 testing images

Test images and nearest neighbors

Alex Krizhevsky, “Learning Multiple Layers of Features from Tiny Images”, Technical Report, 2009.

17 of 68

Nearest Neighbor classifier

for two dimensional data

18 of 68

Nearest Neighbor classifier

for two dimensional data

19 of 68

Distance Metric to compare images

19

Fei-Fei Li & Justin Johnson & Serena Yeung

L1 distance:

Lecture 2 -

add

We generally use L2 distance as distance metric

20 of 68

Build nearest neighbor classifier

in numpy

  • Make a copy of the python notebook file in your Google Colab editor
  • Monon (RA CCDS) will guide through the codes

21 of 68

TO DO (10 minutes)

  • Write code to print the predicted labels for  first 15 validation images using the nearest_neighbor() function.

22 of 68

K - Nearest Neighbors classifier

for two dimensional data

23 of 68

TO DO (20 minutes)

  • Write code snippet to calculate the predicted label based on k nearest neighbors.
  • first 15 validation images using the k_nearest_neighbor() function.
  • See the difference

24 of 68

Evaluation in Test images

25 of 68

Evaluation in Test images

26 of 68

Finally ..

Putting everything together in OOP structure (class, methods)

27 of 68

April 5, 2018

27

Lecture 2 -

KNN is not good at all….

28 of 68

Linear Classification

f(x,W) = Wx + b

x is an Array of 32x32x3 numbers (3072 numbers total)

W is 10 x 3072

b is 10 x 1

29 of 68

Linear Classifier: Three Viewpoints

29

Fei-Fei Li & Justin Johnson & Serena Yeung

f(x,W) = Wx + b

Algebraic Viewpoint

Visual Viewpoint

Geometric Viewpoint

One template per class

Hyperplanes cutting up space

Lecture 2 -

30 of 68

Softmax Classifier

(Multinomial Logistic Regression)

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

31 of 68

Softmax Classifier

(Multinomial Logistic Regression)

Want to interpret raw classifier scores as probabilities

Softmax Function

cat car frog

3.2

5.1

-1.7

32 of 68

Softmax Classifier

(Multinomial Logistic Regression)

32

Fei-Fei Li & Justin Johnson & Serena Yeung

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

Softmax Function

Probabilities must be >= 0

24.5

164.0

0.18

exp

April 10, 2018

unnormalized probabilities

33 of 68

Softmax Classifier

(Multinomial Logistic Regression)

33

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

Softmax Function

24.5

164.0

0.18

0.13

0.87

0.00

exp

normalize

unnormalized

probabilities

Fei-Fei Li & Justin Johnson & Serena Yeung

probabilities

April 10, 2018

Probabilities must be >= 0

Probabilities must sum to 1

34 of 68

Softmax Classifier

(Multinomial Logistic Regression)

34

Fei-Fei Li & Justin Johnson & Serena Yeung

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

Softmax Function

24.5

164.0

0.18

0.13

0.87

0.00

exp

normalize

log-probabilities / logits

probabilities

April 10, 2018

unnormalized

Probabilities must be >= 0

Probabilities must sum to 1

probabilities

Unnormalized

35 of 68

Softmax Classifier

(Multinomial Logistic Regression)

35

Fei-Fei Li & Justin Johnson & Serena Yeung

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

Softmax Function

24.5

164.0

0.18

0.13

0.87

0.00

exp

normalize

Probabilities must be >= 0

Probabilities must sum to 1

1.00

0.00

0.00

compare

Unnormalized

log-probabilities / logits

unnormalized

probabilities

probabilities

Correct

probs

April 10, 2018

36 of 68

Softmax Classifier

(Multinomial Logistic Regression)

36

cat car frog

3.2

5.1

-1.7

Want to interpret raw classifier scores as probabilities

Softmax Function

Maximize probability of correct class

Putting it all together:

37 of 68

Whole Picture & CE Loss Calculation

Total loss:

38 of 68

Optimization

Lecture 3 -

38

April 10, 2018

Find the W and b that will give the total loss minimum.

How?

39 of 68

Follow the slope

Lecture 3 -

39

April 10, 2018

Gradient Descent

40 of 68

Gradient Calculation

41 of 68

Parameter updates

42 of 68

Training Procedure

(Learning the parameters)

Mini-batch Stochastic Gradient Descent (SGD) Loop:

    • Sample a batch of data
    • Forward propagate it through the network
    • Calculate loss
    • Backprop to calculate the gradients
    • Update the parameters using the gradient

43 of 68

Visualize the weights

44 of 68

TODO

  • Change the following hyper parameter for better average f1 score
    • Epoch size: 100
    • Batch size: 64, 128, 256, 512
    • Learning rate: 0.05, 0.001, 0.005
  • Visualize the weights now. Are you seeing some patterns in the weights? If so, your weights are perfectly learned.

45 of 68

Convolutional Neural Network

Day 1, Session 2

46 of 68

Big Picture

Illustration of LeCun et al. 1998 from CS231n 2017 Lecture 1

47 of 68

Convolution Layer

32

3

32

depth

32x32x3 image -> preserve spatial structure

width

height

48 of 68

Convolution Layer

32x32x3 image

5x5x3 filter

32

Convolve the filter with the image

i.e. “slide over the image spatially, computing dot products”

32

3

49 of 68

Convolution Layer

32x32x3 image

5x5x3 filter

32

Convolve the filter with the image

i.e. “slide over the image spatially, computing dot products”

Filters always extend the full depth of the input volume

32

3

50 of 68

Convolution Layer

32

32x32x3 image 5x5x3 filter

32

1 number:

the result of taking a dot product between the filter and a small 5x5x3 chunk of the image

(i.e. 5*5*3 = 75-dimensional dot product + bias)

3

51 of 68

consider a second, green filter

Convolution Layer

32

32

3

32x32x3 image 5x5x3 filter

convolve (slide) over all spatial locations

activation maps

1

28

28

52 of 68

Convolution Demo

53 of 68

32

32

3

Convolution Layer

activation maps

6

28

28

We stack these up to get a “new image” of size 28x28x6!

We have six filters

54 of 68

Summary of convolutional layer

55 of 68

Sub-sampling/Pooling Layer

  • makes the representations smaller and more manageable
  • operates over each activation map independently:

56 of 68

Max Pooling

1

1

2

4

5

6

7

8

3

2

1

0

1

2

3

4

Single depth slice

x

y

max pool with 2x2 filters and stride 2

6

8

3

4

57 of 68

Summary of maxpool layer

Common settings:

F = 2, S = 2

F = 3, S = 2

58 of 68

Convolutional Net

ConvNet is a sequence of Convolution Layers, interspersed with activation functions

32

32

3

28

28

6

CONV, ReLU

e.g. 6 5x5x3 filters

59 of 68

Convolutional Net

32

32

3

CONV, ReLU

e.g. 6 5x5x3 filters

28

28

6

CONV, ReLU

e.g. 10 5x5x6 filters

CONV, ReLU

….

10

24

24

60 of 68

ReLu Activation Function

61 of 68

VGG 16

62 of 68

Fully Connected Layer (FC layer)

- Contains neurons that connect to the entire input volume, as in ordinary Neural Networks

Demo: http://cs231n.stanford.edu/

63 of 68

Developing a CNN

for image classification

  • Conv layer, Relu
  • Conv layer, Relu
  • Max pool
  • Conv layer, Relu
  • Conv layer, Relu
  • Max pool
  • Flatten
  • FC1, Relu
  • FC2, Relu
  • FC3, Relu
  • Softmax

64 of 68

Experiments

  • Monon will guide you through the CNN codes for above architecture
  • Run your code
  • Produce the train and validation loss curves
  • Test and produce your confusion matrix and evaluation metrics

65 of 68

Loss curve investigation

66 of 68

Loss curve investigation

67 of 68

TODO

  • Generally, fully connected layers are followed by sigmoid function. Add sigmoid activation function after each linear layer. Remove the Relu function for linear layers
    • Run your code
    • Produce the train and validation loss curves
    • Test and produce your confusion matrix and evaluation metrics
  • Add one convolution layer and Relu layer just before the flattening.
    • Run your code
    • Produce the train and validation loss curves
    • Test and produce your confusion matrix and evaluation metrics
  • Add one batch normalization layer after sigmoid activation function
    • For each fc layers
    • Run your code
    • Produce the train and validation loss curves
    • Test and produce your confusion matrix and evaluation metrics

68 of 68

CIFAR10 Image classification

using pretrained VGG16

  • VGG 16 weights are trained in Image nets
  • We will finetune the weights for CIFAR10
  • Performance Evaluation