1 of 109

Convolutional Neural Networks�(CNN) 

2 of 109

Convolution

2

3 of 109

Convolution

  • Integral (or sum) of the product of the two signals after one is reversed and shifted
  • Cross-correlation and convolution

3

4 of 109

1D Convolution

  • (actually cross-correlation)

4

Source: Dr. Francois Fleuret at EPFL

1

3

2

3

0

-1

1

2

2

1

1

3

0

-1

Input

Kernel

Output

L = W-w+1

7

W

w

5 of 109

1D Convolution

  • (actually cross-correlation)

5

Source: Dr. Francois Fleuret at EPFL

1

3

2

3

0

-1

1

2

2

1

1

3

0

-1

Input

Kernel

Output

L = W-w+1

7

9

W

w

6 of 109

1D Convolution

  • (actually cross-correlation)

6

Source: Dr. Francois Fleuret at EPFL

1

3

2

3

0

-1

1

2

2

1

1

3

0

-1

w

Input

Kernel

Output

L = W-w+1

7

9

12

2

-1

0

6

W

7 of 109

Example: 1D Convolution

7

8 of 109

De-noising a Piecewise Smooth Signal

  •  

8

9 of 109

De-noising a Piecewise Smooth Signal

9

10 of 109

Edge Detection

10

11 of 109

Smoothing and Detection of Abrupt Changes

11

12 of 109

Images

12

13 of 109

Images Are Numbers

  • Grayscale image as a matrix

13

Source: 6.S191 Intro. to Deep Learning at MIT

14 of 109

Colored Images

  • Colored images are represented as 3D tensors

14

Original image

R

G

B

Grayscale image

15 of 109

2D Convolution

15

16 of 109

Convolution on Image (= Convolution in 2D)

  • Filter (or Kernel)
    • Discrete convolution can be viewed as element-wise multiplication by a matrix
    • Modify or enhance an image by filtering
    • Filter images to emphasize certain features or remove other features
    • Filtering includes smoothing, sharpening and edge enhancement

16

Image

Kernel

Output

17 of 109

Convolution on Image (= Convolution in 2D)

  • The filter slides (or “convolves”) across the image, repeating this operation for each overlapping region.

17

image

Kernel

Feature Map

18 of 109

Convolution on Image

18

Kernel

19 of 109

Convolution on Image

19

Kernel

20 of 109

Convolution on Image

20

Kernel

21 of 109

Convolution on Image

21

Kernel

22 of 109

Convolution on Image

22

Kernel

23 of 109

Convolution on Image

23

Kernel

24 of 109

Convolution on Image

24

Kernel

Kernel

25 of 109

Convolution on Image

25

Kernel

Kernel

26 of 109

Convolution on Image

26

27 of 109

How to Find the Right Kernels

  • We learn many different kernels that make specific effect on images

  • Let’s apply an opposite approach

  • We are not designing the kernel, but are learning the kernel from data

  • Can learn feature extractor from data using a deep learning framework

27

28 of 109

Learning Visual Features

28

29 of 109

Image Classification

  • Motivation
    • The bird occupies a local area and looks the same in different parts of an image.
    • We should construct neural networks which exploit these properties.

29

30 of 109

ANN for Object Classification in Image

  • Flattening does not seem the best
  • Did not make use of the fact that we are dealing with images

30

bird

31 of 109

Fully Connected Neural Network (= ANN)

  • Input
    • 2D image
    • Vector of pixel values

  • Fully connected
    • Connect neuron in hidden layer to all neurons in input layer
    • No spatial information
    • Spatial organization of the input is destroyed by flatten
    • And many, many parameters !

  • How can we use spatial structure in the input to inform the architecture of the network?

31

Source: 6.S191 Intro. to Deep Learning at MIT

32 of 109

Convolution Mask + Neural Network

  • Naïve idea
    • 2D filter (or Kernel) to the entire image
    • To maintain spatial organization

32

Fully connected

33 of 109

Locality

  •  

33

Locality

Fully connected

Locally connected

34 of 109

Locality

  •  

34

Locality

Fully connected

Locally connected

Weight Sharing

Convolution

35 of 109

Locality

  •  

35

Locality

Fully connected

Locally connected

Weight Sharing

Convolution

36 of 109

Locality

  •  

36

Locality

Fully connected

Locally connected

Weight Sharing

Convolution

37 of 109

Locality

  •  

37

Locality

Fully connected

Locally connected

Weight Sharing

Convolution

38 of 109

Convolution for Image Classification

  • Convolution, originally defined in the field of mathematics, is naturally well-suited for image classification
    • Locality: kernel focuses on local regions of the input, allowing the model to capture spatially localized patterns
    • Weight Sharing: The same set of weights (kernel values) is applied across different spatial locations, which greatly reduces the number of parameters

    • Scan the image locally while preserving the spatial structure of the data
    • By convolving a small kernel over the image and sharing the same weights across locations, the model can efficiently detect patterns (e.g., edges, corners, textures) regardless of their position in the image.

38

Convolution

39 of 109

Convolution + Neural Network

  • Unlike traditional image processing where kernels are manually designed, the kernel is learned directly from the data

  • Model automatically discovers the most relevant features for a given task, effectively learning its own feature extractor.

39

Convolution

 

 

 

 

 

 

 

 

 

40 of 109

More on Convolution of CNN

  • Multiple Channels

  • Multiple Kernels

40

41 of 109

Multiple Filters (or Kernels)

  • Multiple filters (or kernels) are applied to extract a diverse set of features from images
  • Each filter produces one output channel (also called a feature map)
  • Channel refers to a distinct layer of information in the input or feature maps

41

42 of 109

Multiple Channels

  • Channel refers to a distinct layer of information in the input or feature maps
  • Examples of multiple channels
    • Grayscale images have 1 channel
    • RGB images have 3 channels (Red, Green, Blue)

42

Source: Dr. Francois Fleuret at EPFL

43 of 109

Multi-channel 2D Convolution

43

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

44 of 109

Multi-channel 2D Convolution

44

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Kernel

w

h

c

45 of 109

Multi-channel 2D Convolution

45

Kernel

w

h

c

Output

Input

W

H

C

46 of 109

Multi-channel 2D Convolution

46

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Output

Kernel

w

h

c

47 of 109

Multi-channel 2D Convolution

47

Source: Dr. Francois Fleuret at EPFL

Kernel

w

h

c

Input

W

H

C

Output

48 of 109

Multi-channel 2D Convolution

48

Source: Dr. Francois Fleuret at EPFL

Kernel

w

h

c

Input

W

H

C

Output

49 of 109

Multi-channel 2D Convolution

49

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Kernel

w

h

c

Output

50 of 109

Multi-channel 2D Convolution

50

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Kernel

w

h

c

Output

51 of 109

Multi-channel 2D Convolution

51

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Kernel

w

h

c

Output

52 of 109

Multi-channel 2D Convolution

52

Source: Dr. Francois Fleuret at EPFL

Input

W

H

C

Kernel

w

h

c

Output

53 of 109

Multi-channel 2D Convolution

53

Output

Input

W

H

C

Kernel

w

h

c

Source: Dr. Francois Fleuret at EPFL

54 of 109

Multi-channel 2D Convolution

54

Output

H – h + 1

W – w + 1

1

Input

W

H

C

Kernel

w

h

c

Source: Dr. Francois Fleuret at EPFL

55 of 109

Multi-channel and Multi-kernel 2D Convolution

55

Source: Dr. Francois Fleuret at EPFL

Output

H – h + 1

W – w + 1

D

w

h

c

Kernels

Input

W

H

C

D

56 of 109

Dealing with Shapes

  •  

56

Source: Dr. Francois Fleuret at EPFL

57 of 109

Multi-channel 2D Convolution

  • The kernel is not swiped across channels, just across rows and columns.
  • Note that a convolution preserves the signal support structure.
    • A 1D signal is converted into a 1D signal, a 2D signal into a 2D, and neighboring parts of the input signal influence neighboring parts of the output signal.

  • We usually refer to one of the channels generated by a convolution layer as an activation map.
  • The sub-area of an input map that influences a component of the output as the receptive field of the latter.

57

58 of 109

Padding and Stride

58

59 of 109

Strides

  • Strides: increment step size for the convolution operator
  • Reduces the size of the output map

59

Example with kernel size 3×3 and a stride of 2 (image in blue)

Source: https://github.com/vdumoulin/conv_arithmetic

60 of 109

Padding

  • Padding: artificially fill borders of image
  • Useful to keep spatial dimension constant across filters
  • Useful with strides and large receptive fields
  • Usually fill with 0s

60

Source: https://github.com/vdumoulin/conv_arithmetic

61 of 109

Padding and Stride

  •  

61

Source: Dr. Francois Fleuret at EPFL

 

Input

 

62 of 109

Padding and Stride

  •  

62

Source: Dr. Francois Fleuret at EPFL

Input

Output

63 of 109

Padding and Stride

  •  

63

Source: Dr. Francois Fleuret at EPFL

Input

Output

64 of 109

Padding and Stride

  •  

64

Source: Dr. Francois Fleuret at EPFL

Input

Output

65 of 109

Padding and Stride

  •  

65

Source: Dr. Francois Fleuret at EPFL

Input

Output

66 of 109

Padding and Stride

  •  

66

Source: Dr. Francois Fleuret at EPFL

Input

Output

67 of 109

Padding and Stride

  •  

67

Source: Dr. Francois Fleuret at EPFL

Input

Output

68 of 109

Padding and Stride

  •  

68

Source: Dr. Francois Fleuret at EPFL

Input

Output

69 of 109

Padding and Stride

  •  

69

Source: Dr. Francois Fleuret at EPFL

Input

Output

70 of 109

Padding and Stride

  •  

70

Source: Dr. Francois Fleuret at EPFL

Input

Output

71 of 109

Nonlinear Activation Function

  • Convolution is inherently a linear operation, meaning it can only represent linear relationships
  • To enable the model to learn nonlinear and more complex patterns, a nonlinear activation function (such as ReLU, sigmoid, or tanh) is applied element-wise to the output of the convolution

71

72 of 109

Pooling

72

73 of 109

Max Pooling

  • Max pooling is a down-sampling operation to reduce the spatial dimensions (height and width) of feature maps while preserving the most important features
    • A small window slides over the input feature map
    • For each window, the maximum value is selected and retained
    • The result is a smaller, more compact representation of the input
    • The most common pooling method

73

1

2

4

1

6

7

8

5

2

1

0

3

2

3

4

1

6

8

3

4

 

74 of 109

Average Pooling

  • Average pooling is a down-sampling technique to reduce the spatial dimensions (height and width) of feature maps by computing the average value within a sliding window
    • A fixed-size window slides over the input feature map
    • For each window, it computes the average of all values in that region
    • The output is a reduced-size feature map that retains overall information in a smoothed form

74

3.25

5.25

2

2

1

2

4

1

6

7

8

5

2

1

0

3

2

3

4

1

Avg pool with 2×2 filters and stride 2

75 of 109

Global Average Pooling (GAP)

  • Takes the average of each entire feature map and produces a single value

75

1

2

4

1

6

7

8

5

2

1

0

3

2

3

4

1

Global Avg pool

3.125

76 of 109

Max Pooling in 1D

  • The most standard type of pooling is the max-pooling
  • For instance, in 1D with a window of size 2

76

Source: Dr. Francois Fleuret at EPFL

1

3

2

3

0

-1

1

2

2

1

w

3

3

0

2

2

Input

r w

Output

r

77 of 109

Pooling: Invariance to Small Deformations

  • Offer a form of pseudo-invariance to small deformations in the input, such as local translations, which may occur due to slight shifts, distortions, or noise in the data
  • Reduce sensitivity to small spatial variations in the input

77

Source: Dr. Francois Fleuret at EPFL

Input

Output

78 of 109

Pooling: Invariance to Small Deformations

  • Offer a form of pseudo-invariance to small deformations in the input, such as local translations, which may occur due to slight shifts, distortions, or noise in the data
  • Reduce sensitivity to small spatial variations in the input

78

Source: Dr. Francois Fleuret at EPFL

Input

Output

79 of 109

Pooling Performs Separately on Each of Channels

79

Source: Dr. Francois Fleuret at EPFL

r w

s h

C

Input

80 of 109

Multi-channel Pooling

80

Input

r w

s h

C

Output

81 of 109

Multi-channel Pooling

81

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

82 of 109

Multi-channel Pooling

82

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

83 of 109

Multi-channel Pooling

83

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

84 of 109

Multi-channel Pooling

84

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

85 of 109

Multi-channel Pooling

85

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

86 of 109

Multi-channel Pooling

86

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

87 of 109

Multi-channel Pooling

87

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

88 of 109

Multi-channel Pooling

88

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

89 of 109

Multi-channel Pooling

89

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

90 of 109

Multi-channel Pooling

90

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

91 of 109

Multi-channel Pooling

91

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

92 of 109

Multi-channel Pooling

92

Source: Dr. Francois Fleuret at EPFL

Output

Input

r w

s h

C

93 of 109

Multi-channel Pooling

93

Source: Dr. Francois Fleuret at EPFL

Input

r w

s h

C

Output

r

s

C

94 of 109

Convolution Layer Block

94

95 of 109

Inside the Convolution Layer Block

95

Convolution layer block

Input to layer

Next layer

96 of 109

Convolution Neural Network

  • Input

  • Conv blocks
    • Convolution + activation (relu)
    • Convolution + activation (relu)
    • ...
    • Maxpooling

  • Output
    • Fully connected layers
    • Softmax
  •  

96

97 of 109

Artificial Neural Networks

  • Universal function approximator
    • Linear connected networks
    • Simple nonlinear neurons

  • Hidden layers
    • Autonomous feature learning

97

Class 2

Class 1

98 of 109

Convolutional Neural Networks

  • Structure
    • Weight sharing
    • Local connectivity
    • Typically have sparse interactions

  • Optimization
    • Smaller searching space of unknown parameters

  • Note:
    • Used an ANN-style illustration, as CNN architectures are often challenging to visualize

98

Class 2

Class 1

99 of 109

CNNs for Classification

99

Source: 6.S191 Intro. to Deep Learning at MIT

100 of 109

CNNs for Classification: Feature Learning

  • Learn features in input image through convolution
  • Introduce non-linearity through activation function (real-world data is non-linear!)
  • Reduce dimensionality and preserve spatial invariance with pooling

100

Source: 6.S191 Intro. to Deep Learning at MIT

101 of 109

CNNs for Classification: Class Probabilities

  • CONV and POOL layers output high-level features of input
  • Fully connected layer uses these features for classifying input image
  • Express output as probability of image belonging to a particular class

101

Source: 6.S191 Intro. to Deep Learning at MIT

102 of 109

CNN in TensorFlow

102

103 of 109

Lab: CNN with TensorFlow

  • MNIST example
  • To classify handwritten digits

103

104 of 109

CNN Structure

104

105 of 109

Loss and Optimizer

  • Loss
    • Classification: Cross entropy
    • Equivalent to applying logistic regression
  • Optimizer
    • GradientDescentOptimizer
    • AdamOptimizer: the most popular optimizer

105

106 of 109

Test or Evaluation

106

107 of 109

CNN for Steel Surface Defects

107

108 of 109

Steel Surface Defects

  • NEU steel surface defects example

108

109 of 109

CNN with TensorFlow

109