1 of 93

Convolutional Neural Network

Lecture 18

Introduction to convolutional architecture for image analysis

EECS 189/289, Fall 2025 @ UC Berkeley

Joseph E. Gonzalez and Narges Norouzi

EECS 189/289, Fall 2025 @ UC Berkeley

Joseph E. Gonzalez and Narges Norouzi

2 of 93

Roadmap

  • Network Design Considerations for Images
  • Element of the Convolutional Neural Network
  • CNN Design
  • CNN History

2499499

3 of 93

Network Design Considerations for Images

  • Network Design Considerations for Images
  • Element of the Convolutional Neural Network
  • CNN Design
  • CNN History

2499499

4 of 93

Story So Far

  • NNs are universal function approximators.
  • NNs can be trained through variations of gradient descent.
  • Can recognize patterns in data.

 

 

 

 

 

 

2499499

5 of 93

A Problem

  • Will a NN that recognizes the left image as a cat also recognize the one on the right one as a cat?

 

 

 

 

 

 

2499499

6 of 93

A Problem

  • Will a NN that recognizes the left image as a cat also recognize the one on the right one as a cat?

 

 

 

 

 

 

A network that recognizes "cat" feature regardless of the location and scale of the target object.

2499499

7 of 93

A Problem

Network must be shift invariant. 

2499499

8 of 93

Solution: Scanning

  • Scan for the presence of the cat in the image.

2499499

9 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

10 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

11 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

12 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

13 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

14 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

15 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

16 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

17 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

18 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

19 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

2499499

20 of 93

Scanning Is Referred to as Convolution

  • We need to apply a feature extractor in each region of the image.

max

Maximum of all the outputs (Equivalent of Boolean OR)

The strength of having the feature anywhere in the image

Same shared weights

2499499

21 of 93

Scanning Is Referred to as Convolution

  • The entire operation can be viewed as one giant network.
    • With many subnetworks, one per window.
  • Restriction: All subnetworks are identical.
  • The network is shift-invariant!

max

2499499

22 of 93

Which of the following is correct?

The Slido app must be installed on every computer you’re presenting from

2499499

23 of 93

Regular Network vs. Scanning Network

  • In a FC-NN, each neuron is connected by a unique weight to every neuron in the previous layer.
    • Entries in the weight matrix are unique.
  • In a scanning NN, each neuron is connected to a subset of neurons in the previous layer
    • The weights matrix is sparse.
    • The weights matrix is block structured with identical blocks.

max

 

2499499

24 of 93

Element of the Convolutional Neural Network

  • Network Design Considerations for Images
  • Element of the Convolutional Neural Network
  • CNN Design
  • CNN History

2499499

25 of 93

Back to Scanning Network

  • Patches are referred to as "receptive fields".
  • The weights form a small two-dimensional grid known as a filter or a kernel.
  • Assume we unroll the 2D structures into a vector: output from each receptive field is calculated as:

max

 

2499499

26 of 93

How to Identify a Feature in a Receptive Field?

  •  

 

2499499

27 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

2499499

28 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

2499499

29 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

1

2499499

30 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

1

3

2499499

31 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

1

3

2

2499499

32 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

1

3

2

1

2499499

33 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

Stride = 1

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

1

3

2

1

1

2499499

34 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

2499499

35 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

2499499

36 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

2499499

37 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

2499499

38 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

2499499

39 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

1

2499499

40 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

1

0

2499499

41 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

1

0

2

2499499

42 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

1

0

2

-2

2499499

43 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

0

1

0

2

-2

1

2499499

44 of 93

How to Implement Scanning of Features with Weight-Sharing?

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

1

1

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

0

1

1

1

0

0

0

0

0

1

1

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Horizontal line detector

 

-1

0

Stride = 1

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

3

-3

-2

1

1

0

3

-3

-1

-1

2

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

Vertical line detector

Feature Map

2499499

45 of 93

Convolution Example

2499499

46 of 93

If we convolve a kernel of size 5x5 with an image of size 8x8 with the stride of 1, what are the output dimensions?

The Slido app must be installed on every computer you’re presenting from

2499499

47 of 93

If we convolve 4 kernels of size 5x5 with an image of size 8x8 with the stride of 1, what are the output dimensions?

The Slido app must be installed on every computer you’re presenting from

2499499

48 of 93

 

 

 

2499499

49 of 93

 

 

 

 

The dimension of features decreases but the depth usually increases.

2499499

50 of 93

What If We Want to Keep the Same Image Dimensions?

  • We can use padding.

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

0

2499499

51 of 93

What If We Want to Keep the Same Image Dimensions?

  • We can use padding.

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

0

0

Feature Map

 

2499499

52 of 93

Avoiding Redundancy

  • We do not need to move one pixel at a time.

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

0

Stride = 2

2499499

53 of 93

Avoiding Redundancy

  • We do not need to move one pixel at a time.

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

1

1

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

1

1

1

0

0

0

0

0

0

0

1

1

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

-1

-1

-1

2

2

2

-1

-1

-1

Feature extractor

Or kernel

Horizontal line detector

 

0

-1

 

Stride = 2

……

2499499

54 of 93

Characteristics of Image Analysis

Patterns are usually smaller than high-resolution images.

Patterns can appear anywhere in the image.

Addressed with convolution

2499499

55 of 93

Characteristics of Image Analysis

Patterns are usually smaller than high-resolution images.

Patterns can appear anywhere in the image.

Patterns might appear at different scales.

2499499

56 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

-1

0

1

3

2

1

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

Down-sampling features

 

This is called pooling.

2499499

57 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

1

3

2

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

This is called pooling.

 

Down-sampling features

2499499

58 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

1

3

2

1

1

1

0

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

 

Down-sampling features

2499499

59 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

 

Down-sampling features

2499499

60 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

0

0

1

1

1

0

0

0

-1

-1

-1

1

1

1

-1

-1

-1

-1

-1

-1

1

1

2

1

1

0

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

 

Down-sampling features

2499499

61 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

1

1

1

2

1

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

 

Down-sampling features

2499499

62 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

1

1

1

2

1

-1

0

1

0

2

-2

1

1

1

0

-3

-2

1

1

3

-3

-1

1

3

-2

-1

-1

2

2

-1

-1

-2

1

2

1

-2

0

-1

2

-1

-1

2

-1

-1

2

-1

3

0

2

-1

2

-1

 

Down-sampling features

2499499

63 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

1

1

1

2

1

1

1

1

3

-1

2

1

0

-1

2

-1

-1

2

-1

-1

2

-1

3

2

2

 

Down-sampling features

2499499

64 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

1

1

1

2

1

1

1

1

3

2

1

0

-1

2

-1

-1

2

-1

-1

2

-1

3

2

 

Down-sampling features

2499499

65 of 93

How Can We Detect Features at Different Scale?

-1

-1

-1

2

2

2

-1

-1

-1

3

2

1

1

1

1

1

2

1

1

1

1

3

2

1

0

-1

2

-1

-1

2

-1

-1

2

-1

3

2

 

Down-sampling features

2499499

66 of 93

CNN Design

  • Network Design Considerations for Images
  • Element of the Convolutional Neural Network
  • CNN Design
  • CNN History

2499499

67 of 93

Convolution

Pooling

Good features for image understanding tasks

Flatten

Cat?

This network can also be deep.

2499499

68 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

2499499

69 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

def __init__(self, num_classes=10):

super().__init__()

 

 

self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)

nn.Conv2d(

in_channels, # # of input feature maps (e.g., 3 for RGB)

out_channels, # # of filters you learn → # of output feature maps

kernel_size, # filter size (int or (kh, kw))

stride=1, # step size of the filter (int or (sh, sw))

padding=0, # add zeros around the image (int, tuple, or 'same')

bias=True, # learn an additive bias per output channel

padding_mode='zeros' # 'zeros' (default), 'reflect', 'replicate'

)

 

2499499

70 of 93

How many parameters are trained in the first convolution layer?

The Slido app must be installed on every computer you’re presenting from

2499499

71 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

def __init__(self, num_classes=10):

super().__init__()

 

self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)

self.pool = nn.MaxPool2d(2, 2)

 

 

 

2499499

72 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

def __init__(self, num_classes=10):

super().__init__()

 

self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)

self.pool = nn.MaxPool2d(2, 2)

 

 

 

self.conv2 = nn.Conv2d(8, 64, kernel_size=3, padding=1)

 

 

 

2499499

73 of 93

How many parameters are trained in the second convolution layer?

The Slido app must be installed on every computer you’re presenting from

2499499

74 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

def __init__(self, num_classes=10):

super().__init__()

 

self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)

self.pool = nn.MaxPool2d(2, 2)

 

 

 

self.conv2 = nn.Conv2d(8, 64, kernel_size=3, padding=1)

 

 

self.fc1 = nn.Linear(64 * 7 * 7, 256)

self.fc2 = nn.Linear(256, num_classes)

self.dropout = nn.Dropout(0.25)

 

2499499

75 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

def forward(self, x):

x = self.pool(F.relu(self.conv1(x))) # 8x14x14

x = self.pool(F.relu(self.conv2(x))) # 64x7x7

x = x.view(x.size(0), -1) # flatten: 64*7*7

x = self.dropout(F.relu(self.fc1(x)))

return self.fc2(x)

 

 

 

 

 

 

 

2499499

76 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

 

 

 

 

 

 

 

2499499

77 of 93

Convolution

Flatten

Pooling

Convolution

Pooling

Output

 

 

 

 

 

 

 

class SmallCNN(nn.Module):

def __init__(self):

super().__init__()

self.features = nn.Sequential(

nn.Conv2d(1, 8, 3, padding=1),

nn.ReLU(),

nn.MaxPool2d(2),

nn.Conv2d(8, 64, 3, padding=1),

nn.ReLU(),

nn.MaxPool2d(2), )

self.classifier = nn.Sequential(

nn.Flatten(),

nn.Linear(64*7*7, 256),

nn.ReLU(),

nn.Dropout(0.25),

nn.Linear(256, 10))

def forward(self, x):

x = self.features(x)

return self.classifier(x)

2499499

78 of 93

What Do the Filters Learn?

2499499

79 of 93

CNN History

  • Network Design Considerations for Images
  • Element of the Convolutional Neural Network
  • CNN Design
  • CNN History

2499499

80 of 93

History

  • How do animals see?
    • What is the neural process from eye to recognition?
  • Hubel and Wiesel 1959
    • First study on neural correlates of vision.
    • Receptive Fields in Cat's visual cortex
    • 24 cats, anaesthetized, immobilized, on artificial respirators
      • Electrodes in brain
    • Beamed light of different patterns into the eyes and measured neural responses

2499499

81 of 93

Gabor Filters

  • Hubel and Wiesel study introduced two types of cells in brain:
    • Simple cells: Strong response to visual inputs with a simple edge oriented at a particular angle and located at a particular position.
      • More studies showed that the responses from simple cells can be modeled as Gabor filters:

    • Complex cells: Respond to more complex stimuli derived by combining and processing the output of simple cells.

 

 

 

2499499

82 of 93

1980s and 1990s …

  • By the end of the 1980’s, several papers were produced that considerably advanced the field.
  • The idea of backpropagation was first published in French by Yann LeCun in 1985.
  • November 1998, LeCun published one of his most recognized papers describing a “modern” CNN architecture for document recognition, called LeNet-5.

LeCun, Yann, et al. “Gradient-based learning applied to document recognition.” Proceedings of the IEEE 86.11 (1998): 2278–2324.

2499499

83 of 93

The ImageNet Task (2009)

  • 14 million images in more than 20,000 hierarchical categories with hand-annotation
    • Indicate what objects are pictured
    • In at least one million of the images, bounding boxes are also provided

2499499

84 of 93

ImageNet Larg-Scale Visual Recognition Challenge (ILSVRC) (2010)

  • ILSVRC evaluates algorithms for object detection and image classification at large scale.
  • The classification task:
    • Get the “correct” class in your top 5 bets. There are 1000 classes.
  • The localization task:
    • For each bet, put a box around the object. Your box must have at least 50% overlap with the correct box.

2499499

85 of 93

Examples from the Test Set

2499499

86 of 93

ImageNet Larg-Scale Visual Recognition Challenge (ILSVRC)

The first column shows the ground truth labeling on an example image, and the next three show three sample outputs with the corresponding evaluation score.

2499499

87 of 93

The ILSVRC-2012 Competition on ImageNet

Some of the best existing computer vision methods were tried on this dataset by leading computer vision groups from Oxford, INRIA, XRCE

2499499

88 of 93

AlexNet (Link)

  • Alex Krizhevsky (NeurIPS 2012)'s AlexNet architecture is shown here.
  • The paper has 185,000 citations.
  • ReLU activation is used in every hidden layer. These train much faster and are more expressive than logistic units.

2499499

89 of 93

AlexNet's Tricks to Improve Generalization

  •  

Slides adapted from Geofrey Hinton, University of Toronto

2499499

90 of 93

AlexNet's First Convolution Layer

96 convolutional kernels of size 11×11×3 learned by the first convolutional layer on the 224×224×3 input images.

The top 48 kernels were learned on GPU 1 while the bottom 48 kernels were learned on GPU 2.

2499499

91 of 93

How About Higher Layers?

  • Which images make a specific neuron activate?

Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, CVPR, 2014

2499499

92 of 93

How About Higher Layers?

  • Which images make a specific neuron activate?

Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, CVPR, 2014

Top regions for six POOL5 units: Maximally activating images for some POOL5 (5th pool layer) neurons of an AlexNet. The activation values and the receptive field of the particular neuron are shown in white. Some neurons are responsive to upper bodies, text, or specular highlights.

2499499

93 of 93

Convolutional Neural Network

Lecture 18

Reading: Chapter 10 in Bishop Deep Learning Textbook