Convolutional Neural Network
Lecture 18
Introduction to convolutional architecture for image analysis
EECS 189/289, Fall 2025 @ UC Berkeley
Joseph E. Gonzalez and Narges Norouzi
EECS 189/289, Fall 2025 @ UC Berkeley
Joseph E. Gonzalez and Narges Norouzi
Roadmap
2499499
Network Design Considerations for Images
2499499
Story So Far
2499499
A Problem
2499499
A Problem
A network that recognizes "cat" feature regardless of the location and scale of the target object.
2499499
A Problem
Network must be shift invariant.
2499499
Solution: Scanning
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
2499499
Scanning Is Referred to as Convolution
max
Maximum of all the outputs (Equivalent of Boolean OR)
The strength of having the feature anywhere in the image
Same shared weights
2499499
Scanning Is Referred to as Convolution
max
2499499
Which of the following is correct?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
2499499
Regular Network vs. Scanning Network
max
2499499
Element of the Convolutional Neural Network
2499499
Back to Scanning Network
max
2499499
How to Identify a Feature in a Receptive Field?
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
1
3
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
1
3
2
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
1
3
2
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
Stride = 1
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
1
3
2
1
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
1
0
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
1
0
2
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
1
0
2
-2
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
0
1
0
2
-2
1
2499499
How to Implement Scanning of Features with Weight-Sharing?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Horizontal line detector
-1
0
Stride = 1
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
3
-3
-2
1
1
0
3
-3
-1
-1
2
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
Vertical line detector
Feature Map
2499499
Convolution Example
2499499
If we convolve a kernel of size 5x5 with an image of size 8x8 with the stride of 1, what are the output dimensions?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
2499499
If we convolve 4 kernels of size 5x5 with an image of size 8x8 with the stride of 1, what are the output dimensions?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
2499499
2499499
The dimension of features decreases but the depth usually increases.
2499499
What If We Want to Keep the Same Image Dimensions?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
0
2499499
What If We Want to Keep the Same Image Dimensions?
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
0
0
Feature Map
2499499
Avoiding Redundancy
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
0
Stride = 2
2499499
Avoiding Redundancy
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
Feature extractor
Or kernel
Horizontal line detector
0
-1
Stride = 2
……
2499499
Characteristics of Image Analysis
Patterns are usually smaller than high-resolution images.
Patterns can appear anywhere in the image.
Addressed with convolution
2499499
Characteristics of Image Analysis
Patterns are usually smaller than high-resolution images.
Patterns can appear anywhere in the image.
Patterns might appear at different scales.
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
-1
0
1
3
2
1
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
This is called pooling.
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
1
3
2
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
This is called pooling.
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
1
3
2
1
1
1
0
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
0
0
1
1
1
0
0
0
-1
-1
-1
1
1
1
-1
-1
-1
-1
-1
-1
1
1
2
1
1
0
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
1
1
1
2
1
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
1
1
1
2
1
-1
0
1
0
2
-2
1
1
1
0
-3
-2
1
1
3
-3
-1
1
3
-2
-1
-1
2
2
-1
-1
-2
1
2
1
-2
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
0
2
-1
2
-1
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
1
1
1
2
1
1
1
1
3
-1
2
1
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
2
2
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
1
1
1
2
1
1
1
1
3
2
1
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
2
Down-sampling features
2499499
How Can We Detect Features at Different Scale?
-1 | -1 | -1 |
2 | 2 | 2 |
-1 | -1 | -1 |
3
2
1
1
1
1
1
2
1
1
1
1
3
2
1
0
-1 | 2 | -1 |
-1 | 2 | -1 |
-1 | 2 | -1 |
3
2
Down-sampling features
2499499
CNN Design
2499499
Convolution
Pooling
Good features for image understanding tasks
Flatten
Cat?
This network can also be deep.
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
def __init__(self, num_classes=10):
super().__init__()
self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)
nn.Conv2d(
in_channels, # # of input feature maps (e.g., 3 for RGB)
out_channels, # # of filters you learn → # of output feature maps
kernel_size, # filter size (int or (kh, kw))
stride=1, # step size of the filter (int or (sh, sw))
padding=0, # add zeros around the image (int, tuple, or 'same')
bias=True, # learn an additive bias per output channel
padding_mode='zeros' # 'zeros' (default), 'reflect', 'replicate'
)
2499499
How many parameters are trained in the first convolution layer?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
def __init__(self, num_classes=10):
super().__init__()
self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)
self.pool = nn.MaxPool2d(2, 2)
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
def __init__(self, num_classes=10):
super().__init__()
self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)
self.pool = nn.MaxPool2d(2, 2)
self.conv2 = nn.Conv2d(8, 64, kernel_size=3, padding=1)
2499499
How many parameters are trained in the second convolution layer?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
def __init__(self, num_classes=10):
super().__init__()
self.conv1 = nn.Conv2d(3, 8, kernel_size=3, padding=1)
self.pool = nn.MaxPool2d(2, 2)
self.conv2 = nn.Conv2d(8, 64, kernel_size=3, padding=1)
self.fc1 = nn.Linear(64 * 7 * 7, 256)
self.fc2 = nn.Linear(256, num_classes)
self.dropout = nn.Dropout(0.25)
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
def forward(self, x):
x = self.pool(F.relu(self.conv1(x))) # 8x14x14
x = self.pool(F.relu(self.conv2(x))) # 64x7x7
x = x.view(x.size(0), -1) # flatten: 64*7*7
x = self.dropout(F.relu(self.fc1(x)))
return self.fc2(x)
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
2499499
Convolution
Flatten
Pooling
Convolution
Pooling
Output
class SmallCNN(nn.Module):
def __init__(self):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(1, 8, 3, padding=1),
nn.ReLU(),
nn.MaxPool2d(2),
nn.Conv2d(8, 64, 3, padding=1),
nn.ReLU(),
nn.MaxPool2d(2), )
self.classifier = nn.Sequential(
nn.Flatten(),
nn.Linear(64*7*7, 256),
nn.ReLU(),
nn.Dropout(0.25),
nn.Linear(256, 10))
def forward(self, x):
x = self.features(x)
return self.classifier(x)
2499499
What Do the Filters Learn?
2499499
CNN History
2499499
History
2499499
Gabor Filters
2499499
1980s and 1990s …
LeCun, Yann, et al. “Gradient-based learning applied to document recognition.” Proceedings of the IEEE 86.11 (1998): 2278–2324.
Image from https://medium.com/
2499499
The ImageNet Task (2009)
2499499
ImageNet Larg-Scale Visual Recognition Challenge (ILSVRC) (2010)
2499499
Examples from the Test Set
2499499
ImageNet Larg-Scale Visual Recognition Challenge (ILSVRC)
The first column shows the ground truth labeling on an example image, and the next three show three sample outputs with the corresponding evaluation score.
2499499
The ILSVRC-2012 Competition on ImageNet
Some of the best existing computer vision methods were tried on this dataset by leading computer vision groups from Oxford, INRIA, XRCE
2499499
AlexNet (Link)
Credit: https://medium.com/data-science/advanced-topics-in-deep-convolutional-neural-networks-71ef1190522d
2499499
AlexNet's Tricks to Improve Generalization
Slides adapted from Geofrey Hinton, University of Toronto
2499499
AlexNet's First Convolution Layer
96 convolutional kernels of size 11×11×3 learned by the first convolutional layer on the 224×224×3 input images.
The top 48 kernels were learned on GPU 1 while the bottom 48 kernels were learned on GPU 2.
2499499
How About Higher Layers?
Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, CVPR, 2014
2499499
How About Higher Layers?
Ross Girshick, Jeff Donahue, Trevor Darrell, Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, CVPR, 2014
Top regions for six POOL5 units: Maximally activating images for some POOL5 (5th pool layer) neurons of an AlexNet. The activation values and the receptive field of the particular neuron are shown in white. Some neurons are responsive to upper bodies, text, or specular highlights.
2499499
Convolutional Neural Network
Lecture 18
Reading: Chapter 10 in Bishop Deep Learning Textbook