1
EC 500 A1: �Foundations of Computer Vision
Course information
https://sites.google.com/view/bu-ec-500-au26-chao
(for course information, weekly schedule, and reading assignment updates)
Dr. Wei-Lun (Harry) Chao (chao209@bu.edu), Office: PHO437
Associate professor in ECE (PhD: USC; Postdoc: Cornell)
Sanjana Sanjeev Kumar (sanjask@bu.edu), CS MS student
2
Machine learning, computer vision, and applications in
3
Machine learning, computer vision, and applications in
4
Application-inspired ML & CV
ML & CV for applications
Learning with “imperfect” data
5
[Zhu et al., 2014]
KITTI
(Germany)
Argoverse
(USA)
nuScenes
(USA, Singapore)
Lyft
(USA)
Waymo
(USA)
[Wang et al., 2020]
Course information
Course information
7
Communications
8
9
Questions?
Grading and homework (tentative)
Grading (subject to slight change)
Guidelines
Final project first glance (subject to change)
Final project first glance (subject to change)
Schedule
In-class quizzes (linear algebra)
In-class quizzes
Schedule
Homework
Exams & final project presentation
Policy
Academic integrity
(Re-)grading
AI Policy for the course
16
17
Questions?
Pre-requisites & what to expect?
18
Review materials
19
Caution!
20
This is a 500-level course!
Caution!
21
One of the aims is to provide students with a strong foundational background, enabling them to pursue computer-vision-centered or machine-learning-centered MS/PhD paths or explore future opportunities in the computer vision, machine learning, and artificial intelligence industries.
This is a 500-level course!
Caution!
22
The course is not simply knowledge feeding, and I will leave space for you to read, think, and explore!
This is a 500-level course!
23
Questions?
Course descriptions & goals
24
Course descriptions & goals
25
Textbook
26
Foundations of Computer Vision
First Day®
ACCESS
CONVENIENCE
AFFORDABILITY
27
Benefits of Inclusive Access:
• Lower price than traditional purchase
• Guaranteed to get the right materials for your course
• Seamless digital access
• Option to opt out before deadline
28
Exclusive Preferred Pricing Through First Day®
Your First Day Price | $54.85 |
ENG EC 500
FOUNDATIONS OF COMPUTER VISION
*This course material charge will be applied to your student account in October.
Opt Out Deadline: September 22nd
29
If you do not wish to participate in the program, you can choose to opt out within Blackboard by September 22nd. Opting out is not recommended, and you will be responsible for purchasing your required materials without preferred pricing. To opt out, use the “Course Materials” link in Blackboard.
If you do not opt out by the deadline, you have agreed to purchase these materials, and the cost will be charged to your student account at the price listed.
Suggested References
30
Generative Deep Learning:
Teaching Machines To Paint, Write, Compose, and Play
(second edition)
Computer Vision: Algorithms and Applications
(second edition)
Other great textbooks
31
Deep Learning:
Foundations and Concepts
Understanding Deep Learning
Dive into Deep Learning
PDF accessible for the 1st and 3rd books – check their websites
Other excellent CV courses
32
Other excellent CV courses
33
How to do/learn well?
34
Semi-flipped classroom approach
35
How to read a textbook?
My vision for this course
My suggestions …
Important for this week
38
Important dates
39
40
Questions?
About using AI tools
41
Writing is thinking
42
Writing is thinking
43
About math
44
About math
45
About linear algebra
46
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
47
Questions?
Today
Introduction
Course overview
48
49
What is computer vision?
What is computer vision?
Human vision is capable of extracting information about the world around us using only the light that reflects off surfaces in the direction of our eyes.
Our eyes are sensors. Our brains have to translate the information collected by millions of photoreceptors in our retinas into an interpretation of the world in front of us.
Computer vision studies how to reproduce in a computer the ability to see
Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.
50
Input: the structure of ambient light
51
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Output: measuring lights vs. scene properties
52
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
The study of vision is interdisciplinary
53
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
The study of vision is interdisciplinary
54
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Visual pathways
55
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
56
Questions?
What is computer vision?
Computer vision studies how to reproduce in a computer the ability to see
Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.
57
Vision is the process of discovering from images what is presented in the world, and where it is.
David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.
Computer vision
58
[Source: Detectron2]
[Source: Graham Murdoch/Popular Science]
Computer vision: data
59
Image (s)
Video (s) = sequence of images
RGB image (s): Three matrices
Computer vision: data
60
0 | 0 | 124 | 255 | 125 |
0 | 0 | 125 | 126 | 60 |
0 | 0 | 126 | 60 | 126 |
0 | 0 | 0 | 127 | 60 |
0 | 0 | 0 | 0 | 128 |
0 | 0 | 124 | 255 | 125 |
0 | 0 | 125 | 126 | 60 |
0 | 0 | 126 | 60 | 126 |
0 | 0 | 0 | 127 | 60 |
0 | 0 | 0 | 0 | 128 |
0 | 0 | 124 | 255 | 125 |
0 | 0 | 125 | 126 | 60 |
0 | 0 | 126 | 60 | 126 |
0 | 0 | 0 | 127 | 60 |
0 | 0 | 0 | 0 | 128 |
Computer vision: data
61
Image (s)
Video (s) = sequence of images
RGBD image (s): Four matrices
Entry value
= depth
Computer vision: data
62
Point cloud
A collection of 3D (or 4D) points
x coordinate |
y coordinate |
z coordinate |
reflectance |
N points = 3-by-N or 4-by-N matrix
Computer vision: data
63
Image aligned with point cloud
LiDAR-based vision
64
[Source: Graham Murdoch/Popular Science]
LiDAR:
LiDAR-based vision
65
[Credits: Lisa Wu’s presentation]
66
Questions?
What is computer vision?
Computer vision studies how to reproduce in a computer the ability to see
Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.
67
Vision is the process of discovering from images what is presented in the world, and where it is.
David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.
Three representation directions
68
S: scene
I: image
2: Reconstruction
1: Recognition
tree
3: Generation
tree
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Computer vision: representative tasks
69
Computer vision: representative tasks
Computer vision: representative tasks
71
Retrieval, image-to-image search
Computer vision: representative tasks
72
Depth estimation and 3D reconstruction
Computer vision: representative tasks
73
Computer vision: representative tasks
74
Style transfer
[Figure credit: CycleGAN, ICCV 2017]
75
Questions?
What is computer vision?
Computer vision studies how to reproduce in a computer the ability to see
Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.
76
Vision is the process of discovering from images what is presented in the world, and where it is.
David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.
How to let computers recognize objects?
A cat?
A lion?
A car?
Percept:
See a picture
Action:
Tell the object class
Human design vs. machine-learning-based
cat
Design
cat
cat
cat
Data
collection
“Learn”
“Coding” the rules:
Can you list the rules of recognizing a cat?
Underlying idea:
Humans sometimes are good at “making decisions” BUT are not good at “explaining decisions”.
Learning-based computer vision
79
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
What is machine learning?
This book is about learning from data.
Sergios Theodoridis. Machine learning: a Bayesian and optimization perspective.
We choose the title “learning from data” that faithfully describes what the subject is about.
Y. Abu-Mostafa, M. Magdon-Ismail, H-T Lin. Learning from data.
Machine Learning Overview
Learning from Data
81
Machine Learning Overview
Learning from Data
Algorithm
Data
Evaluation
82
Machine Learning Overview
Learning from Data
Algorithm
Data
Evaluation
Goal
83
Example: coin classifier
84
Machine learning algorithms
Training data
Learned models
Test data
[Figure credit: Y. Abu-Mostafa, M. Magdon-Ismail, H-T Lin. Learning from data.]
What is deep learning (deep neural networks)?
Image
Label (e.g., dog or cat)
Classifier
See a picture
Tell the object class
A sequence of “learnable” computation!
Example: image classification
86
[Gif credits: Gradient descent 3Blue1Brown series S3 E2]
A sequence of “learnable” computation!
The progress of deep learning
[Simonyan et al., 2015]
[Szegedy et al., 2015]
[Huang et al., 2017]
[He et al., 2016]
[Krizhevsky et al., 2012]
The progress of deep learning
Visual transformers
[Liu et al., 2021]
[Battaglia et al., 2018]
Graph neural networks
[Qi et al., 2017]
PointNet
[Zoph et al., 2017]
Neural architecture search
89
Questions?
Today
Introduction
Course overview
90
Topics
1. Introduction to computer vision
a. Introduction to the course
b. A simple vision system
91
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
2. Image formation
a. Concepts of imaging and lenses
b. Images and 3D geometry
c. Camera modeling
d. Cameras as linear systems
92
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
3. Foundations of image processing
a. Linear filtering and convolution
b. Fourier analysis
c. Blur filters, image derivatives, and filter banks
d. (Up/down) sampling
e. Image pyramids
93
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
4. Foundations of learning
a. Introduction to learning
b. Gradient-based learning algorithms
c. Generalization
d. Neural networks as distribution transformers
94
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
5. Probabilistic models of images
a. Color
b. Statistical image models
c. Textures
95
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
6. Neural architectures for vision
a. Convolutional neural nets
b. Transformers
96
Visual transformers
[Liu et al., 2021]
[Simonyan et al., 2015]
ConvNet (VGG Net)
Topics
7. Generative image models and representation learning
a. Representation learning
b. Generative models
97
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
8. Understanding vision with semantics and language
a. Visual recognition
b. Vision and language
98
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
9. Challenges in learning-based vision
a. Data bias and shift
b. Robustness and generality
c. Transfer learning and adaptation
99
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Topics
10. Understanding geometry
a. Stereo vision
b. Homographies
c. Depth estimation from single images
d. Feature detection and matching
e. Multi-view geometry and structure from motion
f. Radiance fields
100
Topics
11. Understanding motion
a. Motion estimation
b. Optical flow estimation
c. Object tracking
101
Topics
12. Advanced topics
a. Exciting 3D & 4D models
b. Computer vision for biodiversity
102
[Figure credit: Jianyuan Wang, VGGT, CVPR 2025.]
TODO
103
Today
104
Neural networks for image classification
105
[Gif credits: Gradient descent 3Blue1Brown series S3 E2]
A sequence of “learnable” computation!
Learning neural networks for image classification
Pre-trained network (from GitHub, Huggingface)
[Very Deep Convolutional Networks for Large-Scale Image Recognition, ICLR 2015]
Other topics
108