1 of 73

CSE 5524: �Computer Vision

2 of 73

Course information

  • Course website:

https://sites.google.com/view/osu-cse-5524-sp25-chao/home

(for course information, weekly schedule, and reading update)

  • Instructor:

Dr. Wei-Lun (Harry) Chao (chao.209), Office: DL 587

Assistant professor in CSE (PhD: USC; Postdoc: Cornell)

  • TA:

Amin Karimi Monsefi (karimimonsefi.1), CSE PhD student

2

3 of 73

A bit about me

Machine learning and its applications to

  • Autonomous driving
  • Computer vision
  • Natural language processing
  • Health care
  • Imageomics

3

Pancreatic

cancer

4 of 73

A bit about me

4

5 of 73

Learning with “imperfect” data

  • Limited data and supervision
  • Imbalanced data
  • Inaccessible data
  • Domain shifts

5

[Zhu et al., 2014]

KITTI

(Germany)

Argoverse

(USA)

nuScenes

(USA, Singapore)

Lyft

(USA)

Waymo

(USA)

[Wang et al., 2020]

6 of 73

Course information

  • Lecture time: Tuesday and Thursday, 2:20 PM - 3:40 PM

  • Office hours: TBD, DL587
    • No office hours the first week

  • TA Office hours: TBD, BE406
    • No office hours the first week

7 of 73

Course information

  • Carmen/GitHub:
    • To set up by the end of this week
    • For announcement, posting course materials (slides), and homework submission

  • Piazza:
    • For discussion. Please register!
    • Link: TBA
    • Please use name.#@osu.edu
    • Access code: osu-cse-5524-SP25-chao

  • Detailed syllabus (pdf):
    • can be found on Carmen and the course website

7

8 of 73

Communications

  • Schedule and reading will be updated on the website
  • Announcements will be made through Carmen
  • Discussions and questions must be posted in Piazza
  • Please only use email to contact me or the TA for urgent or personal issues. Please include the tag "[OSU-CSE-5524]" in the subject line.
  • More details: See website, Carmen, and the syllabus

8

9 of 73

Questions?

10 of 73

Grading and homework (tentative)

 Grading (tentative)

  • Quiz – 10%
  • Homework – 50%
  • Midterm (March 4) – 20%
  • Final (April 23, 2:20 pm) – 20%
    • The final exam is cumulative.
  • The final exam may be replaced by a project. If so, the presentations are on:
    • April 22, 2:20 pm (please reserve)
    • April 23, 2:00 pm

Guidelines

  •  Expect >= 6 homework assignments (including problem and programming sets)
    • Solutions may involve derivations. Grading is based on correctness and clarity. Be concise and show your reasoning in a clear and precise way.
    • Homework completion and submissions are individual but feel free to discuss. You must strictly follow the submission instructions.
    • NOT ALLOWED: ask/search for solutions
    • No late days are accepted.

11 of 73

Tentative schedule

Homework

  • Dates: TBA
  • You will have 1~2 weeks to complete each homework
  • Due is at 23:59 ET

Exams & final project presentation

  • Midterm date(s): 3/4/2025
  • Final exam: 4/23/2025
  • Final project presentation:
    • 4/22/2025 (study day before the final)
    • 4/23/2025

12 of 73

Policy

Academic integrity

  • Plagiarism and other unacceptable violations
    • Zero tolerance
    • I MUST report incidents
  • Please study the related sections in the syllabus (pdf) on academic integrity.
  • Please read OAA’s message on large language models: https://oaa.osu.edu/artificial-intelligence-and-academic-integrity

(Re-)grading

  • Only factual errors will be corrected.
  • Request: one week within the release of your homework and exam grade
  • Format: TBA

13 of 73

Pre-requisites & what to expect?

  • Pre-requisites
    • Data structures and algorithms: 2331
    • Statistics and probability: 5522, Stat 3460, or 3470
    • Decent degree of mathematical sophistication
    • Knowledge of programming, algorithm design, and data structures
  • Suggested backgrounds
    • Linear algebra: Math 2568, 2174, 4568, or 5520H
    • Artificial intelligence: 3521, 5521, or 5243
  • Extensive math and programming-related homework
    • Multivariate calculus, linear algebra, and probability
    • Python 3
    • PyTorch & Hugging Face
  • CV algorithms are often difficult to debug
    • We strongly recommend that you start early.

13

14 of 73

Review materials

  • Please see the reading list in the spreadsheet on the website

  • 4 points quizzes related to linear algebra: completed by 1/30/2024

14

15 of 73

Questions?

16 of 73

Course descriptions & goals

  • Course Description:  Computer vision algorithms for use in human-computer interactive systems; image formation, image features, segmentation, shape analysis, object tracking, motion calculation, and applications. 

  • Course Goals / Objectives: 
    • Master fundamental and recent computer vision algorithms
    • Be competent with computer vision application design and evaluation
    • Be familiar with the Python/PyTorch programming environment
    • Be exposed to original research and applications in computer vision

16

17 of 73

Textbook

  • Required

17

Foundations of Computer Vision

18 of 73

Suggested References

18

Generative Deep Learning:

Teaching Machines To Paint, Write, Compose, and Play

(second edition)

Computer Vision: Algorithms and Applications

(second edition)

PDF accessible through the OSU Library website

19 of 73

Other great textbooks

19

Deep Learning:

Foundations and Concepts

Understanding Deep Learning

Dive into Deep Learning

PDF accessible for the 1st and 3rd books – check their websites

20 of 73

Other excellent CV courses

  • Brown CV: https://browncsci1430.github.io/

20

21 of 73

Other excellent CV courses

  • Computer vision courses are hard to be comprehensive and unified
    • 3D vision, generative vision, deep learning for vision, robotic vision, etc.
    • Even the basic CV courses can be very different

21

22 of 73

Important for this week

  • Register. If you are on the waitlist, you might or might not get in depending on how many empty seats or how many students drop.
  • Register for the class on piazza (see Carmen for the link) --- our main platform for discussion and communication
  • Math review/self-diagnostic: do “CSE5523” Homework #0 and check suggested materials on the website (e.g., linear algebra slides)--- extremely important to check your readiness for the course
  • Python: check suggested tutorials on the website
  • Decision: stay or drop

  • Office hours: start next week

22

23 of 73

How to do/learn well?

  • Lecture and lecture slides for basics
    • Describe basic concepts, tools
    • Describe algorithms and their development with intuition and rigor

  • Textbook reading for completeness and extension

  • Homework for practice, generalization, and implementation

  • Discussion (Piazza, office hours) for further understanding

23

24 of 73

Important dates

  • Midterm: in class (date: 3/4/2025, subject to change)

  • Final: in class (date: 4/23/2025, following the university’s schedule)

  • Final project presentation: in class
    • 4/22/2025 – study day before the final exam week
    • 4/23/2025

  • Online (or pre-recorded) teaching or guest lectures
    • For some weeks, I may be traveling, such as 1/23, 3/4, and 4/17

24

25 of 73

Questions?

26 of 73

Today

Introduction

  • What is computer vision?

Course overview

26

27 of 73

What is computer vision?

28 of 73

What is computer vision?

Human vision is capable of extracting information about the world around us using only the light that reflects off surfaces in the direction of our eyes.

Our eyes are sensors. Our brains have to translate the information collected by millions of photoreceptors in our retinas into an interpretation of the world in front of us.

Computer vision studies how to reproduce in a computer the ability to see

Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.

28

29 of 73

Input: the structure of ambient light

29

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

30 of 73

Output: measuring lights vs. scene properties

30

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

31 of 73

The study of vision is interdisciplinary

  • involving many disciplines (physics, phycology, biology, neuroscience, art, and computer science)

31

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

32 of 73

The study of vision is interdisciplinary

  • Gestalt phycology grouping rules for perceptual organization

32

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

33 of 73

Visual pathways

33

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

34 of 73

Questions?

35 of 73

What is computer vision?

Computer vision studies how to reproduce in a computer the ability to see

Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.

35

Vision is the process of discovering from images what is presented in the world, and where it is.

David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.

36 of 73

Computer vision

36

[Source: Detectron2]

  • A computer sees the world through sensors, which generate images, videos, point cloud, etc.

[Source: Graham Murdoch/Popular Science]

37 of 73

Computer vision: data

37

Image (s)

Video (s) = sequence of images

RGB image (s): Three matrices

38 of 73

Computer vision: data

  • RGB images:

  • What is inside each matrix?
    • {0,1,……,255}
    • Interval: [0, 1]

38

0

0

124

255

125

0

0

125

126

60

0

0

126

60

126

0

0

0

127

60

0

0

0

0

128

0

0

124

255

125

0

0

125

126

60

0

0

126

60

126

0

0

0

127

60

0

0

0

0

128

0

0

124

255

125

0

0

125

126

60

0

0

126

60

126

0

0

0

127

60

0

0

0

0

128

39 of 73

Computer vision: data

39

Image (s)

Video (s) = sequence of images

RGBD image (s): Four matrices

Entry value

= depth

40 of 73

Computer vision: data

40

Point cloud

A collection of 3D (or 4D) points

x coordinate

y coordinate

z coordinate

reflectance

N points = 3-by-N or 4-by-N matrix

41 of 73

Computer vision: data

41

Image aligned with point cloud

42 of 73

LiDAR-based vision

42

[Source: Graham Murdoch/Popular Science]

LiDAR:

  • Light Detection and Ranging sensor
  • accurate 3D point clouds of the environment, centered at the ego-car

43 of 73

LiDAR-based vision

  • A point cloud is formed by LiDAR responses within a short time period

43

[Credits: Lisa Wu’s presentation]

44 of 73

Questions?

45 of 73

What is computer vision?

Computer vision studies how to reproduce in a computer the ability to see

Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.

45

Vision is the process of discovering from images what is presented in the world, and where it is.

David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.

46 of 73

Three representation directions

46

S: scene

I: image

2: Reconstruction

1: Recognition

tree

3: Generation

tree

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

47 of 73

Computer vision: representative tasks

47

48 of 73

Computer vision: representative tasks

49 of 73

Computer vision: representative tasks

49

Retrieval, image-to-image search

50 of 73

Computer vision: representative tasks

50

Depth estimation and 3D reconstruction

51 of 73

Computer vision: representative tasks

51

52 of 73

Computer vision: representative tasks

52

Style transfer

[Figure credit: CycleGAN, ICCV 2017]

53 of 73

Questions?

54 of 73

What is computer vision?

Computer vision studies how to reproduce in a computer the ability to see

Antonio Torralba, Phillip Isola, and William T. Freeman, Foundations of Computer Vision, 2024.

54

Vision is the process of discovering from images what is presented in the world, and where it is.

David Marr, Vision A Computational Investigation into the Human Representation and Processing of Visual Information, 1982.

55 of 73

How to let computers recognize objects?

A cat?

A lion?

A car?

Percept:

See a picture

Action:

Tell the object class

56 of 73

Human design vs. machine-learning-based

cat

Design

cat

cat

cat

Data

collection

“Learn”

“Coding” the rules:

Can you list the rules of recognizing a cat?

Underlying idea:

Humans sometimes are good at “making decisions” BUT are not good at “explaining decisions”.

57 of 73

Learning-based computer vision

57

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

58 of 73

What is machine learning?

This book is about learning from data.

Sergios Theodoridis. Machine learning: a Bayesian and optimization perspective.

We choose the title “learning from data” that faithfully describes what the subject is about.

Y. Abu-Mostafa, M. Magdon-Ismail, H-T Lin. Learning from data.

59 of 73

Machine Learning Overview

  • What is machine learning?

Learning from Data

59

60 of 73

Machine Learning Overview

  • What is machine learning?

Learning from Data

Algorithm

Data

Evaluation

60

61 of 73

Machine Learning Overview

  • What is machine learning?

Learning from Data

Algorithm

Data

Evaluation

Goal

61

62 of 73

Example: coin classifier

62

Machine learning algorithms

Training data

Learned models

Test data

[Figure credit: Y. Abu-Mostafa, M. Magdon-Ismail, H-T Lin. Learning from data.]

63 of 73

What is deep learning (deep neural networks)?

Image

Label (e.g., dog or cat)

Classifier

See a picture

Tell the object class

A sequence of “learnable” computation!

64 of 73

Example: image classification

64

[Gif credits: Gradient descent 3Blue1Brown series S3 E2]

A sequence of “learnable” computation!

65 of 73

The progress of deep learning

[Simonyan et al., 2015]

[Szegedy et al., 2015]

[Huang et al., 2017]

[He et al., 2016]

[Krizhevsky et al., 2012]

66 of 73

The progress of deep learning

Visual transformers

[Liu et al., 2021]

[Battaglia et al., 2018]

Graph neural networks

[Qi et al., 2017]

PointNet

[Zoph et al., 2017]

Neural architecture search

67 of 73

Questions?

68 of 73

Today

Introduction

  • What is computer vision?

Course overview

68

69 of 73

Topics

  • Image formation

  • Image processing, filtering, sampling, etc.

69

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

70 of 73

Topics

  • Foundation of learning
  • Neural network architectures (CNN, Transformers)
  • Visual recognition
  • Vision & language
  • Challenges in learning-based vision

70

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

71 of 73

Topics

  • Probabilistic image modeling
  • Representation learning
  • Generative models

71

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

72 of 73

Topics

  • Geometry & camera modeling
  • 3D from single & stereo
  • Multi-view & SFM
  • Radiance fields
  • Motion & tracking

72

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

73 of 73

TODO

  • See the beginning slides and the course website for suggested reading
  • Background review: math and programming

73