1 of 86

CS5670: Intro to Computer Vision (Cornell Tech)

Instructor: Noah Snavely

2 of 86

Instructor

  • Noah Snavely (snavely@cs.cornell.edu)�
  • Research interests:
    • Computer vision and graphics
    • 3D reconstruction and visualization of Internet photo collections
    • Deep learning for computer graphics

3 of 86

Noah’s work

  • Automatic 3D reconstruction from Internet photo collections

“Statue of Liberty”

3D model

Flickr photos

“Half Dome, Yosemite”

“Colosseum, Rome”

4 of 86

City-scale 3D reconstruction

Reconstruction of Dubrovnik, Croatia, from ~40,000 images

5 of 86

Depth from a single image

6 of 86

Visualizing scenes from tourist photos

7 of 86

Reconstructing dynamic 3D scenes

DynIBaR: Neural Dynamic Image-Based Rendering [https://dynibar.github.io/]

Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, Noah SnavelyCVPR 2023

8 of 86

Teaching assistants

  • Please check back on course webpage for office hours

https://www.cs.cornell.edu/courses/cs5670/2025sp/

Amritansh Kwatra

ak2244@cornell.edu

9 of 86

Important information

  • Textbook:

Rick Szeliski, Computer Vision: Algorithms and Applications online at: http://szeliski.org/Book/

  • Announcements/discussion via Ed Discussions (via Canvas)

  • Assignment turnin via GitHub Classroom and CMSX: https://cmsx.cs.cornell.edu

10 of 86

Today

  1. What is computer vision?

  • Why study computer vision?

  • Course overview

  • Images & image filtering [time permitting]

11 of 86

Today

  • Readings
    • Szeliski, Chapter 1 (Introduction)

12 of 86

Every image tells a story

  • Goal of computer vision: perceive the “story” behind the picture
  • Compute properties of the world
    • 3D shape
    • Names of people or objects
    • What happened?

13 of 86

The goal of computer vision

14 of 86

Can computers match human perception?

  • Yes and no
    • humans are better at “hard” things, are more robust, and make inferences quickly and cheaply

  • But huge progress
    • Accelerating in the last 10 years due to deep learning, large vision-language models
    • What is considered “hard” keeps changing

15 of 86

16 of 86

17 of 86

18 of 86

19 of 86

Current models still make very silly mistakes

[Tomer Ullmann, The Illusion-Illusion: Vision Language Models See Illusions Where There are None, arXiv 2024]

20 of 86

Human perception has its shortcomings

https://twitter.com/pickover/status/1460275132958662657/

21 of 86

But humans can tell a lot about a scene from a little information…

Source: “80 million tiny images” by Torralba, et al.

22 of 86

23 of 86

The goal of computer vision

24 of 86

The goal of computer vision

  • Compute the 3D shape of the world

ZED 2i Camera

25 of 86

The goal of computer vision

  • Recognize objects and people

Terminator 2, 1991

26 of 86

slide credit: Fei-Fei, Fergus & Torralba

27 of 86

sky

building

flag

wall

banner

bus

cars

bus

face

street lamp

slide credit: Fei-Fei, Fergus & Torralba

28 of 86

The goal of computer vision

  • “Enhance” images

29 of 86

30 of 86

The goal of computer vision

  • Forensics

Source: Nayar and Nishino, “Eyes for Relighting”

31 of 86

Source: Nayar and Nishino, “Eyes for Relighting”

32 of 86

Source: Nayar and Nishino, “Eyes for Relighting”

33 of 86

The goal of computer vision

  • Improve photos (“Computational Photography”)

Super-resolution (source: 2d3)

Low-light photography

(credit: Hasinoff et al., SIGGRAPH ASIA 2016)

Depth of field on cell phone camera (source: Google Research Blog)

Removing objects (Google Magic Eraser)

34 of 86

April 10, 2019

35 of 86

Why study computer vision?

  • Billions of images/videos captured per day
  • Huge number of potential applications
  • The next slides show the current state of the art

36 of 86

Optical character recognition (OCR)

Digit recognition, AT&T labs (1990’s)

http://yann.lecun.com/exdb/lenet/

    • If you have a scanner, it probably came with OCR software

Automatic check processing

37 of 86

Face detection

  • Nearly all cameras detect faces in real time
    • (Why?)

38 of 86

Face analysis and recognition

39 of 86

Vision-based biometrics

Who is she?

Source: S. Seitz

40 of 86

Vision-based biometrics

How the Afghan Girl was Identified by Her Iris Patterns” Read the story

Source: S. Seitz

41 of 86

Login without a password

Fingerprint scanners on many new smartphones and other devices

Face unlock on Apple iPhone X�See also http://www.sensiblevision.com/

42 of 86

New York Times, Jan. 18, 2020

by Kashmir Hill

43 of 86

44 of 86

Bird identification

Merlin Bird ID (based on Cornell Tech technology!)

45 of 86

Special effects: shape capture

The Matrix movies, ESC Entertainment, XYZRGB, NRC

Source: S. Seitz

46 of 86

Special effects: motion capture

Pirates of the Carribean, Industrial Light and Magic

Source: S. Seitz

47 of 86

48 of 86

49 of 86

3D face tracking w/ consumer cameras

Snapchat Lenses

Face2Face system (Thies et al.)

50 of 86

Image synthesis

Karras, et al., Progressive Growing of GANs for Improved Quality, Stability, and Variation, ICLR 2018

51 of 86

Which face is real?

52 of 86

Image synthesis

“An astronaut riding a horse in a photorealistic style” – DALL-E 2

“A photo of a Corgi dog riding a bike in Times Square. It is wearing sunglasses and a beach hat” – Imagen

53 of 86

Sports

Sportvision first down line

Explanation on www.howstuffworks.com

Source: S. Seitz

54 of 86

Smart cars

  • Mobileye
  • Tesla Autopilot
  • Safety features in many cars

55 of 86

Self-driving cars

Waymo

56 of 86

Robotics

Amazon Prime Air

Amazon Scout

57 of 86

Medical imaging

3D imaging (MRI, CT)

Skin cancer classification with deep learning https://cs.stanford.edu/people/esteva/nature/

58 of 86

59 of 86

Virtual & Augmented Reality

6DoF head tracking

Hand & body tracking

3D-360 video capture

3D scene understanding

60 of 86

Current state of the art

  • You just saw many examples of current systems.
    • Many of these are less than 10 years old

  • Computer vision is an active research area, and rapidly changing
    • Many new apps in the next 5 years
    • Deep learning and generative methods powering many modern applications

  • Many startups across a dizzying array of areas
    • Generative AI, robotics, autonomous vehicles, medical imaging, construction, inspection, VR/AR, …

61 of 86

Why is computer vision difficult?

Viewpoint variation

Illumination

Scale

62 of 86

Why is computer vision difficult?

Intra-class variation

Background clutter

Motion (Source: S. Lazebnik)

Occlusion

63 of 86

Challenges: local ambiguity

slide credit: Fei-Fei, Fergus & Torralba

64 of 86

But there are lots of visual cues we can use…

Source: S. Lazebnik

65 of 86

Bottom line

  • Perception is an inherently ambiguous problem
    • Many different 3D scenes could have given rise to a given 2D image

    • We often must use prior knowledge about the world’s structure

Image source: F. Durand

Artist Julian Beever with his anamorphic Coke bottle

66 of 86

67 of 86

68 of 86

69 of 86

CS5670: Introduction to Computer Vision

  • Project-based course whose goal is to teach you the fundamentals of computer vision – image processing, geometry, recognition – in a hands-on way

  • This course covers fundamentals, including mathematical fundamentals. It is not a course specifically on deep learning or generative AI, though those topics will be covered.

70 of 86

Course requirements

  • Prerequisites
    • Data structures
    • Good working knowledge of Python programming
    • Linear algebra
    • Vector calculus

  • Course does not assume prior imaging experience
    • computer vision, image processing, graphics, etc.

71 of 86

Course overview (tentative)

  1. Low-level vision
    • image processing, edge detection, feature detection, cameras, image formation

  • Geometry & appearance
    • projective geometry, stereo, structure from motion, optimization, lighting & materials

  • Recognition & generative models
    • object classification, deep learning, diffusion models

72 of 86

1. Low-level vision

  • Basic image processing and image formation

Filtering, edge detection

*

=

Feature extraction

Image formation

73 of 86

Project: Hybrid images

74 of 86

75 of 86

76 of 86

Project: Feature detection and matching

77 of 86

2. Geometry & appearance

Projective geometry

Multi-view stereo

Structure from motion

Stereo vision

Image credit: IDS Imaging

78 of 86

Project: Creating panoramas

79 of 86

Project: Neural Radiance Fields (NeRFs)

80 of 86

3. Recognition, Deep Learning & Generative Models

“dog”

Image classification

Convolutional Neural Networks

Image generation

“a class watching a computer vision lecture at Cornell Tech”

81 of 86

Project: Image diffusion models

82 of 86

Lectures

  • Lectures will be held in person in Bloomberg 131
  • If there is an instance where you need to attend lecture remotely, please reach out to the instructor for approval
  • No conversations during class, please!

83 of 86

Grading

  • Approximately weekly short quizzes (typically at the beginning of class on Thursdays)
  • One midterm (take-home), one final exam (in class)

  • Grade breakdown (subject to minor tweaks):
    • Quizzes: 5% (lowest quiz grade dropped)
    • Midterm: 16%
    • Programming projects: 63%
    • Final exam: 16%

84 of 86

Late policy

  • Four free “slip days” will be available for the semester

  • A late project will be penalized by 10% for each day it is late (excepting slip days), and no extra credit will be awarded

85 of 86

Academic Integrity

  • Assignments will be done solo or in pairs (we’ll let you know for each project)
  • Please do not leave any code public on GitHub (or the like) at the end of the semester!
  • We will follow the Cornell Code of Academic Integrity (http://cuinfo.cornell.edu/aic.cfm)
  • If you use ChatGPT (or CoPilot, or similar) on coding assignments, you must disclose that with your submission
    • BUT: We advise you to do all coding yourself, unassisted. You will learn less, and become less capable experts in vision, if you rely on LLMs.

86 of 86

Questions?