CS5670: Intro to Computer Vision (Cornell Tech)
Instructor: Noah Snavely
Instructor
Noah’s work
“Statue of Liberty”
3D model
Flickr photos
“Half Dome, Yosemite”
“Colosseum, Rome”
City-scale 3D reconstruction
Reconstruction of Dubrovnik, Croatia, from ~40,000 images
Depth from a single image
Visualizing scenes from tourist photos
Reconstructing dynamic 3D scenes
DynIBaR: Neural Dynamic Image-Based Rendering [https://dynibar.github.io/]
Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, Noah Snavely�CVPR 2023
Teaching assistants
Gene Chou
Ruojin Cai
Haian Jin
Amritansh Kwatra
Important information
Rick Szeliski, Computer Vision: Algorithms and Applications online at: http://szeliski.org/Book/
Today
Today
Every image tells a story
The goal of computer vision
Can computers match human perception?
Current models still make very silly mistakes
[Tomer Ullmann, The Illusion-Illusion: Vision Language Models See Illusions Where There are None, arXiv 2024]
Human perception has its shortcomings
https://twitter.com/pickover/status/1460275132958662657/
But humans can tell a lot about a scene from a little information…
Source: “80 million tiny images” by Torralba, et al.
The goal of computer vision
The goal of computer vision
ZED 2i Camera
The goal of computer vision
Terminator 2, 1991
slide credit: Fei-Fei, Fergus & Torralba
sky
building
flag
wall
banner
bus
cars
bus
face
street lamp
slide credit: Fei-Fei, Fergus & Torralba
The goal of computer vision
The goal of computer vision
Source: Nayar and Nishino, “Eyes for Relighting”
Source: Nayar and Nishino, “Eyes for Relighting”
Source: Nayar and Nishino, “Eyes for Relighting”
The goal of computer vision
Super-resolution (source: 2d3)
Low-light photography
(credit: Hasinoff et al., SIGGRAPH ASIA 2016)
Depth of field on cell phone camera (source: Google Research Blog)
Removing objects (Google Magic Eraser)
April 10, 2019
Why study computer vision?
Optical character recognition (OCR)
Digit recognition, AT&T labs (1990’s)
Automatic check processing
Sudoku grabber
Face detection
Face analysis and recognition
Vision-based biometrics
Who is she?
Source: S. Seitz
Vision-based biometrics
“How the Afghan Girl was Identified by Her Iris Patterns” Read the story
Source: S. Seitz
Login without a password
Fingerprint scanners on many new smartphones and other devices
Face unlock on Apple iPhone X�See also http://www.sensiblevision.com/
New York Times, Jan. 18, 2020
by Kashmir Hill
Bird identification
Merlin Bird ID (based on Cornell Tech technology!)
Special effects: shape capture
The Matrix movies, ESC Entertainment, XYZRGB, NRC
Source: S. Seitz
Special effects: motion capture
Pirates of the Carribean, Industrial Light and Magic
Source: S. Seitz
3D face tracking w/ consumer cameras
Snapchat Lenses
Face2Face system (Thies et al.)
Image synthesis
Karras, et al., Progressive Growing of GANs for Improved Quality, Stability, and Variation, ICLR 2018
Which face is real?
Image synthesis
“An astronaut riding a horse in a photorealistic style” – DALL-E 2
“A photo of a Corgi dog riding a bike in Times Square. It is wearing sunglasses and a beach hat” – Imagen
Sports
Sportvision first down line
Explanation on www.howstuffworks.com
Source: S. Seitz
Smart cars
Self-driving cars
Waymo
Robotics
NASA’s Mars Curiosity Rover
Amazon Picking Challenge
http://www.robocup2016.org/en/events/amazon-picking-challenge/
Amazon Prime Air
Amazon Scout
Medical imaging
3D imaging (MRI, CT)
Skin cancer classification with deep learning https://cs.stanford.edu/people/esteva/nature/
Virtual & Augmented Reality
6DoF head tracking
Hand & body tracking
3D-360 video capture
3D scene understanding
Current state of the art
Why is computer vision difficult?
Viewpoint variation
Illumination
Scale
Why is computer vision difficult?
Intra-class variation
Background clutter
Motion (Source: S. Lazebnik)
Occlusion
Challenges: local ambiguity
slide credit: Fei-Fei, Fergus & Torralba
But there are lots of visual cues we can use…
Source: S. Lazebnik
Bottom line
Image source: F. Durand
Artist Julian Beever with his anamorphic Coke bottle
CS5670: Introduction to Computer Vision
Course requirements
Course overview (tentative)
1. Low-level vision
Filtering, edge detection
*
=
Feature extraction
Image formation
Project: Hybrid images
Project: Feature detection and matching
2. Geometry & appearance
Projective geometry
Multi-view stereo
Structure from motion
Stereo vision
Image credit: IDS Imaging
Project: Creating panoramas
Project: Neural Radiance Fields (NeRFs)
3. Recognition, Deep Learning & Generative Models
“dog”
Image classification
Convolutional Neural Networks
Image generation
“a class watching a computer vision lecture at Cornell Tech”
Project: Image diffusion models
Lectures
Grading
Late policy
Academic Integrity
Questions?