1 of 56

Introduction to computer vision 11

Jean Ponce

jean.ponce@ens.fr

Zuhaib Akhtar za2023@nyu.edu

Ayush Jain aj3152@nyu.edu

Slides will be available after classes

2 of 56

Self-calibration

3 of 56

Types of ambiguity

Projective

15dof

Affine

12dof

Similarity

7dof

Euclidean

6dof

Preserves intersection and tangency

Preserves parallellism, volume ratios

Preserves angles, ratios of length

Preserves angles, lengths

  • With no constraints on the camera calibration matrix or on the scene, we get a projective reconstruction
  • Need additional information to upgrade the reconstruction to affine, similarity, or Euclidean

4 of 56

From uncalibrated to calibrated (affine) cameras

Weak-perspective camera:

Calibrated camera:

Problem: what is Q ?

Assume k and s are known (k=1, s=0)

5 of 56

From uncalibrated to calibrated (affine) cameras

Weak-perspective camera:

Calibrated camera:

Problem: what is Q ?

Assume k and s are known (k=1, s=0)

6 of 56

Reconstruction Results (Tomasi and Kanade, 1992)

Reprinted from “Factoring Image Sequences into Shape and Motion,” by C. Tomasi and

T. Kanade, Proc. IEEE Workshop on Visual Motion (1991). © 1991 IEEE.

7 of 56

What is some parameters are known?

Weak perspective camera:

Zero skew:

Problem: what is Q ?

(the Euclidean upgrade)

0

Self calibration!

^

8 of 56

Projective case, known intrinsic parameters

Euclidean (perspective) camera:

Take then

If K is known take it to be the identity, then:

where

9 of 56

Euclidean upgrade from projective bundle adjustment.

Mean relative error: 1.2%

10 of 56

Projective case, self calibration (Pollefeys, 1998)

Euclidean (perspective) camera:

Take then

If u0=v0=0 (or equivalently are known):

thus

11 of 56

Multi-view object models

12 of 56

Visual Hulls

[Baumgart, 1974]

13 of 56

Visual Hulls

[Baumgart, 1974]

14 of 56

15 of 56

Experimental setup. Calibration courtesy of Jean-Marc Lavest.

16 of 56

[Lazebnik, Furukawa & Ponce, IJCV’07] [Franco & Boyer, PAMI’09]

17 of 56

3

4

8

12

18 of 56

Multi-view stereo:

PMVS

Feature detection

Matching

Expansion

Filtering

Yasutaka Furukawa and Jean Ponce,

Accurate, Dense and Robust Stereopsis, CVPR’07, PAMI’10

19 of 56

A1

B1

C1

D1

A2

B2

C2

b

c

d

e

f

g

I1

I2

Expansion

20 of 56

A1

B1

C1

D1

A2

B2

C2

b

c

d

e

f

g

I1

I2

b

Expansion

21 of 56

A1

B1

C1

D1

A2

B2

C2

a

b

c

d

e

f

g

I1

I2

b

Expansion

22 of 56

A1

B1

C1

D1

A2

B2

C2

a

b

c

d

e

f

g

I1

I2

p

q

Filtering

23 of 56

24 of 56

A 3D survey of Palazzo Ducale in Venice, Italy

Courtesy of Yves Ubelmann, Iconem – Exhibit at the Grand Palais https://www.grandpalais.fr/fr/evenement/venise-revelee

25 of 56

Outline:

  • Texture
    • Textons
    • Bags of words

  • Segmentation
    • K-means
    • Mean-shift algorithm

26 of 56

Texture Classification

  • Profound observation: Grass and sea pictures don’t look the same!
  • Basic idea: Model the distribution of “texture” over the image (or over a region) and classify in different classes based on the texture models learned from training examples.

Grass

Sea

27 of 56

Image categorization

  • Profound observation: Cows and buildings don’t look the same!
  • Basic idea: Model the distribution of “texture” over the image (or over a region) and classify in different classes based on the texture models learned from training examples.

Cow

Building

28 of 56

The Concept of “Texton”

Multiple training images of the same texture

Multiple training images of the same texture

Filter responses over a bank of filters

Clustering

Texton Dictionary

Question: How do we

perform clustering?

29 of 56

Example of Filter Banks

Isotropic Gabor

Gaussian derivatives at different scales and orientations

‘S’

‘LM’

‘MR8’

30 of 56

Example Textons (LM)

(Linear combinations of filters corresponding to cluster centers)

31 of 56

Example: Visual words in photographs

Images

Word maps

Visual dictionary

Visual word = texton for « objects »

32 of 56

Modeling Texton Distributions

Training

image

Filter Responses

Texton Map

Model = Histogram of textons in the image

33 of 56

Analogy with Text Analysis

Political observers say that the government of Zorgia does not control the political situation. The government will not hold elections …

Analogy:

Text fragment 🡨🡪 Image region

Word 🡨🡪 Texton

Government

Political

Gigabyte

Observers

Election

Memory

Gigahertz

Bus

Word from vocabulary

Frequency of occurrence

« Bag of words »

34 of 56

Analogy with Text Analysis

The ZH-20 unit is a 200Gigahertz processor with 2Gigabyte memory. Its strength is its bus and high-speed memory……

Political

Government

Gigabyte

Observers

Election

Memory

Gigahertz

Bus

Word from vocabulary

Frequency of occurrence

Government

Observers

Histogram from input fragment

Political

Government

Gigabyte

Observers

Election

Memory

Gigahertz

Bus

Frequency of occurrence

Histogram from training “computer” fragments

Political

Gigabyte

Election

Memory

Gigahertz

Bus

Frequency of occurrence

Histogram from training “political” fragments

Compare

35 of 56

Classification

Input Image (or Region of an Input Image)

Model

Compare with Stored Models from Training Images

Models of Plastic

Models of Grass

36 of 56

Example Classification

Input Region

Textons

37 of 56

Examples

38 of 56

  • Sources:
    • J. Winn, A. Criminisi and T. Minka. Object Categorization by Learned Universal Visual Dictionary. Proc. IEEE Intern. Conf. Comp. Vision. 2005. (Also Csurka et al., 2004)
    • M. Varma and A. Zisserman. A statistical approach to texture classification from single images. IJCV, 62(1–2):61–81, April 2005. (Also Lazebnik et al., 2003)

  • Questions:
    • How many textons/words?
    • What filters?
    • How to construct clusters?
    • How to compare histogram distributions?
    • How to exploit the spatial distribution of textons (these examples completely ignore the relative positions of textons in the image)?

  • Will be revisited for object recognition

39 of 56

Segmentation and clustering

40 of 56

From images to objects

What Defines an Object?

    • Cues: color, texture, regions, contours….
    • Subjective problem, but has been well-studied in psychology and computer vision

41 of 56

The goals of segmentation

  • Group together similar-looking pixels for efficiency of further processing
    • “Bottom-up” process
    • Unsupervised

“superpixels”

42 of 56

The goals of segmentation

  • Separate image into coherent “objects”
    • “Bottom-up” or “top-down” process?
    • Supervised or unsupervised?

image

human segmentation

43 of 56

The goals of segmentation

  • Separate image into coherent “objects”
    • “Bottom-up” or “top-down” process?
    • Supervised or unsupervised?

image

human segmentation

44 of 56

The goals of segmentation

  • Separate image into coherent “objects”
    • Segmentation is a SUBJECTIVE task

image

human segmentation

45 of 56

Segmentation as clustering

Source: K. Grauman

46 of 56

Segmentation as clustering

Simplest methods:

Agglomerative clustering (Merge):

  • grouping stuff that belongs together, or
  • iteratively merging the closest clusters

Divisive clustering (Split):

  • split clusters recursively, or
  • iteratively split the cluster that yields the most

diverse cluster

Split and merge

47 of 56

Clustering

How to choose the representative colors?

    • This is a clustering problem!

Objective

    • Each point should be as close as possible to a cluster center
      • Minimize sum squared distance of each point to closest center

R

G

R

G

48 of 56

Solution: Break it down into subproblems

Suppose I tell you the cluster centers ci

    • Q: how to determine which points to associate with each ci?
    • A: for each point p, choose closest ci

Suppose I tell you the points in each cluster

    • Q: how to determine the cluster centers?
    • A: choose ci to be the mean of all points in the cluster

49 of 56

K-means clustering

K-means clustering algorithm

    • Randomly initialize the cluster centers, c1, ..., cK
    • Given cluster centers, determine points in each cluster
      • For each point p, find the closest ci. Put p into cluster i
    • Given points in each cluster, solve for ci
      • Set ci to be the mean of points in cluster i
    • If ci have changed, repeat Step 2

Java demo: http://home.dei.polimi.it/matteucc/Clustering/tutorial_html/AppletKM.html

Properties

    • Will always converge to some solution
    • Can be a “local minimum”
      • does not always find the global minimum of objective function:

50 of 56

Segmentation as clustering

  • K-means clustering based on intensity or color is essentially vector quantization of the image attributes
    • Clusters don’t have to be spatially coherent

Image

Intensity-based clusters

Color-based clusters

51 of 56

Segmentation as clustering

Source: K. Grauman

(But apples and oranges)

52 of 56

Segmentation as clustering

  • Clustering based on (r,g,b,x,y) values enforces more spatial coherence

53 of 56

K-Means for segmentation

  • Pros
    • Very simple method
    • Converges to a local minimum of the error function
  • Cons
    • Memory-intensive
    • Need to pick K
    • Sensitive to initialization
    • Sensitive to outliers
    • Only finds “spherical” �clusters

54 of 56

Histogram-based segmentation

Goal

    • Break the image into K regions (segments)
    • Solve this by reducing the number of colors to K and mapping each pixel to the closest color

55 of 56

Histogram-based segmentation

Goal

    • Break the image into K regions (segments)
    • Solve this by reducing the number of colors to K and mapping each pixel to the closest color

Here’s what it looks like if we use two colors

56 of 56

Finding Modes in a Histogram

How Many Modes Are There? What are they? Which points belong with which modes?

    • Easy to see, hard to compute