1 of 106

Lecture 5

Multi-Scale Pyramids

6.8300/6.8301 Advances in Computer Vision

Spring 2024

Vincent Sitzmann, Sara Beery, Mina Konaković Luković, Kaiming He

2 of 106

Detecting a pattern in an image

2

Pattern to detect

3 of 106

Template�matching

Template�matching

Template�matching

Motivation for translation invariance!

4 of 106

4

We need translation and scale invariance

“Training”�example

5 of 106

5

Consider all patch location and sizes

No match

No match

Match

“Training”�example

Downsample

6 of 106

Idea: Build Multi-Scale Pyramids

6

Template

7 of 106

Idea: Build Multi-Scale Pyramids

7

Template

Multiscale image pyramid

A fast & efficient way to provide scale & translation invariance!

8 of 106

Gaussian Pyramid

8

1/2

1/2

1/2

1/2

9 of 106

Subsampling and aliasing

1/2

1/2

10 of 106

The Gaussian pyramid

10

For each level

1. Blur input image with a Gaussian filter

[1, 4, 6, 4, 1]

[1, 4, 6, 4, 1]

[1

4

6

4

1]

(The Gaussian filter is approximated with a binomial filter)

11 of 106

The Gaussian pyramid

11

For each level

1. Blur input image with a Gaussian filter

2. Downsample image

[1, 4, 6, 4, 1]

[1

4

6

4

1]

12 of 106

The Gaussian pyramid

12

256×256

128×128

64×64

32×32

2

2

2

13 of 106

The Gaussian pyramid

13

512×512

256×256

128×128

64×64

32×32

(original image)

14 of 106

Convolution as matrix multiplication

In the 1D case, it helps to make explicit the structure of the matrix:

O

=

In the 1D case, it helps to make explicit the structure of the matrix:

/3

/3

/3

1/3

1/3

1/3

1/3

1/3

1/3

90

90

90

60

60

30

30

0

0

0

0

90

90

90

90

90

0

0

0

30

60

90

90

90

60

30

Reminder from Lecture 3

15 of 106

The Gaussian pyramid

[1, 4, 6, 4, 1]

[1, 4, 6, 4, 1]

(The arrays shown here are for 1D signals)

16 of 106

The Gaussian pyramid

[1, 4, 6, 4, 1]

[1, 4, 6, 4, 1]

17 of 106

The Gaussian pyramid

[1, 4, 6, 4, 1]

[1, 4, 6, 4, 1]

=

18 of 106

The Gaussian pyramid

18

For each level

1. Blur input image with a Gaussian filter

2. Downsample image

19 of 106

What about the opposite of blurring?

19

Gaussian filter

-

+

-

Laplacian filter

20 of 106

The Laplacian Pyramid

Compute the difference between upsampled Gaussian pyramid level k+1 and Gaussian pyramid level k.

20

+

-

21 of 106

The Laplacian Pyramid

21

Gaussian pyramid

22 of 106

The Laplacian Pyramid

22

Gaussian pyramid

Laplacian pyramid

23 of 106

The Laplacian Pyramid

23

Blurring and downsampling:

Upsampling and blurring:

(blur)

(Downsampling by 2)

24 of 106

Upsampling

24

=

Insert zeros

64x64

128x128

128x128

25 of 106

Upsampling

25

1

0

1

0

1

0

0

0

0

0

1

0

1

0

1

0

0

0

0

0

1

0

1

0

1

= ?

26 of 106

Upsampling

26

1

0

1

0

1

0

0

0

0

0

1

0

1

0

1

0

0

0

0

0

1

0

1

0

1

=

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

1

27 of 106

The Laplacian Pyramid

27

Blurring and downsampling:

Upsampling and blurring:

(blur)

(Downsampling by 2)

(Upsampling by 2)

(blur)

28 of 106

The Laplacian Pyramid

28

29 of 106

The Laplacian Pyramid

29

Gaussian pyramid

Laplacian pyramid

=

30 of 106

The Laplacian Pyramid

30

Laplacian pyramid

31 of 106

The Laplacian Pyramid

31

Gaussian pyramid

Laplacian pyramid

32 of 106

The Laplacian Pyramid

32

Gaussian pyramid

Laplacian pyramid

Analysis/Encoder

Synthesis/Decoder

33 of 106

The Laplacian Pyramid

33

Laplacian pyramid

Gaussian residual

Laplacian pyramid

Gaussian residual

34 of 106

Laplacian pyramid applications

  • Image compression
  • Noise removal
  • Computing image features (e.g., SIFT)
  • Connections to common neural network architectures: ResNets, U-Nets, skip connections in transformers…

34

35 of 106

Image Blending

35

36 of 106

Image Blending

36

37 of 106

Image Blending

37

38 of 106

Image Blending with the Laplacian Pyramid

38

39 of 106

Image Blending with the Laplacian Pyramid

39

40 of 106

Image Blending with the Laplacian Pyramid

  • Build Laplacian pyramid for both images: LA, LB
  • Build Gaussian pyramid for mask: G
  • Build a combined Laplacian pyramid:
  • Collapse L to obtain the blended image

40

http://persci.mit.edu/pub_pdfs/pyramid83.pdf

41 of 106

Image pyramids

41

Gaussian Pyr

Laplacian Pyr

And many more: QMF, steerable, …

Convnets!

42 of 106

Orientations

42

43 of 106

Orientations

43

44 of 106

Steerable Pyramid

44

Low pass residual

45 of 106

Steerable Pyramid

45

=

Oriented filters

Blur

Downsampling

46 of 106

Steerable Pyramid

3 scales, 4 orientations

47 of 106

Steerable Pyramid

47

Analysis/Encoder

Synthesis/Decoder

48 of 106

Visual summary: Linear Image Transforms

48

1D

1D

1D

2D

Fourier Transform

Gaussian Pyramid

Laplacian Pyramid

Steerable Pyramid

49 of 106

Steerable pyramid applications

  • Texture recognition
  • Image compression
  • Noise removal
  • Computing image features (e.g., SIFT)
  • Object recognition

49

50 of 106

Dall-E 2

https://openai.com/dall-e-2/

An astronaut riding a horse�in a photorealistic style

51 of 106

Making textures

52 of 106

Textures

Stationary

Stochastic

53 of 106

54 of 106

Pre-attentive texture discrimination

Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.

55 of 106

Pre-attentive texture discrimination

Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.

56 of 106

Pre-attentive texture discrimination

Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.

57 of 106

Pre-attentive texture discrimination

Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.

This texture pair is pre-attentively indistinguishable. Why?

58 of 106

Julesz - Textons

Textons: fundamental texture elements.

Textons might be represented by features such as terminators, corners, and intersections within the patterns…

59 of 106

“We note here that simpler, lower-level mechanisms tuned for size may be �sufficient to explain this discrimination.”

Observation: the Xs�look smaller than the Ls.

60 of 106

Ls 25% larger �contrast adjusted to keep mean constant

Ls 25% shorter

61 of 106

SIGGRAPH 1995

https://www.cns.nyu.edu/heegerlab/content/publications/Heeger-siggraph95.pdf

62 of 106

Steerable Pyramid

62

Analysis/Encoder

Synthesis/Decoder

Representation

Can we use this representation to generate more samples from the same texture?

63 of 106

Steerable Pyramid

63

Analysis/Encoder

Representation

Representation

Representation =

Histograms of pixel values�for each subband

64 of 106

Transferring Image Statistics via Representation Matching

64

Analysis/Encoder

Representation

Analysis/Encoder

Synthesis/Decoder

65 of 106

Two main tools:

1- steerable pyramid

2- matching histograms

Transferring Image Statistics via Representation Matching

66 of 106

Input texture

Steerable pyr

1-The steerable pyramid

67 of 106

Overview of the algorithm

Two main tools:

1- steerable pyramid

2- matching histograms

68 of 106

2-Matching histograms

Cumulative distribution

9% of pixels have an intensity value

within the range[0.37, 0.41]

75% of pixels have an intensity value

smaller than 0.5

5% of pixels have an intensity value

within the range[0.37, 0.41]

Histogram

69 of 106

2-Matching histograms

?

Z[n,m]

Y[n,m]

We look for a transformation

of the image Y

Y’ = f (Y)

Such that

Hist(Y) = Hist(f(Z))

There are infinitely many functions

that can do this transformation.

A natural choice is to use f being:

  • point-wise non linearity
  • stationary
  • monotonic (most of the time invertible)

70 of 106

2-Matching histograms

Y’ = f (Y)

The function f is just a look up table: it says, change all the pixels of value Y into a value f(Y).

Y’= 0.5

New

intensity

Y= 0.8

Original

intensity

Y[n,m]

Target histogram

71 of 106

2-Matching histograms

Y’ = f (Y)

72 of 106

Another example: Matching histograms

10% of pixels are black

and 90% are white

5% of pixels have an intensity value

within the range[0.37, 0.41]

?

?

73 of 106

Another example: Matching histograms

Y= 0.8

Y’ = f (Y)

The function f is just a look up table: it says, change all the pixels of

value Y into a value f(Y).

Original

intensity

Y(x,y)

Y’= 1

New

intensity

74 of 106

Another example: Matching histograms

Y’ = f (Y)

In this example, f is a step function.

Cumulative distribution

75 of 106

Matching histograms of a subband

76 of 106

Matching histograms of a subband

Y’ = f (Y)

77 of 106

Texture analysis

Input texture

(histogram)

Steerable pyramid

(histogram)

The texture is represented as a collection of

marginal histograms.

78 of 106

Texture synthesis

Input texture

(histogram)

Heeger and Bergen, 1995

(histogram)

79 of 106

Why does it work? (sort of)

80 of 106

Why does it work? (sort of)

The black and white

blocks appear by

thresholding (f) a

blobby image

Iteration 0

Filter bank

81 of 106

Why does it work? (sort of)

The black and white blocks appear by

thresholding (f) a blobby image

82 of 106

Why does it work? (sort of)

After 6 iterations

Histograms match ok

red = target histogram, blue = current iteration

83 of 106

Color textures

R

G

B

Three textures

84 of 106

Color textures

R

G

B

85 of 106

Color textures

R

G

B

This does not work

86 of 106

Color textures

Problem: we create new colors not present in the original image.

Why? Color channels are not independent.

R

G

B

87 of 106

PCA and decorrelation

R

G

B

R

G

In the original image, R and G are correlated, but, after synthesis,…

R

G

88 of 106

PCA and decorrelation

R

G

The texture synthesis algorithm assumes that the channels

are independent.

What we want to do is some rotation

See that in this rotated space,

if I specify one coordinate the

other remains unconstrained.

U1

U2

Rotation

89 of 106

PCA and decorrelation

R

G

1.0000 0.9303 0.6034

0.9303 0.9438 0.6620

0.6034 0.6620 0.5569

C =

correlation(R,G)

C = D D’

0.6347 0.6072 0.4779

0.6306 -0.0496 -0.7745

0.4466 -0.7930 0.4144

D =

=

3 x Npixels

3 x Npixels

3 x 3

D’

R

G

B

U1

U2

U3

PCA finds the principal directions of variation of the data.

It gives a decomposition of the covariance matrix as:

By transforming the original data (RGB) using D we get:

The new components (U1,U2,U3) are decorrelated.

90 of 106

Color textures

R

G

B

Rotation

Matrix

(3x3)

These three textures

look similar

(high dependency)

These three textures

Look less similar

(lower dependency)

D’

91 of 106

Color textures

Inverse

Rotation

Matrix

R

G

B

D

92 of 106

Color textures

Inverse

Rotation

R

G

B

D

93 of 106

Color channels

Without PCA

With PCA

94 of 106

Color channels

95 of 106

Color channels

96 of 106

Examples from the paper

Heeger and Bergen, 1995

97 of 106

Examples from the paper

98 of 106

Examples not from the paper

Input

texture

Synthetic

texture

But, does it really work even when it seems to work?

99 of 106

The main idea: it works by ‘kind of’ projecting a random image into the set of equivalent textures

Space of all images

Set of equivalent textures

Set of perceptually

equivalent textures

100 of 106

But, does it really work???�How to measure how well the representation constraints the set of equivalent textures?

?

All the textures in this

set have the same

parameters.

?

?

?

?

101 of 106

How to identify the set of equivalent textures?

?

This does not reveal how poor

the representation actually is.

102 of 106

How to identify the set of equivalent textures?

These trajectories are

more perceptually

salient

This set is huge

103 of 106

Portilla and Simoncelli

  • Parametric representation, based on Gaussian scale mixture prior model for images.
  • About 1000 numbers to describe a texture.
  • Ok results; maybe as good as DeBonet.

104 of 106

Portilla and Simoncelli

105 of 106

Portilla & Simoncelli

Heeger & Bergen

Portilla & Simoncelli

106 of 106

How to identify the set of equivalent textures?

Now they look good, but maybe

they look too good…

Portilla & Simoncelli