Lecture 5
Multi-Scale Pyramids
6.8300/6.8301 Advances in Computer Vision
Spring 2024
Vincent Sitzmann, Sara Beery, Mina Konaković Luković, Kaiming He
Detecting a pattern in an image
2
Pattern to detect
Template�matching
Template�matching
Template�matching
Motivation for translation invariance!
4
We need translation and scale invariance
“Training”�example
…
5
Consider all patch location and sizes
No match
No match
Match
“Training”�example
…
Downsample
Idea: Build Multi-Scale Pyramids
6
Template
Idea: Build Multi-Scale Pyramids
7
Template
Multiscale image pyramid
A fast & efficient way to provide scale & translation invariance!
Gaussian Pyramid
8
1/2
1/2
1/2
1/2
Subsampling and aliasing
1/2
1/2
The Gaussian pyramid
10
For each level
1. Blur input image with a Gaussian filter
[1, 4, 6, 4, 1]
[1, 4, 6, 4, 1]
[1
4
6
4
1]
(The Gaussian filter is approximated with a binomial filter)
The Gaussian pyramid
11
For each level
1. Blur input image with a Gaussian filter
2. Downsample image
[1, 4, 6, 4, 1]
[1
4
6
4
1]
The Gaussian pyramid
12
256×256
128×128
64×64
32×32
2
2
2
The Gaussian pyramid
13
512×512
256×256
128×128
64×64
32×32
(original image)
Convolution as matrix multiplication
In the 1D case, it helps to make explicit the structure of the matrix:
O
=
In the 1D case, it helps to make explicit the structure of the matrix:
/3
/3
/3
1/3 | 1/3 | 1/3 | | | | | | | |
| 1/3 | 1/3 | 1/3 | | | | | | |
| | | | | | | | | |
| | | | | | | | | |
| | | | | | | | | |
| | | | | | | | | |
| | | | | | | | | |
| | | | | | | | | |
90
90
90
60
60
30
30
0
0 |
0 |
0 |
90 |
90 |
90 |
90 |
90 |
0 |
0 |
0 |
30 |
60 |
90 |
90 |
90 |
60 |
30 |
Reminder from Lecture 3
The Gaussian pyramid
[1, 4, 6, 4, 1]
[1, 4, 6, 4, 1]
(The arrays shown here are for 1D signals)
The Gaussian pyramid
[1, 4, 6, 4, 1]
[1, 4, 6, 4, 1]
The Gaussian pyramid
[1, 4, 6, 4, 1]
[1, 4, 6, 4, 1]
=
The Gaussian pyramid
18
For each level
1. Blur input image with a Gaussian filter
2. Downsample image
What about the opposite of blurring?
19
Gaussian filter
-
+
-
Laplacian filter
The Laplacian Pyramid
Compute the difference between upsampled Gaussian pyramid level k+1 and Gaussian pyramid level k.
20
+
-
The Laplacian Pyramid
21
Gaussian pyramid
The Laplacian Pyramid
22
Gaussian pyramid
Laplacian pyramid
The Laplacian Pyramid
23
Blurring and downsampling:
Upsampling and blurring:
(blur)
(Downsampling by 2)
Upsampling
24
=
Insert zeros
64x64
128x128
128x128
Upsampling
25
1 | 0 | 1 | 0 | 1 |
0 | 0 | 0 | 0 | 0 |
1 | 0 | 1 | 0 | 1 |
0 | 0 | 0 | 0 | 0 |
1 | 0 | 1 | 0 | 1 |
= ?
Upsampling
26
1 | 0 | 1 | 0 | 1 |
0 | 0 | 0 | 0 | 0 |
1 | 0 | 1 | 0 | 1 |
0 | 0 | 0 | 0 | 0 |
1 | 0 | 1 | 0 | 1 |
=
1 | 1 | 1 | 1 | 1 |
1 | 1 | 1 | 1 | 1 |
1 | 1 | 1 | 1 | 1 |
1 | 1 | 1 | 1 | 1 |
1 | 1 | 1 | 1 | 1 |
The Laplacian Pyramid
27
Blurring and downsampling:
Upsampling and blurring:
(blur)
(Downsampling by 2)
(Upsampling by 2)
(blur)
The Laplacian Pyramid
28
The Laplacian Pyramid
29
Gaussian pyramid
Laplacian pyramid
=
The Laplacian Pyramid
30
Laplacian pyramid
The Laplacian Pyramid
31
Gaussian pyramid
Laplacian pyramid
The Laplacian Pyramid
32
Gaussian pyramid
Laplacian pyramid
Analysis/Encoder
Synthesis/Decoder
The Laplacian Pyramid
33
Laplacian pyramid
Gaussian residual
Laplacian pyramid
Gaussian residual
Laplacian pyramid applications
34
Image Blending
35
Image Blending
36
Image Blending
37
Image Blending with the Laplacian Pyramid
38
Image Blending with the Laplacian Pyramid
39
Image Blending with the Laplacian Pyramid
40
http://persci.mit.edu/pub_pdfs/pyramid83.pdf
Image pyramids
41
Gaussian Pyr
Laplacian Pyr
And many more: QMF, steerable, …
Convnets!
Orientations
42
Orientations
43
Steerable Pyramid
44
Low pass residual
Steerable Pyramid
45
=
Oriented filters
Blur
Downsampling
…
…
Steerable Pyramid
3 scales, 4 orientations
Steerable Pyramid
47
Analysis/Encoder
Synthesis/Decoder
Visual summary: Linear Image Transforms
48
1D
1D
1D
2D
Fourier Transform
Gaussian Pyramid
Laplacian Pyramid
Steerable Pyramid
Steerable pyramid applications
49
Dall-E 2
https://openai.com/dall-e-2/
An astronaut riding a horse�in a photorealistic style
Making textures
Textures
Stationary
Stochastic
Pre-attentive texture discrimination
Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.
Pre-attentive texture discrimination
Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.
Pre-attentive texture discrimination
Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.
Pre-attentive texture discrimination
Bela Julesz, "Textons, the Elements of Texture Perception, and their Interactions". Nature 290: 91-97. March, 1981.
This texture pair is pre-attentively indistinguishable. Why?
Julesz - Textons
Textons: fundamental texture elements.
Textons might be represented by features such as terminators, corners, and intersections within the patterns…
“We note here that simpler, lower-level mechanisms tuned for size may be �sufficient to explain this discrimination.”
Observation: the Xs�look smaller than the Ls.
Ls 25% larger �contrast adjusted to keep mean constant
Ls 25% shorter
SIGGRAPH 1995
https://www.cns.nyu.edu/heegerlab/content/publications/Heeger-siggraph95.pdf
Steerable Pyramid
62
Analysis/Encoder
Synthesis/Decoder
Representation
Can we use this representation to generate more samples from the same texture?
Steerable Pyramid
63
Analysis/Encoder
Representation
Representation
Representation =
Histograms of pixel values�for each subband
Transferring Image Statistics via Representation Matching
64
Analysis/Encoder
Representation
Analysis/Encoder
Synthesis/Decoder
Two main tools:
1- steerable pyramid
2- matching histograms
Transferring Image Statistics via Representation Matching
Input texture
Steerable pyr
1-The steerable pyramid
Overview of the algorithm
Two main tools:
1- steerable pyramid
2- matching histograms
2-Matching histograms
Cumulative distribution
9% of pixels have an intensity value
within the range[0.37, 0.41]
75% of pixels have an intensity value
smaller than 0.5
5% of pixels have an intensity value
within the range[0.37, 0.41]
Histogram
2-Matching histograms
?
Z[n,m]
Y[n,m]
We look for a transformation
of the image Y
Y’ = f (Y)
Such that
Hist(Y) = Hist(f(Z))
There are infinitely many functions
that can do this transformation.
A natural choice is to use f being:
2-Matching histograms
Y’ = f (Y)
The function f is just a look up table: it says, change all the pixels of value Y into a value f(Y).
Y’= 0.5
New
intensity
Y= 0.8
Original
intensity
Y[n,m]
Target histogram
2-Matching histograms
Y’ = f (Y)
Another example: Matching histograms
10% of pixels are black
and 90% are white
5% of pixels have an intensity value
within the range[0.37, 0.41]
?
?
Another example: Matching histograms
Y= 0.8
Y’ = f (Y)
The function f is just a look up table: it says, change all the pixels of
value Y into a value f(Y).
Original
intensity
Y(x,y)
Y’= 1
New
intensity
Another example: Matching histograms
Y’ = f (Y)
In this example, f is a step function.
Cumulative distribution
Matching histograms of a subband
Matching histograms of a subband
Y’ = f (Y)
Texture analysis
Input texture
(histogram)
Steerable pyramid
(histogram)
The texture is represented as a collection of
marginal histograms.
Texture synthesis
Input texture
(histogram)
Heeger and Bergen, 1995
(histogram)
Why does it work? (sort of)
Why does it work? (sort of)
The black and white
blocks appear by
thresholding (f) a
blobby image
Iteration 0
…
Filter bank
Why does it work? (sort of)
The black and white blocks appear by
thresholding (f) a blobby image
Why does it work? (sort of)
After 6 iterations
Histograms match ok
red = target histogram, blue = current iteration
Color textures
R
G
B
Three textures
Color textures
R
G
B
Color textures
R
G
B
This does not work
Color textures
Problem: we create new colors not present in the original image.
Why? Color channels are not independent.
R
G
B
PCA and decorrelation
R
G
B
R
G
In the original image, R and G are correlated, but, after synthesis,…
R
G
PCA and decorrelation
R
G
The texture synthesis algorithm assumes that the channels
are independent.
What we want to do is some rotation
See that in this rotated space,
if I specify one coordinate the
other remains unconstrained.
U1
U2
Rotation
PCA and decorrelation
R
G
1.0000 0.9303 0.6034
0.9303 0.9438 0.6620
0.6034 0.6620 0.5569
C =
correlation(R,G)
C = D D’
0.6347 0.6072 0.4779
0.6306 -0.0496 -0.7745
0.4466 -0.7930 0.4144
D =
=
3 x Npixels
3 x Npixels
3 x 3
D’
R
G
B
U1
U2
U3
PCA finds the principal directions of variation of the data.
It gives a decomposition of the covariance matrix as:
By transforming the original data (RGB) using D we get:
The new components (U1,U2,U3) are decorrelated.
Color textures
R
G
B
Rotation
Matrix
(3x3)
These three textures
look similar
(high dependency)
These three textures
Look less similar
(lower dependency)
D’
Color textures
Inverse
Rotation
Matrix
R
G
B
D
Color textures
Inverse
Rotation
R
G
B
D
Color channels
Without PCA
With PCA
Color channels
Color channels
Examples from the paper
Heeger and Bergen, 1995
Examples from the paper
Examples not from the paper
Input
texture
Synthetic
texture
But, does it really work even when it seems to work?
The main idea: it works by ‘kind of’ projecting a random image into the set of equivalent textures
Space of all images
Set of equivalent textures
Set of perceptually
equivalent textures
But, does it really work???�How to measure how well the representation constraints the set of equivalent textures?
?
All the textures in this
set have the same
parameters.
?
?
?
?
How to identify the set of equivalent textures?
?
This does not reveal how poor
the representation actually is.
How to identify the set of equivalent textures?
These trajectories are
more perceptually
salient
This set is huge
Portilla and Simoncelli
Portilla and Simoncelli
Portilla & Simoncelli
Heeger & Bergen
Portilla & Simoncelli
How to identify the set of equivalent textures?
Now they look good, but maybe
they look too good…
Portilla & Simoncelli