1 of 100

From Pixels to Percepts

CS280: Computer Vision

  1. Efros, UC Berkeley, Spring 2026

2 of 100

From Images to Entities

3 of 100

"I stand at the window and see a house, trees, sky. Theoretically I might say there were 327 brightnesses and nuances of colour. Do I have "327"? No. I have sky, house, and trees.”

— Max Wertheimer, 1923

4 of 100

From Pixels to Perception

Tiger

Grass

Water

Sand

outdoor

wildlife

Tiger

tail

eye

legs

head

back

shadow

mouth

5 of 100

Need to handle complexity!

6 of 100

Ways of Reducing Complexity

Representation

(e.g. texture, blur, small scale)

Raw

Image

pixels

Segmentation

(partition the input)

Categorization

(partition the world)

7 of 100

Two aspects of object recognition

  • Shape
  • Texture

8 of 100

Attneave’s Cat (1954)�Line drawings convey most of the information

8

9 of 100

Modeling shape variation in a category

  • D’Arcy Thompson: On Growth and Form, 1917
    • studied transformations between shapes of organisms

10 of 100

Matching� Example

model

target

11 of 100

Example Natural Materials

11

Terrycloth

Rough Plastic

Plaster-b

Sponge

Rug-a

Painted Spheres

Columbia-Utrecht Database (http://www.cs.columbia.edu/CAVE)

12 of 100

Texture Recognition

12

Felt?

Polyester?

Terrycloth?

Rough Plaster?

Leather?

Plaster?

Concrete?

Crumpled Paper?

Sponge?

Limestone?

Brick?

?

?

13 of 100

When are two textures similar?

14 of 100

14

15 of 100

Preattentive vs Attentive Vision (Julesz)

Human vision operates in two distinct modes:

1. Preattentive vision

parallel, instantaneous (~100--200ms), without scrutiny,

independent of the number of patterns, covering a large visual field.

2. Attentive vision

serial search by focal attention in 50ms steps limited to small aperture.

16 of 100

Evidence for Pre-attentive Recognition

  • On a task of judging animal vs no animal, humans can make mostly correct saccades in 150 ms (Kirchner & Thorpe, 2006)

    • Comparable to synaptic delay in the retina, LGN, V1, V2, V4, IT pathway.
    • Doesn’t rule out feed back but shows feed forward only is very powerful
  • Detection and categorization are practically simultaneous (Grill-Spector & Kanwisher, 2005)

17 of 100

Object

Bag of ‘words’

18 of 100

19 of 100

Clustered Image Patches (“Bag of Visual Words”)

Fei-Fei et al. 2005

20 of 100

Image representation

…..

frequency

codewords

21 of 100

Scene Classification (Renninger & Malik)

kitchen

livingroom

bedroom

bathroom

city

street

farm

beach

mountain

forest

Vision Science &

Computer Vision Groups

University of California Berkeley

22 of 100

Texton Histogram Matching

Vision Science &

Computer Vision Groups

University of California Berkeley

23 of 100

Discrimination of Basic Categories

texture model

Vision Science &

Computer Vision Groups

University of California Berkeley

24 of 100

Discrimination of Basic Categories

texture model

chance

Vision Science &

Computer Vision Groups

University of California Berkeley

25 of 100

Discrimination of Basic Categories

texture model

chance

37 ms

Vision Science &

Computer Vision Groups

University of California Berkeley

26 of 100

Discrimination of Basic Categories

texture model

chance

50 ms

Vision Science &

Computer Vision Groups

University of California Berkeley

27 of 100

Discrimination of Basic Categories

texture model

chance

69 ms

Vision Science &

Computer Vision Groups

University of California Berkeley

28 of 100

Discrimination of Basic Categories

texture model

chance

37 ms

50 ms

69 ms

Vision Science &

Computer Vision Groups

University of California Berkeley

29 of 100

When is Object Recognition �Just Texture Recognition?

30 of 100

31 of 100

“Collie”

image X

label Y

Convolutional Neural Network

When Object Recognition

is just Texture Recognition

32 of 100

“Collie”

image X

label Y

Convolutional Neural Network

When Object Recognition

is just Texture Recognition

33 of 100

From Images to Entities?

Is it even necessary?

34 of 100

Spatial Support

  • Spatial Support
    • Which pixels to include?
  • Similarity Metric
    • Which statistics to compute?

Surprising Result: second will get easier, if we make progress on the first

model

35 of 100

Does Spatial Support Matter?

Classify

Ground-Truth Segment

Bounding Box

Classify

vs.

36 of 100

Why segmentation/grouping is useful?

Separate image into objects

    • Reduction in complexity, tokenization, compression, reasoning

Improved spatial support for object detection

    • Super-pixels, self-attention, etc.

Help occlusion reasoning

    • Some segment boundaries encode occlusion

Help reason about object shape

    • Shape is largely missing in texture processing

37 of 100

From Images to Entities

Some History

38 of 100

�Chickening out: �“Semantic Segmentation”

Input

Label

Input

Label

Each pixel has label, inc. background, and unknown

Usually visualized by colors.

Note: don’t distinguish between object instances

Image Credit: Everingham et al. Pascal VOC 2012.

Slide by David Fouhey

39 of 100

Structuralism

© Stephen E. Palmer, 2002

Structuralism:

Perception results from the association

of basic sensory atoms in memory via

repeated, prior joint occurrences.

Derived from philosophy of

British Empiricists (e.g., Locke,

Berkeley, Hume, and Mills).

Proposed by Wilhelm Wundt,

the father of modern Psychology.

40 of 100

Structuralism

© Stephen E. Palmer, 2002

Sensory Atoms

Retinal mosaic

Greenness

at (x3,y3)

Yellowness

at (x2,y2)

Redness

at (x1,y1)

41 of 100

Structuralism

© Stephen E. Palmer, 2002

Perceptual Complexes

Retinal mosaic

42 of 100

Structuralism

© Stephen E. Palmer, 2002

Perceptual Complexes

Retinal mosaic

Red apple

at (x0,y0)

43 of 100

Structuralism

© Stephen E. Palmer, 2002

Perceptual Complexes

Retinal mosaic

Red apple

at (x0,y0)

Association

44 of 100

Structuralism

© Stephen E. Palmer, 2002

Chemical Analogy

Perceptions are made of basic sensory experiences

just as molecules are made of basic atoms.

45 of 100

Gestaltism

© Stephen E. Palmer, 2002

Gestaltism:

Perception results from the interaction

between the intrinsic structure of the stimulus

and the intrinsic structure of the brain.

Max

Wertheimer

Wolfgang

Köhler

Kurt

Koffka

46 of 100

Gestaltism

© Stephen E. Palmer, 2002

The Gestalt movement in perceptual theory

was primarily a reaction against Structuralism:

Successful in arguing against Structuralism,

but less successful in promoting its own

theoretical agenda.

Rejected atomism

Rejected empiricism

Rejected associationism

47 of 100

Gestaltism

© Stephen E. Palmer, 2002

Holism: The whole is different from the sum of its parts.

Emergent properties:

Features of a configuration

that are not features of

its components, e.g.:

  • length
  • orientation
  • curvature
  • closure
  • connectedness

48 of 100

“The whole is different

from its parts”

-- Kurt Koffka

49 of 100

© Stephen E. Palmer, 2002

Wertheimer’s “laws” of grouping

Rows

Perceptual Grouping

Columns

50 of 100

© Stephen E. Palmer, 2002

14.20

Proximity

Rows

Columns

Perceptual Grouping

51 of 100

© Stephen E. Palmer, 2002

14.21

Color Similarity

Rows

Columns

Perceptual Grouping

52 of 100

© Stephen E. Palmer, 2002

14.22

Size Similarity

Rows

Columns

Perceptual Grouping

53 of 100

© Stephen E. Palmer, 2002

14.23

Orientation Similarity

Rows

Columns

Perceptual Grouping

54 of 100

© Stephen E. Palmer, 2002

14.24

Similarity of texture

Rows

Perceptual Grouping

Columns

55 of 100

© Stephen E. Palmer, 2002

14.26

Common Fate

Columns

Perceptual Grouping

56 of 100

Common fate

Image credit: Arthus-Bertrand (via F. Durand)

57 of 100

© Stephen E. Palmer, 2002

14.28

Closure

Columns

Rows

Perceptual Grouping

58 of 100

© Stephen E. Palmer, 2002

14.29

Common Region

Rows

Perceptual Grouping

Columns

59 of 100

© Stephen E. Palmer, 2002

14.30

Element Connectedness

Rows

Perceptual Grouping

Columns

60 of 100

© Stephen E. Palmer, 2002

14.27

Good Continuation

Columns

Rows

Perceptual Grouping

61 of 100

© Stephen E. Palmer, 2002

14.32

Past Experience

Perceptual Grouping

62 of 100

63 of 100

What’s wrong with classical clustering algorithms?

e.g. K-means or EM?

64 of 100

Similarity and dissimilarity

  • Which colors are similar?

  • Which ones are dissimilar?

64

65 of 100

Similarity and dissimilarity

  • Similarity:
  • comparable properties:
    • brightness, color, texture, motion

  • Dissimilarity:
  • changes in properties
    • edges, motion discontinuities

  • May want to both group and separate

Image Segmentation

65

3/3/2003

66 of 100

Graph-based segmentation

  • Node = pixel
  • Edge = pair of neighboring pixels
  • Edge weight = similarity or dissimilarity of the respective nodes

wij

i

j

Source: S. Seitz

67 of 100

  • Idea:
    • Iteratively merge subgraphs (starting with initially nodes) if the edge(s) between them are “weak” enough.
    • Question: How to evaluate the strength of edges between subgraphs

Ci

Cj

 

Don’t merge if:

 

68 of 100

σ = 0.8, k = 300

 

 

69 of 100

Mean shift clustering and segmentation

  • An advanced and versatile technique for clustering-based segmentation

70 of 100

Mean shift algorithm

  • The mean shift algorithm seeks modes or local maxima of density in the feature space

image

Feature space

(L*u*v* color values)

71 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

72 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

73 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

74 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

75 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

76 of 100

Mean shift

Search�window

Center of

mass

Mean Shift

vector

Slide by Y. Ukrainitz & B. Sarel

77 of 100

Mean shift

Search�window

Center of

mass

Slide by Y. Ukrainitz & B. Sarel

78 of 100

Mean shift clustering

  • Cluster: all data points in the attraction basin of a mode
  • Attraction basin: the region for which all trajectories lead to the same mode

Slide by Y. Ukrainitz & B. Sarel

79 of 100

Mean shift clustering/segmentation

  • Find features (color, gradients, texture, etc)
  • Initialize windows at individual feature points
  • Perform mean shift for each window until convergence
  • Merge windows that end up near the same “peak” or mode

80 of 100

Mean shift segmentation results

http://www.caip.rutgers.edu/~comanici/MSPAMI/msPamiResults.html

81 of 100

Mean shift segmentation results

82 of 100

Segmentation: Spectral Graph Techniques

83 of 100

Exploiting global constraints:�Image Segmentation as Graph Partitioning

83

Build a weighted graph G=(V,E) from image

V: image pixels

E: connections between pairs of nearby pixels

Partition graph so that similarity within group is large and similarity between groups is small -- Normalized Cuts [Shi & Malik 97]

84 of 100

Wij small when intervening contour strong, small when weak..� �Cij = max Pb(x,y) for (x,y) on line segment ij; Wij = exp ( - Cij / σ

84

85 of 100

How to partition a graph

  • We can find the minimum cut efficiently, but this tends to break the graph into isolated little pieces

85

86 of 100

Normalized Cut is a better measure ..

  • We normalize by the total volume of connections

86

87 of 100

Solving the Normalized Cut problem

  • Exact discrete solution to Ncut is NP-hard even on regular grid [Papadimitriou’97]
  • We first transform to

  • Drawing on spectral graph theory, good approximation can be obtained by solving a generalized eigenvalue problem.

87

88 of 100

Normalized Cuts as a Spring-Mass system

  • Each pixel is a point mass; each connection is a spring:

  • Fundamental modes are generalized eigenvectors of

(D - W) y = λDy

88

89 of 100

Eigenvectors carry contour information

89

90 of 100

Temporal NCuts [Shi & Malik, 98]

91 of 100

in (Marr’s) Theory

Input Image

Boundaries

Segmentation

Recognition

?

in Practice

Input Image

Edges

Segmentation

Recognition

Person

Car#1

Car#2

Road

...

92 of 100

What is a “good” segmentation??

93 of 100

Compare to human segmentation or to “ground truth”

  • http://www.eecs.berkeley.edu/Research/Projects/CS/vision/grouping/resources.html

No objective definition of segmentation!

Subject 1

Subject 2

Subject 3

94 of 100

No objective definition of segmentation!�

  • http://www.eecs.berkeley.edu/Research/Projects/CS/vision/bsds/BSDS300/html/dataset/images/color/317080.html

95 of 100

Evaluation: Boundary agreement

True boundary

Detected boundary

Correct if

D < T

Precision = % of detected boundary

pixels that are correct

Recall = % of boundary pixels that are detected

Varying T

96 of 100

Evaluation: Region overlap with ground truth

97 of 100

Evaluation: Region overlap with ground truth

Graph-based

Mean shift

Spectral

Ground truth

98 of 100

Results: Berkeley Segmentation Engine

99 of 100

Segmentation is not an aim in itself �– it’s a result of image understanding!

input image point process curve process

a color region texture regions objects

100 of 100

Superpixels