1 of 44

Introduction to computer vision 13

Jean Ponce

jean.ponce@ens.fr

Zuhaib Akhtar za2023@nyu.edu

Ayush Jain aj3152@nyu.edu

Slides will be available after classes

2 of 44

Deep learning

  • Representation learning

  • History of neural networks

  • Training

  • CNNs

3 of 44

Image categorization as

representation learning

Image “space”

Feature (Hilbert) space

n

θ

θ

θ

 

4 of 44

Deep learning

Layer 1

Layer 2

Layer n

Linear head

Learned

representation

5 of 44

Traditional Recognition Approach

Hand-designed�feature extraction

Trainable�classifier

Image/ Video

Pixels

Object�Class

6 of 44

What about learning the features?

    • Learn a feature hierarchy all the way from pixels to classifier
    • Each layer extracts features from the output of previous layer
    • Train all layers jointly

Layer 1

Layer 2

Layer 3

Simple �Classifier

Image/ Video

Pixels

7 of 44

“Shallow” vs. “deep” architectures

Hand-designed�feature extraction

Trainable�classifier

Image/ Video

Pixels

Object�Class

Layer 1

Layer N

Simple classifier

Object Class

Image/ Video

Pixels

Traditional recognition: “Shallow” architecture

Deep learning: “Deep” architecture

8 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]

Output:

 

bias

(Source Wikipedia)

9 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers

10 of 44

Inspiration: Neuron cells

11 of 44

Hubel/­Wiesel Architecture

  • D. Hubel and T. Wiesel (1959, 1962, Nobel Prize 1981)
    • Visual cortex consists of a hierarchy of simple, complex, and hyper-complex cells

12 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers
  • Shift-invariant network inspired by Hubel & Wiesel study of visual cortex “Neocognitron” [Fukushima 1980]

13 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers
  • Shift-invariant network inspired by Hubel & Wiesel study of visual cortex “Neocognitron” [Fukushima 1980]
  • Training of multi-layer networks with “backpropagation”

14 of 44

Convolutional Neural Networks (CNN, Convnet)

  • Neural network with specialized connectivity structure
  • Stack multiple stages of feature extractors
  • Higher stages compute more global, more invariant features
  • Classification layer at the end

Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recognition, Proceedings of the IEEE 86(11): 2278–2324, 1998.

15 of 44

Parametric supervised learning

  •  

 

 

(Slides borrowed from S. Lazebnik and M. Trager)

16 of 44

Neural networks

  •  

 

 

 

17 of 44

Feedforward neural network

  •  

 

18 of 44

Nonlinearity

  •  

 

 

19 of 44

Convolutional Neural Networks (CNN, Convnet)

  • Feed-forward feature extraction:
    1. Convolve input with learned filters
    2. Non-linearity
    3. Spatial pooling
    4. Normalization
  • Supervised training of convolutional �filters by back-propagating �classification error

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Normalization

Feature maps

20 of 44

Feedforward neural networks for images

image

Fully connected layer

Source: S. Lazebnik

21 of 44

CNNs: Neural networks for images

image

Convolutional layer

Source: S. Lazebnik

22 of 44

CNNs: Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

23 of 44

CNNs: Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

24 of 44

Convolution as feature extraction

Input

Feature Map

...

Source: S. Lazebnik

25 of 44

Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

26 of 44

Neural networks for images

image

next layer

Convolutional layer

+ ReLU

Source: S. Lazebnik

27 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Input

Feature Map

...

Source: R. Fergus, Y. LeCun

28 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Source: R. Fergus, Y. LeCun

Rectified Linear Unit (ReLU)

29 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Max

Source: R. Fergus, Y. LeCun

30 of 44

Key operations in a CNN

Softmax layer:

Source: S. Lazebnik

 

31 of 44

Loss functions

  •  

 

 

 

32 of 44

Gradient descent

  •  

 

 

33 of 44

Stochastic gradient descent

  • Idea: instead of computing the gradient, compute an approximation

Note that in expectation

 

 

 

34 of 44

GD

SGD

35 of 44

How do we optimize the very nonlinear and

nonconvex empirical risk?

Many possible choices:

  • Second-order (Newton-like) methods

  • Gradient descent

  • Stochastic gradient descent and variants

36 of 44

Léon Bottou on large-scale learning..

37 of 44

Léon Bottou on large-scale learning..

38 of 44

Léon Bottou on large-scale learning..

39 of 44

Léon Bottou on large-scale learning..

40 of 44

Léon Bottou on large-scale learning..

41 of 44

Léon Bottou on large-scale learning..

42 of 44

Léon Bottou on large-scale learning..

43 of 44

Bottom line:

  • “Bad” optimization algorithms make sense in the large-scale setting, especially since we minimize expectations

  • Second-order (Newton-like) methods

  • Gradient descent

  • Stochastic gradient descent and variants

44 of 44

Batch SGD

  •