1 of 44

Introduction to computer vision 13

Jean Ponce

jean.ponce@ens.fr

​

Zuhaib Akhtar za2023@nyu.edu

Ayush Jain aj3152@nyu.edu

​

​

​

Slides will be available after classes

2 of 44

Deep learning

​

  • Representation learning

​

  • History of neural networks

​

  • Training

​

  • CNNs

​

3 of 44

Image categorization as

representation learning

Image “space”

Feature (Hilbert) space

n

θ

θ

θ

 

4 of 44

Deep learning

Layer 1

Layer 2

Layer n

Linear head

Learned

representation

5 of 44

Traditional Recognition Approach

Hand-designed�feature extraction

Trainable�classifier

Image/ Video

Pixels

Object�Class

6 of 44

What about learning the features?

    • Learn a feature hierarchy all the way from pixels to classifier
    • Each layer extracts features from the output of previous layer
    • Train all layers jointly

Layer 1

Layer 2

Layer 3

Simple �Classifier

Image/ Video

Pixels

7 of 44

“Shallow” vs. “deep” architectures

Hand-designed�feature extraction

Trainable�classifier

Image/ Video

Pixels

Object�Class

Layer 1

Layer N

Simple classifier

Object Class

Image/ Video

Pixels

Traditional recognition: “Shallow” architecture

Deep learning: “Deep” architecture

…

8 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]

Output:

 

bias

(Source Wikipedia)

9 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers

10 of 44

Inspiration: Neuron cells

11 of 44

Hubel/­Wiesel Architecture

  • D. Hubel and T. Wiesel (1959, 1962, Nobel Prize 1981)
    • Visual cortex consists of a hierarchy of simple, complex, and hyper-complex cells

12 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers
  • Shift-invariant network inspired by Hubel & Wiesel study of visual cortex “Neocognitron” [Fukushima 1980]

13 of 44

Brief history of Neural Networks

  • Perceptron algorithm [Rosenblatt, 57]
  • Multi-layer perceptron with hidden layers
  • Shift-invariant network inspired by Hubel & Wiesel study of visual cortex “Neocognitron” [Fukushima 1980]
  • Training of multi-layer networks with “backpropagation”

14 of 44

Convolutional Neural Networks (CNN, Convnet)

  • Neural network with specialized connectivity structure
  • Stack multiple stages of feature extractors
  • Higher stages compute more global, more invariant features
  • Classification layer at the end

​

Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recognition, Proceedings of the IEEE 86(11): 2278–2324, 1998.

15 of 44

Parametric supervised learning

  •  

 

 

(Slides borrowed from S. Lazebnik and M. Trager)

16 of 44

Neural networks

  •  

 

 

 

17 of 44

Feedforward neural network

  •  

 

18 of 44

Nonlinearity

  •  

 

 

19 of 44

Convolutional Neural Networks (CNN, Convnet)

  • Feed-forward feature extraction:
    1. Convolve input with learned filters
    2. Non-linearity
    3. Spatial pooling
    4. Normalization
  • Supervised training of convolutional �filters by back-propagating �classification error

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Normalization

Feature maps

20 of 44

Feedforward neural networks for images

image

Fully connected layer

Source: S. Lazebnik

21 of 44

CNNs: Neural networks for images

image

Convolutional layer

Source: S. Lazebnik

22 of 44

CNNs: Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

23 of 44

CNNs: Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

24 of 44

Convolution as feature extraction

Input

Feature Map

...

Source: S. Lazebnik

25 of 44

Neural networks for images

image

feature map

learned weights

Convolutional layer

Source: S. Lazebnik

26 of 44

Neural networks for images

image

next layer

Convolutional layer

+ ReLU

Source: S. Lazebnik

27 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Input

Feature Map

...

Source: R. Fergus, Y. LeCun

28 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Source: R. Fergus, Y. LeCun

Rectified Linear Unit (ReLU)

29 of 44

Key operations in a CNN

Input Image

Convolution (Learned)

Non-linearity

Spatial pooling

Feature maps

Max

Source: R. Fergus, Y. LeCun

30 of 44

Key operations in a CNN

Softmax layer:

Source: S. Lazebnik

 

31 of 44

Loss functions

  •  

 

 

 

32 of 44

Gradient descent

  •  

 

 

33 of 44

Stochastic gradient descent

  • Idea: instead of computing the gradient, compute an approximation

​

​

​

​

​

​

Note that in expectation

 

 

 

34 of 44

GD

SGD

35 of 44

How do we optimize the very nonlinear and

nonconvex empirical risk?

​

Many possible choices:

​

  • Second-order (Newton-like) methods

​

  • Gradient descent

​

  • Stochastic gradient descent and variants

36 of 44

Léon Bottou on large-scale learning..

37 of 44

Léon Bottou on large-scale learning..

38 of 44

Léon Bottou on large-scale learning..

39 of 44

Léon Bottou on large-scale learning..

40 of 44

Léon Bottou on large-scale learning..

41 of 44

Léon Bottou on large-scale learning..

42 of 44

Léon Bottou on large-scale learning..

43 of 44

Bottom line:

​

  • “Bad” optimization algorithms make sense in the large-scale setting, especially since we minimize expectations

​

  • Second-order (Newton-like) methods

​

  • Gradient descent

​

  • Stochastic gradient descent and variants

44 of 44

Batch SGD

  •