1 of 120

Lecture 7

Neural Networks

6.8300/1 Advances in Computer Vision

Spring 2024

Sara Beery, Kaiming He, Vincent Sitzmann, Mina Konaković Luković

2 of 120

7. Introduction to Deep Learning

  • Brief history
  • Basic formulation (hierarchical processing)
  • Optimization via gradient descent
  • Layer types (Linear, Pointwise non-linearity)
  • Everything is a tensor
  • Deep nets as data transformers

3 of 120

Deep learning

In the past, we didn’t have enough data to fit these models. But now we do!

Modeling the visual world is incredibly complicated. We need high capacity models.

We want a class of high capacity models that are easy to optimize.

Deep neural networks!

4 of 120

A brief history of Neural Networks

time

enthusiasm

5 of 120

Perceptrons, 1958

Rosenblatt

6 of 120

Perceptrons, 1958

7 of 120

time

enthusiasm

Perceptrons,

1958

8 of 120

Minsky and Papert, Perceptrons, 1972

9 of 120

time

enthusiasm

Perceptrons,

1958

Minsky and Papert,

1972

10 of 120

Parallel Distributed Processing (PDP), 1986

11 of 120

XOR problem

Inputs

Output

0 0 0

1 0 1

0 1 1

1 1 0

PDP authors pointed to the backpropagation algorithm

as a breakthrough, allowing multi-layer neural networks to be

trained. Among the functions that a multi-layer network can represent but a single-layer network cannot: the XOR function.

0 1

0 1

12 of 120

time

enthusiasm

Perceptrons,

1958

Minsky and Papert,

1972

PDP book,

1986

13 of 120

LeCun conv nets, 1998

Demos:

14 of 120

14

15 of 120

Neural networks to recognize handwritten digits? yes

Neural networks for tougher problems? not really

16 of 120

Neural Information Processing Systems 2000

  • Neural Information Processing Systems, is the premier conference on machine learning. Evolved from an interdisciplinary conference to a machine learning conference.
  • For the 2000 conference:
    • title words predictive of paper acceptance: “Belief Propagation” and “Gaussian”.
    • title words predictive of paper rejection: “Neural” and “Network”.

17 of 120

time

enthusiasm

Perceptrons,

1958

Minsky and Papert,

1972

PDP book,

1986

AI winter,

2000

18 of 120

Krizhevsky, Sutskever, and Hinton, NeurIPS 2012

“Alexnet”

19 of 120

Slide from Rob Fergus, NYU

20 of 120

Krizhevsky, Sutskever, and Hinton, NeurIPS 2012

21 of 120

time

enthusiasm

Perceptrons,

1958

Minsky and Papert,

1972

PDP book,

1986

AI winter,

2000

Krizhevsky, Sutskever,

Hinton, 2012

28 years

28 years

22 of 120

What comes next?

time

enthusiasm

Perceptrons,

1958

Minsky and Papert,

1972

PDP book,

1986

AI winter,

2000

Krizhevsky, Sutskever,

Hinton, 2012

28 years

28 years

2028 ?

23 of 120

What comes next?

Perceptrons,

1958

Minsky and Papert,

1972

PDP book,

1986

AI winter,

2000

time

enthusiasm

28 years

28 years

Krizhevsky, Sutskever,

Hinton, 2012

2028 ?

24 of 120

[“Mask RCNN”, He et al. 2017]

25 of 120

[“Neural module networks”, Andreas et al. 2017]

26 of 120

Ivy Tasi @ivymyt

Vitaly Vidmirov @vvid

[“pix2pix”, Isola et al. 2017]

27 of 120

Serre, 2014

28 of 120

Image classification

Edges

Texture

Colors

Segments

Parts

“clown fish”

29 of 120

“clown fish”

Edges

Texture

Colors

Segments

Parts

Learned

“clown fish”

Classifier

Image classification

30 of 120

“clown fish”

Learned

Image classification

31 of 120

“clown fish”

Learned

Neural net

Image classification

32 of 120

“clown fish”

Learned

Deep neural net

Image classification

33 of 120

“clown fish”

Loss

Learned

Deep learning

Training data

“Fish”

“Grizzly”

“Chameleon”

34 of 120

Gradient descent

35 of 120

Gradient descent

x

36 of 120

Gradient descent

One iteration of gradient descent:

37 of 120

Gradient descent

For large N, computing J in every iteration can be expensive

38 of 120

Stochastic gradient descent (SGD)

  • Want to minimize overall loss function J, which is sum of individual losses over each example.
  • In Stochastic gradient descent, compute gradient on sub-set (batch) of data.

If batchsize=1 then θ is updated after each example.

If batchsize=N (full set) then this is standard gradient descent.

  • Gradient direction is noisy, relative to average over all examples (standard gradient descent).
  • Advantages
    • Faster: approximate total gradient with small sample
    • Implicit regularizer
  • Disadvantages
    • High variance, unstable updates

39 of 120

Input representation

Output representation

Computation in a neural net

40 of 120

Computation in a neural net

Input representation

Output representation

Linear layer

41 of 120

Computation in a neural net

Input representation

Output representation

Linear layer

weights

bias

42 of 120

Computation in a neural net

Input representation

Output representation

Linear layer

weights

bias

parameters of the model

43 of 120

Input representation

Output representation

Computation in a neural net

“Perceptron”

44 of 120

Example: linear classification with a perceptron

One layer neural net (perceptron) can

perform linear classification!

45 of 120

Training data

46 of 120

Non-differentiable non-linearity

Input representation

Output representation

47 of 120

Non-linearity with soft activation

Input representation

Output representation

Tanh

48 of 120

Tanh

  • Bounded between [-1,+1]
  • Saturation for large +/- inputs
  • Gradients go to zero (vanishing gradients)
  • Outputs centered at 0
  • tanh(z) = 2 sigmoid(2z) −1

Computation in a neural net — non-linearity

49 of 120

Sigmoid

  • Interpretation as firing rate of neuron
  • Bounded between [0,1]
  • Saturation for large +/- inputs
  • Gradients go to zero
  • Outputs centered at 0.5 � (poor conditioning)
  • Not used in practice

Computation in a neural net — non-linearity

50 of 120

Rectified linear unit (ReLU)

  • Unbounded output (on positive side)
  • Efficient to implement:
  • Also seems to help convergence (see 6x speedup vs tanh in [Krizhevsky et al.])
  • Drawback: if strongly in negative region, unit is dead forever (no gradient).
  • Default choice: widely used in current models.

Computation in a neural net — non-linearity

51 of 120

Leaky ReLU

  • where α is small (e.g. 0.02)
  • Efficient to implement:
  • Also known as probabilistic ReLU (PReLU)
  • Has non-zero gradients everywhere (unlike ReLU)
  • α can also be learned (see Kaiming He et al. 2015).

Computation in a neural net — non-linearity

52 of 120

Output representation

Input representation

Intermediate representation

Stacking layers - Multi-layer Perceptron (MLP)

= “hidden units

= “pre-activation hidden layer”

= “post-activation hidden layer”

53 of 120

Input representation

Intermediate representation

Stacking layers - fully connected layers

Output representation

54 of 120

Input representation

Intermediate representation

Example: how signal evolves

Output representation

positive

negative

55 of 120

Input representation

Intermediate representation

Example: how signal evolves

Output representation

positive

negative

56 of 120

Input representation

Intermediate representation

Example: how signal evolves

Output representation

positive

negative

57 of 120

Input representation

Intermediate representation

Example: how signal evolves

Output representation

positive

negative

58 of 120

Connectivity patterns

Input representation

Output representation

Fully connected layer

Locally connected layer

(Sparse W)

Input representation

Output representation

59 of 120

“clown fish”

Linear

Non-linearity

Deep nets

60 of 120

Example: linear classification with a perceptron

One layer neural net (perceptron) can

perform linear classification!

61 of 120

Example: nonlinear classification with a deep net net

62 of 120

  • 1 layer? Linear decision surface.
  • 2+ layers? In theory, can represent any function. Assuming non-trivial non-linearity.
    • Bengio 2009,

http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf

    • Bengio, Courville, Goodfellow book

http://www.deeplearningbook.org/contents/mlp.html

    • Simple proof by M. Neilsen

http://neuralnetworksanddeeplearning.com/chap4.html

    • D. Mackay book

http://www.inference.phy.cam.ac.uk/mackay/itprnn/ps/482.491.pdf

  • But issue is efficiency: very wide two layers vs narrow deep model? In practice, more layers helps.

Representational power

63 of 120

“clown fish”

Linear

Non-linearity

Deep nets

64 of 120

Last layer

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

“clown fish”

argmax

Classifier layer

65 of 120

“clown fish”

Loss

error

Network output

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

Ground truth label

Loss function

66 of 120

“clown fish”

Loss

small

Network output

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

Ground truth label

Loss function

67 of 120

“grizzly bear”

Loss

large

Network output

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

Ground truth label

Loss function

68 of 120

Network output

dolphin

cat

grizzly bear

angel fish

chameleon

clown fish

Ground truth label

iguana

elephant

Probability of the observed data under the model

69 of 120

Prediction

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

Ground truth label

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

70 of 120

Prediction

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

Ground truth label

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

Loss

0

1

71 of 120

Prediction

Ground truth label

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

Likelihood

0

1

0

1

0

1

Likelihood of observed true data under predictive model

72 of 120

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

0

1

0

1

Prediction

Ground truth label

Loss

73 of 120

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

0

1

0

1

Prediction

Ground truth label

Loss

74 of 120

Softmax regression (a.k.a. multinomial logistic regression)

logits: vector of K scores, one for each class

squash into a non-negative vector that sums to 1 — i.e. a probability mass function!

dolphin

cat

grizzly bear

angel fish

chameleon

iguana

elephant

clown fish

0

1

75 of 120

“clown fish”

Loss

Learned

Deep learning

76 of 120

“grizzly bear”

Loss

Learned

Deep learning

77 of 120

“chameleon”

Loss

Learned

Deep learning

78 of 120

Batch (parallel) processing

Loss

Loss

Loss

Features

Images

79 of 120

Tensors

(multi-dimensional arrays)

Furry?

Is a fish?

Size

# Stripes

Each layer is a representation of the data

80 of 120

Tensors

(multi-dimensional arrays)

# neurons

# features

# units

# “channels”

81 of 120

Everything is a tensor

82 of 120

Regularizing deep nets

Deep nets have millions of parameters!

On many datasets, it is easy to overfit — we may have more free parameters than data points to constrain them.

How can we prevent the network from overfitting?

  1. Fewer neurons, fewer layers
  2. Weight decay and other regularizers
  3. Normalization layers

83 of 120

Recall: regularized least squares

Only use polynomial terms if you really need them! Most terms should be zero

ridge regression, a.k.a., Tikhonov regularization

84 of 120

Regularizing the weights in a neural net

weight decay

“We prefer to keep weights small.”

85 of 120

Dropout

Input representation

Intermediate representation

Output representation

86 of 120

Dropout

Input representation

Intermediate representation

Output representation

87 of 120

Dropout

Input representation

Intermediate representation

Output representation

88 of 120

Dropout

Input representation

Intermediate representation

Output representation

89 of 120

Dropout

Randomly zero out hidden units.

Prevents network from relying too much on spurious correlations between different hidden units.

Can be understood as averaging over an exponential ensemble of subnetworks. This averaging smooths the function, thereby reducing the effective capacity of the network.

90 of 120

Normalization layers

ReLU

Norm

91 of 120

Normalization layers

ReLU

Norm

92 of 120

Normalization layers

ReLU

Norm

93 of 120

Normalization layers

ReLU

Norm

94 of 120

Keep track of mean and variance of a unit (or a population of units) over time.

Standardize unit activations by subtracting mean and dividing by variance.

Squashes units into a standard range, avoiding overflow.

Normalization layers

Also achieves invariance to mean and variance of the training signal.

Both these properties reduce the effective capacity of the model, i.e. regularize the model.

95 of 120

Why do deep nets generalize?

  • Deep nets have so many parameters they could just act like look up tables, regurgitating their training data
  • Instead, they learn rules that generalize
  • Defies classical theory!

96 of 120

The more parameters, the simpler the learned function

[Double-descent: Belkin, Hsu, Ma, Mandal, PNAS 2019]

More features —> smoother solutions

97 of 120

The simplicity hypothesis

Emerging theory:

deep nets learn simple functions that fit the data

Classical theory:

big models learn complicated functions, and overfit the data

98 of 120

Tensors

(multi-dimensional arrays)

# neurons

# features

# units

# “channels”

99 of 120

“Tensor flow”

100 of 120

101 of 120

z

z

Layer L

Input

Deep nets are data transformers

  • Deep nets transform datapoints, layer by layer
  • Each layer is a different representation of the data
  • We call these representations embeddings

102 of 120

Two different ways to represent a function

103 of 120

Two different ways to represent a function

104 of 120

Data transformations for a variety of neural net layers

105 of 120

Mapping 2D

Wiring graph

Equation

Mapping 1D

Activations

Parameters

106 of 120

Wiring graph

Equation

Mapping

Matrix

N+1

M

Activations

Parameters

107 of 120

z

Training iteration

logits

class probabilites

relu

softmax

108 of 120

logits

class probabilites

relu

softmax

z

Training data

Training iteration

109 of 120

z

Training iteration

logits

class probabilites

relu

softmax

110 of 120

z

Training iteration

111 of 120

z

112 of 120

113 of 120

114 of 120

115 of 120

116 of 120

117 of 120

Training iteration

118 of 120

119 of 120

Layer 1 representation

[DeCAF, Donahue, Jia, et al. 2013]

[Visualization technique : t-sne, van der Maaten & Hinton, 2008]

120 of 120

7. Introduction to Deep Learning

  • Brief history
  • Basic formulation (hierarchical processing)
  • Optimization via gradient descent
  • Layer types (Linear, Pointwise non-linearity)
  • Everything is a tensor
  • Deep nets as data transformers