Lecture 7
Neural Networks
6.8300/1 Advances in Computer Vision
Spring 2024
Sara Beery, Kaiming He, Vincent Sitzmann, Mina Konaković Luković
7. Introduction to Deep Learning
Deep learning
In the past, we didn’t have enough data to fit these models. But now we do!
Modeling the visual world is incredibly complicated. We need high capacity models.
We want a class of high capacity models that are easy to optimize.
Deep neural networks!
A brief history of Neural Networks
time
enthusiasm
Perceptrons, 1958
Rosenblatt
Perceptrons, 1958
time
enthusiasm
Perceptrons,
1958
Minsky and Papert, Perceptrons, 1972
time
enthusiasm
Perceptrons,
1958
Minsky and Papert,
1972
Parallel Distributed Processing (PDP), 1986
XOR problem
Inputs
Output
0 0 0
1 0 1
0 1 1
1 1 0
PDP authors pointed to the backpropagation algorithm
as a breakthrough, allowing multi-layer neural networks to be
trained. Among the functions that a multi-layer network can represent but a single-layer network cannot: the XOR function.
0 1
0 1
time
enthusiasm
Perceptrons,
1958
Minsky and Papert,
1972
PDP book,
1986
LeCun conv nets, 1998
Demos:
14
Neural networks to recognize handwritten digits? yes
Neural networks for tougher problems? not really
Neural Information Processing Systems 2000
time
enthusiasm
Perceptrons,
1958
Minsky and Papert,
1972
PDP book,
1986
AI winter,
2000
Krizhevsky, Sutskever, and Hinton, NeurIPS 2012
“Alexnet”
Slide from Rob Fergus, NYU
Krizhevsky, Sutskever, and Hinton, NeurIPS 2012
time
enthusiasm
Perceptrons,
1958
Minsky and Papert,
1972
PDP book,
1986
AI winter,
2000
Krizhevsky, Sutskever,
Hinton, 2012
28 years
28 years
What comes next?
time
enthusiasm
Perceptrons,
1958
Minsky and Papert,
1972
PDP book,
1986
AI winter,
2000
Krizhevsky, Sutskever,
Hinton, 2012
28 years
28 years
2028 ?
What comes next?
Perceptrons,
1958
Minsky and Papert,
1972
PDP book,
1986
AI winter,
2000
time
enthusiasm
28 years
28 years
Krizhevsky, Sutskever,
Hinton, 2012
2028 ?
[“Mask RCNN”, He et al. 2017]
[“Neural module networks”, Andreas et al. 2017]
Ivy Tasi @ivymyt
Vitaly Vidmirov @vvid
[“pix2pix”, Isola et al. 2017]
Serre, 2014
Image classification
Edges
Texture
Colors
Segments
Parts
“clown fish”
“clown fish”
Edges
Texture
Colors
Segments
Parts
Learned
“clown fish”
Classifier
Image classification
“clown fish”
Learned
Image classification
“clown fish”
Learned
Neural net
Image classification
“clown fish”
Learned
Deep neural net
Image classification
“clown fish”
Loss
Learned
Deep learning
Training data
…
“Fish”
“Grizzly”
“Chameleon”
Gradient descent
Gradient descent
x
Gradient descent
One iteration of gradient descent:
Gradient descent
For large N, computing J in every iteration can be expensive
Stochastic gradient descent (SGD)
If batchsize=1 then θ is updated after each example.
If batchsize=N (full set) then this is standard gradient descent.
Input representation
Output representation
Computation in a neural net
Computation in a neural net
Input representation
Output representation
Linear layer
Computation in a neural net
Input representation
Output representation
Linear layer
weights
bias
Computation in a neural net
Input representation
Output representation
Linear layer
weights
bias
parameters of the model
Input representation
Output representation
Computation in a neural net
“Perceptron”
Example: linear classification with a perceptron
One layer neural net (perceptron) can
perform linear classification!
Training data
Non-differentiable non-linearity
Input representation
Output representation
Non-linearity with soft activation
Input representation
Output representation
Tanh
Tanh
Computation in a neural net — non-linearity
Sigmoid
Computation in a neural net — non-linearity
Rectified linear unit (ReLU)
Computation in a neural net — non-linearity
Leaky ReLU
Computation in a neural net — non-linearity
Output representation
Input representation
Intermediate representation
Stacking layers - Multi-layer Perceptron (MLP)
= “hidden units”
= “pre-activation hidden layer”
= “post-activation hidden layer”
Input representation
Intermediate representation
Stacking layers - fully connected layers
Output representation
Input representation
Intermediate representation
Example: how signal evolves
Output representation
positive
negative
Input representation
Intermediate representation
Example: how signal evolves
Output representation
positive
negative
Input representation
Intermediate representation
Example: how signal evolves
Output representation
positive
negative
Input representation
Intermediate representation
Example: how signal evolves
Output representation
positive
negative
Connectivity patterns
Input representation
Output representation
Fully connected layer
Locally connected layer
(Sparse W)
Input representation
Output representation
“clown fish”
Linear
Non-linearity
…
Deep nets
Example: linear classification with a perceptron
One layer neural net (perceptron) can
perform linear classification!
Example: nonlinear classification with a deep net net
http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf
http://www.deeplearningbook.org/contents/mlp.html
http://neuralnetworksanddeeplearning.com/chap4.html
http://www.inference.phy.cam.ac.uk/mackay/itprnn/ps/482.491.pdf
Representational power
“clown fish”
Linear
Non-linearity
…
Deep nets
Last layer
…
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
“clown fish”
argmax
Classifier layer
“clown fish”
Loss
error
Network output
…
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
Ground truth label
Loss function
“clown fish”
Loss
small
Network output
…
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
Ground truth label
Loss function
“grizzly bear”
Loss
large
Network output
…
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
Ground truth label
Loss function
Network output
…
dolphin
cat
grizzly bear
angel fish
chameleon
clown fish
Ground truth label
…
iguana
elephant
Probability of the observed data under the model
Prediction
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
Ground truth label
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
…
Prediction
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
Ground truth label
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
…
Loss
0
1
Prediction
Ground truth label
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
Likelihood
0
1
0
1
0
1
Likelihood of observed true data under predictive model
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
0
1
0
1
Prediction
Ground truth label
Loss
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
…
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
0
1
0
1
Prediction
Ground truth label
Loss
Softmax regression (a.k.a. multinomial logistic regression)
logits: vector of K scores, one for each class
squash into a non-negative vector that sums to 1 — i.e. a probability mass function!
dolphin
cat
grizzly bear
angel fish
chameleon
iguana
elephant
clown fish
0
1
…
…
“clown fish”
Loss
Learned
Deep learning
“grizzly bear”
Loss
Learned
Deep learning
“chameleon”
Loss
Learned
Deep learning
Batch (parallel) processing
Loss
Loss
Loss
…
Features
Images
Tensors
(multi-dimensional arrays)
…
Furry?
Is a fish?
Size
# Stripes
…
Each layer is a representation of the data
…
Tensors
(multi-dimensional arrays)
# neurons
# features
# units
# “channels”
Everything is a tensor
Regularizing deep nets
Deep nets have millions of parameters!
On many datasets, it is easy to overfit — we may have more free parameters than data points to constrain them.
How can we prevent the network from overfitting?
Recall: regularized least squares
Only use polynomial terms if you really need them! Most terms should be zero
ridge regression, a.k.a., Tikhonov regularization
Regularizing the weights in a neural net
weight decay
“We prefer to keep weights small.”
Dropout
Input representation
Intermediate representation
Output representation
Dropout
Input representation
Intermediate representation
Output representation
Dropout
Input representation
Intermediate representation
Output representation
Dropout
Input representation
Intermediate representation
Output representation
Dropout
Randomly zero out hidden units.
Prevents network from relying too much on spurious correlations between different hidden units.
Can be understood as averaging over an exponential ensemble of subnetworks. This averaging smooths the function, thereby reducing the effective capacity of the network.
Normalization layers
ReLU
Norm
Normalization layers
ReLU
Norm
Normalization layers
ReLU
Norm
Normalization layers
ReLU
Norm
Keep track of mean and variance of a unit (or a population of units) over time.
Standardize unit activations by subtracting mean and dividing by variance.
Squashes units into a standard range, avoiding overflow.
Normalization layers
Also achieves invariance to mean and variance of the training signal.
Both these properties reduce the effective capacity of the model, i.e. regularize the model.
Why do deep nets generalize?
The more parameters, the simpler the learned function
[Double-descent: Belkin, Hsu, Ma, Mandal, PNAS 2019]
More features —> smoother solutions
The simplicity hypothesis
Emerging theory:
deep nets learn simple functions that fit the data
Classical theory:
big models learn complicated functions, and overfit the data
Tensors
(multi-dimensional arrays)
# neurons
# features
# units
# “channels”
“Tensor flow”
z
z
Layer L
Input
Deep nets are data transformers
Two different ways to represent a function
Two different ways to represent a function
Data transformations for a variety of neural net layers
Mapping 2D
Wiring graph
Equation
Mapping 1D
Activations
Parameters
Wiring graph
Equation
Mapping
Matrix
N+1
M
Activations
Parameters
z
Training iteration
logits
class probabilites
relu
softmax
logits
class probabilites
relu
softmax
z
Training data
Training iteration
z
Training iteration
logits
class probabilites
relu
softmax
z
Training iteration
z
Training iteration
Layer 1 representation
[DeCAF, Donahue, Jia, et al. 2013]
[Visualization technique : t-sne, van der Maaten & Hinton, 2008]
7. Introduction to Deep Learning