Deep Learning
Chris Gregg
CS109, Stanford University
Summer 2026
Innovations in deep learning
2
AlphaGO (2016)
Notes:
Self Driving Cars
Computers making art
4
The Next Rembrandt
A Neural Algorithm of Artistic Style
Detecting skin cancer
5
Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.
Our Little Buddy
6
http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo.html
7
x1
x2
x3
x4
θ1
θ2
θ3
θ4
y
x1
x2
x3
x4
Biological Basis for Neural Networks
Actually, it’s probably someone else’s brain
Review
8
Classification Task
9
Heart 1
Heart 2
Heart n
ROI 1
ROI 2
ROI m
Output
0
1
1
0
1
1
1
0
0
0
0
1
…
…
…
Machine Learning
10
(inputs)
(prediction)
(model)
The Training / Testing Paradigm
11
Deployment
The Training / Testing Paradigm
12
Deployment
If your model passes testing…
Learn your parameters
Make sure that they work
13
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
14
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
15
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
16
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
17
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0.817
0
x:
1
1
18
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0.817
0
x:
1
1
19
Math for Logistic Regression
1
2
3
Make logistic regression assumption
Calculate the log likelihood for all data
Get derivative of log likelihood with respect to thetas
Often call this
1
Initialize: θj = 0 for all 0 ≤ j ≤ m
20
Logistic Regression Training
Repeat many times:
For each parameter j
For each training example (x, y):
𝜃j += η * gradient[j] for all 0 ≤ j ≤ m
gradient[j] = 0 for all 0 ≤ j ≤ m
gradient[j]
21
+
Training
Dataset likelihood:
Training iterations
Likelihood
22
Artificial Neurons
Last Class: Comparing Classifiers
Comparing Classifiers: Test Accuracy
24
Model Train Accuracy Test Accuracy
-------------------------------------------------------------
Baseline 0.6031 0.5887
Naive Bayes 0.7909 0.8067
Logistic Regression 0.8169 0.8307
Decision Tree 0.8514 0.8307
Random Forest 0.8726 0.8500
Gradient Boosting 0.8611 0.8440
AdaBoost 0.8334 0.8353
BayesNet 0.8320 0.8507
Comparing Classifiers: Calibration
25
Comparing Classifiers: Precision / Recall Curve
26
Dog classifier. Image credit:
wikipedia
End Review
27
Predicting a Categorical
28
Logistic Regression to Predict a Categorical?
29
+
z = 2.1
σ(z) = 0.7
Logistic Regression to Predict a Categorical?
30
Standard Logistic Regression
Multi-Class Logistic Regression
Logistic Regression to Predict a Categorical
31
32
Softmax is a generalization of the sigmoid function that squashes a K -dimensional vector z of arbitrary real values to a K-dimensional vector softmax(z) of real values in the range [0, 1] that add up to 1.
Categorical Classification?
Sigmoid is to Bernoulli as Softmax is to Categorical
Understanding Softmax
33
List: 7 7 7 7
[0.25, 0.25, 0.25, 0.25]
List: 2 8 5 6
[0.0, 0.84, 0.04, 0.11]
Why not this?
34
Softmax
Normalized Logistic
(oh…not convex)
Softmax is Related to Sigma Function!
35
In softmax with two classes,
What is Log Likelihood?
36
What is Log Likelihood?
37
We are ready…
39
x1
x2
x3
x4
θ1
θ2
θ3
θ4
y
x1
x2
x3
x4
Biological Basis for Neural Networks
Core idea behind the revolution in AI
41
Deep learning is (at its core) many logistic regression pieces stacked on top of each other.
(aka Neural Networks)
42
Digit Recognition Example
Let’s make feature vectors from pictures of numbers
Computer Vision
Hundreds of millions of neurons [1]
Visual neurons make up up 30% of your cortex [1]
Vision in your Brain
[1] http://discovermagazine.com/1993/jun/thevisionthingma227
45
Logistic Regression
This means it predicts a 0
…
This is one neuron…your brain has hundreds of millions
46
Logistic Regression
This means it predicts a 0
Indicates logistic regression connection
…
47
Logistic Regression
This means it predicts a 1
…
48
Not So Good
This means it predicts a 1
…
49
We Can Put Neurons Together
This means it predicts a 0
…
50
We Can Put Neurons Together
Look at a single “hidden” neuron
This means it predicts a 0
There is a parameter for every connection
…
51
We Can Put Neurons Together
Look at another “hidden” neuron
This means it predicts a 0
There is a parameter for every connection
…
52
We Can Put Neurons Together
This means it predicts a 0
…
53
We Can Put Neurons Together
This means it predicts a 0
…
Look at another neuron
There is a parameter for every connection
54
We Can Put Neurons Together
This means it predicts a 0
…
55
We Can Put Neurons Together
…
*lots
+
56
We Can Put Neurons Together
…
+
Single neuron in a hidden layer
57
We Can Put Neurons Together
…
*lots
+
58
We Can Put Neurons Together
…
*lots
*logistic
regression
Deep learning
def Deep learning is
maximum likelihood estimation with neural networks.
59
> 0.5?
Yes.�Predict 1
Lots of Logistic (regressions)
LOL
60
Demonstration
https://adamharley.com/nn_vis/cnn/2d.html
61
Deep learning gets its intelligence from its thetas (aka its parameters)
Lets Build a Neural Network!
62
How do we train?
MLE of Thetas!
First: Learning Goals…
1. Understand Chain Rule
as ♡ of Deep Learning
2. Demystify:
Deep Learning is MLE
3. Become experts of
logistic regression
Math worth knowing:
70
New Notation
Layer x
Layer h
Layer
71
New Notation
Layer x
Layer h
Layer
72
New Notation
Layer x
Layer h
Layer
73
Forward Pass
…
Layer x
Layer h
Layer
74
Forward Pass
…
Layer x
Layer h
Layer
75
Forward Pass
…
Layer x
Layer h
Layer
76
Forward Pass
…
Layer x
Layer h
Layer
77
All Together
Neural Network
x
h
…
…
78
Smoke Check 1
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in ?
b) 20
c) 40
d) 800
a) 2
79
Smoke Check 2
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in ?
b) 20
c) 40
d) 800
a) 2
80
Smoke Check 3
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in total?
b) 20
c) 820
d) 16000
a) 800
Today: Do Something Brave
82
Forward Pass
…
Layer x
Layer h
Layer
20 parameters need setting
800 parameters need setting
83
Only Have to Do Three Things
2
3
Calculate the log probability for all data
Get partial derivative of log likelihood with respect to each theta
1
Make deep learning assumption
84
Smoke Check
3
Get partial derivative of log likelihood with respect to each theta
Why?
85
Smoke Check
3
Get partial derivative of log likelihood with respect to each theta
Why?
We need to do gradient ascent
86
A deep learning model gets its intelligence by having useful thetas.
We can find useful thetas, by searching for ones that maximize likelihood of our training data
We can maximize likelihood using optimization techniques (such as gradient ascent).
In order to use optimization techniques, we need to calculate the partial derivative of likelihood with respect to thetas.
Why We Calculate Partial Derivatives
Basically MLE is hard because
it has so many details
87
Thanks to Keith Eicher
88
89
Only Have to Do Three Things
2
Calculate the log probability for all data
1
Make deep learning assumption
90
Same Assumption, Same LL
For one datum
For IID data
Take the log
Feel the Bern!
91
Only Have to Do Three Things
2
3
Calculate the log probability for all data
Get partial derivative of log likelihood with respect to each theta
1
Make deep learning assumption
92
Derivative Goals
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
93
Bad Approach
Neural Network
x
h
…
…
Math bug
94
Derivatives Without Tears
95
Woah Ms. Forster, you were right.
Chain rule is useful!
First use:
Big Idea #1: Chain Rule
96
Big Idea #2: Sigmoid Derivative
True fact about sigmoid functions
97
Big Idea #3: Derivative of Sum
We only need to calculate the gradient for one training example!
We will pretend we only have one example
We can sum up the gradients of each example to get the correct answer
Recall
99
Sigmoid has a Beautiful Slope
Sigmoid, you should be a ski hill
Chain rule!
Plug and chug
100
Derivative Goals
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
101
Neural Network
x
h
…
…
Chain Rule Example 1
Goal
Network
Decomposition
102
Neural Network
x
h
…
…
Chain Rule Example 2
Goal
Network
Decomposition
Decomposition
104
Gradient of output layer params
Neural Network
x
h
…
…
105
Gradient of output layer params
106
Gradient of output layer params
What! That’s not scary!
where
107
Make it Simple
Boom!
109
Neural Network
x
h
…
…
110
Gradient of hidden layer params
Neural Network
x
h
…
…
111
Gradient of hidden layer params
Wait is it over?
112
Gradient of hidden layer params
That one too?
113
Make it Simple
Congrats. You now know
Backpropagation
Moment of silence
117
Summary: Simple Calculations For
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
118
Bigger Neural Network
x
h(1)
…
…
h(2)
…
What Would You Do Here?
Chain rule:
Game changer for
artificial intelligence
120
Neural Networks Can Learn Complex Functions
a
x
b
c
d
e
f
g
LL
Works for any number of layers
1 Trillion Artificial Neurons
GoogLeNet Brain
Piech
123
GoogLeNet Brain
22 layers deep
Multiple,
Multi class output
Piech
Optimal stimulus
by numerical optimization
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Top stimuli from the test set
The Cat Neuron
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 1
Neuron 2
Neuron 3
Neuron 4
Neuron 5
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 7
Neuron 8
Neuron 6
Neuron 9
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 11
Neuron 10
Neuron 12
Neuron 13
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
22,000 categories
14,000,000 images
Hand-engineered features (SIFT, HOG, LBP),
Spatial pyramid, SparseCoding/Compression
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
ImageNet Classification
…
smoothhound, smoothhound shark, Mustelus mustelus
American smooth dogfish, Mustelus canis
Florida smoothhound, Mustelus norrisi
whitetip shark, reef whitetip shark, Triaenodon obseus
Atlantic spiny dogfish, Squalus acanthias
Pacific spiny dogfish, Squalus suckleyi
hammerhead, hammerhead shark
smooth hammerhead, Sphyrna zygaena
smalleye hammerhead, Sphyrna tudes
shovelhead, bonnethead, bonnet shark, Sphyrna tiburo
angel shark, angelfish, Squatina squatina, monkfish
electric ray, crampfish, numbfish, torpedo
smalltooth sawfish, Pristis pectinatus
guitarfish
roughtail stingray, Dasyatis centroura
butterfly ray
eagle ray
spotted eagle ray, spotted ray, Aetobatus narinari
cownose ray, cow-nosed ray, Rhinoptera bonasus
manta, manta ray, devilfish
Atlantic manta, Manta birostris
devil ray, Mobula hypostoma
grey skate, gray skate, Raja batis
little skate, Raja erinacea
…
Stingray
Mantaray
22,000 is a lot!
0.005%
Random guess
1.5%
?
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Pre Neural Networks
GoogLeNet
0.005%
Random guess
1.5%
Pre Neural Networks
43.9%
GoogLeNet
Szegedy et al, Going Deeper With Convolutions, CVPR 2015
0.005%
Random guess
1.5%
Pre Neural Networks
95.1%
SE-ResNet
How many parameters
is too many?
135
Good ML = Generalization
136
Dropout when your model is training, randomly turn off your neurons with probability 0.5. It will make your network more robust.
Prevent Overfitting?
Piech, CS106A, Stanford University
137
Convolution it turns out if you want to force some of your weights to be shared for different neurons, the math isn’t that much harder. This is used a lot for vision (CNN).
Shared Weights?
Piech, CS106A, Stanford University
Not everything is classification
Come to Monday’s Lecture!
Piech, CS106A, Stanford University
https://www.nytimes.com/2021/07/20/technology/ai-education-neural-networks.html
Human-Level Accuracy
Rubric Level Accuracy on Few-Shot Grading a Novel Question
General Exam Grading Model
Student answer (as text)
Question
(as text)
Rubric
(as text)
Education Tasks
General Exam Embedding
Code embedding
Question embedding
Rubric embedding
Invented the Proto-Transformer
Human-Level Accuracy
Rubric Level Accuracy on Few-Shot Grading a Novel Question
Gave Feedback to 3,500 Real Students
Do you agree? AI feedback 97.9%. Human feedback 96.7%
But what about interactive, creative assignments?
1M ungraded code.org assignments.
The AI is shown a brand new student game. Does it work?
Simultaneously learn to grade and play to grade.
Majority class: 50%
Code-as-text: 67%
Play-to-grade: 94%
Piech
Detecting skin cancer
148
Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.
Piech
Piech