Deep Learning
Chris Gregg
CS109, Stanford University
Summer 2026
Innovations in deep learning
2
AlphaGO (2016)
Notes:
Self Driving Cars
Computers making art
4
The Next Rembrandt
A Neural Algorithm of Artistic Style
Detecting skin cancer
5
Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.
Our Little Buddy
6
http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo.html
7
x1
x2
x3
x4
θ1
θ2
θ3
θ4
y
x1
x2
x3
x4
Biological Basis for Neural Networks
Actually, it’s probably someone else’s brain
Logistics
8
End of Class Schedule
9
Week | Monday | Wed | Fri |
Next (May 25th) | Memorial Day (no class) | Ethics | Regression |
Last Week (June 1st) | Application / Practice | Beyond CS109 | |
Final PEP
ML PSet
Review PSet
Final PEP Sign Ups
10
Review
11
Classification Task
12
Heart 1
Heart 2
Heart n
ROI 1
ROI 2
ROI m
Output
0
1
1
0
1
1
1
0
0
0
0
1
…
…
…
Machine Learning
13
(inputs)
(prediction)
(model)
The Training / Testing Paradigm
14
Deployment
The Training / Testing Paradigm
15
Deployment
If your model passes testing…
Learn your parameters
Make sure that they work
16
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
17
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
18
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
19
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0
x:
1
1
20
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0.817
0
x:
1
1
21
+
z = 2.1
σ(z) = 0.7
Logistic Regression
0.817
0
x:
1
1
22
Math for Logistic Regression
1
2
3
Make logistic regression assumption
Calculate the log likelihood for all data
Get derivative of log likelihood with respect to thetas
Often call this
1
Initialize: θj = 0 for all 0 ≤ j ≤ m
23
Logistic Regression Training
Repeat many times:
For each parameter j
For each training example (x, y):
𝜃j += η * gradient[j] for all 0 ≤ j ≤ m
gradient[j] = 0 for all 0 ≤ j ≤ m
gradient[j]
24
+
Training
Dataset likelihood:
Training iterations
Likelihood
25
Artificial Neurons
Last Class: Comparing Classifiers
Comparing Classifiers: Test Accuracy
27
Model Train Accuracy Test Accuracy
-------------------------------------------------------------
Baseline 0.6031 0.5887
Naive Bayes 0.7909 0.8067
Logistic Regression 0.8169 0.8307
Decision Tree 0.8514 0.8307
Random Forest 0.8726 0.8500
Gradient Boosting 0.8611 0.8440
AdaBoost 0.8334 0.8353
BayesNet 0.8320 0.8507
Comparing Classifiers: Calibration
28
Comparing Classifiers: Precision / Recall Curve
29
Dog classifier. Image credit:
wikipedia
End Review
30
Predicting a Categorical
31
Logistic Regression to Predict a Categorical?
32
+
z = 2.1
σ(z) = 0.7
Logistic Regression to Predict a Categorical?
33
Standard Logistic Regression
Multi-Class Logistic Regression
Logistic Regression to Predict a Categorical
34
35
Softmax is a generalization of the sigmoid function that squashes a K -dimensional vector z of arbitrary real values to a K-dimensional vector softmax(z) of real values in the range [0, 1] that add up to 1.
Categorical Classification?
Sigmoid is to Bernoulli as Softmax is to Categorical
Understanding Softmax
36
List: 7 7 7 7
[0.25, 0.25, 0.25, 0.25]
List: 2 8 5 6
[0.0, 0.84, 0.04, 0.11]
Why not this?
37
Softmax
Normalized Logistic
(oh…not convex)
Softmax is Related to Sigma Function!
38
In softmax with two classes,
What is Log Likelihood?
39
What is Log Likelihood?
40
We are ready…
42
x1
x2
x3
x4
θ1
θ2
θ3
θ4
y
x1
x2
x3
x4
Biological Basis for Neural Networks
Core idea behind the revolution in AI
44
Deep learning is (at its core) many logistic regression pieces stacked on top of each other.
(aka Neural Networks)
45
Digit Recognition Example
Let’s make feature vectors from pictures of numbers
Computer Vision
Hundreds of millions of neurons [1]
Visual neurons make up up 30% of your cortex [1]
Vision in your Brain
[1] http://discovermagazine.com/1993/jun/thevisionthingma227
48
Logistic Regression
This means it predicts a 0
…
This is one neuron…your brain has hundreds of millions
49
Logistic Regression
This means it predicts a 0
Indicates logistic regression connection
…
50
Logistic Regression
This means it predicts a 1
…
51
Not So Good
This means it predicts a 1
…
52
We Can Put Neurons Together
This means it predicts a 0
…
53
We Can Put Neurons Together
Look at a single “hidden” neuron
This means it predicts a 0
There is a parameter for every connection
…
54
We Can Put Neurons Together
Look at another “hidden” neuron
This means it predicts a 0
There is a parameter for every connection
…
55
We Can Put Neurons Together
This means it predicts a 0
…
56
We Can Put Neurons Together
This means it predicts a 0
…
Look at another neuron
There is a parameter for every connection
57
We Can Put Neurons Together
This means it predicts a 0
…
58
We Can Put Neurons Together
…
*lots
+
59
We Can Put Neurons Together
…
+
Single neuron in a hidden layer
60
We Can Put Neurons Together
…
*lots
+
61
We Can Put Neurons Together
…
*lots
*logistic
regression
Deep learning
def Deep learning is
maximum likelihood estimation with neural networks.
62
> 0.5?
Yes.�Predict 1
Lots of Logistic (regressions)
LOL
63
Demonstration
https://adamharley.com/nn_vis/cnn/2d.html
64
Deep learning gets its intelligence from its thetas (aka its parameters)
Lets Build a Neural Network!
65
How do we train?
MLE of Thetas!
First: Learning Goals…
1. Understand Chain Rule
as ♡ of Deep Learning
2. Demystify:
Deep Learning is MLE
3. Become experts of
logistic regression
Math worth knowing:
73
New Notation
Layer x
Layer h
Layer
74
New Notation
Layer x
Layer h
Layer
75
New Notation
Layer x
Layer h
Layer
76
Forward Pass
…
Layer x
Layer h
Layer
77
Forward Pass
…
Layer x
Layer h
Layer
78
Forward Pass
…
Layer x
Layer h
Layer
79
Forward Pass
…
Layer x
Layer h
Layer
80
All Together
Neural Network
x
h
…
…
81
Smoke Check 1
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in ?
b) 20
c) 40
d) 800
a) 2
82
Smoke Check 2
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in ?
b) 20
c) 40
d) 800
a) 2
83
Smoke Check 3
Neural Network
x
h
…
…
|x| = 40
|h| = 20
How many parameters in total?
b) 20
c) 820
d) 16000
a) 800
Today: Do Something Brave
85
Forward Pass
…
Layer x
Layer h
Layer
20 parameters need setting
800 parameters need setting
86
Only Have to Do Three Things
2
3
Calculate the log probability for all data
Get partial derivative of log likelihood with respect to each theta
1
Make deep learning assumption
87
Smoke Check
3
Get partial derivative of log likelihood with respect to each theta
Why?
88
Smoke Check
3
Get partial derivative of log likelihood with respect to each theta
Why?
We need to do gradient ascent
89
A deep learning model gets its intelligence by having useful thetas.
We can find useful thetas, by searching for ones that maximize likelihood of our training data
We can maximize likelihood using optimization techniques (such as gradient ascent).
In order to use optimization techniques, we need to calculate the partial derivative of likelihood with respect to thetas.
Why We Calculate Partial Derivatives
Basically MLE is hard because
it has so many details
90
Thanks to Keith Eicher
91
92
Only Have to Do Three Things
2
Calculate the log probability for all data
1
Make deep learning assumption
93
Same Assumption, Same LL
For one datum
For IID data
Take the log
Feel the Bern!
94
Only Have to Do Three Things
2
3
Calculate the log probability for all data
Get partial derivative of log likelihood with respect to each theta
1
Make deep learning assumption
95
Derivative Goals
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
96
Bad Approach
Neural Network
x
h
…
…
Math bug
97
Derivatives Without Tears
98
Woah Ms. Forster, you were right.
Chain rule is useful!
First use:
Big Idea #1: Chain Rule
99
Big Idea #2: Sigmoid Derivative
True fact about sigmoid functions
100
Big Idea #3: Derivative of Sum
We only need to calculate the gradient for one training example!
We will pretend we only have one example
We can sum up the gradients of each example to get the correct answer
Recall
102
Sigmoid has a Beautiful Slope
Sigmoid, you should be a ski hill
Chain rule!
Plug and chug
103
Derivative Goals
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
104
Neural Network
x
h
…
…
Chain Rule Example 1
Goal
Network
Decomposition
105
Neural Network
x
h
…
…
Chain Rule Example 2
Goal
Network
Decomposition
Decomposition
107
Gradient of output layer params
Neural Network
x
h
…
…
108
Gradient of output layer params
109
Gradient of output layer params
What! That’s not scary!
where
110
Make it Simple
Boom!
112
Neural Network
x
h
…
…
113
Gradient of hidden layer params
Neural Network
x
h
…
…
114
Gradient of hidden layer params
Wait is it over?
115
Gradient of hidden layer params
That one too?
116
Make it Simple
Congrats. You now know
Backpropagation
Moment of silence
120
Summary: Simple Calculations For
Neural Network
x
h
…
…
Loss with respect to
output layer params
Loss with respect to
hidden layer params
121
Bigger Neural Network
x
h(1)
…
…
h(2)
…
What Would You Do Here?
Chain rule:
Game changer for
artificial intelligence
123
Neural Networks Can Learn Complex Functions
a
x
b
c
d
e
f
g
LL
Works for any number of layers
1 Trillion Artificial Neurons
GoogLeNet Brain
Piech
126
GoogLeNet Brain
22 layers deep
Multiple,
Multi class output
Piech
Optimal stimulus
by numerical optimization
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Top stimuli from the test set
The Cat Neuron
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 1
Neuron 2
Neuron 3
Neuron 4
Neuron 5
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 7
Neuron 8
Neuron 6
Neuron 9
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
Pooling Size = 5
Number
of maps = 8
Image Size = 200
Number of output
channels = 8
Number of input
channels = 3
One layer
RF size = 18
Input to another layer above
(image with 8 channels)
W
H
LCN Size = 5
Neuron 11
Neuron 10
Neuron 12
Neuron 13
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Best Neuron Stimuli
22,000 categories
14,000,000 images
Hand-engineered features (SIFT, HOG, LBP),
Spatial pyramid, SparseCoding/Compression
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
ImageNet Classification
…
smoothhound, smoothhound shark, Mustelus mustelus
American smooth dogfish, Mustelus canis
Florida smoothhound, Mustelus norrisi
whitetip shark, reef whitetip shark, Triaenodon obseus
Atlantic spiny dogfish, Squalus acanthias
Pacific spiny dogfish, Squalus suckleyi
hammerhead, hammerhead shark
smooth hammerhead, Sphyrna zygaena
smalleye hammerhead, Sphyrna tudes
shovelhead, bonnethead, bonnet shark, Sphyrna tiburo
angel shark, angelfish, Squatina squatina, monkfish
electric ray, crampfish, numbfish, torpedo
smalltooth sawfish, Pristis pectinatus
guitarfish
roughtail stingray, Dasyatis centroura
butterfly ray
eagle ray
spotted eagle ray, spotted ray, Aetobatus narinari
cownose ray, cow-nosed ray, Rhinoptera bonasus
manta, manta ray, devilfish
Atlantic manta, Manta birostris
devil ray, Mobula hypostoma
grey skate, gray skate, Raja batis
little skate, Raja erinacea
…
Stingray
Mantaray
22,000 is a lot!
0.005%
Random guess
1.5%
?
Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012
Pre Neural Networks
GoogLeNet
0.005%
Random guess
1.5%
Pre Neural Networks
43.9%
GoogLeNet
Szegedy et al, Going Deeper With Convolutions, CVPR 2015
0.005%
Random guess
1.5%
Pre Neural Networks
95.1%
SE-ResNet
How many parameters
is too many?
138
Good ML = Generalization
139
Dropout when your model is training, randomly turn off your neurons with probability 0.5. It will make your network more robust.
Prevent Overfitting?
Piech, CS106A, Stanford University
140
Convolution it turns out if you want to force some of your weights to be shared for different neurons, the math isn’t that much harder. This is used a lot for vision (CNN).
Shared Weights?
Piech, CS106A, Stanford University
Not everything is classification
Come to Monday’s Lecture!
Piech, CS106A, Stanford University
https://www.nytimes.com/2021/07/20/technology/ai-education-neural-networks.html
Human-Level Accuracy
Rubric Level Accuracy on Few-Shot Grading a Novel Question
General Exam Grading Model
Student answer (as text)
Question
(as text)
Rubric
(as text)
Education Tasks
General Exam Embedding
Code embedding
Question embedding
Rubric embedding
Invented the Proto-Transformer
Human-Level Accuracy
Rubric Level Accuracy on Few-Shot Grading a Novel Question
Gave Feedback to 3,500 Real Students
Do you agree? AI feedback 97.9%. Human feedback 96.7%
But what about interactive, creative assignments?
1M ungraded code.org assignments.
The AI is shown a brand new student game. Does it work?
Simultaneously learn to grade and play to grade.
Majority class: 50%
Code-as-text: 67%
Play-to-grade: 94%
Piech
Detecting skin cancer
151
Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.
Piech
Piech