1 of 149

Deep Learning

Chris Gregg

CS109, Stanford University

Summer 2026

2 of 149

Innovations in deep learning

  • Deep learning (neural networks) is the core idea driving the current revolution in AI.

2

AlphaGO (2016)

Notes:

3 of 149

Self Driving Cars

4 of 149

Computers making art

4

5 of 149

Detecting skin cancer

5

Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.

6 of 149

Our Little Buddy

6

http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo.html

7 of 149

  • A neuron

  • Your brain

7

x1

x2

x3

x4

θ1

θ2

θ3

θ4

y

x1

x2

x3

x4

Biological Basis for Neural Networks

Actually, it’s probably someone else’s brain

8 of 149

Review

8

9 of 149

Classification Task

9

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

0

1

1

0

1

1

1

0

0

0

0

1

10 of 149

Machine Learning

10

(inputs)

(prediction)

(model)

11 of 149

The Training / Testing Paradigm

11

Deployment

12 of 149

The Training / Testing Paradigm

12

Deployment

If your model passes testing…

Learn your parameters

Make sure that they work

13 of 149

13

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

14 of 149

14

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

15 of 149

15

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

16 of 149

16

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

17 of 149

17

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0.817

0

x:

1

1

18 of 149

18

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0.817

0

x:

1

1

19 of 149

19

Math for Logistic Regression

1

2

3

Make logistic regression assumption

Calculate the log likelihood for all data

Get derivative of log likelihood with respect to thetas

Often call this

1

20 of 149

Initialize: θj = 0 for all 0 ≤ jm

20

Logistic Regression Training

Repeat many times:

For each parameter j

For each training example (x, y):

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

gradient[j]

21 of 149

21

+

Training

Dataset likelihood:

Training iterations

Likelihood

22 of 149

22

Artificial Neurons

23 of 149

Last Class: Comparing Classifiers

24 of 149

Comparing Classifiers: Test Accuracy

24

Model Train Accuracy Test Accuracy

-------------------------------------------------------------

Baseline 0.6031 0.5887

Naive Bayes 0.7909 0.8067

Logistic Regression 0.8169 0.8307

Decision Tree 0.8514 0.8307

Random Forest 0.8726 0.8500

Gradient Boosting 0.8611 0.8440

AdaBoost 0.8334 0.8353

BayesNet 0.8320 0.8507

25 of 149

Comparing Classifiers: Calibration

25

26 of 149

Comparing Classifiers: Precision / Recall Curve

26

Dog classifier. Image credit:

wikipedia

27 of 149

End Review

27

28 of 149

Predicting a Categorical

28

29 of 149

Logistic Regression to Predict a Categorical?

29

+

z = 2.1

σ(z) = 0.7

30 of 149

Logistic Regression to Predict a Categorical?

30

Standard Logistic Regression

Multi-Class Logistic Regression

31 of 149

Logistic Regression to Predict a Categorical

31

32 of 149

32

Softmax is a generalization of the sigmoid function that squashes a K -dimensional vector z of arbitrary real values to a K-dimensional vector softmax(z) of real values in the range [0, 1] that add up to 1.

Categorical Classification?

Sigmoid is to Bernoulli as Softmax is to Categorical

33 of 149

Understanding Softmax

33

List: 7 7 7 7

[0.25, 0.25, 0.25, 0.25]

List: 2 8 5 6

[0.0, 0.84, 0.04, 0.11]

34 of 149

Why not this?

34

Softmax

Normalized Logistic

(oh…not convex)

35 of 149

Softmax is Related to Sigma Function!

35

In softmax with two classes,

36 of 149

What is Log Likelihood?

36

37 of 149

What is Log Likelihood?

37

38 of 149

We are ready…

39 of 149

  • A neuron

  • Your brain

39

x1

x2

x3

x4

θ1

θ2

θ3

θ4

y

x1

x2

x3

x4

Biological Basis for Neural Networks

40 of 149

Core idea behind the revolution in AI

41 of 149

41

Deep learning is (at its core) many logistic regression pieces stacked on top of each other.

(aka Neural Networks)

42 of 149

42

Digit Recognition Example

Let’s make feature vectors from pictures of numbers

43 of 149

Computer Vision

44 of 149

Hundreds of millions of neurons [1]

Visual neurons make up up 30% of your cortex [1]

Vision in your Brain

[1] http://discovermagazine.com/1993/jun/thevisionthingma227

45 of 149

45

Logistic Regression

This means it predicts a 0

This is one neuron…your brain has hundreds of millions

46 of 149

46

Logistic Regression

This means it predicts a 0

Indicates logistic regression connection

47 of 149

47

Logistic Regression

This means it predicts a 1

48 of 149

48

Not So Good

This means it predicts a 1

49 of 149

49

We Can Put Neurons Together

This means it predicts a 0

50 of 149

50

We Can Put Neurons Together

Look at a single “hidden” neuron

This means it predicts a 0

There is a parameter for every connection

51 of 149

51

We Can Put Neurons Together

Look at another “hidden” neuron

This means it predicts a 0

There is a parameter for every connection

52 of 149

52

We Can Put Neurons Together

This means it predicts a 0

53 of 149

53

We Can Put Neurons Together

This means it predicts a 0

Look at another neuron

There is a parameter for every connection

54 of 149

54

We Can Put Neurons Together

This means it predicts a 0

55 of 149

55

We Can Put Neurons Together

*lots

+

56 of 149

56

We Can Put Neurons Together

+

Single neuron in a hidden layer

57 of 149

57

We Can Put Neurons Together

*lots

+

58 of 149

58

We Can Put Neurons Together

*lots

*logistic

regression

59 of 149

Deep learning

def Deep learning is

maximum likelihood estimation with neural networks.

59

 

 

> 0.5?

Yes.�Predict 1

Lots of Logistic (regressions)

LOL

  • def A neural network is
  • (at its core) many logistic regression pieces stacked on top of each other.

60 of 149

60

Demonstration

https://adamharley.com/nn_vis/cnn/2d.html

61 of 149

61

Deep learning gets its intelligence from its thetas (aka its parameters)

62 of 149

Lets Build a Neural Network!

62

63 of 149

How do we train?

64 of 149

MLE of Thetas!

65 of 149

First: Learning Goals…

66 of 149

1. Understand Chain Rule

as ♡ of Deep Learning

67 of 149

2. Demystify:

Deep Learning is MLE

68 of 149

3. Become experts of

logistic regression

69 of 149

Math worth knowing:

70 of 149

70

New Notation

Layer x

Layer h

Layer

71 of 149

71

New Notation

Layer x

Layer h

Layer

72 of 149

72

New Notation

Layer x

Layer h

Layer

73 of 149

73

Forward Pass

Layer x

Layer h

Layer

74 of 149

74

Forward Pass

Layer x

Layer h

Layer

75 of 149

75

Forward Pass

Layer x

Layer h

Layer

76 of 149

76

Forward Pass

Layer x

Layer h

Layer

77 of 149

77

All Together

Neural Network

x

h

78 of 149

78

Smoke Check 1

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in ?

b) 20

c) 40

d) 800

a) 2

79 of 149

79

Smoke Check 2

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in ?

b) 20

c) 40

d) 800

a) 2

80 of 149

80

Smoke Check 3

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in total?

b) 20

c) 820

d) 16000

a) 800

81 of 149

Today: Do Something Brave

82 of 149

82

Forward Pass

Layer x

Layer h

Layer

20 parameters need setting

800 parameters need setting

83 of 149

83

Only Have to Do Three Things

2

3

Calculate the log probability for all data

Get partial derivative of log likelihood with respect to each theta

1

Make deep learning assumption

84 of 149

84

Smoke Check

3

Get partial derivative of log likelihood with respect to each theta

Why?

85 of 149

85

Smoke Check

3

Get partial derivative of log likelihood with respect to each theta

Why?

We need to do gradient ascent

86 of 149

86

A deep learning model gets its intelligence by having useful thetas.

We can find useful thetas, by searching for ones that maximize likelihood of our training data

We can maximize likelihood using optimization techniques (such as gradient ascent).

In order to use optimization techniques, we need to calculate the partial derivative of likelihood with respect to thetas.

Why We Calculate Partial Derivatives

Basically MLE is hard because

it has so many details

87 of 149

87

Thanks to Keith Eicher

88 of 149

88

89 of 149

89

Only Have to Do Three Things

2

Calculate the log probability for all data

1

Make deep learning assumption

90 of 149

90

Same Assumption, Same LL

For one datum

For IID data

Take the log

Feel the Bern!

91 of 149

91

Only Have to Do Three Things

2

3

Calculate the log probability for all data

Get partial derivative of log likelihood with respect to each theta

1

Make deep learning assumption

92 of 149

92

Derivative Goals

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

93 of 149

93

Bad Approach

Neural Network

x

h

Math bug

94 of 149

94

Derivatives Without Tears

95 of 149

95

Woah Ms. Forster, you were right.

Chain rule is useful!

First use:

Big Idea #1: Chain Rule

96 of 149

96

Big Idea #2: Sigmoid Derivative

True fact about sigmoid functions

97 of 149

97

Big Idea #3: Derivative of Sum

We only need to calculate the gradient for one training example!

We will pretend we only have one example

We can sum up the gradients of each example to get the correct answer

98 of 149

Recall

99 of 149

99

Sigmoid has a Beautiful Slope

Sigmoid, you should be a ski hill

Chain rule!

Plug and chug

100 of 149

100

Derivative Goals

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

101 of 149

101

Neural Network

x

h

Chain Rule Example 1

Goal

Network

Decomposition

102 of 149

102

Neural Network

x

h

Chain Rule Example 2

Goal

Network

Decomposition

103 of 149

Decomposition

104 of 149

104

Gradient of output layer params

Neural Network

x

h

105 of 149

105

Gradient of output layer params

106 of 149

106

Gradient of output layer params

What! That’s not scary!

where

107 of 149

107

Make it Simple

108 of 149

Boom!

109 of 149

109

Neural Network

x

h

110 of 149

110

Gradient of hidden layer params

Neural Network

x

h

111 of 149

111

Gradient of hidden layer params

Wait is it over?

112 of 149

112

Gradient of hidden layer params

That one too?

113 of 149

113

Make it Simple

114 of 149

115 of 149

Congrats. You now know

Backpropagation

116 of 149

Moment of silence

117 of 149

117

Summary: Simple Calculations For

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

118 of 149

118

Bigger Neural Network

x

h(1)

h(2)

What Would You Do Here?

119 of 149

Chain rule:

Game changer for

artificial intelligence

120 of 149

  • Some data sets/functions are not separable

    • These are classifiers learned by neural networks

120

Neural Networks Can Learn Complex Functions

121 of 149

a

x

b

c

d

e

f

g

LL

Works for any number of layers

122 of 149

1 Trillion Artificial Neurons

GoogLeNet Brain

Piech

123 of 149

123

GoogLeNet Brain

22 layers deep

Multiple,

Multi class output

Piech

124 of 149

Optimal stimulus

by numerical optimization

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Top stimuli from the test set

The Cat Neuron

125 of 149

126 of 149

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 1

Neuron 2

Neuron 3

Neuron 4

Neuron 5

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

127 of 149

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 7

Neuron 8

Neuron 6

Neuron 9

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

128 of 149

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 11

Neuron 10

Neuron 12

Neuron 13

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

129 of 149

22,000 categories

14,000,000 images

Hand-engineered features (SIFT, HOG, LBP),

Spatial pyramid, SparseCoding/Compression

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

ImageNet Classification

130 of 149

smoothhound, smoothhound shark, Mustelus mustelus

American smooth dogfish, Mustelus canis

Florida smoothhound, Mustelus norrisi

whitetip shark, reef whitetip shark, Triaenodon obseus

Atlantic spiny dogfish, Squalus acanthias

Pacific spiny dogfish, Squalus suckleyi

hammerhead, hammerhead shark

smooth hammerhead, Sphyrna zygaena

smalleye hammerhead, Sphyrna tudes

shovelhead, bonnethead, bonnet shark, Sphyrna tiburo

angel shark, angelfish, Squatina squatina, monkfish

electric ray, crampfish, numbfish, torpedo

smalltooth sawfish, Pristis pectinatus

guitarfish

roughtail stingray, Dasyatis centroura

butterfly ray

eagle ray

spotted eagle ray, spotted ray, Aetobatus narinari

cownose ray, cow-nosed ray, Rhinoptera bonasus

manta, manta ray, devilfish

Atlantic manta, Manta birostris

devil ray, Mobula hypostoma

grey skate, gray skate, Raja batis

little skate, Raja erinacea

Stingray

Mantaray

22,000 is a lot!

131 of 149

0.005%

Random guess

1.5%

?

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Pre Neural Networks

GoogLeNet

132 of 149

0.005%

Random guess

1.5%

Pre Neural Networks

43.9%

GoogLeNet

Szegedy et al, Going Deeper With Convolutions, CVPR 2015

133 of 149

0.005%

Random guess

1.5%

Pre Neural Networks

95.1%

SE-ResNet

134 of 149

How many parameters

is too many?

135 of 149

  • Goal of machine learning: build models that generalize well to predicting new data
    • “Overfitting”: fitting the training data too well, so we lose generality of model

    • Polynomial on the right fits training data perfectly!
    • Which would you rather use to predict a new data point?

135

Good ML = Generalization

136 of 149

136

Dropout when your model is training, randomly turn off your neurons with probability 0.5. It will make your network more robust.

Prevent Overfitting?

Piech, CS106A, Stanford University

137 of 149

137

Convolution it turns out if you want to force some of your weights to be shared for different neurons, the math isn’t that much harder. This is used a lot for vision (CNN).

Shared Weights?

Piech, CS106A, Stanford University

138 of 149

Not everything is classification

Come to Monday’s Lecture!

Piech, CS106A, Stanford University

139 of 149

https://www.nytimes.com/2021/07/20/technology/ai-education-neural-networks.html

140 of 149

Human-Level Accuracy

Rubric Level Accuracy on Few-Shot Grading a Novel Question

141 of 149

General Exam Grading Model

Student answer (as text)

Question

(as text)

Rubric

(as text)

Education Tasks

General Exam Embedding

Code embedding

Question embedding

Rubric embedding

142 of 149

Invented the Proto-Transformer

143 of 149

Human-Level Accuracy

Rubric Level Accuracy on Few-Shot Grading a Novel Question

144 of 149

Gave Feedback to 3,500 Real Students

Do you agree? AI feedback 97.9%. Human feedback 96.7%

145 of 149

146 of 149

But what about interactive, creative assignments?

1M ungraded code.org assignments.

The AI is shown a brand new student game. Does it work?

Simultaneously learn to grade and play to grade.

Majority class: 50%

Code-as-text: 67%

Play-to-grade: 94%

147 of 149

Piech

148 of 149

Detecting skin cancer

148

Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.

Piech

149 of 149

Piech