1 of 152

Deep Learning

Chris Gregg

CS109, Stanford University

Summer 2026

2 of 152

Innovations in deep learning

  • Deep learning (neural networks) is the core idea driving the current revolution in AI.

2

AlphaGO (2016)

Notes:

3 of 152

Self Driving Cars

4 of 152

Computers making art

4

5 of 152

Detecting skin cancer

5

Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.

6 of 152

Our Little Buddy

6

http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo.html

7 of 152

  • A neuron

  • Your brain

7

x1

x2

x3

x4

θ1

θ2

θ3

θ4

y

x1

x2

x3

x4

Biological Basis for Neural Networks

Actually, it’s probably someone else’s brain

8 of 152

Logistics

8

9 of 152

End of Class Schedule

9

Week

Monday

Wed

Fri

Next

(May 25th)

Memorial Day

(no class)

Ethics

Regression

Last Week

(June 1st)

Application / Practice

Beyond CS109

Final PEP

ML PSet

Review PSet

10 of 152

Final PEP Sign Ups

10

  • Personal Exam Prep (PEP) for the final is going to be during the beginning of week 10 (by Wednesday)
  • Signups will be live next week (look for an Ed Announcement)
  • Similar to the midterm, with one homework question, which will be from either pset5 or pset6
  • You must do a final PEP as part of your grade
  • We will assume that you are caught up on the material through today. You don’t need to solve everything perfectly for a good grade on the PEP, but you need to be able to engage meaningfully.

11 of 152

Review

11

12 of 152

Classification Task

12

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

0

1

1

0

1

1

1

0

0

0

0

1

13 of 152

Machine Learning

13

(inputs)

(prediction)

(model)

14 of 152

The Training / Testing Paradigm

14

Deployment

15 of 152

The Training / Testing Paradigm

15

Deployment

If your model passes testing…

Learn your parameters

Make sure that they work

16 of 152

16

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

17 of 152

17

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

18 of 152

18

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

19 of 152

19

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0

x:

1

1

20 of 152

20

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0.817

0

x:

1

1

21 of 152

21

+

z = 2.1

σ(z) = 0.7

Logistic Regression

0.817

0

x:

1

1

22 of 152

22

Math for Logistic Regression

1

2

3

Make logistic regression assumption

Calculate the log likelihood for all data

Get derivative of log likelihood with respect to thetas

Often call this

1

23 of 152

Initialize: θj = 0 for all 0 ≤ jm

23

Logistic Regression Training

Repeat many times:

For each parameter j

For each training example (x, y):

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

gradient[j]

24 of 152

24

+

Training

Dataset likelihood:

Training iterations

Likelihood

25 of 152

25

Artificial Neurons

26 of 152

Last Class: Comparing Classifiers

27 of 152

Comparing Classifiers: Test Accuracy

27

Model Train Accuracy Test Accuracy

-------------------------------------------------------------

Baseline 0.6031 0.5887

Naive Bayes 0.7909 0.8067

Logistic Regression 0.8169 0.8307

Decision Tree 0.8514 0.8307

Random Forest 0.8726 0.8500

Gradient Boosting 0.8611 0.8440

AdaBoost 0.8334 0.8353

BayesNet 0.8320 0.8507

28 of 152

Comparing Classifiers: Calibration

28

29 of 152

Comparing Classifiers: Precision / Recall Curve

29

Dog classifier. Image credit:

wikipedia

30 of 152

End Review

30

31 of 152

Predicting a Categorical

31

32 of 152

Logistic Regression to Predict a Categorical?

32

+

z = 2.1

σ(z) = 0.7

33 of 152

Logistic Regression to Predict a Categorical?

33

Standard Logistic Regression

Multi-Class Logistic Regression

34 of 152

Logistic Regression to Predict a Categorical

34

35 of 152

35

Softmax is a generalization of the sigmoid function that squashes a K -dimensional vector z of arbitrary real values to a K-dimensional vector softmax(z) of real values in the range [0, 1] that add up to 1.

Categorical Classification?

Sigmoid is to Bernoulli as Softmax is to Categorical

36 of 152

Understanding Softmax

36

List: 7 7 7 7

[0.25, 0.25, 0.25, 0.25]

List: 2 8 5 6

[0.0, 0.84, 0.04, 0.11]

37 of 152

Why not this?

37

Softmax

Normalized Logistic

(oh…not convex)

38 of 152

Softmax is Related to Sigma Function!

38

In softmax with two classes,

39 of 152

What is Log Likelihood?

39

40 of 152

What is Log Likelihood?

40

41 of 152

We are ready…

42 of 152

  • A neuron

  • Your brain

42

x1

x2

x3

x4

θ1

θ2

θ3

θ4

y

x1

x2

x3

x4

Biological Basis for Neural Networks

43 of 152

Core idea behind the revolution in AI

44 of 152

44

Deep learning is (at its core) many logistic regression pieces stacked on top of each other.

(aka Neural Networks)

45 of 152

45

Digit Recognition Example

Let’s make feature vectors from pictures of numbers

46 of 152

Computer Vision

47 of 152

Hundreds of millions of neurons [1]

Visual neurons make up up 30% of your cortex [1]

Vision in your Brain

[1] http://discovermagazine.com/1993/jun/thevisionthingma227

48 of 152

48

Logistic Regression

This means it predicts a 0

This is one neuron…your brain has hundreds of millions

49 of 152

49

Logistic Regression

This means it predicts a 0

Indicates logistic regression connection

50 of 152

50

Logistic Regression

This means it predicts a 1

51 of 152

51

Not So Good

This means it predicts a 1

52 of 152

52

We Can Put Neurons Together

This means it predicts a 0

53 of 152

53

We Can Put Neurons Together

Look at a single “hidden” neuron

This means it predicts a 0

There is a parameter for every connection

54 of 152

54

We Can Put Neurons Together

Look at another “hidden” neuron

This means it predicts a 0

There is a parameter for every connection

55 of 152

55

We Can Put Neurons Together

This means it predicts a 0

56 of 152

56

We Can Put Neurons Together

This means it predicts a 0

Look at another neuron

There is a parameter for every connection

57 of 152

57

We Can Put Neurons Together

This means it predicts a 0

58 of 152

58

We Can Put Neurons Together

*lots

+

59 of 152

59

We Can Put Neurons Together

+

Single neuron in a hidden layer

60 of 152

60

We Can Put Neurons Together

*lots

+

61 of 152

61

We Can Put Neurons Together

*lots

*logistic

regression

62 of 152

Deep learning

def Deep learning is

maximum likelihood estimation with neural networks.

62

 

 

> 0.5?

Yes.�Predict 1

Lots of Logistic (regressions)

LOL

  • def A neural network is
  • (at its core) many logistic regression pieces stacked on top of each other.

63 of 152

63

Demonstration

https://adamharley.com/nn_vis/cnn/2d.html

64 of 152

64

Deep learning gets its intelligence from its thetas (aka its parameters)

65 of 152

Lets Build a Neural Network!

65

66 of 152

How do we train?

67 of 152

MLE of Thetas!

68 of 152

First: Learning Goals…

69 of 152

1. Understand Chain Rule

as ♡ of Deep Learning

70 of 152

2. Demystify:

Deep Learning is MLE

71 of 152

3. Become experts of

logistic regression

72 of 152

Math worth knowing:

73 of 152

73

New Notation

Layer x

Layer h

Layer

74 of 152

74

New Notation

Layer x

Layer h

Layer

75 of 152

75

New Notation

Layer x

Layer h

Layer

76 of 152

76

Forward Pass

Layer x

Layer h

Layer

77 of 152

77

Forward Pass

Layer x

Layer h

Layer

78 of 152

78

Forward Pass

Layer x

Layer h

Layer

79 of 152

79

Forward Pass

Layer x

Layer h

Layer

80 of 152

80

All Together

Neural Network

x

h

81 of 152

81

Smoke Check 1

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in ?

b) 20

c) 40

d) 800

a) 2

82 of 152

82

Smoke Check 2

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in ?

b) 20

c) 40

d) 800

a) 2

83 of 152

83

Smoke Check 3

Neural Network

x

h

|x| = 40

|h| = 20

How many parameters in total?

b) 20

c) 820

d) 16000

a) 800

84 of 152

Today: Do Something Brave

85 of 152

85

Forward Pass

Layer x

Layer h

Layer

20 parameters need setting

800 parameters need setting

86 of 152

86

Only Have to Do Three Things

2

3

Calculate the log probability for all data

Get partial derivative of log likelihood with respect to each theta

1

Make deep learning assumption

87 of 152

87

Smoke Check

3

Get partial derivative of log likelihood with respect to each theta

Why?

88 of 152

88

Smoke Check

3

Get partial derivative of log likelihood with respect to each theta

Why?

We need to do gradient ascent

89 of 152

89

A deep learning model gets its intelligence by having useful thetas.

We can find useful thetas, by searching for ones that maximize likelihood of our training data

We can maximize likelihood using optimization techniques (such as gradient ascent).

In order to use optimization techniques, we need to calculate the partial derivative of likelihood with respect to thetas.

Why We Calculate Partial Derivatives

Basically MLE is hard because

it has so many details

90 of 152

90

Thanks to Keith Eicher

91 of 152

91

92 of 152

92

Only Have to Do Three Things

2

Calculate the log probability for all data

1

Make deep learning assumption

93 of 152

93

Same Assumption, Same LL

For one datum

For IID data

Take the log

Feel the Bern!

94 of 152

94

Only Have to Do Three Things

2

3

Calculate the log probability for all data

Get partial derivative of log likelihood with respect to each theta

1

Make deep learning assumption

95 of 152

95

Derivative Goals

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

96 of 152

96

Bad Approach

Neural Network

x

h

Math bug

97 of 152

97

Derivatives Without Tears

98 of 152

98

Woah Ms. Forster, you were right.

Chain rule is useful!

First use:

Big Idea #1: Chain Rule

99 of 152

99

Big Idea #2: Sigmoid Derivative

True fact about sigmoid functions

100 of 152

100

Big Idea #3: Derivative of Sum

We only need to calculate the gradient for one training example!

We will pretend we only have one example

We can sum up the gradients of each example to get the correct answer

101 of 152

Recall

102 of 152

102

Sigmoid has a Beautiful Slope

Sigmoid, you should be a ski hill

Chain rule!

Plug and chug

103 of 152

103

Derivative Goals

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

104 of 152

104

Neural Network

x

h

Chain Rule Example 1

Goal

Network

Decomposition

105 of 152

105

Neural Network

x

h

Chain Rule Example 2

Goal

Network

Decomposition

106 of 152

Decomposition

107 of 152

107

Gradient of output layer params

Neural Network

x

h

108 of 152

108

Gradient of output layer params

109 of 152

109

Gradient of output layer params

What! That’s not scary!

where

110 of 152

110

Make it Simple

111 of 152

Boom!

112 of 152

112

Neural Network

x

h

113 of 152

113

Gradient of hidden layer params

Neural Network

x

h

114 of 152

114

Gradient of hidden layer params

Wait is it over?

115 of 152

115

Gradient of hidden layer params

That one too?

116 of 152

116

Make it Simple

117 of 152

118 of 152

Congrats. You now know

Backpropagation

119 of 152

Moment of silence

120 of 152

120

Summary: Simple Calculations For

Neural Network

x

h

Loss with respect to

output layer params

Loss with respect to

hidden layer params

121 of 152

121

Bigger Neural Network

x

h(1)

h(2)

What Would You Do Here?

122 of 152

Chain rule:

Game changer for

artificial intelligence

123 of 152

  • Some data sets/functions are not separable

    • These are classifiers learned by neural networks

123

Neural Networks Can Learn Complex Functions

124 of 152

a

x

b

c

d

e

f

g

LL

Works for any number of layers

125 of 152

1 Trillion Artificial Neurons

GoogLeNet Brain

Piech

126 of 152

126

GoogLeNet Brain

22 layers deep

Multiple,

Multi class output

Piech

127 of 152

Optimal stimulus

by numerical optimization

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Top stimuli from the test set

The Cat Neuron

128 of 152

129 of 152

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 1

Neuron 2

Neuron 3

Neuron 4

Neuron 5

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

130 of 152

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 7

Neuron 8

Neuron 6

Neuron 9

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

131 of 152

Pooling Size = 5

Number

of maps = 8

Image Size = 200

Number of output

channels = 8

Number of input

channels = 3

One layer

RF size = 18

Input to another layer above

(image with 8 channels)

W

H

LCN Size = 5

Neuron 11

Neuron 10

Neuron 12

Neuron 13

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Best Neuron Stimuli

132 of 152

22,000 categories

14,000,000 images

Hand-engineered features (SIFT, HOG, LBP),

Spatial pyramid, SparseCoding/Compression

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

ImageNet Classification

133 of 152

smoothhound, smoothhound shark, Mustelus mustelus

American smooth dogfish, Mustelus canis

Florida smoothhound, Mustelus norrisi

whitetip shark, reef whitetip shark, Triaenodon obseus

Atlantic spiny dogfish, Squalus acanthias

Pacific spiny dogfish, Squalus suckleyi

hammerhead, hammerhead shark

smooth hammerhead, Sphyrna zygaena

smalleye hammerhead, Sphyrna tudes

shovelhead, bonnethead, bonnet shark, Sphyrna tiburo

angel shark, angelfish, Squatina squatina, monkfish

electric ray, crampfish, numbfish, torpedo

smalltooth sawfish, Pristis pectinatus

guitarfish

roughtail stingray, Dasyatis centroura

butterfly ray

eagle ray

spotted eagle ray, spotted ray, Aetobatus narinari

cownose ray, cow-nosed ray, Rhinoptera bonasus

manta, manta ray, devilfish

Atlantic manta, Manta birostris

devil ray, Mobula hypostoma

grey skate, gray skate, Raja batis

little skate, Raja erinacea

Stingray

Mantaray

22,000 is a lot!

134 of 152

0.005%

Random guess

1.5%

?

Le, et al., Building high-level features using large-scale unsupervised learning. ICML 2012

Pre Neural Networks

GoogLeNet

135 of 152

0.005%

Random guess

1.5%

Pre Neural Networks

43.9%

GoogLeNet

Szegedy et al, Going Deeper With Convolutions, CVPR 2015

136 of 152

0.005%

Random guess

1.5%

Pre Neural Networks

95.1%

SE-ResNet

137 of 152

How many parameters

is too many?

138 of 152

  • Goal of machine learning: build models that generalize well to predicting new data
    • “Overfitting”: fitting the training data too well, so we lose generality of model

    • Polynomial on the right fits training data perfectly!
    • Which would you rather use to predict a new data point?

138

Good ML = Generalization

139 of 152

139

Dropout when your model is training, randomly turn off your neurons with probability 0.5. It will make your network more robust.

Prevent Overfitting?

Piech, CS106A, Stanford University

140 of 152

140

Convolution it turns out if you want to force some of your weights to be shared for different neurons, the math isn’t that much harder. This is used a lot for vision (CNN).

Shared Weights?

Piech, CS106A, Stanford University

141 of 152

Not everything is classification

Come to Monday’s Lecture!

Piech, CS106A, Stanford University

142 of 152

https://www.nytimes.com/2021/07/20/technology/ai-education-neural-networks.html

143 of 152

Human-Level Accuracy

Rubric Level Accuracy on Few-Shot Grading a Novel Question

144 of 152

General Exam Grading Model

Student answer (as text)

Question

(as text)

Rubric

(as text)

Education Tasks

General Exam Embedding

Code embedding

Question embedding

Rubric embedding

145 of 152

Invented the Proto-Transformer

146 of 152

Human-Level Accuracy

Rubric Level Accuracy on Few-Shot Grading a Novel Question

147 of 152

Gave Feedback to 3,500 Real Students

Do you agree? AI feedback 97.9%. Human feedback 96.7%

148 of 152

149 of 152

But what about interactive, creative assignments?

1M ungraded code.org assignments.

The AI is shown a brand new student game. Does it work?

Simultaneously learn to grade and play to grade.

Majority class: 50%

Code-as-text: 67%

Play-to-grade: 94%

150 of 152

Piech

151 of 152

Detecting skin cancer

151

Esteva, Andre, et al. "Dermatologist-level classification of skin cancer with deep neural networks." Nature 542.7639 (2017): 115-118.

Piech

152 of 152

Piech