1 of 178

Logistic Regression

Chris Gregg

CS109, Stanford University

Summer, 2026

2 of 178

2

3 of 178

3

4 of 178

Where are we in CS109?

5 of 178

Where are we in CS109?

6 of 178

Earlier the start, the better, for PSet #6

6

7 of 178

Earlier the start, the better, for PSet #6

7

8 of 178

Earlier the start, the better, for PSet #6

8

9 of 178

What will be on Pset #7?

Review and Applications!

10 of 178

Background

11 of 178

11

Background: Sigmoid Function

The sigmoid function squashes z to be a number between 0 and 1

Caution:

different use of sigma than normal std

12 of 178

12

Background: Key Notation

Sigmoid function

Weighted sum

(aka dot product)

Sigmoid function of

weighted sum

13 of 178

13

Background: Chain Rule

Who knew calculus would be so useful?

Aka decomposition of composed functions

14 of 178

Machine Learning

(aka Applied Probability)

15 of 178

15

Machine Learning in CS109

Parameter Estimation

Logistic

Regression

Neural Networks

Theory

Core Algorithms

Great Idea

MLE

16 of 178

MLE Idea: Chose params that make the data look likely

16

17 of 178

Likelihood Definition

17

A generalized term for “PDF / PMF / Joint”

of data as a function of parameters

Wikipedia:

18 of 178

18

Define likelihood, use independence.

Define the loglikelihood

Use LL to chose params

Chose the params

That maximize likelihood

Maximum Likelihood

19 of 178

MLE for a Pareto

19

We know sand is distributed as a pareto with PDF

20 of 178

20

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Xi ~ Pareto(α). Use Maximum Likelihood to estimate α.
    • Likelihood:

    • Log-likelihood:

    • Chose α to be the argmax of LL:

MLE for a Pareto

3. Find the value of α which maximizes log likelihood

1. What is the likelihood of all the data

2. What is the log-likelihood all the data

21 of 178

21

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Xi ~ Pareto(α). Use Maximum Likelihood to estimate α.
    • Likelihood:

    • Log-likelihood:

    • Chose α to be the argmax of LL:

MLE for a Pareto

3. Find the value of α which maximizes log likelihood

2. What is the log-likelihood all the data

22 of 178

22

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Xi ~ Pareto(α). Use Maximum Likelihood to estimate α.
    • Likelihood:

    • Log-likelihood:

    • Chose α to be the argmax of LL:

MLE for a Pareto

3. Find the value of α which maximizes log likelihood

23 of 178

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Xi ~ Pareto(α). Use Maximum Likelihood to estimate α.
    • Likelihood:

    • Log-likelihood:

    • Chose α to be the argmax of LL:

23

MLE for a Pareto

Argmax Option #1: set the derivative to 0, and solve for alpha

24 of 178

 

25 of 178

25

Gradient Ascent

Walk uphill and you will find a local maxima

(if your step size is small enough)

Especially good if function is convex

26 of 178

26

Gradient Ascent

Repeat many times

Walk uphill and you will find a local maxima

(if your step size is small enough)

This is some profound life philosophy

Step size constant

27 of 178

Initialize: θj = random for all 0 ≤ jm

27

Gradient Ascent

Repeat many times:

Calculate all gradient[j]s based on data

𝜃j += η * gradient[j] for all 0 ≤ jm

28 of 178

MLE of Erlang

28

Time to finish Medical Diagnosis in Seconds

29 of 178

End Review

Core Algorithms

30 of 178

Warmup

Core Algorithms

31 of 178

MLE for Bernoulli

31

32 of 178

Don’t we already have the Beta?

Yes! But this example is critical for developing towards deep learning.

33 of 178

Maximum Likelihood with Bernoulli

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Use Maximum Likelihood to estimate p
    • Probability mass function can be written as:

33

34 of 178

Differentiable PMF for Bernoulli

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Xi ~ Ber(p)
    • Probability mass function,

34

0

1

p

1 - p

PMF of Bernoulli

PMF of Bernoulli (p = 0.2)

35 of 178

35

Bernoulli PMF

36 of 178

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Use Maximum Likelihood to estimate p
    • Probability mass function can be written as:

    • Likelihood:
    • Log-likelihood:

36

36

1. What is the likelihood of one Xi

2. What is the likelihood of all the data

3. What is the log-likelihood all the data

4. Find the value of p which maximizes log likelihood

Maximum Likelihood For Bernoulli

37 of 178

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Use Maximum Likelihood to estimate p
    • Probability mass function can be written as:

    • Likelihood:
    • Log-likelihood:

37

2. What is the likelihood of all the data

3. What is the log-likelihood all the data

4. Find the value of p which maximizes log likelihood

Maximum Likelihood For Bernoulli

38 of 178

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Use Maximum Likelihood to estimate p
    • Probability mass function can be written as:

    • Likelihood:
    • Log-likelihood:

38

38

3. What is the log-likelihood all the data

4. Find the value of p which maximizes log likelihood

Maximum Likelihood For Bernoulli

39 of 178

  • Consider I.I.D. random variables X1, X2, ..., Xn
    • Use Maximum Likelihood to estimate p
    • Probability mass function can be written as:

    • Likelihood:
    • Log-likelihood:

39

39

4. Find the value of p which maximizes log likelihood

Maximum Likelihood For Bernoulli

40 of 178

Take Derivative:

40

41 of 178

Take Derivative:

41

Take the derivative wrt p

42 of 178

Take Derivative:

42

Take the derivative wrt p

Derivative of a sum!

43 of 178

Take Derivative:

43

Take the derivative wrt p

Derivative of a sum!

Derivative of a sum!

44 of 178

Take Derivative:

44

Take the derivative wrt p

Derivative of a sum!

Derivative of a sum!

45 of 178

Take Derivative:

45

Take the derivative wrt p

Derivative of a sum!

Derivative of log p

Derivative of a sum!

46 of 178

Take Derivative:

46

Take the derivative wrt p

Derivative of a sum!

Derivative of log p

Derivative of a sum!

47 of 178

Take Derivative:

47

Take the derivative wrt p

Derivative of a sum!

Derivative of log p

Derivative of a sum!

Derivative of log (1-p)

48 of 178

Set to Zero:

48

49 of 178

Set to Zero:

49

50 of 178

Set to Zero:

50

Let

To make life easier

51 of 178

Set to Zero:

51

Let

And

To make life easier

52 of 178

Set to Zero:

52

Let

And

To make life easier

53 of 178

Set to Zero:

53

Let

And

To make life easier

54 of 178

Set to Zero:

54

Let

And

To make life easier

55 of 178

Set to Zero:

55

Let

And

To make life easier

56 of 178

Set to Zero:

56

Let

And

To make life easier

57 of 178

Set to Zero:

57

Let

And

To make life easier

58 of 178

Isn’t that the same as

unbiased estimator?

59 of 178

Yes. For Bernoulli.

The ML estimate of p of a Bernoulli is the same as the sample mean.

60 of 178

MLE vs Beta

60

The medicine is tried on 20 patients. It “works” for 14 and “doesn’t work” for 6. What is your new belief that the drug works?

In other words I have 20 IID samples from a Bernoulli. Estimate p. The data is [1,1,1,1,1,1,1,1,1,1,1,1,1,1,0,0,0,0,0,0]

MLE estimate:

Prior

Posterior

mode

Beta estimate:

61 of 178

61

Think about the difference between a point estimate and a distribution

p = 0.75

p =

62 of 178

Param Estimation With a Prior == Inference

62

We know sand is distributed as a pareto with PDF

Prior:

63 of 178

63

Maximum A Posteriori (MAP)

1

Infer the posterior belief in params

2

Return the argmax

(aka the mode)

64 of 178

You need to know MLE and Inference.

MAP is something you should recognize!

65 of 178

It is time….

66 of 178

Today a very special MLE problem

67 of 178

Classification

68 of 178

68

Example Datasets

Heart

Ancestry

Netflix

69 of 178

69

Training Data

Assume IID data:

n training datapoints

Each datapoint has m features and a single output

Training Data: assignments all random variables X and Y

70 of 178

70

Single Feature Value

User 1

User 2

User n

Movie 1

Movie 2

Movie m

Output

1

0

1

1

1

1

0

0

0

0

1

1

71 of 178

71

Single Feature Value

1

0

1

1

1

1

0

0

0

0

1

1

User 1

User 2

User n

Movie 1

Movie 2

Movie m

Output

72 of 178

72

Single Feature Value

1

0

1

1

1

1

0

0

0

0

1

1

User 1

User 2

User n

Movie 1

Movie 2

Movie m

Output

73 of 178

73

Single Feature Value

1

0

1

1

1

1

0

0

0

0

1

1

User 1

User 2

User n

Movie 1

Movie 2

Movie m

Output

74 of 178

74

Single Feature Value

1

0

1

1

1

1

0

0

0

0

1

1

In general:

In this case:

User 1

User 2

User n

Movie 1

Movie 2

Movie m

Output

75 of 178

75

Healthy Heart Classifier

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

0

1

1

0

1

1

1

0

0

0

0

1

76 of 178

76

Healthy Heart Classifier

0

1

1

0

1

1

1

0

0

0

0

1

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

77 of 178

77

Healthy Heart Classifier

0

1

1

0

1

1

1

0

0

0

0

1

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

78 of 178

78

Healthy Heart Classifier

0

1

1

0

1

1

1

0

0

0

0

1

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

79 of 178

79

Healthy Heart Classifier

0

1

1

0

1

1

1

0

0

0

0

1

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

80 of 178

80

Ancestry Classifier

User 1

User 2

User n

SNP 1

SNP 2

SNP m

Output

1

0

1

0

0

0

1

1

1

1

0

1

81 of 178

81

Aside Predicting Real Numbers is Regression

Game 1

Game 2

Game n

Opposing team

ELO

Points in last game

At Home?

Output

84

105

1

120

90

102

0

95

74

120

0

115

# Points

82 of 178

82

Ancestry Classifier

User 1

User 2

User n

SNP 1

SNP 2

SNP m

Output

1

0

1

0

0

0

1

1

1

1

0

1

83 of 178

83

Still Classification

User 1

User 2

User n

SNP 1

SNP 2

SNP m

Output

15

0

2

0

0

0

1

1

7

20

0

1

84 of 178

Classification

85 of 178

85

Healthy Heart Classifier

Heart 1

Heart 2

Heart n

ROI 1

ROI 2

ROI m

Output

0

1

1

0

1

1

1

0

0

0

0

1

86 of 178

86

Classification is Building a Harry Potter Hat

87 of 178

Machine Learning for Classification

87

(inputs)

(prediction)

(model)

88 of 178

Machine Learning for Classification

88

(inputs)

(prediction)

(model)

89 of 178

89

Logistic Regression

90 of 178

Chapter 1: Big Picture

91 of 178

  • In classification we care about

91

From Naïve Bayes to Logistic Regression

Let’s build a machine that can you can put x into, which then spits out

92 of 178

  • Could we compute via a machine?
  • Welcome our friend: logistic regression!

92

0.81

[1, 1, 0]

Logistic Regression Assumption

93 of 178

  • Could we compute via a machine?
  • Welcome our friend: logistic regression!

93

0.81

[1, 1, 0]

Logistic Regression Assumption

94 of 178

94

Logistic Regression Assumption

95 of 178

95

Logistic Regression

+

z = 2.1

σ(z) = 0.7

96 of 178

96

Logistic Regression

97 of 178

97

+

z = 2.1

σ(z) = 0.7

Logistic Regression Cartoon

98 of 178

98

+

Inputs x = [0, 1, 1]

z = 2.1

σ(z) = 0.7

99 of 178

99

+

Inputs

z = 2.1

σ(z) = 0.7

100 of 178

100

+

Inputs + Output

z = 2.1

σ(z) = 0.7

101 of 178

101

+

Inputs

z = 2.1

σ(z) = 0.7

102 of 178

102

+

Weights

z = 2.1

σ(z) = 0.7

103 of 178

103

+

Weighed Sum

z = 2.1

σ(z) = 0.7

104 of 178

104

+

Squashing Function

z = 2.1

σ(z) = 0.7

105 of 178

105

+

Prediction

z = 2.1

σ(z) = 0.7

106 of 178

106

+

z = 2.1

σ(z) = 0.7

Parameters Affect Prediction

107 of 178

107

+

z = -1.5

σ(z) = 0.4

Parameters Affect Prediction

108 of 178

108

+

z = 2.1

σ(z) = 0.7

Parameters Affect Prediction

109 of 178

109

+

Different Predictions for Different Inputs

z = 2.1

σ(z) = 0.7

110 of 178

110

+

Different Predictions for Different Inputs

z = 2.1

σ(z) = 0.7

111 of 178

111

+

Different Predictions for Different Inputs

z = -1.9

σ(z) = 0.3

112 of 178

  • Model conditional likelihood
  • with logistic function:

    • For simplicity define so
    • Since P(Y = 0 | X = x) + P(Y = 1 | X = x) = 1:

112

Recall:

Sigmoid function

Handling the Intercept

113 of 178

113

Logistic Regression Assumption:

Big Assumption

114 of 178

Note: inflection point at z = 0. f (0) = 0.5

114

Want to distinguish y = 1 (blue) points from y = 0 (red) points

The Sigmoid Function

115 of 178

115

What is in a Name

Classification Algorithms

Regression Algorithms

Linear Regression

Logistic Regression

If Chris could rename it he would call it: Sigmoidal Classification

Awesome classifier, terrible name

Decision Tree Classifier

116 of 178

116

What makes for a “smart”

logistic regression algorithm?

117 of 178

117

Logistic regression gets its intelligence from its thetas (aka its parameters)

118 of 178

118

+

z = -1.5

σ(z) = 0.4

Let’s say that:

y = 1

Data looks unlikely

How Do We Learn Parameters?

119 of 178

119

+

z = -1.5

σ(z) = 0.4

How Do We Learn Parameters?

Let’s say that:

y = 1

Data looks unlikely

120 of 178

120

+

z = 2.1

σ(z) = 0.9

How Do We Learn Parameters?

y = 1

Data is much more likely!

Let’s say that:

121 of 178

121

Maximum Likelihood Estimation

Remember this?

122 of 178

122

Let’s show you the big picture,

then we can derive it!

123 of 178

123

Math for Logistic Regression

1

2

3

Make logistic regression assumption

Calculate the log likelihood for all data

Get derivative of log likelihood with respect to thetas

Often call this

1

124 of 178

124

Gradient Ascent

Walk uphill and you will find a local maxima

(if your step size is small enough)

Logistic regression LL function is convex

125 of 178

125

Gradient Ascent Step

Do this

for all

thetas!

𝜃1

𝜃2

LL(𝜃)

126 of 178

What does this look like in code?

127 of 178

Initialize: θj = 0 for all 0 ≤ jm

127

Real Code!!!

Repeat many times:

For each training example (x, y):

For each parameter j:

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

gradient[j]

128 of 178

Step by Step

129 of 178

Initialize: θj = 0 for all 0 ≤ jm

129

Logistic Regression Training

Calculate all θj

130 of 178

Initialize: θj = 0 for all 0 ≤ jm

130

Logistic Regression Training

Repeat many times:

Calculate all gradient[j]s based on data

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

131 of 178

Initialize: θj = 0 for all 0 ≤ jm

131

Logistic Regression Training

Repeat many times:

For each training example (x, y):

For each parameter j:

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

Update gradient[j] for current training example (x, y)

132 of 178

Initialize: θj = 0 for all 0 ≤ jm

132

Logistic Regression Training

Repeat many times:

For each training example (x, y):

For each parameter j:

𝜃j += η * gradient[j] for all 0 ≤ jm

gradient[j] = 0 for all 0 ≤ jm

gradient[j]

133 of 178

133

+

Training

Dataset likelihood:

Likelihood

Training iterations

134 of 178

134

+

Training

Dataset likelihood:

Training iterations

Likelihood

135 of 178

135

+

Training

Dataset likelihood:

Training iterations

Likelihood

136 of 178

136

+

Training

Dataset likelihood:

Training iterations

Likelihood

137 of 178

137

+

Training

Dataset likelihood:

Training iterations

Likelihood

138 of 178

138

+

Training

Dataset likelihood:

Training iterations

Likelihood

139 of 178

139

Don’t forget:

xj is j-th input variable and x0 = 1.

Allows for θ0 to be an intercept.

140 of 178

140

+

z = 2.1

σ(z) = 0.7

Prediction

0

x:

1

1

141 of 178

141

+

z = 2.1

σ(z) = 0.7

0

x:

1

1

Prediction

142 of 178

142

+

z = 2.1

σ(z) = 0.7

0

x:

1

1

Prediction

143 of 178

143

+

z = 2.1

σ(z) = 0.7

0

x:

1

1

Prediction

0.817

144 of 178

144

+

z = 2.1

σ(z) = 0.7

0.817

0

x:

1

1

Prediction

145 of 178

145

Live Demo!!

146 of 178

Chapter 2: How Come?

147 of 178

147

Logistic Regression

1

2

3

Make logistic regression assumption

Calculate the log probability for all data

Get derivative of log probability with respect to thetas

Often call this

1

148 of 178

148

How did we get that LL function?

149 of 178

Probability mass function:

149

Recall: PMF of Bernoulli

0

1

p

1 - p

PMF of Bernoulli

PMF of Bernoulli (p = 0.2)

150 of 178

150

Implies

For IID data

Take the log

1

Recall:

151 of 178

151

How did we get that gradient?

152 of 178

152

Sigmoid has a Beautiful Slope

True fact about sigmoid functions

153 of 178

153

Sigmoid has a Beautiful Slope

Sigmoid, you should be a ski hill

Chain rule!

Plug and chug

154 of 178

154

Sigmoid has a Beautiful Slope

155 of 178

ARE YOU READY???

156 of 178

156

I think I’m Ready…

Where

157 of 178

157

Think About Only One Training Instance

We only need to calculate the gradient for one training example!

We will pretend we only have one example

We can sum up the gradients of each example to get the correct answer

158 of 178

158

First, imagine only one example

Where

CHAIN RULE

Already did that one

Derive this one

Simplify

159 of 178

159

Now, all the data

See last slide

Derivative of sum…

Some people don’t like hats…

160 of 178

160

Logistic Regression

1

2

3

Make logistic regression assumption

Calculate the log probability for all data

Get derivative of log probability with respect to thetas

1

161 of 178

161

The Hard Way

162 of 178

Phew!

163 of 178

Chapter 3: Philosophy

(if time)

164 of 178

Neuron

165 of 178

Neuron

166 of 178

Neuron

167 of 178

Neuron

168 of 178

Neuron

169 of 178

Some inputs are more important

170 of 178

Artificial Neurons

171 of 178

  • A neuron

  • Your brain

x1

x2

x3

x4

θ1

θ2

θ3

θ4

y

x1

x2

x3

x4

Biological Basis for Neural Networks

172 of 178

Deep learning is (at its core) many logistic regression pieces stacked on top of each other.

(aka Neural Networks)

173 of 178

Computer Vision

174 of 178

Alpha GO

175 of 178

Revolution in AI

176 of 178

Computers Making Art

177 of 178

177

178 of 178

Basically just many logistic regression cells

And lots of chain rule…