1
Parameter Estimation
Chris Gregg
CS109, Stanford University
Summer 2026
Where are we in CS109?
2
You are here
3
4
5
General “Inference”
6
Probabilistic Model
Fever
Tired
Flu
Undergrad
7
Probabilistic Model
Fever
Tired
Flu
Undergrad
If you know the probability of each random variables given the ones that directly cause it, you can joint sample!
8
Four Prototypical Trajectories
But where do those numbers come from?
9
Four Prototypical Trajectories
Suspense
10
Four Prototypical Trajectories
At this point, if you are given a model, with all the involved probabilities, you can make predictions
11
Four Prototypical Trajectories
But what if you want to learn the probabilities in the model?
12
Four Prototypical Trajectories
Machine Learning
13
AI and Machine Learning
ML: Rooted in probability theory
Artificial
Intelligence
Machine
Learning
Deep
Learning
Gen AI
14
Our Path
Parameter Estimation
Deep Learning
Core Algorithms
15
Our Path
Deep Learning
Core Algorithms
Unbiased
estimators
Maximizing
likelihood
16
17
Four Prototypical Trajectories
Review
Shorthand for Equality Events
18
Our shorthand notation
Is shorthand for the event
Is shorthand for the event
Full Notation
19
Four Prototypical Trajectories
End Review
20
21
Once upon a time…
22
…there was parameter estimation
23
What are Parameters?
24
What are Parameters?
25
What are Parameters?
26
What are Parameters?
27
θ = p
θ = λ
θ = (α, β)
θ = (μ, σ2)
θ = (m, b)
What are Parameters?
28
Fever
Tired
Flu
Undergrad
Parameters
What are Parameters?
29
Why Do We Care?
Real World Problem
Formal Model 𝜃
Prediction
Function 𝜃*
Model the problem
Learning Algorithm
Testing
Data
Training Data
Evaluation
score
30
Modelling
Real World Problem
Formal Model 𝜃
Prediction
Function 𝜃*
Model the problem
Learning Algorithm
Testing
Data
Training Data
Evaluation
score
31
Real World Problem
Model the problem
Parameter Estimation (aka Training)
Formal Model 𝜃
Prediction
Function 𝜃*
Learning Algorithm
Testing
Data
Training Data
Evaluation
score
32
We have already seen some
parameter estimators!
33
9 Heads out of 10 Flips. What is your Belief in p?
34
Unbiased Estimators of Mean and Variance
35
Unbiased Estimators of Mean and Variance
36
Unbiased Estimators of Mean and Variance
37
Sample variance:
Unbiased Estimators of Mean and Variance
38
Sample variance:
Unbiased Estimators of Mean and Variance
39
Our Path
Deep Learning
Core Algorithms
Unbiased
estimators
Maximizing
likelihood
40
Unbiased Estimation is a limited tool:
how could we use that for fitting WebMD?
41
Great idea in Machine Learning
42
To the Course Reader!
43
We want to choose the parameter value that maximizes the probability of the data.
How to Choose the “Best” Parameters: MLE
44
We want to choose the parameter value that maximizes the probability of the data.
Maximum
Likelihood
Estimation!
How to Choose the “Best” Parameters: MLE
45
“I feel seen”
46
47
We want to choose the parameter value that maximizes the probability of the data.
Maximum
Likelihood
Estimation!
How do we quantify “probability of the data”?
How to Choose the “Best” Parameters: MLE
48
A generalized term for “PDF / PMF / Joint”
of data as a function of parameters
Wikipedia:
Likelihood Definition
49
θ is shorthand for parameter(s)
(if we have a Poisson, θ = λ)
Definition: The probability of our observed data if our parameters were θ.
The Likelihood Function
50
(in our example, this would just be the Poisson PMF)
If we had a single observation, X = x:
The Likelihood Function
Definition: The probability of our observed data if our parameters were θ.
51
For a list of observations, [x1, x2, …, xn]:
The Likelihood Function
Definition: The joint probability of our observed data if our parameters were θ.
52
For a list of observations, [x1, x2, …, xn]:
The Likelihood Function
Definition: The joint probability of our observed data if our parameters were θ.
We assume that data points are I.I.D.
53
For a list of observations, [x1, x2, …, xn]:
The Likelihood Function
Definition: The joint probability of our observed data if our parameters were θ.
We assume that data points are I.I.D.
54
For a list of observations, [x1, x2, …, xn]:
The Likelihood Function
Definition: The joint probability of our observed data if our parameters were θ.
We assume that data points are I.I.D.
We always use f for likelihood in MLE (even for discrete) 🙃
55
The Likelihood Function
n I.I.D. data points
We explicitly specify parameter θ of distribution
This is just a product since Xi are I.I.D.
56
Likelihood (of data given parameters):
Either the
PDF (continuous) or
PMF (discrete), or
joint if multiple variables per datapoint
57
We want to choose the parameter value that maximizes the probability of the data:
How to Choose the “Best” Parameters: MLE
Likelihood
58
We want to choose the parameter value that maximizes the probability of the data:
How to Choose the “Best” Parameters: MLE
Likelihood
To put words into math:
Our best estimate
59
We want to choose the parameter value that maximizes the probability of the data:
How to Choose the “Best” Parameters: MLE
Likelihood
To put words into math:
Our best estimate
Log Likelihood
60
Sidequest: What is argmax?
61
Argmax
62
63
64
But how do we compute argmax?
65
Option #1: Straight optimization
66
Finding the Argmax with Calculus
67
Differentiate w.r.t.�argmax’s argument
Finding the Argmax with Calculus
68
Differentiate w.r.t.�argmax’s argument
Finding the Argmax with Calculus
69
Differentiate w.r.t.�argmax’s argument
Finding the Argmax with Calculus
70
Differentiate w.r.t.�argmax’s argument
Set to 0 and solve
Finding the Argmax with Calculus
71
Differentiate w.r.t.�argmax’s argument
Set to 0 and solve
Finding the Argmax with Calculus
72
Differentiate w.r.t.�argmax’s argument
Set to 0 and solve
Finding the Argmax with Calculus
73
Argmax of Log
Claim:
Log is monotonic
x ≤ y ⇔ log(x) ≤ log(y) for all x, y > 0
74
Argmax of Log
75
Log I Love You
76
Natural Log
77
End Sidequest
78
Let’s go deeper!
79
MLE For Poisson
80
We observed the following samples:
[6, 1, 2, 1, 2, 3, 3, 2, 1, 3, 1, 3]
What is lambda, ?
λ
MLE For Poisson
81
1. What is the likelihood of one Xi
2. What is the likelihood of all the data
3. What is the log-likelihood all the data
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
82
2. What is the likelihood of all the data
3. What is the log-likelihood all the data
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
83
3. What is the log-likelihood all the data
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
84
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
85
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
86
4. Find the value of λ which maximizes log likelihood
MLE For Poisson
87
MLE For Poisson
88
MLE For Poisson
Differentiate w.r.t. λ, and set to 0:
89
MLE For Poisson
90
MLE For Poisson
91
MLE For Poisson
92
Isn’t that the same as
the sample mean?
93
Yes. For Poisson.
94
MLE For Poisson
95
MLE For Poisson
96
We know sand is distributed as a pareto with PDF
MLE For Pareto
97
MLE for a Pareto
3. Find the value of α which maximizes log likelihood
1. What is the likelihood of all the data
2. What is the log-likelihood all the data
98
MLE for a Pareto
3. Find the value of α which maximizes log likelihood
2. What is the log-likelihood all the data
99
MLE for a Pareto
3. Find the value of α which maximizes log likelihood
100
MLE for a Pareto
101
Time to finish Medical Diagnosis in Seconds
MLE for Erlang
102
MLE for Erlang
103
104
Gradient Ascent
Walk uphill and you will find a local maxima
(if your step size is small enough)
105
Gradient Ascent
Walk uphill and you will find a local maxima
(if your step size is small enough)
Especially good if function is convex
106
Gradient Ascent
Repeat many times
Walk uphill and you will find a local maxima
(if your step size is small enough)
This is some profound life philosophy
Step size constant
Initialize: θj = random for all 0 ≤ j ≤ m
107
Gradient Ascent
Repeat many times:
Calculate all gradient[j]’s based on data
𝜃j += η * gradient[j] for all 0 ≤ j ≤ m
108
To the code!
109
Gradient Ascent for MLE of Erlang
110
Derived these partial derivatives of log likelihood
Gradient Ascent for MLE of Erlang
111
MLE For Bernoulli
112
Don’t we already have the Beta?
Yes! But this example is critical for developing towards deep learning.
113
MLE For Bernoulli
114
MLE For Bernoulli
115
Differentiable PMF for Bernoulli
116
PMF of Bernoulli
Differentiable PMF for Bernoulli
117
0
1
PMF of Bernoulli
Differentiable PMF for Bernoulli
118
0
1
p
PMF of Bernoulli
Differentiable PMF for Bernoulli
119
0
1
p
1 - p
PMF of Bernoulli
Differentiable PMF for Bernoulli
120
0
1
p
1 - p
PMF of Bernoulli
PMF of Bernoulli (p = 0.2)
Differentiable PMF for Bernoulli
121
0
1
p
1 - p
PMF of Bernoulli
PMF of Bernoulli (p = 0.2)
Differentiable PMF for Bernoulli
122
Bernoulli PMF
123
1. What is the likelihood of one Xi
2. What is the likelihood of all the data
3. What is the log-likelihood all the data
4. Find the value of p which maximizes log likelihood
Maximum Likelihood For Bernoulli
124
2. What is the likelihood of all the data
3. What is the log-likelihood all the data
4. Find the value of p which maximizes log likelihood
Maximum Likelihood For Bernoulli
125
3. What is the log-likelihood all the data
4. Find the value of p which maximizes log likelihood
Maximum Likelihood For Bernoulli
126
4. Find the value of p which maximizes log likelihood
Maximum Likelihood For Bernoulli
Take Derivative:
127
Take Derivative:
128
Take the derivative wrt p
Take Derivative:
129
Take the derivative wrt p
Derivative of a sum!
Take Derivative:
130
Take the derivative wrt p
Derivative of a sum!
Derivative of a sum!
Take Derivative:
131
Take the derivative wrt p
Derivative of a sum!
Derivative of a sum!
Take Derivative:
132
Take the derivative wrt p
Derivative of a sum!
Derivative of log p
Derivative of a sum!
Take Derivative:
133
Take the derivative wrt p
Derivative of a sum!
Derivative of log p
Derivative of a sum!
Take Derivative:
134
Take the derivative wrt p
Derivative of a sum!
Derivative of log p
Derivative of a sum!
Derivative of log (1-p)
Set to Zero:
135
Set to Zero:
136
Set to Zero:
137
Let
To make life easier
Set to Zero:
138
Let
And
To make life easier
Set to Zero:
139
Let
And
To make life easier
Set to Zero:
140
Let
And
To make life easier
Set to Zero:
141
Let
And
To make life easier
Set to Zero:
142
Let
And
To make life easier
Set to Zero:
143
Let
And
To make life easier
Set to Zero:
144
Let
And
To make life easier
145
Isn’t that the same as
unbiased estimator?
146
Yes. For Bernoulli.
147
MLE for Bernoulli is Sample Mean
148
The medicine is tried on 20 patients. It “works” for 14 and “doesn’t work” for 6. What is your new belief that the drug works?
In other words I have 20 IID samples from a Bernoulli. Estimate p. The data is [1,1,1,1,1,1,1,1,1,1,1,1,1,1,0,0,0,0,0,0]
MLE estimate:
Prior
Posterior
mode
Beta estimate:
MLE vs. Beta
149
Think about the difference between a point estimate and a distribution
p = 0.75
p =