Join at slido.com�#9624861
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
9624861
Density Estimation and �Gaussian Mixture Models
Lecture 5
Introducing Maximum Likelihood Estimation, and �Gaussian Mixture Models
EECS 189/289, Fall 2025 @ UC Berkeley
Joseph E. Gonzalez and Narges Norouzi
EECS 189/289, Fall 2026 @ UC Berkeley
Joseph E. Gonzalez and Narges Norouzi
Recap
In the last lecture we introduced basic probability rules and derived the Maximum Likelihood Estimate (MLE) and started to derive the Maximum a Posteriori (MAP) Estimate.
Today we will:
9624861
Estimating the Parameters of a Distribution
9624861
The Bernoulli MLE Is Just Counting
9624861
The Parameter as a Random Variable
9624861
Modeling the Prior
9624861
The Beta Distribution
9624861
Deriving the Posterior for Bernoulli + Beta
9624861
Computing the MAP
9624861
The Prior as Pseudo-Counts
9624861
9624861
For a,b > 2 which has higher variance?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
9624861
Starting Lecture 5
More Practice with MLE
Normal (Gaussian) Distribution
9624861
The MLE for IID Gaussian Samples (Part 1)
9624861
The MLE for IID Gaussian Samples (Part 2)
9624861
The MLE for IID Gaussian Samples (Part 3)
9624861
The MLE for IID Gaussian Samples (Part 3)
9624861
Recap of the MLE for IID Gaussian Samples
Sufficient Statistics
9624861
Bias
Questions
9624861
Bias of an Estimate
9624861
Unbiased
9624861
Is the maximum likelihood variance estimate biased?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
9624861
Variance Identity
9624861
Variance Identity
Sample Mean Variance
9624861
9624861
Sometimes called the
unbiased �sample variance.
9624861
Demo
Why np.var has
a ddof argument
9624861
Fitting Data with Multiple Modes
Often data has multiple modes, and a single Gaussian is not a great model for the distribution.
9624861
Gaussian Mixture Models
Questions
9624861
Gaussian Mixture Model (GMM)
The Gaussian mixture model defines the probability of sampling data from a weighted combination of Gaussians
9624861
Gaussian Mixture Model (GMM)
pi = np.array([0.2, 0.5, 0.3])
mu = np.array([-1, 2, 5])
sigma = np.array([0.2, 0.5, .1])
Parameters
9624861
Demo
Gaussian Mixture Model
9624861
The GMM is a Latent Variable Model
| |
1.3 | 1 |
9.6 | 3 |
2.4 | 1 |
5.7 | 2 |
… | … |
Think of this as the color in the earlier visualization
Density for is based on it’s Gaussian.
9624861
The GMM is a Generative Model
Example
9624861
Demo
Sampling from a GMM
9624861
The Connection to K-Means
9624861
Latent Variable Posteriors
9624861
Trying the Stationary Point Method
N
K
9624861
Estimating the GMM Parameters
N
K
9624861
Quick Recap
If we knew the model parameters, we could easily compute the cluster assignments.
If we knew the cluster assignments, we could easily estimate the model parameters.
How can we solve this cyclic dependency?
Model Parameters
Cluster Assignments
9624861
Intuition Behind �Expectation-maximization (EM)
9624861
The EM Algorithm: E-step
N
K
Cluster “pseudo counts”
1
1
1
1
1
1
9624861
The EM Algorithm: E-step
N
K
1
1
1
1
1
1
9624861
Method of Lagrange Multipliers (Lecture 3) �for the normalization constraint
N
K
1
1
1
1
1
1
9624861
N
K
1
1
1
1
1
1
Need figure out the Lagrange Multiplier
9624861
9624861
9624861
The EM Algorithm for GMMs
9624861
The EM Algorithm
Easy to optimize joint probability
Current distribution �over the latent Z
Updates distribution over Z.
9624861
Convergence and Local Optima
9624861
Convergence, Initialization, and Local Optima
9624861
Demo
Implementing EM
for GMMs
9624861
What We Did Today
9624861
Density Estimation and �Gaussian Mixture Models
Lecture 5
Credit: Joseph E. Gonzalez and Narges Norouzi
Reference Book Chapters: