1 of 56

Join at slido.com�#9624861

The Slido app must be installed on every computer you’re presenting from

9624861

2 of 56

Density Estimation and �Gaussian Mixture Models

Lecture 5

Introducing Maximum Likelihood Estimation, and �Gaussian Mixture Models

EECS 189/289, Fall 2025 @ UC Berkeley

Joseph E. Gonzalez and Narges Norouzi

EECS 189/289, Fall 2026 @ UC Berkeley

Joseph E. Gonzalez and Narges Norouzi

3 of 56

Recap

In the last lecture we introduced basic probability rules and derived the Maximum Likelihood Estimate (MLE) and started to derive the Maximum a Posteriori (MAP) Estimate.

Today we will:

  • Finish deriving the MAP estimate
  • Introduce the Multivariate Gaussian and how to compute Bias
  • Introduce the Mixture of Gaussians model that combines many of these ideas.

9624861

4 of 56

Estimating the Parameters of a Distribution

  •  

9624861

5 of 56

The Bernoulli MLE Is Just Counting

  •  

9624861

6 of 56

The Parameter as a Random Variable

  •  

9624861

7 of 56

Modeling the Prior

  •  

9624861

8 of 56

The Beta Distribution

  •  

 

9624861

9 of 56

Deriving the Posterior for Bernoulli + Beta

  •  

9624861

10 of 56

Computing the MAP

  •  

9624861

11 of 56

The Prior as Pseudo-Counts

  •  

9624861

12 of 56

 

  •  

9624861

13 of 56

For a,b > 2 which has higher variance?

The Slido app must be installed on every computer you’re presenting from

9624861

14 of 56

Starting Lecture 5

More Practice with MLE

15 of 56

Normal (Gaussian) Distribution

  •  

9624861

16 of 56

The MLE for IID Gaussian Samples (Part 1)

  •  

9624861

17 of 56

The MLE for IID Gaussian Samples (Part 2)

  •  

9624861

18 of 56

The MLE for IID Gaussian Samples (Part 3)

  •  

 

 

 

9624861

19 of 56

The MLE for IID Gaussian Samples (Part 3)

  •  

 

 

 

 

 

 

9624861

20 of 56

Recap of the MLE for IID Gaussian Samples

  •  

Sufficient Statistics

9624861

21 of 56

Bias

  • MAP Estimation
  • Gaussians
    • Maximum Likelihood
    • Bias
  • Gaussian Mixture Models

Questions

9624861

22 of 56

Bias of an Estimate

  •  

9624861

23 of 56

 

  •  

Unbiased

9624861

24 of 56

Is the maximum likelihood variance estimate biased?

The Slido app must be installed on every computer you’re presenting from

9624861

25 of 56

 

  •  

 

Variance Identity

 

 

 

 

9624861

26 of 56

 

  •  

 

Variance Identity

 

 

Sample Mean Variance

 

 

9624861

27 of 56

 

  •  

 

 

9624861

28 of 56

 

  •  

Sometimes called the

unbiasedsample variance.

9624861

29 of 56

Demo

Why np.var has

a ddof argument

9624861

30 of 56

Fitting Data with Multiple Modes

Often data has multiple modes, and a single Gaussian is not a great model for the distribution.

9624861

31 of 56

Gaussian Mixture Models

  • MAP Estimation
  • Gaussians
    • Maximum Likelihood
    • Bias
  • Gaussian Mixture Models

Questions

9624861

32 of 56

Gaussian Mixture Model (GMM)

The Gaussian mixture model defines the probability of sampling data from a weighted combination of Gaussians

  • Example: K=3

  • Note the data doesn’t have a color or a y-value: both were added for visualization purposes.

 

 

 

9624861

33 of 56

Gaussian Mixture Model (GMM)

  •  

 

 

 

pi = np.array([0.2, 0.5, 0.3])

mu = np.array([-1, 2, 5])

sigma = np.array([0.2, 0.5, .1])

Parameters

9624861

34 of 56

Demo

Gaussian Mixture Model

9624861

35 of 56

The GMM is a Latent Variable Model

  •  

1.3

1

9.6

3

2.4

1

5.7

2

Think of this as the color in the earlier visualization

Density for is based on it’s Gaussian.

9624861

36 of 56

The GMM is a Generative Model

  •  

Example

 

 

 

 

9624861

37 of 56

Demo

Sampling from a GMM

9624861

38 of 56

The Connection to K-Means

  •  

9624861

39 of 56

Latent Variable Posteriors

  •  

9624861

40 of 56

Trying the Stationary Point Method

  •  

 

N

K

9624861

41 of 56

Estimating the GMM Parameters

  •  

 

N

K

9624861

42 of 56

Quick Recap

If we knew the model parameters, we could easily compute the cluster assignments.

If we knew the cluster assignments, we could easily estimate the model parameters.

How can we solve this cyclic dependency?

 

Model Parameters

 

Cluster Assignments

9624861

43 of 56

Intuition Behind �Expectation-maximization (EM)

  •  

9624861

44 of 56

The EM Algorithm: E-step

  •  

 

N

K

 

Cluster “pseudo counts”

1

1

1

1

1

1

9624861

45 of 56

The EM Algorithm: E-step

  •  

 

N

K

 

1

1

1

1

1

1

9624861

46 of 56

 

  •  

Method of Lagrange Multipliers (Lecture 3) �for the normalization constraint

 

 

 

N

K

 

1

1

1

1

1

1

9624861

47 of 56

 

  •  

 

 

N

K

 

1

1

1

1

1

1

Need figure out the Lagrange Multiplier

 

 

9624861

48 of 56

 

  •  

9624861

49 of 56

 

  •  

9624861

50 of 56

The EM Algorithm for GMMs

  •  

9624861

51 of 56

The EM Algorithm

  •  

Easy to optimize joint probability

Current distribution �over the latent Z

Updates distribution over Z.

9624861

52 of 56

Convergence and Local Optima

  •  

9624861

53 of 56

Convergence, Initialization, and Local Optima

  •  

9624861

54 of 56

Demo

Implementing EM

for GMMs

9624861

55 of 56

What We Did Today

  •  

9624861

56 of 56

Density Estimation and �Gaussian Mixture Models

Lecture 5

Credit: Joseph E. Gonzalez and Narges Norouzi

Reference Book Chapters:

  • Probability: Chapter 2.[1-2], 2.3.[1-3] (we will return to 2.4 onward later)
  • Distributions: Chapter 3.1, 3.2.1 (the rest of 3 is more advanced Gaussians)
  • Clustering: Chapter 15.1 (k-means), 15.2 (Gaussian Mixture Models)