Join at slido.com�#3862491
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
3862491
Probability and �Density Estimation
Lecture 4
Review of Probability and Fitting Probability Models to Data
EECS 189/289, Fall 2026 @ UC Berkeley
Joseph E. Gonzalez and Narges Norouzi
Alexa, call my mom!
Calling your mom!
*Someone on YouTube just accidentally called their mom.
3862491
Siri, call my mom!
...
*Unlikely that anyone accidentally calls their mom…
3862491
Try It Yourself
Take out your phone or laptop and say one of these out loud.
Now say them again, but replace the first word: “Hey Sardine…”, “Alaska…”, “OK Doodle…”
Some devices woke up. Some did not. Some woke up for the wrong phrase.
3862491
How does a smart assistant �“know when to listen”?
3862491
Wake Words
A wake word is a verbal cue that triggers �voice assistants to start actively listening.
Example: “Siri, set an alarm for 6:00AM”
Most voice assistants continuously run a wake word detector model on every sound they hear.
1 day
Rare wake word events.
3862491
Wake Words
A wake word is a verbal cue that triggers �voice assistants to start actively listening.
Example: “Siri, set an alarm for 6:00AM”
Most voice assistants continuously run a wake word detector model on every sound they hear.
Streamed to the cloud for processing.
3862491
Building a Wake-Word Detector
Is this a learning problem?
What kind of learning problem?
Streamed to the cloud for processing.
3862491
What kind of learning problem is wake word detection.
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
3862491
Building a Wake-Word Detector
Is this a learning problem?
What kind of learning problem is this?
What kind of model should we use?
Streamed to the cloud for processing.
How accurate does the detector need to be?
3862491
Building a Wake-Word Detector
Streamed to the cloud for processing.
Introduce Random Variables to �Model this Process
3862491
Building a Wake-Word Detector
Streamed to the cloud for processing.
3862491
Today's Plan
Probability. Joint, marginal, conditional, independence, Bayes.
Expectations. Turning probabilities into quantities with units.
Density estimation. Combining probability and optimization to fit models
3862491
Basics of Probability
A brief review of the
The Joint Probability Distribution
| 0 | 1 |
| 0.2 | 0.1 |
| 0.25 | 0.45 |
1
The joint probability satisfies the �following two properties:
0.45 | 0.55 |
0.3 |
0.7 |
3862491
The Joint Probability Distribution
The Sum Rule (Marginalization): defines the distribution over a subset of the random variables.
| 0 | 1 |
| 0.2 | 0.1 |
| 0.25 | 0.45 |
1
0.45 | 0.55 |
0.3 |
0.7 |
3862491
Conditional Probability
| 0.25 | 0.45 |
0.7 |
0.36 | 0.64 |
=
| 0 | 1 |
| 0.2 | 0.1 |
| 0.25 | 0.45 |
1
0.45 | 0.55 |
0.3 |
0.7 |
3862491
Product Rule: Chain Rule of Probability
3862491
Bayes’ Theorem
3862491
Independent Random Variables
3862491
Summarizing the Four Rules
3862491
Analyzing the �Wake Word Detector
What happens with rare events?
Building a Wake-Word Detector
3862491
Building a Wake-Word Detector
3862491
Building a Wake-Word Detector
3862491
How could we improve the Wake Word Detector?
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
3862491
Building a Wake-Word Detector
Ultimately want high precision and recall.
3862491
Analyzing the Wake Word Detector
Bayes’ Theorem
3862491
Analyzing the Wake Word Detector
3862491
Bayesian Updates: Wake Word Detector
3862491
Revisiting the Math with Counts
Imagine 1,000,000 one-second segments.
The detector fires 99+999 = 1,098 times. Only 99 of those are real.
The base rate dominates. There are 10,000 times more silent segments than wake-word segments, so a small error rate on a huge population becomes a real challenge.
3862491
Demo
Which knob
matters?
3862491
What Is 9% Precision Costing Us?
3862491
Expectations
3862491
Functions of Random Variables
3862491
For any two correlated random variables 𝑋 and 𝑌
The Slido app must be installed on every computer you’re presenting from
Do not edit�How to change the design
3862491
Linearity of Expectation
3862491
Linearity of Expectation
3862491
Variance
3862491
Covariance
3862491
What Is 9% Precision Costing Us?
3862491
Demo
What 9% precision
costs per day
3862491
Modeling Distributions
From counting to
Modeling Distributions
Many machine learning models attempt to model the (joint) probability distribution of the data.
There are many ways to model a (joint) probability distribution.
3862491
Tabular Representations
| | |
| 0.2 | 0.1 |
| 0.1 | 0.1 |
| 0.15 | 0.35 |
| | |
| | |
| | |
| | |
3862491
Bernoulli Distribution
3862491
Continuous Random Variables
3862491
Normal (Gaussian) Distribution
3862491
Density Estimation and MLE
Fitting distributions to data with
Density Estimation
3862491
Empirical Probability Distributions
3862491
The Empirical Distribution Is a Model
3862491
Estimating the Parameters of a Distribution
3862491
The Likelihood Function
Independent
Identically
3862491
The Log Likelihood Function
3862491
Maximum Likelihood Estimation
3862491
The MLE for ChatGPT
3862491
Demo
Maximum likelihood,
in one picture
3862491
The MLE for IID Bernoulli Samples (Part 1)
3862491
The MLE for IID Bernoulli Samples (Part 2)
Left as an exercise �(for your AI).
3862491
The MLE for IID Bernoulli Samples (Part 3)
3862491
Stopped Here
The Bernoulli MLE Is Just Counting
3862491
Your Turn: Let’s flip a coin
3862491
The Issue with Maximum Likelihood �and Rare Events
3862491
The Parameter as a Random Variable
3862491
Modeling the Prior
3862491
The Beta Distribution
3862491
Deriving the Posterior for Bernoulli + Beta
3862491
Computing the MAP
3862491
The Prior as Pseudo-Counts
3862491
3862491
What We Did Today
Reviewed basic probability, Bayes Theorem, Expectations, and Variance
Maximum likelihood determine model parameters by maximizing the likelihood of the data under the model.
Maximum a posteriori determine model parameters that maximize the posterior distribution.
Next Lecture: we will combine ideas learned today to develop Gaussian Mixture Models.
3862491
Probability and Density Estimation
Lecture 4
Credit: Joseph E. Gonzalez and Narges Norouzi
Reference Book Chapters: