1 of 75

Join at slido.com�#3862491

The Slido app must be installed on every computer you’re presenting from

3862491

2 of 75

Probability and �Density Estimation

Lecture 4

Review of Probability and Fitting Probability Models to Data

EECS 189/289, Fall 2026 @ UC Berkeley

Joseph E. Gonzalez and Narges Norouzi

3 of 75

Alexa, call my mom!

Calling your mom!

*Someone on YouTube just accidentally called their mom.

3862491

4 of 75

Siri, call my mom!

...

*Unlikely that anyone accidentally calls their mom…

3862491

5 of 75

Try It Yourself

Take out your phone or laptop and say one of these out loud.

  • “Hey Siri, what is the airspeed velocity of an unladen swallow?”
  • “Alexa, set a timer for one second.”
  • “OK Google, tell me a joke about Berkeley.”

​

Now say them again, but replace the first word: “Hey Sardine…”, “Alaska…”, “OK Doodle…”

​

Some devices woke up. Some did not. Some woke up for the wrong phrase.

  • Every one of those devices was running a detector on every second of sound in this room.

3862491

6 of 75

How does a smart assistant �“know when to listen”?

3862491

7 of 75

Wake Words

A wake word is a verbal cue that triggers �voice assistants to start actively listening.

Example: “Siri, set an alarm for 6:00AM”

Most voice assistants continuously run a wake word detector model on every sound they hear.

1 day

Rare wake word events.

3862491

8 of 75

Wake Words

A wake word is a verbal cue that triggers �voice assistants to start actively listening.

Example: “Siri, set an alarm for 6:00AM”

Most voice assistants continuously run a wake word detector model on every sound they hear.

Streamed to the cloud for processing.

3862491

9 of 75

Building a Wake-Word Detector

Is this a learning problem?

  • Yes – difficult to describe easy to demonstrate

What kind of learning problem?

Streamed to the cloud for processing.

3862491

10 of 75

What kind of learning problem is wake word detection.

The Slido app must be installed on every computer you’re presenting from

3862491

11 of 75

Building a Wake-Word Detector

Is this a learning problem?

  • Yes – difficult to describe easy to demonstrate

What kind of learning problem is this?

  • Supervised (we will collect labeled examples)
  • Binary classification (is this sound a wake word? Yes/No)

What kind of model should we use?

  • Many possible solutions, … probably a small neural network.

Streamed to the cloud for processing.

How accurate does the detector need to be?

3862491

12 of 75

Building a Wake-Word Detector

  •  

Streamed to the cloud for processing.

Introduce Random Variables to �Model this Process

3862491

13 of 75

Building a Wake-Word Detector

  •  

Streamed to the cloud for processing.

 

 

 

3862491

14 of 75

Today's Plan

Probability. Joint, marginal, conditional, independence, Bayes.

  • The detector fires. What is the chance a wake word was said?

​

Expectations. Turning probabilities into quantities with units.

  • What that accuracy costs per day.

​

Density estimation. Combining probability and optimization to fit models

  • Maximum likelihood: the objective behind most models in this course.
  • Maximum a Posteriori: accounting for priors.

​

3862491

15 of 75

Basics of Probability

A brief review of the

16 of 75

The Joint Probability Distribution

  •  

​

0

1

​

0.2

0.1

​

0.25

0.45

 

 

 

1

The joint probability satisfies the �following two properties:

0.45

0.55

0.3

0.7

3862491

17 of 75

The Joint Probability Distribution

  •  

 

 

The Sum Rule (Marginalization): defines the distribution over a subset of the random variables.

​

 

​

0

1

​

0.2

0.1

​

0.25

0.45

 

1

0.45

0.55

0.3

0.7

3862491

18 of 75

Conditional Probability

  •  

 

​

0.25

0.45

0.7

0.36

0.64

=

 

 

 

​

0

1

​

0.2

0.1

​

0.25

0.45

 

1

0.45

0.55

0.3

0.7

3862491

19 of 75

Product Rule: Chain Rule of Probability

  •  

3862491

20 of 75

Bayes’ Theorem

  •  

3862491

21 of 75

Independent Random Variables

  •  

3862491

22 of 75

Summarizing the Four Rules

  •  

3862491

23 of 75

Analyzing the �Wake Word Detector

What happens with rare events?

24 of 75

Building a Wake-Word Detector

  •  

 

 

 

 

3862491

25 of 75

Building a Wake-Word Detector

  •  

 

 

 

 

3862491

26 of 75

Building a Wake-Word Detector

  •  

3862491

27 of 75

How could we improve the Wake Word Detector?

The Slido app must be installed on every computer you’re presenting from

3862491

28 of 75

Building a Wake-Word Detector

  •  

Ultimately want high precision and recall.

3862491

29 of 75

Analyzing the Wake Word Detector

  •  

Bayes’ Theorem

3862491

30 of 75

Analyzing the Wake Word Detector

  •  

3862491

31 of 75

Bayesian Updates: Wake Word Detector

  •  

3862491

32 of 75

Revisiting the Math with Counts

Imagine 1,000,000 one-second segments.

  • Wake words are said in 0.01% of them: 100 segments contain a wake word, 999,900 do not.
  • The detector catches 99% of the real ones: 99 true detections.
  • The detector fires on 0.1% of the silent ones: 999 false alarms.

​

The detector fires 99+999 = 1,098 times. Only 99 of those are real.

  • 99 / 1,098 ≈ 9%

​

The base rate dominates. There are 10,000 times more silent segments than wake-word segments, so a small error rate on a huge population becomes a real challenge.

3862491

33 of 75

Demo

Which knob

matters?

3862491

34 of 75

What Is 9% Precision Costing Us?

  •  

3862491

35 of 75

Expectations

  •  

3862491

36 of 75

Functions of Random Variables

  •  

3862491

37 of 75

For any two correlated random variables 𝑋 and 𝑌

The Slido app must be installed on every computer you’re presenting from

3862491

38 of 75

Linearity of Expectation

  •  

3862491

39 of 75

Linearity of Expectation

  •  

 

 

 

 

 

3862491

40 of 75

Variance

  •  

 

 

 

 

 

3862491

41 of 75

Covariance

  •  

3862491

42 of 75

What Is 9% Precision Costing Us?

  •  

 

 

 

 

3862491

43 of 75

Demo

What 9% precision

costs per day

3862491

44 of 75

Modeling Distributions

From counting to

45 of 75

Modeling Distributions

Many machine learning models attempt to model the (joint) probability distribution of the data.

  • Instead of modeling “Does this sound contain a wake word” they model “The probability that this sound is a wake word”.

​

There are many ways to model a (joint) probability distribution.

  • Tabular representations
  • Classic probability models
  • Nonparametric models (histograms, kernel density estimation)

​

3862491

46 of 75

Tabular Representations

  •  

​

​

​

​

0.2

0.1

​

0.1

0.1

​

0.15

0.35

 

​

​

​

​

​

​

​

​

​

​

​

​

 

 

3862491

47 of 75

Bernoulli Distribution

  •  

 

 

 

 

 

 

 

 

 

3862491

48 of 75

Continuous Random Variables

  •  

 

3862491

49 of 75

Normal (Gaussian) Distribution

  •  

3862491

50 of 75

Density Estimation and MLE

Fitting distributions to data with

51 of 75

Density Estimation

  •  

3862491

52 of 75

Empirical Probability Distributions

  •  

3862491

53 of 75

The Empirical Distribution Is a Model

  •  

3862491

54 of 75

Estimating the Parameters of a Distribution

  •  

3862491

55 of 75

The Likelihood Function

  •  

Independent

Identically

3862491

56 of 75

The Log Likelihood Function

  •  

3862491

57 of 75

Maximum Likelihood Estimation

  •  

3862491

58 of 75

The MLE for ChatGPT

  •  

3862491

59 of 75

Demo

Maximum likelihood,

in one picture

3862491

60 of 75

The MLE for IID Bernoulli Samples (Part 1)

  •  

3862491

61 of 75

The MLE for IID Bernoulli Samples (Part 2)

  •  

Left as an exercise �(for your AI).

3862491

62 of 75

The MLE for IID Bernoulli Samples (Part 3)

  •  

3862491

63 of 75

Stopped Here

64 of 75

The Bernoulli MLE Is Just Counting

  •  

3862491

65 of 75

Your Turn: Let’s flip a coin

  •  

3862491

66 of 75

The Issue with Maximum Likelihood �and Rare Events

  •  

3862491

67 of 75

The Parameter as a Random Variable

  •  

3862491

68 of 75

Modeling the Prior

  •  

3862491

69 of 75

The Beta Distribution

  •  

 

3862491

70 of 75

Deriving the Posterior for Bernoulli + Beta

  •  

3862491

71 of 75

Computing the MAP

  •  

3862491

72 of 75

The Prior as Pseudo-Counts

  •  

3862491

73 of 75

 

  •  

3862491

74 of 75

What We Did Today

Reviewed basic probability, Bayes Theorem, Expectations, and Variance

Maximum likelihood determine model parameters by maximizing the likelihood of the data under the model.

Maximum a posteriori determine model parameters that maximize the posterior distribution.

Next Lecture: we will combine ideas learned today to develop Gaussian Mixture Models.

3862491

75 of 75

Probability and Density Estimation

Lecture 4

Credit: Joseph E. Gonzalez and Narges Norouzi

Reference Book Chapters:

  • Probability: Chapter 2.[1-2]
  • Density Estimation and MLE: Chapter 2.3