Beyond CS109
Chris Gregg
Summer, 2026
CS109, Stanford University
2
Final Exam Logistics
What have you
learned?
What is a probability?
4
The event we care about
How many times does it occur?
Out of (close to) infinite trials
Lecture 1
Mehran Sahami, Chris Piech, Lisa Yan, Jerry Cain, Juliette Woodrow, and Chris Gregg, Spring 2026
Time to Start Flippin Coins
Bayes’ Theorem
6
Lecture 2
Program the General Version
Learning Goals of Today
11
Mutually Exclusive
Independent
Makes AND easy:
Makes OR easy:
Lecture 3
Sort semi-distinct objects
12
Also:
Lecture 4
Sahami, Piech, Yang, Cain, Woodrow, Gregg
Counting Cards
Declaring a Random Variable to be Binomial
14
Our random variable
Is distributed as a
Binomial
With these parameters
Num trials
Probability of success on each trial
Piech & Cain, CS109, Stanford University
Lecture 5
Expected Value
16
Loop over all values x that X can take on
The value
The probability of that value
Lecture 6
Poisson Random Variable
17
Lecture 7
19
Truth 2:
Truth 3:
Truths of Probability For Continuous Random Variables
Truth 1:
Truth 4:
Know why!
That’s all possible values (Axiom 2)
Since the integral is a probability (Axiom 1)
Area under the curve!
Lecture 8
Normal Probability Density Function
23
x
f(x)
Lecture 9
Joint table: mutually exclusive and covers sample space.
24
| Single | Relationship | Complicated |
Frosh | 0.07 ⋅ k | 0.04 ⋅ k | 0.01 ⋅ k |
Soph | 0.09 ⋅ k | 0.05 ⋅ k | 0.01 ⋅ k |
Junior | 0.05 ⋅ k | 0.05 ⋅ k | 0.01 ⋅ k |
Senior | 0.01 ⋅ k | 0.03 ⋅ k | 0.01 ⋅ k |
5+ | 0.03 ⋅ k | 0.03 ⋅ k | 0.02 ⋅ k |
X is dating status.
Y is year.
Each combination is mutually exclusive, and they span the sample space
Lecture 10
Chris Piech, CS109
25
Beer Lambert Law
Rate of muons depends on x, amount of limestone
Lecture 11
Chris Piech, CS109
Inference
26
Age from C14
Updated Delivery Prob
Age from Name
Hidden Chambers
Stanford Eye Test
Updating Lidar Belief
Inference
Inference noun
Updating one’s belief about a random variable (or multiple) based on conditional knowledge regarding another random variable (or multiple) in a probabilistic model.
TLDR: conditional probability with random variables.
28
def update_belief_carbon_dating(m = 900):
# pr_A[i] is P(Age = i| m = 900).
pr_A = {}
for i in range(100,10000+1):
prior = 1 / n_years # P(A = i)
likelihood = calc_likelihood(m, i) #P(M=m|A=i)
pr_A[i] = likelihood * prior
# implicitly computes the normalization constant
normalize(pr_A)
return pr_A
def update_belief_baby(prior, today = 10):
# pr_D[i] is P(D = i| No Baby Yet).
pr_D = {}
for i in range(-50,25):
# P(NoBaby | D = i)
likelihood = 0 if i < today else 1
pr_D[i] = likelihood * prior[i]
# implicitly computes the LOTP
normalize(pr_D)
return pr_D
def update_belief_name_to_age(name = 'Laura'):
# pr_age[i] is P(Age = i| name).
# prob_name_and_age is just a counting from the US
# Social Security database.
pr_age = {}
for i in range(10,110):
pr_age[i] = calc_prob_name_and_age(name, i)
# implicitly computes the normalization constant
normalize(pr_age)
return pr_age
What do you notice is the same. What is different?
Chris Piech, CS109
Four Prototypical Trajectories
Lecture 12
Chris Piech, CS109, 2021
Multinomial Random Variable?
30
Joint PMF
where
and
Multinomial # of ways of ordering the outcomes
Probability of each ordering is equal + mutually exclusive
Lecture 13
Chris Piech, CS109, 2021
Beta is the Random Variable for Probabilities
31
Used to represent a distributed belief of a probability
Lecture 14
Chris Piech, CS109, 2021
Central Limit Theorem
32
Lecture 16
Chris Piech, CS109, 2021
Four Prototypical Trajectories
Bootstraping allows you to:
Lecture 17
Chris Piech, CS109, 2021
Conditional Expectation
34
X = units in fall quarter
Y = year in school
E[X | Y] ?
Lecture 18
Chris Piech, CS109, 2021
Which Question is Better?
35
Lecture 19
Chris Piech, CS109, 2021
36
We want to choose the parameter value that maximizes the probability of the data:
How to Choose the “Best” Parameters: MLE
Likelihood
To put words into math:
Our best estimate
Log Likelihood
Lecture 20
37
Logistic Regression
+
z = 2.1
σ(z) = 0.7
Lecture 21
Gradient Ascent
Walk uphill and you will find a local maxima
(if your step size is small enough)
argmax
Results
39
overfitting
Model Train Accuracy Test Accuracy
-------------------------------------------------------------
Baseline 0.6031 0.5887
Naive Bayes 0.7909 0.8067
Logistic Regression 0.8169 0.8307
Decision Tree 0.8514 0.8307
Random Forest 0.8726 0.8500
Gradient Boosting 0.8611 0.8440
AdaBoost 0.8334 0.8353
BayesNet 0.8320 0.8507
Lecture 22
40
Calibration
New!
41
We Can Put Neurons Together
…
+
Single neuron in a hidden layer
Lecture 23
Lets start training a Critter
42
Lecture 25
What should you do next?
Go solve amongst the abundance of important problems
Think about intersectionality
Your side passion
Data that you have access to
Your lived experience
Thompson sampling
CS109
Probability Fundamentals
Single Random Variables
Probabilistic Models
Uncertainty Theory
AI
I had an important job!
I hope you think I did it justice
Thank you so much!