1 of 31

CMSC 320

Basic Probability

& Distributions

Part 1 • Core principles, distributions, and why they matter in data science

person

Fardina Alam

UNCERTAINTY → DECISION

0

0%

Impossible

0.5

50%

Even chance

1

100%

Certain

Probability Bounds

0 ≤ P(A) ≤ 1

2 of 31

WHY THIS MATTERS

Start with Data Science: Where Does Probability Enter?

CMSC 320 • Basic Probability & Distributions

2

Predict Will a customer churn?

P(churn | behavior)

Infer Is an observed difference real?

Sampling uncertainty

Model Which class is most likely?

Probabilistic assumptions

Decide Which action has best payoff?

Expected value

Probability turns uncertainty into quantities we can compare.

3 of 31

Why Does a Data Scientist Need Probability & Distributions?

  • Fraud model: P(Fraud | transaction features) = 0.82
  • Churn model: P(Churn | customer history) = 0.31
  • A/B test: 12.4% conversion vs. 10.8%
  • Service metric: 95th-percentile latency = 420 ms
  • Housing data: mean price = $510K, median price = $390K

All of these scenarios require us to reason about uncertainty and distributions.

We need to start with the data question. Then learn the statistical concept needed to answer it.

CMSC 320 • Basic Probability & Distributions

4 of 31

TOPICS

Today’s Roadmap

CMSC 320 • Basic Probability & Distributions

3

1

Events &�probability

2

Conditional�reasoning

3

Bayes &�total probability

4

Independence &�expected value

5

Common�distributions

6

Normal, Z�& CLT

Goal: connect each mathematical idea to a data-science decision.

Textbook reference: Chapter 6

5 of 31

DS Scenario 1: Which Customers Are Likely to Churn?

Dataset: 10,000 customers

Features: tenure, monthly charge, contract type, support calls

Target: churn = Yes / No

Overall churn rate = 30%, but customer groups may behave differently.

We want P(Churn | observed customer information), not only P(Churn).

Needed concepts → events, probability axioms, conditional probability, independence, Bayes' theorem.

6 of 31

CORE CONCEPTS

Probability Basics: Outcomes → Events → Probability

CMSC 320 • Basic Probability & Distributions

4

Sample space Ω

All possible outcomes��Die: {1, 2, 3, 4, 5, 6}

Event A

A subset of outcomes of interest��A = “roll an even number”

Probability P(A)

Likelihood of the event��P(A) = 3/6 = 0.5

0 ≤ P(A) ≤ 1

Data science translation: define the population → define the event → quantify uncertainty.

7 of 31

CORE CONCEPTS

From Events to Random Variables

CMSC 320 • Basic Probability & Distributions

5

A random variable assigns a number to an uncertain outcome.

2

3

4

5

6

7

8

9

10

11

12

X = sum of two dice

Distribution

Each possible X value has an associated probability.��P(X=7)=6/36

Focus on the pattern across outcomes—not one roll.

8 of 31

CONDITIONAL REASONING

Conditional Probability: Information Changes the Sample Space

CMSC 320 • Basic Probability & Distributions

6

P(A | B) = P(A ∩ B) / P(B)

Marble example

10 marbles total�5 red • 3 of the red are shiny��Given RED → only 5 marbles remain relevant

3 shiny red

P(shiny | red) = 3/5 = 0.60

Most real-world predictions are conditional.

What is the probability that a marble is shiny, given that it is red?

9 of 31

9

Conditional probability from a table

Always inspect the denominator

Observed users

Churn

Stay

Total

Premium

80

120

200

Standard

240

560

800

Total

320

680

1000

Question 1: Among Premium users, what proportion churned?

P(Churn | Premium) = 80/200 = 0.40

Question 2: Among users who churned, what proportion were Premium users?

P(Premium | Churn) = 80/320 = 0.25

Important

P(A|B) and P(B|A) are generally different.

CMSC320 • Introduction to Data Science

Basic Probability and Distributions — Part 1

10 of 31

RELATIONSHIPS

Independence: Does Knowing B Change A?

CMSC 320 • Basic Probability & Distributions

10

Independent ⇔ P(A | B) = P(A)

Equivalent: P(A ∩ B) = P(A)P(B)

Independent

Two separate fair coin tosses

Not independent

Rain and umbrella use

Modeling question

Does one feature add information about another?

Independence is an assumption to check—not simply assume.

In data science, independence assumptions appear in models and statistical tests

Knowing B does not change the probability of A

11 of 31

RELATIONSHIPS

Conditional Independence: Context Can Remove a Relationship

CMSC 320 • Basic Probability & Distributions

11

Rain

Dog barks

Cat runs

Given Dog barks: P(Cat runs | Rain, Dog barks) = P(Cat runs | Dog barks)

Once the middle variable is known, Rain adds no extra information about Cat runs.

Why DS cares: conditional independence simplifies models such as Naive Bayes.

12 of 31

CONDITIONAL REASONING

Bayes’ Rule: Update a Belief with Evidence

P(A | B)=P(B | A)P(A)/P(B)

Prior

P(A)

What we believed before

Likelihood

P(B | A)

How compatible is evidence?

Evidence

P(B)

How common is the evidence?

Posterior

P(A | B)

Updated belief

Example:update P(spam) after seeing the word “discount”.

CMSC 320 • Basic Probability & Distributions

7

13 of 31

Example: Picnic Day

13

What is the chance of rain during the day?

Scenario: Find the chance of rain given that the morning is cloudy.

1. P(Rain) = 10%

Only 3 out of 30 days are rainy.

2. P(Cloud | Rain) = 50%

50% of rainy days start off cloudy.

3. P(Cloud) = 40%

40% of all days start off cloudy.

Applying Bayes' Theorem:

P(Rain | Cloud) = 12.5%

Only a 12.5% chance of rain.

Not too bad, let's have a picnic! 🧺☀️

14 of 31

PROBABILITY LAWS

Law of Total Probability: Add Across Possible Causes

When an event can happen through several non-overlapping cases:

P(A) = Σ P(A | Bᵢ) P(Bᵢ)

Case 1: Fair Die

Selection Chance: 70% (0.7)

Conditional Prob: P(6 | Fair) = 1/6

Contribution: 0.7 × (1/6) = 0.1167

Case 2: Loaded Die

Selection Chance: 30% (0.3)

Conditional Prob: P(6 | Loaded) = 1/2

Contribution: 0.3 × (1/2) = 0.1500

Total Probability P(6)

Sum of all weighted cases:

0.7(1/6) + 0.3(1/2)

P(6) = 0.2667 (26.67%)

Key Takeaway: Useful when the underlying source/cause is hidden.

CMSC 320 • Basic Probability & Distributions

9

15 of 31

Example: Two Bags of Marbles

Setup Scenario

Bag 1: 6 white, 4 black    ➔    P(Black|B1) = 4/10 | Bag 2: 3 white, 7 black    ➔    P(Black|B2) = 7/10

Case 1: Equal Bag Selection

If we put both bags in a box, close our eyes, grab a bag, and then grab a marble: what is the probability it is black?

SOLUTION:

P(B1) = 1/2,   P(B2) = 1/2

P(Black) = P(B1)·P(Black|B1) + P(B2)·P(Black|B2)

Final Calculation:

1/2·4/10 + 1/2·7/10 = 11/20 = 0.55

Case 2: Unequal Bag Selection

If Bag 1 is much larger, making us twice as likely to grab Bag 1 as Bag 2: what is the probability of grabbing a black marble?

SOLUTION:

P(B1) = 2/3,   P(B2) = 1/3

P(Black) = P(B1)·P(Black|B1) + P(B2)·P(Black|B2)

Final Calculation:

2/3·4/10 + 1/3·7/10 = 15/30 = 1/2 (0.50)

16 of 31

Law of Total Probability: back to the example

Sample Space S: The scenario of choosing a bag and drawing a marble.

  1. Mutually Exclusive Events B1​ and B2​:
    • B1​: Selecting Bag 1.
    • B2​: Selecting Bag 2.

These events are mutually exclusive because you can only choose one bag at a time.

  1. Event A: Drawing a black marble.

According to the law of total probability, the probability of drawing a black marble (event A) is:

P(A)=P(B1)⋅P(A | B1)+P(B2)⋅P(A |B2)P(A)

16

17 of 31

Try Yourself

Scenario: Imagine you are organizing a charity event, and there are three possible venues (A, B, and C) where you can hold the event. The probability of each venue being available on a given day is as follows:

Venue A: 40% chance of being available.

Venue B: 30% chance of being available.

Venue C: 30% chance of being available.

You also know that if Venue A is available, there's a 70% chance of raising a large amount of money for the charity, while if Venue B or C is available, there's a 50% chance of raising a large amount of money.

What is the overall probability L of raising a large amount of money at the charity event?

17

18 of 31

DECISION MAKING

Expected Value: Average Behavior Over Repetition

CMSC 320 • Basic Probability & Distributions

12

E[X] = Σ x · P(X=x)

What it means

Long-run average outcome�Not a guaranteed single outcome�Can represent reward, cost, clicks, loss, or risk

20%

$0

30%

$10

40%

$20

10%

$10

EV = $12 per spin

In ML, optimization often means minimizing expected loss.

19 of 31

Expected Value (Average Behavior)

Expected value summarizes long-run behavior.

  • Not a guaranteed outcome
  • The average result over many repetitions
  • Can be positive or negative

CONCEPT & PROPERTIES

FORMULA & EXAMPLE

In Data Science:

  • Helps predict average outcomes like revenue, user clicks, or risk.
  • Forms the foundational basis for machine learning decision rules.

Remember:

Expected value directly drives loss, reward, and algorithmic optimization.

20 of 31

20

Expected Value: Examples

EXAMPLE 1: GAME SHOW

Someone offers for you to go on a game show with two possible outcomes:

  • 5% chance: Win $1,000,000
  • 95% chance: Hit with sticks, causing $10,000 in medical bills

Should you go on the game show?

Calculation:

EV = (1,000,000 × .05) + (-10,000 × .95)�   = 50,000 - 9,500 = $40,500

Conclusion: Net positive!

EXAMPLE 2: WHEEL SPIN

You play a game spinning a wheel with the following reward probabilities:

  • 20% chance: Win $0
  • 30% chance: Win $10 gift card
  • 40% chance: Win $20 gift card
  • 10% chance: Win $10 gift card

Calculation:

EV = (.2 × 0) + (.3 × 10) + (.4 × 20) + (.1 × 10)�   = 0 + 3 + 8 + 1 = $12

On average, expect to win $12 per spin over many spins.

21 of 31

Probability Distributions and Their Types

Distributions describe how values are spread.

Discrete

Countable values (0, 1, 2, …)

  • Bernoulli Distribution
  • Binomial Distribution
  • Poisson Distribution
  • Zero-Inflated Poisson Distribution

Continuous

Smooth ranges of values

  • Uniform Distribution
  • Normal (Gaussian) Distribution
  • Many more..

Key Insights

Understanding how your data is distributed can tell you a lot about the process generating the data.

  • The nature of your distribution affects which statistical tools you use

Shape tells us about the data-generating process

Remember: Distribution choice comes before modeling.

22 of 31

1(a,b). Bernoulli and Binomial Distribution

Bernoulli Distribution

Binary Outcome: Success (1) or Failure (0)

  • Models a single trial
  • Success probability = p
  • P(X = 1) = p
  • P(X = 0) = 1 − p
  • Example: Single coin flip, spam vs. not spam

Binomial Distribution

Number of successes (k) in n independent Bernoulli trials.

X ~ Binomial(n, p)

Example: Heads in 10 flips, quiz scores

Key Idea: Binomial = sum of independent Bernoulli trials

Worked Example

Tossing a fair coin 10 times:

n = 10, p = 0.5

23 of 31

1a. Poisson Distribution

What it Models

Counts of events occurring in a fixed time or space interval, where events occur randomly at a constant average rate.

Key Assumptions

  • Events occur independently of one another
  • Average rate λ remains constant
  • Simultaneous events cannot occur at exact same instant

Probability Mass Function (PMF)

Where:

  • P(X=k): Probability of observing exactly k events
  • λ: Average rate parameter per interval
  • e: Euler's constant (≈ 2.71828)

X ~ Poisson(λ)

Distribution Shape

Right-skewed for small values of λ.

Becomes increasingly symmetric as λ increases.

Worked Example

Emails arrive at an average rate of 3 per hour. What is the probability of receiving exactly 2 emails in 1 hour?

Given Parameters:

  • λ = 3 events/hour | t = 1 hour | k = 2 emails

24 of 31

Zero-Inflated Poisson (ZIP) Distribution

Used for count data with more zeros than expected.

Components

  • Poisson component: models the count process
  • Inflation component: models excess zeros (e.g., Bernoulli)

Why ZIP?

  • Standard Poisson underestimates zeros
  • ZIP captures structural zeros + counts�

Examples

  • Insurance claims
  • Hospital visits
  • Product defects

Oftentimes, poisson distributions can have a spike at zero.

25 of 31

25

2a. The Uniform Distribution

Key Characteristics

A continuous or discrete probability distribution where all outcomes within a specified boundary are equally likely.

  • Equally Likely: Every outcome has the exact same probability of occurrence.
  • Fixed Interval: Values strictly occur within a defined minimum and maximum range.

Classic Example: Fair Six-Sided Die

Rolling a standard, balanced die gives 6 discrete outcomes {1, 2, 3, 4, 5, 6}.

Probability Mass Function:

  • P(X = k) = 1/n = 1/6 ≈ 16.67% for each face
  • Flat probability density across all valid outcomes.

Distribution Visualization

26 of 31

2b. Normal (Gaussian) Distribution

Key Characteristics

  • Symmetric & Bell-Shaped: Data is symmetrically distributed with no skew.
  • Central Tendency: Mean = Median = Mode = μ.
  • Parameters: Defined by mean (μ) and standard deviation (σ).

"Averages are common, extremes are rare."

Real-World Examples

  • Heights of people: Human physical characteristics across populations.
  • Measurement errors: Random variations in scientific experiments.
  • Standardized test scores: Large-scale educational evaluations.

Empirical Rule: 68%–95%–99.7%1σ, 2σ, 3σ bounds

Frequency Distribution & Bell CurveSymmetric around Mean

27 of 31

Parameters of the Normal Distribution: μ & σ

Mean (μ)

Controls the position of the curve

  • Shifts the curve left or right
  • Controls the center of the distribution

Standard Deviation (σ)

Controls the shape and dispersion of the curve

  • Controls the spread of the curve
  • Smaller σ → narrower, taller
  • Larger σ → wider, shorter

28 of 31

Problem Solving: Finding percentages

28

Example 1 (Solved)

The travel time between two regional cities is approximately normally distributed with a mean of 70 minutes and a standard deviation of 2 minutes.

Q: What is the percentage of travel times that are between 66 minutes and 72 minutes?

Try Yourself

The volume of soup served via machine is normally distributed with a mean of 240 mL and a standard deviation of 5 mL. A store serves 160 cups of soup.

Q: What number of these cups of soup are expected to contain less than 230 ml?

Solution & Parameters

  • Mean (μ) = 70 minutes
  • Standard Dev (σ) = 2 minutes
  • Target Interval = [66, 72] → [μ - 2σ, μ + 1σ]

Resulting Percentage

81.5% (13.5% + 34% + 34%)

Time (min):

64

66

68

70

72

74

76

29 of 31

Standard Normal (Z) Distribution

Key Properties & Definition

A special normal distribution characterized by:

  • Mean μ = 0
  • Standard deviation σ = 1

Used to compare values from different distributions by converting raw data into standardized z-scores.

Z-Score Formula

Standardize any normal distribution value:

Interpretation & Worked Example

Interpretation Rules: Z = 0 (at mean) | Z > 0 (above mean) | Z < 0 (below mean)

Example Calculation:

Given: Mean μ = 100, Std Dev σ = 15, Score X = 115

Z = (115 − 100) / 15 = 1.0

Meaning: Score is 1 standard deviation above mean, indicating better-than-average performance.

30 of 31

The Central Limit Theorem (CLT)

For a sufficiently large sample size n, the distribution of the sample mean is approximately normal, regardless of the population’s original distribution.

Population vs. Sampling Distribution

Key Concept:

If we sample a distribution many times, the set of sample means becomes normally distributed.

Key Observations

  • Applies to sample means, not individual values.
  • Larger sample size n leads to a better normal approximation.

CLT Parameters

Standardized Z-Score for CLT

31 of 31

SUMMARY

What Should You Leave With?

CMSC 320 • Basic Probability & Distributions

24

Conditioning

New information changes the relevant probability.

Bayes

Update beliefs using evidence.

Distributions

Match the model to the data-generating process.

Expected value

Compare long-run outcomes and decisions.

Normal + Z

Reason about center, spread, and standardized distance.

CLT

Explains why sample means support statistical inference.

Next step: use these ideas to describe and reason about real datasets.