1 of 17

Part 01

Introduction to Hypothesis Testing

Fardina F Alam

CMSC 320: Introduction to Data Science

2 of 17

Topics We Will Cover

CMSC320 Textbook Chapter: Chapter 8

01

What is Hypothesis Testing

02

Null and Alternative Hypothesis

03

Level of Significance 𝛂

04

Collect Data using Different Sampling Methods

05

Type I and II Error

3 of 17

“Data alone is not interesting.

It is the interpretation of the data that we are really interested in.”

Hypothesis Testing (Making Informed Decisions with Data)

4 of 17

Hypothesis Testing (HT)

Hypothesis testing is a statistical method used to make informed decisions or draw conclusions about a population based on a sample of data.

HOW HYPOTHESIS TESTING WORKS

1. Make assumption about population

2. Use sample to test assumption

Population: The entire group that you are interested in studying.

Sample: A subset of the population used for analysis.

5 of 17

Main Steps of Hypothesis Testing

  1. State the Null Hypothesis
  2. State the Alternative Hypothesis
  3. Pick a Level of Significance 𝛂
  4. Choose a Test
  5. Collect Data
  6. Calculate a test statistic
  7. Calculate P-Value and compare with 𝛂
  8. Draw a Conclusion

6 of 17

Collect Data

Obtain a representative sample from the population.

Remember the importance of recognizing whether data is collected through an experimental design or observational study.

7 of 17

Sampling Method

Random sampling� • Every unit has equal chance of selection.

A sampling method is a process by which individual items/event (observational units) are selected from the population to be included in the sample. Common sampling methods:

Stratified sampling� • Population divided into meaningful groups (strata).� • Random samples taken from each group.

Cluster sampling� • Population divided into clusters (e.g., locations).� • Randomly select clusters and include all units inside.

Systematic sampling� • Random start → select every kth unit.

Convenience sampling (non-probability)� • Select easiest or most available participants.� • Fast but may introduce bias.

8 of 17

Examples: Sampling Methods

​

  • Random sampling: Choosing names randomly from a list for a survey
  • Stratified sampling: Surveying transportation preferences involves dividing the population into three income strata—low, middle, and high-income households—and then randomly selecting households from each stratum in a city.
  • Cluster sampling: Estimating the average income of households in a city by dividing the city into clusters based on neighborhoods or districts, then randomly selecting specific neighborhoods as clusters and surveying all households within the selected neighborhoods.
  • Systematic sampling: Selecting every 10th person from a list of customers
  • Convenience sampling: Surveying people in a shopping mall (ease of access, availability), Recruiting volunteers from a specific organization or community etc.

9 of 17

Two Hypotheses in Hypothesis Testing

Hypothesis: a claim we test using data (about a population).

  • Null hypothesis (H₀)

no effect / no difference (current assumption)

  • Alternative hypothesis (Hₐ)

real effect / difference (opposite of H₀)

​

Possible Outcomes of a Test

  1. Reject H₀ Evidence suggests a real effect or difference (support Hₐ)
  2. Fail to reject H₀ Not enough evidence to support Hₐ

(We never “prove” H₀ true.)

Example: Candy Machine H₀ & Hₐ are mutually exclusive

Let’s say, it is believed that a candy making machine makes chocolate bars that are on average 5 gram in weight. A worker claims that the machine after maintenance no longer makes 5 gram bar. Write down H₀ and Hₐ.

​

Practice: Doctors believe that the average teen sleeps on average of no longer than 10 hours per day. A researcher believes that teens on average sleep longer. Write down H₀ and Hₐ.

10 of 17

Significance Level (α): How Much Evidence Do We Require?

Significance Level (α): The probability of rejecting the null hypothesis when it is actually true (Type I error).

Common Choices: α = 0.05 (5% Type I error rate) or 0.01 (1% Type I error rate) |  Key Rule: Smaller α → stronger evidence required to reject H₀ ; always choose α before analyzing the data.

Example: New Headache Treatment

H₀ (Null): The treatment has no effect.

H₁ (Alt): The treatment reduces headaches.

Set α = 0.05�Accepts a 5% Type I error rate (rejecting H₀ falsely).

After the study: p-value = 0.03

Decision & Conclusion

p-value (0.03) < α (0.05)

→ Reject H₀

CONCLUSION

The result is statistically significant. There is sufficient evidence to conclude that the treatment reduces headaches.

A threshold chosen before analyzing results to decide how much evidence is required against H₀.

11 of 17

Collect Data → Choose a Test → Compute the Test Statistic

Compute the Test Statistic From sample data, calculate a number that measures how far the sample result is from what H₀ predicts

Common statistics: z, t

​

Interpretation: Larger ∣test statistic∣| sample is far from H₀ |→ stronger evidence against H₀

Example: Compare two means → compute a t-statistic

Choose a Statistical Test Choose based on the research question, data type, and study design:

  • Compare 2 means → t-test
  • Compare 3+ means → ANOVA
  • Compare proportions → Chi-square/Proportion test
  • Relationship between variables → Correlation / Regression

12 of 17

​

After computing the test statistic, decide whether to reject H₀. Two equivalent ways to make a decision:

1. Critical Value Approach

Compare the calculated test statistic directly to a predetermined threshold critical value (cutoff).

2. p-value Approach

Determine the probability of obtaining sample results as extreme as the observed data.

Calculate p-value and compare it to significance level α.

cheBoth methods always yield the same conclusion when using the same test and α.

LATER TOPIC Decision Methods in Hypothesis Testing

13 of 17

LATER TOPIC Decision Methods in Hypothesis Testing

Rejection Region

(Critical Value Approach)

Focuses on values that are too extreme if H₀ were true.

If test statistic falls in:

→ Rejection Region: Reject H₀

→ Otherwise: Fail to Reject H₀

P-value Approach

(Probability Probability)

Probability of observing results this extreme if H₀ is true.

Decision Rule:

→ p ≤ α : Reject H₀

→ p > α : Fail to Reject H₀

Key Idea More extreme result → stronger evidence against H₀

14 of 17

How Confident Are You With Your Decision?

Hypotheses under test

H₀: μ = 5g | Hₐ: μ ≠ 5g

Possible Outcomes

Reject H₀ or Fail to Reject H₀

1. The Experiment

We sample 50 chocolate bars to calculate the average sample mass. Next, we calculated:

  • Test statistic: depends on what type of problem you have
    • Why? Significance: Evaluates if data is extreme enough to reject H₀.

2. Observed Daily Trials

Suppose we receive these average values:

Monday → 5.12 grams

Wednesday → 5.75 grams

Friday → 7.82 grams

How confident can we be that we should reject H₀?

To answer this, we need a CONCRETE METHOD to evaluate the strength of our evidence.

15 of 17

Confidence Level (C) & Significance Level (α)

How confident are we in our decision?

Remember: Significance Level (α) is the risk of being wrong (rejecting a true H₀)

Confidence Level (C)

  • Degree of certainty in our decision
  • Common values: 95%, 99%
  • Higher confidence → more certainty in rejecting H₀

Relationship: α = 1 − C

Example: 95% confidence → α = 0.05

Interpretation

  • Higher confidence → stricter test → harder to reject H₀
  • Lower confidence → easier to reject H₀

How to decide C or Alpha value?

The choice depends on several critical factors:

  • Nature of the research (exploratory vs. clinical)
  • Consequences of making Type I & Type II errors
  • Standard practices established in the field

Use these factors as a general guide for choosing the most appropriate threshold for your hypothesis test.

16 of 17

Type I and Type II Errors

Common mistakes we can make in hypothesis testing

Type I Error (False Positive)

Reject H₀ when it is actually true

  • Think something changed, but it didn't
  • Probability of occurrence = α (alpha)
  • Acts as a false alarm

Type II Error (False Negative)

Fail to reject H₀ when it is actually false

  • Miss a real change or effect
  • Acts as a missed detection

​

​

These errors are crucial to consider because they directly affect the validity of conclusions drawn from statistical tests.

Practice Scenario

A company states their machine makes straws of 4mm diameter. A worker believes this is no longer true, sampling 100 straws at 99% confidence.

Task: Write down H₀, Hₐ, N, C, and α.

17 of 17

Example: What is Type I and II error?

Scenario: Let's say you're testing whether a new drug is effective in reducing blood pressure.

Your null hypothesis (H0): the drug has no effect on blood pressure

Your alternative hypothesis (H1): the drug reduces blood pressure on avg.

Collect data from a sample of patients and conduct a statistical test.