CMSC 320 • PART 03
Hypothesis Testing:
Different Types of Statistical Tests
Instructor: Fardina F Alam
Course Overview
Topics We Will Cover
Primary Focus
Parametric Tests for Means
These tests are used to compare means. They assume that your data follows a normal distribution and that you have a sufficient sample size.
1 Z-Test
2 T-Test
3 One vs Two Sample Test
4 Paired Sample T-Test
5 ANOVA
Alternative
Non-Parametric Tests
Used when distribution assumptions are not met.
1 Chi-squared test
More …..
Follow-Up
Post-Hoc Analysis
Determines specific group differences after overall ANOVA testing.
Key Objective:
Identify which specific means differ from each other.
Hypothesis Testing Overview
Parametric vs. Non-Parametric Tests
A statistical hypothesis test is a method of statistical inference used to decide whether the data at hand sufficiently support a particular hypothesis.
ASSUMES DISTRIBUTION
Parametric Tests
Examples: t-test, ANOVA
DISTRIBUTION-FREE
Non-Parametric Tests
Examples: Mann–Whitney U, Wilcoxon, Kruskal–Wallis
Key Idea: Choose the test based on your data, research question, and assumptions.
In hypothesis testing, a test statistic reduces sample data to a single value (such as a Z, t, or F value). This statistic is then compared to a critical value or p-value to decide whether to reject the null hypothesis.
Statistical Hypothesis Tests Overview
Test | What are we comparing? | Key Condition | When to Use — Example |
Z-test | Sample mean vs. population mean | Population SD σ known; data approx. normal or sample large (typically n>30) | Is average package weight different from 500 g? |
One-sample t-test | Sample mean vs. reference mean | Numerical data; Population σ unknown | Is average sleep time different from 7 hours? |
Two-sample t-test | Means of 2 independent groups | Independent groups, σ unknown | Do men and women have different average heights? |
Paired t-test | Means of 2 related measurements | Same/matched subjects | Is blood pressure different before vs. after treatment? |
ANOVA | Means of 3+ groups | Multiple independent groups | Do 3 teaching methods produce different mean scores? |
Chi-square (χ²) | Categorical counts | Categorical variables | Is product preference related to age group? |
Mann–Whitney U | 2 independent groups | Non-parametric | Compare satisfaction ratings from two independent groups |
Wilcoxon signed-rank | 2 related measurements | Non-parametric | Compare ratings before vs. after an intervention |
Kruskal–Wallis | 3+ independent groups | Non-parametric | Compare satisfaction ratings across 3+ groups |
Means + assumptions reasonable → Z / t / ANOVA | Rank-based/non-parametric approach → Mann–Whitney / Wilcoxon / Kruskal–Wallis | Categorical counts → χ²
Z-Test vs. T-Test: When to Use
Z-Test
Use when population standard deviation (σ) is known.
Z = (x̄ - μ) / (σ / √n)
Think: σ known → Z-test
T-Test
Use when population standard deviation (σ) is unknown.
t = (x̄ - μ) / (s / √n)
Think: σ unknown → t-test
x̄ = sample mean | μ = population mean | σ = population std. deviation | s = sample std. deviation | n = sample size
KEY TAKEAWAY: Main distinction is σ known vs. unknown, not sample size alone.
Hypothesis Testing Context
Comparing Male vs. Female Heights
We have noticed most humans fall into one of two distinct categories—male or female.
We would like to know if our sample of males is taller than our sample of females.
Key Research Question
Can we just take the average of the two samples?
Types of T-Tests
Choose the t-test based on what you are comparing
One-Sample T-Test
Compare one sample mean with a known/reference value.
Example: Is average sleep time different from 7 hours?
Two-Sample T-Test (Independent)
Compare the means of two independent groups.
Example: Is average height different between Group A and Group B?
Paired T-Test
Compare two related measurements from the same/matched subjects.
Example: Blood pressure before vs. after treatment.
KEY DECISION GUIDE
• One group vs. a value? → One-Sample T-Test | • Two independent groups? → Two-Sample T-Test
• Same subjects twice? → Paired T-Test
If the two groups have very different spreads, use Welch’s t-test.
Simply taking the average of two samples is not sufficient to determine if groups differ (doesn't account for variability). Conduct a statistical test to determine significant differences.
How to run a one-sample t-test:
import numpy as np
from scipy import stats
stats.ttest_1samp(your_data, popmean=0.5)
>>> TtestResult(statistic=2.456, pvalue=0.0176, df=49)
T-Test Assumptions
Core statistical requirements for valid hypothesis testing
Continuous Outcome
Outcome is numerical / continuous.
Independence of Observations
Observations are appropriately independent (except paired t-test).
Normality Assumption
Data are approximately normal, especially important for small samples.
Non-Parametric Alternative
If normality is not reasonable → consider a non-parametric alternative.
A difference in sample means alone does not tell us whether the difference is statistically significant—we must also consider variability and sample size.
Why Simple Averages Fail
Taking the average of two samples is not sufficient to determine if one group is taller than the other, because it fails to account for variability within each group.
The Statistical Solution
We conduct a Two-Sample T-test to evaluate whether there is a statistically significant difference between the distributions of the two groups.
Why Comparing Means Alone Is Not Enough
Why Simple Comparison Is Not Enough
Two groups can have different sample means, but that alone does not tell us whether the difference is statistically significant.
We also need to consider:
The Statistical Solution
A Two-Sample T-Test helps determine whether the difference between the means of two independent groups is statistically significant.
Example:� Is the average height different between Group A and Group B?
Key Takeaway
Difference in means + variability + sample size → Statistical evidence
Two Sample T-Test
Q: What sort of p value would we see if men and women had different heights?
Paired Sample t test
The paired sample t test is used to compare the means of two related groups of samples.
Paired Sample t test: Example
The aliens monitor a bunch of humans, test them for intelligence, and then run one half of them through a machine to make them smarter. Afterwards, they want to know if their machine worked.
This would be called a paired t-test.
Null Hypothesis: ?
Alternative Hypothesis: ?
Paired Sample t test: Example
The aliens monitor a bunch of humans, test them for intelligence, and then run one half of them through a machine to make them smarter. Afterwards, they want to know if their machine worked.
Null Hypothesis: The machine did nothing
Alternative Hypothesis: The machine came from a different distribution
Ques: The aliens get a p-value of .05. What can they conclude?
Multiple Groups
The Aliens decide to kidnap humans to study, but we don’t know what humans eat! We have five different food mixes we want to try. We split the humans up into five groups and feed each group a different mix, and then measure how much the humans grow over the next few years.
Ques: How do Aliens know if the mixes have different effects?
Anova (Analysis of Variance) Test
ANOVA is a powerful statistical test for comparing the means of multiple groups (three or more groups (more than two)) to determine if there are significant differences among them.
We use a anova test.
Notes: In t-tests and z-tests, we typically compare means of two groups using individual datasets or assess the mean of a single group against a known value. ANOVA evaluates differences in means across three or more groups as a whole, considering both within-group and between-group variability.
Parametric Tests ( Comparing Means)
Nonparametric Tests ( Comparing Medians)
Nonparametric Hypothesis Tests Used when data do not meet assumptions of parametric tests (e.g., normality, equal variances).
Example:
Rank-based tests → work on ranked data
Frequency-based test → work on counts in categories
Nonparametric Tests: Alternatives to Parametric Tests
Kruskal–Wallis Test
Mann–Whitney U Test
Wilcoxon Signed-Rank Test
Spearman’s Rank Correlation: Measures correlation between two variables based on ranks. Useful when data are ordinal or non-normal
*** Nonparametric tests are mostly based on ranked data instead of raw values, making them more robust when assumptions of parametric tests are not met.
Instead of using the actual numerical values (raw scores), nonparametric tests convert the data into ranks (positions).�
Example: Raw data (exam scores) → 45, 80, 60, 90, 75�Ranked data (from smallest to largest) → 1, 2, 3, 4, 5
So, the test looks at whether groups differ in their rank distributions, not the exact values.
Why? Because ranks are less sensitive to outliers and do not require the assumption that data are normally distributed.
The Chi-squared test (Frequency-based test for categorical data)
Analyze categorical data to check for an association or relationship between two or more categorical variables.�
Type: Nonparametric test (compares frequencies, not raw values)�
When to use: To determine if observed frequencies differ significantly from expected frequencies in a contingency table.�
Example: Is there a relationship between gender and preference for a soda brand (Yes/No)?
What about this?
We are monitoring birds from two different places on the planet, and get the following results:
Bird Type | Location A | Location B |
Grackle | 7 | 13 |
Pigeon | 2 | 7 |
Sea pigeon | 15 | 1 |
One of those big fish-beak things | 13 | 0 |
Big long bird | 22 | 0 |
Bat | 3 | 4 |
Each bird type and location falls into distinct categories, making them categorical variables suitable for analysis using methods like the chi-square test.
We want to find out if two different places on Earth have the same types of birds
Do these locations have the same underlying bird population?
Enter the Chi Square Test! A test for checking if two sets of categorical variables come from the same distribution.
Null hypothesis: ?
Alternative hypothesis: ?
The bird populations observed in Location A and Location B are the same.
The bird populations observed in Location A and Location B are different.
Considerations: (How to decide an appropriate statistical test?)
Post Hoc Tests for ANOVA
When to use: ANOVA tells us that at least one group differs, but not which groups are different.
Post-hoc analysis identifies the specific group differences after ANOVA is significant when more than two groups are compared.
Purpose
Post Hoc Test/ Analysis
Common Post-Hoc Tests
Key Idea Post-hoc tests show which group means differ significantly from each other.
Example
Scenario: We conducted a study to compare the test scores of students from three different schools: School A, School B, and School C.
More than 2 groups (A,B,C) → Apply ANOVA
ANOVA result: p-value < 0.05 (indicating significant overall differences among schools).
Example
Next Step: Applied Tukey's HSD post hoc test.
Interpretation from ANOVA: School A and School B, as well as School A and School C, have significantly different test scores (p < 0.05). But there is no significant difference in test scores between School B and School C.
Summary:
Statistical tests let us reason about one more samples and how they relate to each other and the population.
Summary: Main Steps of Hypothesis Testing