1 of 29

CMSC 320 • PART 03

Hypothesis Testing:

Different Types of Statistical Tests

Instructor: Fardina F Alam

2 of 29

Course Overview

Topics We Will Cover

Primary Focus

Parametric Tests for Means

These tests are used to compare means. They assume that your data follows a normal distribution and that you have a sufficient sample size.

1 Z-Test

2 T-Test

3 One vs Two Sample Test

4 Paired Sample T-Test

5 ANOVA

Alternative

Non-Parametric Tests

Used when distribution assumptions are not met.

1 Chi-squared test

​

More …..

Follow-Up

Post-Hoc Analysis

Determines specific group differences after overall ANOVA testing.

Key Objective:

Identify which specific means differ from each other.

3 of 29

Hypothesis Testing Overview

Parametric vs. Non-Parametric Tests

A statistical hypothesis test is a method of statistical inference used to decide whether the data at hand sufficiently support a particular hypothesis.

ASSUMES DISTRIBUTION

Parametric Tests

  • Assume the data follow a particular distribution (often normal).
  • Usually make assumptions about population parameters.
  • Often more powerful when assumptions are satisfied.

Examples: t-test, ANOVA

DISTRIBUTION-FREE

Non-Parametric Tests

  • Make fewer assumptions about the underlying distribution.
  • Often use ranks rather than original values.
  • Useful for skewed, ordinal, or non-normal data.

Examples: Mann–Whitney U, Wilcoxon, Kruskal–Wallis

Key Idea: Choose the test based on your data, research question, and assumptions.

In hypothesis testing, a test statistic reduces sample data to a single value (such as a Z, t, or F value). This statistic is then compared to a critical value or p-value to decide whether to reject the null hypothesis.

4 of 29

Statistical Hypothesis Tests Overview

Test

What are we comparing?

Key Condition

When to Use — Example

Z-test

Sample mean vs. population mean

Population SD σ known; data approx. normal or sample large (typically n>30)

Is average package weight different from 500 g?

One-sample t-test

Sample mean vs. reference mean

Numerical data; Population σ unknown

Is average sleep time different from 7 hours?

Two-sample t-test

Means of 2 independent groups

Independent groups, σ unknown

Do men and women have different average heights?

Paired t-test

Means of 2 related measurements

Same/matched subjects

Is blood pressure different before vs. after treatment?

ANOVA

Means of 3+ groups

Multiple independent groups

Do 3 teaching methods produce different mean scores?

Chi-square (χ²)

Categorical counts

Categorical variables

Is product preference related to age group?

Mann–Whitney U

2 independent groups

Non-parametric

Compare satisfaction ratings from two independent groups

Wilcoxon signed-rank

2 related measurements

Non-parametric

Compare ratings before vs. after an intervention

Kruskal–Wallis

3+ independent groups

Non-parametric

Compare satisfaction ratings across 3+ groups

Means + assumptions reasonable → Z / t / ANOVA  |  Rank-based/non-parametric approach → Mann–Whitney / Wilcoxon / Kruskal–Wallis  |  Categorical counts → χ²

5 of 29

Z-Test vs. T-Test: When to Use

Z-Test

Use when population standard deviation (σ) is known.

  • Tests a hypothesis about a population mean
  • Population σ is known
  • Data are approximately normal or sample size is large enough for the CLT
  • Uses the standard normal (Z) distribution

Z = (x̄ - μ) / (σ / √n)

Think: σ known → Z-test

T-Test

Use when population standard deviation (σ) is unknown.

  • Tests hypotheses about population mean(s)
  • Population σ is unknown
  • Uses the sample standard deviation (s) instead
  • For small samples, approximate normality is important
  • Can also be used with large samples
  • Uses the t-distribution

t = (x̄ - μ) / (s / √n)

Think: σ unknown → t-test

​

x̄ = sample mean | μ = population mean | σ = population std. deviation | s = sample std. deviation | n = sample size

​

KEY TAKEAWAY: Main distinction is σ known vs. unknown, not sample size alone.

6 of 29

Hypothesis Testing Context

Comparing Male vs. Female Heights

We have noticed most humans fall into one of two distinct categories—male or female.

We would like to know if our sample of males is taller than our sample of females.

Key Research Question

Can we just take the average of the two samples?

​

7 of 29

Types of T-Tests

Choose the t-test based on what you are comparing

One-Sample T-Test

Compare one sample mean with a known/reference value.

Example: Is average sleep time different from 7 hours?

Two-Sample T-Test (Independent)

Compare the means of two independent groups.

Example: Is average height different between Group A and Group B?

Paired T-Test

Compare two related measurements from the same/matched subjects.

Example: Blood pressure before vs. after treatment.

KEY DECISION GUIDE

• One group vs. a value? → One-Sample T-Test   |   • Two independent groups? → Two-Sample T-Test

• Same subjects twice? → Paired T-Test

If the two groups have very different spreads, use Welch’s t-test.

Simply taking the average of two samples is not sufficient to determine if groups differ (doesn't account for variability). Conduct a statistical test to determine significant differences.

How to run a one-sample t-test:

import numpy as np

from scipy import stats

stats.ttest_1samp(your_data, popmean=0.5)

>>> TtestResult(statistic=2.456, pvalue=0.0176, df=49)

8 of 29

T-Test Assumptions

Core statistical requirements for valid hypothesis testing

Continuous Outcome

Outcome is numerical / continuous.

Independence of Observations

Observations are appropriately independent (except paired t-test).

Normality Assumption

Data are approximately normal, especially important for small samples.

Non-Parametric Alternative

If normality is not reasonable → consider a non-parametric alternative.

9 of 29

A difference in sample means alone does not tell us whether the difference is statistically significant—we must also consider variability and sample size.

Why Simple Averages Fail

Taking the average of two samples is not sufficient to determine if one group is taller than the other, because it fails to account for variability within each group.

The Statistical Solution

We conduct a Two-Sample T-test to evaluate whether there is a statistically significant difference between the distributions of the two groups.

10 of 29

Why Comparing Means Alone Is Not Enough

Why Simple Comparison Is Not Enough

Two groups can have different sample means, but that alone does not tell us whether the difference is statistically significant.

We also need to consider:

  • Variability — how spread out the values are
  • Sample size — how much data we have

The Statistical Solution

A Two-Sample T-Test helps determine whether the difference between the means of two independent groups is statistically significant.

Example:� Is the average height different between Group A and Group B?

​

Key Takeaway

Difference in means + variability + sample size → Statistical evidence

11 of 29

Two Sample T-Test

  • Null hypothesis: Men and women are the same height
  • Alternative hypothesis: Men and women are different heights
  • p-value: the probability that we would see these observations if the null hypothesis is true/correct

Q: What sort of p value would we see if men and women had different heights?

12 of 29

Paired Sample t test

The paired sample t test is used to compare the means of two related groups of samples.

  • It is used in a situation where you have two values (i.e., a pair of values) for the same group of samples.
  • Often these two values are measured from the same samples either at two different times, under two different conditions, or after a specific intervention.

13 of 29

Paired Sample t test: Example

The aliens monitor a bunch of humans, test them for intelligence, and then run one half of them through a machine to make them smarter. Afterwards, they want to know if their machine worked.

This would be called a paired t-test.

Null Hypothesis: ?

Alternative Hypothesis: ?

14 of 29

Paired Sample t test: Example

The aliens monitor a bunch of humans, test them for intelligence, and then run one half of them through a machine to make them smarter. Afterwards, they want to know if their machine worked.

Null Hypothesis: The machine did nothing

Alternative Hypothesis: The machine came from a different distribution

Ques: The aliens get a p-value of .05. What can they conclude?

15 of 29

Multiple Groups

The Aliens decide to kidnap humans to study, but we don’t know what humans eat! We have five different food mixes we want to try. We split the humans up into five groups and feed each group a different mix, and then measure how much the humans grow over the next few years.

Ques: How do Aliens know if the mixes have different effects?

16 of 29

Anova (Analysis of Variance) Test

ANOVA is a powerful statistical test for comparing the means of multiple groups (three or more groups (more than two)) to determine if there are significant differences among them.

We use a anova test.

  • Null hypothesis: There is no difference between any of the groups
  • Alternative hypothesis: There is a difference between at least one of the groups

Notes: In t-tests and z-tests, we typically compare means of two groups using individual datasets or assess the mean of a single group against a known value. ANOVA evaluates differences in means across three or more groups as a whole, considering both within-group and between-group variability.

17 of 29

Parametric Tests ( Comparing Means)

18 of 29

Nonparametric Tests ( Comparing Medians)

Nonparametric Hypothesis Tests Used when data do not meet assumptions of parametric tests (e.g., normality, equal variances).

  • Do not rely on population parameters like mean or variance.
  • Often based on ranking data instead of raw values.
  • Suitable for: Nominal or ordinal data, Skewed distributions, Small sample sizes

Example:

Rank-based tests → work on ranked data

  • Mann–Whitney U (2 independent groups)
  • Wilcoxon Signed-Rank (2 paired groups)
  • Kruskal–Wallis (3+ independent groups)
  • Spearman’s Rank Correlation�

Frequency-based test → work on counts in categories

  • Chi-Square (Goodness of Fit, Independence)

19 of 29

Nonparametric Tests: Alternatives to Parametric Tests

​

​

Kruskal–Wallis Test

  • Extension of the one-way ANOVA.
  • Compares medians of 3 or more independent groups.
  • Use when data are not normally distributed or variances are unequal.

​

Mann–Whitney U Test

  • Alternative to independent-samples t-test.
  • Compares two independent groups on median ranks.
  • Useful for ordinal data or non-normal distributions.

​

Wilcoxon Signed-Rank Test

  • Alternative to paired-samples t-test.
  • Compares two related/paired groups (e.g., before vs. after treatment).
  • Assesses differences in median ranks when normality is violated.

Spearman’s Rank Correlation: Measures correlation between two variables based on ranks. Useful when data are ordinal or non-normal

​

*** Nonparametric tests are mostly based on ranked data instead of raw values, making them more robust when assumptions of parametric tests are not met.

Instead of using the actual numerical values (raw scores), nonparametric tests convert the data into ranks (positions).�

Example: Raw data (exam scores) → 45, 80, 60, 90, 75�Ranked data (from smallest to largest) → 1, 2, 3, 4, 5

  • 45 → Rank 1
  • 60 → Rank 2
  • 75 → Rank 3
  • 80 → Rank 4
  • 90 → Rank 5

So, the test looks at whether groups differ in their rank distributions, not the exact values.

Why? Because ranks are less sensitive to outliers and do not require the assumption that data are normally distributed.

20 of 29

The Chi-squared test (Frequency-based test for categorical data)

Analyze categorical data to check for an association or relationship between two or more categorical variables.�

Type: Nonparametric test (compares frequencies, not raw values)�

When to use: To determine if observed frequencies differ significantly from expected frequencies in a contingency table.�

Example: Is there a relationship between gender and preference for a soda brand (Yes/No)?

  • Use Chi-Square to test if soda preference is independent of gender.

​

21 of 29

What about this?

We are monitoring birds from two different places on the planet, and get the following results:

Bird Type

Location A

Location B

Grackle

7

13

Pigeon

2

7

Sea pigeon

15

1

One of those big fish-beak things

13

0

Big long bird

22

0

Bat

3

4

Each bird type and location falls into distinct categories, making them categorical variables suitable for analysis using methods like the chi-square test.

We want to find out if two different places on Earth have the same types of birds

22 of 29

Do these locations have the same underlying bird population?

Enter the Chi Square Test! A test for checking if two sets of categorical variables come from the same distribution.

Null hypothesis: ?

Alternative hypothesis: ?

The bird populations observed in Location A and Location B are the same.

The bird populations observed in Location A and Location B are different.

23 of 29

Considerations: (How to decide an appropriate statistical test?)

  • What are you curious about?
    • Mean? Standard deviation? Frequency?
  • Is your data categorical or continuous?
  • Do you have one or two samples?
  • Is your data normally distributed?
    • If it is, you would use a parametric test. If it is very non-normal, you would use a non-parametric test
  • Is your data paired? Is there a before and after?

​

24 of 29

Post Hoc Tests for ANOVA

25 of 29

When to use: ANOVA tells us that at least one group differs, but not which groups are different.

Post-hoc analysis identifies the specific group differences after ANOVA is significant when more than two groups are compared.

Purpose

  • Compare pairs of groups�Control error from multiple comparisons
  • Identify exactly where differences occur

Post Hoc Test/ Analysis

Common Post-Hoc Tests

  • Tukey’s Honest Significant Difference (HSD)
  • Bonferroni Correction
  • Duncan’s Multiple Range Test�

Key Idea Post-hoc tests show which group means differ significantly from each other.

26 of 29

Example

Scenario: We conducted a study to compare the test scores of students from three different schools: School A, School B, and School C.

More than 2 groups (A,B,C) → Apply ANOVA

ANOVA result: p-value < 0.05 (indicating significant overall differences among schools).

27 of 29

Example

Next Step: Applied Tukey's HSD post hoc test.

Interpretation from ANOVA: School A and School B, as well as School A and School C, have significantly different test scores (p < 0.05). But there is no significant difference in test scores between School B and School C.

28 of 29

Summary:

Statistical tests let us reason about one more samples and how they relate to each other and the population.

  • z-test: A test we can use if your data is normally distributed and we have a large number of samples
  • t-test: A test we can use if your data is normally distributed and we have a small number of samples
  • Paired tests: Used when we follow specific population members through time
  • One tailed vs two tailed tests: Two tailed tests look for any difference in population; one tailed tests require us to pick a direction

29 of 29

Summary: Main Steps of Hypothesis Testing

  1. State the Null Hypothesis: Assumption what you're trying to test
  2. State the Alternative Hypothesis: what you believe might be true
  3. Pick a Level of Significance 𝛂: the probability of rejecting the null hypothesis when it's actually true. Common values for α are 0.05 or 0.01.
  4. Choose a Test: Select the right tool to check your guesses based on your data.
  5. Collect Data: Get the information/data you need through observation or experimentation.
  6. Calculate a test statistic: Using the collected data, calculate the appropriate test statistic.
  7. Calculate P-Value and compare with 𝛂: Based on the comparison, decide if you have enough evidence to believe your guess is right or if you need to keep looking.
  8. Draw a Conclusion: Based on your decision in the previous step, draw a conclusion regarding the null hypothesis. If you reject it, you accept the alternative hypothesis. If you fail to reject it, you do not have enough evidence to support the alternative hypothesis.