1 of 14

Lecture 21

Examples

DATA 8

Spring 2024

2 of 14

Announcements

  • HW 07 due Wednesday 3/6 at 5pm
  • Midterm on Friday 03/08 at 7pm
    • Midterm preparation, Past Exams

3 of 14

Weekly Goals

  • Monday
    • Causation
    • Randomized Control Experiments
  • Today
    • P-Value as an Error
    • Examples
  • Friday
    • Midterm review

4 of 14

Recap

  • Making the wrong decision with a hypothesis test
  • p-value cutoff as an error probability
  • p-value cutoff vs. p-value

5 of 14

Definition of the p-value

Formal name: observed significance level

The p-value is the chance (probability),

  • under the null hypothesis,
  • that the test statistic
  • is equal to the value that was observed in the data
  • or is even further in the direction of the alternative.

6 of 14

Coin Toss Example from Last Lecture

There are 1000 students in Data 8. Each student tests

Null: The coin is fair

Alternative: The coin is unfair

  • based on 10,000 tosses of a coin,
  • the statistic | number of heads - 5,000 |,
  • and the 5% cutoff for the P-value.

Suppose all 10,000 coins are fair. About how many students will conclude that their coins are unfair?

7 of 14

Statistic Simulated Under the Null

About 5% of the area is to the right of the gold line

8 of 14

An Error Probability

  • The cutoff for the P-value is an error probability.

  • If:
    • your cutoff is 5%
    • and the null hypothesis happens to be true

  • then there is about a 5% chance that your test will reject the null hypothesis.

9 of 14

P-value cutoff vs P-value

  • P-value cutoff
    • Does not depend on observed data or simulation
    • Decide on it before seeing the results
    • Conventional values at 5% and 1%
    • Probability of hypothesis testing making an error
  • P-value
    • Depends on the observed data and simulation
    • Probability under the null hypothesis that the test statistic is the observed value or further towards the alternative

10 of 14

P-Value cutoff

Section 11.3.8 from the textbook

The method of statistical testing – choosing between hypotheses based on data in random samples – was developed by Sir Ronald Fisher in the early 20th century. …. About the 5% level, he wrote, “It is convenient to take this point as a limit in judging whether a deviation is to be considered significant or not.”

(Statistical Methods for Research Workers, Ronald Fisher (1925)

“If one in twenty does not seem high enough odds, we may, if we prefer it draw the line at one in fifty (the 2 percent point), or one in a hundred (the 1 percent point). Personally, the author prefers to set a low standard of significance at the 5 percent point …”

11 of 14

More on Hypothesis Tests

12 of 14

Discussion Question

Manufacturers of Super Soda run a taste test. 91 out of 200 tasters prefer Super Soda over its rival

Question: Do fewer people prefer Super Soda than its rival, or is this just chance?

Null hypothesis:

Equal proportions of the population prefer Super Soda and the Rival.

Alternative hypothesis:

Fewer people in the population prefer Super Soda.

Test statistic: observed # of people who prefer Super Soda

p-value: Probability of seeing 91 or fewer people

(Demo)

13 of 14

Hypothesis Test Concerns

The outcome of a hypothesis test can be affected by:

  • The hypotheses you investigate: �How do you define your null distribution?
  • The test statistic you choose: �How do you measure a difference between samples?
  • The empirical distribution of the statistic under the null:�How many times do you simulate under the null distribution?
  • The data you collected:�Did you happen to collect a sample that is similar to the population?
  • The truth:�If the alternative hypothesis is true, how extreme is the difference?

14 of 14

Hypothesis Test Effects

  • Number of simulations:
    • As large as possible.
    • Empirical distribution -> True Distribution
    • No new data needs to be collected.
  • Number of observations:
    • A larger sample will lead you to reject the null more reliably if the alternative is in fact true.
  • Difference from the null:
    • If the truth is similar to the null hypothesis, then even a large sample may not provide enough evidence to reject the null.