1 of 35

Estimating a Population Parameter with�Confidence Intervals

How to answer the questions: (i) what is the value of the population mean? and (ii) what is the value of the population proportion?

2 of 35

�Warm-Up Problem: Coffee Shop Wait Times

Scenario: A popular coffee shop claims that the average wait time for an order is 4.5 minutes, with a standard deviation of 1.2 minutes. During a busy morning, a random sample of 50 customers is taken. Assuming that the population distribution of wait times is not strongly skewed…

  1. What is the probability that the sample mean wait time exceeds 5 minutes?
  2. What is the probability that the sample mean wait time is between 4.2 and 4.8 minutes?
  3. How likely is it that the sample mean wait time is less than 4 minutes?

3 of 35

�Overview

  • A reminder on the Central Limit Theorem and Sampling Distributions
  • Descriptive Statistics versus Inferential Statistics
    • Sample Statistics and Population Parameters
  • Sampling variation and sampling error
    • The distribution of sample statistics
  • The single-sample problem
  • Confidence Intervals
  • A first confidence interval

4 of 35

The Central Limit Theorem and �Sampling Distributions

  •  

5 of 35

�Describing Samples versus Describing Populations

Populations (Inferential Statistics)

  • Population parameters
    • Population mean, μ
    • Population standard deviation, σ
    • Population proportion,  ρ 
  • Want to find or describe these parameters, but we can’t directly

Samples (Descriptive Statistics)

  • Sample statistics
    • Sample mean,
    • Sample standard deviation, s
    • Sample proportion, p or
  • Can find these statistics directly and, under certain conditions, use them to approximate population parameters

6 of 35

�Inference Road Map

Inference On…

Covered?

One Numerical Variable

One Binary Categorical Variable

Associations Between a Numerical Variable and a Binary Categorical Variable

Associations Between Two Binary Categorical Variables

One MultiClass Categorical Variable

Associations Between Two MultiClass Categorical Variables

Associations Between One Numerical Variable and One MultiClass Categorical Variable

Associations Between Two Numerical Variables

7 of 35

�Revisiting the Coffee Shop

Question: How can the coffee shop employees know that their average wait time is 4.5 minutes with a standard deviation of 1.2 minutes?

    • Perhaps their Point of Service (POS) software tracks all of this, they have complete data (a census), and they can compute it directly
    • Maybe they’ve just guessed
    • Perhaps, even without perfect information, they’ve justified this assertion with data.

We’ll investigate how with a hypothetical new coffee shop over the next several slides.

8 of 35

�A Second Location

  •  

9 of 35

�Sampling Wait Times

After their grand opening, the shop employees begin investigating wait times at the new location. They randomly sample 8 customers and observe their wait times.

Wait Times (min): 7.2, 5.8, 6.2, 5.2, 4.7, 6.3, 7.5, 4.5

    • The average wait time is about 5.9 minutes
    • Is the average wait time at the new shop 5.9 minutes?

New Wait Times (min): 5.8, 6.2, 3.8, 6.9, 4.2, 5.5, 5.3, 5.3 (mean ~5.3)

A Third Sample (min): 3.0, 6.2, 4.6, 7.1, 6.9, 3.2, 4.7, 5.8 (mean ~5.2)

10 of 35

�Sampling Wait Times: Sample Variation

We see that the average wait times vary from one sample to the next

This is called sampling variation, and it is unavoidable

So, what can the results of a sample actually tell us?

We’ll take a look at the results visually on the right

11 of 35

�Sampling Wait Times: Sample Variation

We see that the average wait times vary from one sample to the next

This is called sampling variation, and it is unavoidable

So, what can the results of a sample actually tell us?

We’ll take a look at the results visually on the right

12 of 35

�Sampling Wait Times: Sample Variation

We see that the average wait times vary from one sample to the next

This is called sampling variation, and it is unavoidable

So, what can the results of a sample actually tell us?

We’ll take a look at the results visually on the right

13 of 35

�Sampling Wait Times: Sample Variation

We see that the average wait times vary from one sample to the next

This is called sampling variation, and it is unavoidable

So, what can the results of a sample actually tell us?

We’ll take a look at the results visually on the right

Here are the results of 97 more samples

14 of 35

�Sampling Wait Times: Sample Variation

We see that the average wait times vary from one sample to the next

This is called sampling variation, and it is unavoidable

What can we take away from the average wait times from these 100 samples of 8 customers each?

15 of 35

�Sampling Wait Times: Sample Variation

Samples of only 8 customers are quite small, leaving lots of room for uncertainty

Let’s revisit this with samples of 72 customers each

Here are the average wait times resulting from 100 samples of 72 customers each

What do you feel comfortable claiming now?

16 of 35

�The Distribution of Sample Average Wait Times

Let’s take a look at the distributions of the average wait times for our samples

What do you notice?

17 of 35

�The Distribution of Sample Average Wait Times

In case you aren’t convinced, here are collections of 50,000 samples of size 8 and 50,000 samples of size 72

Can you estimate the average wait time at the new location?

18 of 35

�A Serious Problem…

The strategy we’ve just walked through is unreasonable

We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)

Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?

19 of 35

�A Serious Problem…

The strategy we’ve just walked through is unreasonable

We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)

Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?

20 of 35

�A Serious Problem…

The strategy we’ve just walked through is unreasonable

We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)

Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?

21 of 35

�A Serious Problem…

The strategy we’ve just walked through is unreasonable

We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)

Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?

22 of 35

Confidence Intervals:�Estimating Population Parameters

  •  

23 of 35

�Visual Intuition for Confidence Intervals

Before we work through the construction of a confidence interval, let’s build some intuition

What does a confidence interval actually mean?

24 of 35

Confidence Level and Confidence Intervals

How does the confidence level impact the confidence intervals?

Confidence intervals with lower levels of confidence are narrower, but also more likely to miss the parameter

25 of 35

�Summary (So Far)

  •  

26 of 35

�Example: Coffee Shop Wait Times

  •  

27 of 35

�Example: Coffee Shop Wait Times

Scenario: The Coffee Shop owners need to know the average customer wait time at their second location so that they can determine their guarantee. As a reminder, they provide an XX-minute guarantee, where customers who haven’t been served in the guaranteed time are provided a 50%-off coupon for their next visit. The owners are willing to assume that the standard deviation in wait times at this new location will be the same as it is at their existing location, 1.2 minutes. The shop employees monitor a random sample of 33 customer wait times and observe an average wait time of 5 minutes and 48 seconds (5.8 minutes). Construct a 95% confidence interval for the true average customer wait time at the new shop.

We are 95% confident that the true, population mean customer wait time at the new shop is between 5.39 minutes and 6.21 minutes

Note. This interval alone doesn’t address the owners’ needs, but we’ll leave the remaining work up to them.

28 of 35

�Summary (Repeated)

  •  

29 of 35

Example Scenario I: �Heart Rate During Horror Movies

Scenario: A psychology lab is conducting a study on fear responses and wants to estimate the average peak heart rate (in beats per minute) of individuals watching horror movies. They monitor 42 participants during a particularly intense scene and record an average peak heart rate of 117 bpm. Based on prior physiological research, they assume a population standard deviation of 15 bpm for heart rate responses to startling stimuli. Construct and interpret a 90% confidence interval for the average peak heart rate of individuals watching horror movies.

30 of 35

Example Scenario II:�Flight Time of Drones

Scenario: A drone manufacturer is testing a new battery model to estimate the average flight time (in minutes) before a battery dies. From a sample of 30 test flights, the company records an average battery life of 73.4 minutes. Past battery testing has determined that flight time varies with a standard deviation of 6.8 minutes, a value that has been stable across multiple generations of batteries. Find an interval which you are 98% certain contains the average flight time for these drones.

31 of 35

Example Scenario III:�Average Daily Steps of Professional Dancers

Scenario: A research team studying movement patterns of professional dancers wants to estimate the mean number of steps taken per day. They track a sample of 50 professional dancers for a week and find an average of 19,726 steps per day. Based on previous large-scale studies of dancers, the population standard deviation is assumed to be 2,513 steps per day. Estimate, with 95% confidence, the true average daily step count for professional dancers.

32 of 35

Example Scenario IV:�Caffeine Content in Energy Drinks

Scenario: A health organization is concerned about caffeine intake and wants to estimate the average caffeine content (in mg) of a popular energy drink brand with marketing attractive to teens. A random sample of 36 cans is analyzed, revealing a mean caffeine content of 154 mg per can. The manufacturer has provided independent lab results confirming a population standard deviation of 12 mg, as their production process tightly controls caffeine levels. Find and interpret a 98% confidence interval for the average caffeine content in an energy drink from this brand.

33 of 35

Example Scenario V:�Speed of Rollercoasters

Scenario: An amusement park enthusiast group wants to estimate the average top speed (in mph) of roller coasters in the U.S. They collect data from 45 different roller coasters and find a mean top speed of 67.8 mph. Ride engineers have studied coaster speeds for years and have determined the population standard deviation to be 9.5 mph, making it reasonable to assume this value. Construct and interpret a 90% confidence interval for the average top speed of rollercoasters in the U.S.

34 of 35

Scenario VI:�Lifespan of Earbuds on a Single Charge

Scenario: A tech review website is testing a new model of wireless earbuds to estimate the mean battery life (in hours) on a full charge. They run a sample of 35 earbuds through battery drain tests, finding a mean battery life of 6.9 hours. The population standard deviation is assumed to be 1.2 hours, based on extensive previous testing of similar models. Find a 98% confidence interval for the average battery life of this model and provide an interpretation of that interval.

35 of 35

�Next Time…

  •