Estimating a Population Parameter with�Confidence Intervals
How to answer the questions: (i) what is the value of the population mean? and (ii) what is the value of the population proportion?
�Warm-Up Problem: Coffee Shop Wait Times
Scenario: A popular coffee shop claims that the average wait time for an order is 4.5 minutes, with a standard deviation of 1.2 minutes. During a busy morning, a random sample of 50 customers is taken. Assuming that the population distribution of wait times is not strongly skewed…
�Overview
The Central Limit Theorem and �Sampling Distributions
�Describing Samples versus Describing Populations
Populations (Inferential Statistics)
Samples (Descriptive Statistics)
�Inference Road Map
Inference On… | Covered? |
One Numerical Variable | |
One Binary Categorical Variable | |
Associations Between a Numerical Variable and a Binary Categorical Variable | |
Associations Between Two Binary Categorical Variables | |
One MultiClass Categorical Variable | |
Associations Between Two MultiClass Categorical Variables | |
Associations Between One Numerical Variable and One MultiClass Categorical Variable | |
Associations Between Two Numerical Variables | |
�Revisiting the Coffee Shop
Question: How can the coffee shop employees know that their average wait time is 4.5 minutes with a standard deviation of 1.2 minutes?
We’ll investigate how with a hypothetical new coffee shop over the next several slides.
�A Second Location
�Sampling Wait Times
After their grand opening, the shop employees begin investigating wait times at the new location. They randomly sample 8 customers and observe their wait times.
Wait Times (min): 7.2, 5.8, 6.2, 5.2, 4.7, 6.3, 7.5, 4.5
New Wait Times (min): 5.8, 6.2, 3.8, 6.9, 4.2, 5.5, 5.3, 5.3 (mean ~5.3)
A Third Sample (min): 3.0, 6.2, 4.6, 7.1, 6.9, 3.2, 4.7, 5.8 (mean ~5.2)
�Sampling Wait Times: Sample Variation
We see that the average wait times vary from one sample to the next
This is called sampling variation, and it is unavoidable
So, what can the results of a sample actually tell us?
We’ll take a look at the results visually on the right
�Sampling Wait Times: Sample Variation
We see that the average wait times vary from one sample to the next
This is called sampling variation, and it is unavoidable
So, what can the results of a sample actually tell us?
We’ll take a look at the results visually on the right
�Sampling Wait Times: Sample Variation
We see that the average wait times vary from one sample to the next
This is called sampling variation, and it is unavoidable
So, what can the results of a sample actually tell us?
We’ll take a look at the results visually on the right
�Sampling Wait Times: Sample Variation
We see that the average wait times vary from one sample to the next
This is called sampling variation, and it is unavoidable
So, what can the results of a sample actually tell us?
We’ll take a look at the results visually on the right
Here are the results of 97 more samples
�Sampling Wait Times: Sample Variation
We see that the average wait times vary from one sample to the next
This is called sampling variation, and it is unavoidable
What can we take away from the average wait times from these 100 samples of 8 customers each?
�Sampling Wait Times: Sample Variation
Samples of only 8 customers are quite small, leaving lots of room for uncertainty
Let’s revisit this with samples of 72 customers each
Here are the average wait times resulting from 100 samples of 72 customers each
What do you feel comfortable claiming now?
�The Distribution of Sample Average Wait Times
Let’s take a look at the distributions of the average wait times for our samples
What do you notice?
�The Distribution of Sample Average Wait Times
In case you aren’t convinced, here are collections of 50,000 samples of size 8 and 50,000 samples of size 72
Can you estimate the average wait time at the new location?
�A Serious Problem…
The strategy we’ve just walked through is unreasonable
We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)
Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?
�A Serious Problem…
The strategy we’ve just walked through is unreasonable
We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)
Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?
�A Serious Problem…
The strategy we’ve just walked through is unreasonable
We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)
Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?
�A Serious Problem…
The strategy we’ve just walked through is unreasonable
We don’t have the time or resources to conduct 50,000 random samples (or even 100 random samples)
Realistically: We take one sample and need to draw inferences from that one, but which sample do we get?
Confidence Intervals:�Estimating Population Parameters
�Visual Intuition for Confidence Intervals
Before we work through the construction of a confidence interval, let’s build some intuition
What does a confidence interval actually mean?
�Confidence Level and Confidence Intervals
How does the confidence level impact the confidence intervals?
Confidence intervals with lower levels of confidence are narrower, but also more likely to miss the parameter
�Summary (So Far)
�Example: Coffee Shop Wait Times
�Example: Coffee Shop Wait Times
Scenario: The Coffee Shop owners need to know the average customer wait time at their second location so that they can determine their guarantee. As a reminder, they provide an XX-minute guarantee, where customers who haven’t been served in the guaranteed time are provided a 50%-off coupon for their next visit. The owners are willing to assume that the standard deviation in wait times at this new location will be the same as it is at their existing location, 1.2 minutes. The shop employees monitor a random sample of 33 customer wait times and observe an average wait time of 5 minutes and 48 seconds (5.8 minutes). Construct a 95% confidence interval for the true average customer wait time at the new shop.
We are 95% confident that the true, population mean customer wait time at the new shop is between 5.39 minutes and 6.21 minutes
Note. This interval alone doesn’t address the owners’ needs, but we’ll leave the remaining work up to them.
�Summary (Repeated)
Example Scenario I: �Heart Rate During Horror Movies
Scenario: A psychology lab is conducting a study on fear responses and wants to estimate the average peak heart rate (in beats per minute) of individuals watching horror movies. They monitor 42 participants during a particularly intense scene and record an average peak heart rate of 117 bpm. Based on prior physiological research, they assume a population standard deviation of 15 bpm for heart rate responses to startling stimuli. Construct and interpret a 90% confidence interval for the average peak heart rate of individuals watching horror movies.
Example Scenario II:�Flight Time of Drones
Scenario: A drone manufacturer is testing a new battery model to estimate the average flight time (in minutes) before a battery dies. From a sample of 30 test flights, the company records an average battery life of 73.4 minutes. Past battery testing has determined that flight time varies with a standard deviation of 6.8 minutes, a value that has been stable across multiple generations of batteries. Find an interval which you are 98% certain contains the average flight time for these drones.
Example Scenario III:�Average Daily Steps of Professional Dancers
Scenario: A research team studying movement patterns of professional dancers wants to estimate the mean number of steps taken per day. They track a sample of 50 professional dancers for a week and find an average of 19,726 steps per day. Based on previous large-scale studies of dancers, the population standard deviation is assumed to be 2,513 steps per day. Estimate, with 95% confidence, the true average daily step count for professional dancers.
Example Scenario IV:�Caffeine Content in Energy Drinks
Scenario: A health organization is concerned about caffeine intake and wants to estimate the average caffeine content (in mg) of a popular energy drink brand with marketing attractive to teens. A random sample of 36 cans is analyzed, revealing a mean caffeine content of 154 mg per can. The manufacturer has provided independent lab results confirming a population standard deviation of 12 mg, as their production process tightly controls caffeine levels. Find and interpret a 98% confidence interval for the average caffeine content in an energy drink from this brand.
Example Scenario V:�Speed of Rollercoasters
Scenario: An amusement park enthusiast group wants to estimate the average top speed (in mph) of roller coasters in the U.S. They collect data from 45 different roller coasters and find a mean top speed of 67.8 mph. Ride engineers have studied coaster speeds for years and have determined the population standard deviation to be 9.5 mph, making it reasonable to assume this value. Construct and interpret a 90% confidence interval for the average top speed of rollercoasters in the U.S.
Scenario VI:�Lifespan of Earbuds on a Single Charge
Scenario: A tech review website is testing a new model of wireless earbuds to estimate the mean battery life (in hours) on a full charge. They run a sample of 35 earbuds through battery drain tests, finding a mean battery life of 6.9 hours. The population standard deviation is assumed to be 1.2 hours, based on extensive previous testing of similar models. Find a 98% confidence interval for the average battery life of this model and provide an interpretation of that interval.
�Next Time…