MAT 1272 Statistics, CityTech, Mr. Reitz, Spr2013
Vocabulary
-data, data set -parameter vs statistic | -qualitative vs quantitative -sampling -sampling error -random sampling |
Example 1 Identify the population and the sample A. A survey of 1000 U.S. adults found that 59% think buying a home is the best investment a family can make. B. A study of 33,043 infants in Italy was conducted to find a link between a heart rhythm abnormality and sudden infant death syndrome. C. A magazine mails a questionnaires to each company in Fortune magazine’s top 100 best companies to work for and receives responses from 85 of them. |
Example 2 Qualitative or quantitative? A. number of people in a family B. length of a frog’s jump C. color of a car D. how many cars sold by a dealer in a day E. A person’s ethnic background |
Example 3 A magazine wants to know the opinions of New York City residents about the city’s response to Hurricane Sandy. Consider the following methods for choosing a sample. Are there possible sources of bias? A. Contact relief agencies and obtain lists of people affected by the hurricane. Choose randomly from this list. B. Make a list of all street intersections in the city. Choose 10 intersections at random, send reporters to interview 100 people at each intersection. C. Get a list of all phone numbers currently in use in New York City, choose randomly from this list. |
Vocabulary
-sample size n -frequency distribution -frequency f -histogram -polygon -relative frequency -percentage | -classes -lower limit and upper limit of a class -class width -midpoint -cumulative frequency |
Example 1 During a short period on a weekday afternoon, a department store recorded the age of each of its customers: 20 41 33 25 28 21 39 45 32 36 31 30 29 33 38 |
Example 2 The following sample data set lists the prices of 30 portable global positioning system (GPS) navigators. Construct a frequency distribution that has seven classes. Draw a histogram. 90 130 400 200 350 70 325 250 150 250 275 270 150 130 59 200 160 450 300 130 220 100 200 400 200 250 95 180 170 150 |
Example 3 The midpoint of a class can be found using the formula: midpoint = (upper limit - lower limit)/2 Using Example 2 above, find the midpoint of each class. Draw a polygon, using the midpoints to label the x-axis. |
Example 4 The cumulative frequency of a class is the sum (“add up”) of all frequencies of classes up to and including the class. For example, if the frequency of the first class is 3 and the second class is 5, then the cumulative frequency of the first class is 3, and the second class is 8. Using Example 2 above, find the cumulative frequency of each class. Draw a polygon, using the midpoints to label the x-axis. |
Vocabulary
-stem-and-leaf plot |
Example 1 The following are the number of text messages sent last week by the cellular phone users on one floor of a college dormitory. 155 159 144 129 105 145 126 116 130 114 122 112 112 142 126 118 118 108 122 121 109 140 126 119 113 117 118 109 109 119 139 139 122 78 133 126 123 145 121 134 124 119 132 133 124 129 112 126 148 147 Display the data in a stem-and-leaf plot. What two rows contain the bulk of the data? What percentage of cell users in the study do those two rows represent? Based on those two rows, complete the following statement: “Most cell users in the study made between ___ and ___ texts last week.” |
Vocabulary/Symbols
Measures of central tendency: - mean - median - mode - | Measures of Variation / Dispersion - range - variance - standard deviation - |
Example 1 Each semester students fill out an evaluation of their professors, scoring each professor in a number of areas on a scale from 1 (worst) to 5 (best). Suppose that a Professor has an overall average score of 4.2. Would you say that his rating is: A. Very good B. About average C. Poor |
Example 2 Scores on a recent 10-point quiz in a class of 11 students are below. Find the mean. 8 10 2 10 10 9 10 0 9 8 10 |
Example 3 Consider the following two data sets on the ages of all workers in each of two small companies. Company 1: 47 38 36 40 37 42 38 42 Company 2: 70 39 18 46 27
|
Vocabulary
- Quartiles Q1, Q2, Q3, - interquartile range (IQR) | - box-and-whisker plot - standard score (or z-score) |
Example 1 The number of nuclear power plants in the top 15 nuclear power-producing countries in the world are listed. Find the first, second, and third quartiles of the data set. Find the interquartile range (IQR). 7 18 11 6 59 17 18 54 104 20 31 8 10 15 19 |
Example 2 Draw a box-and-whisker plot for the nuclear power plant data in Example 1. |
Example 3 The graph gives the cumulative frequency distribution for SAT test scores of college-bound students in a recent year. What test score represents the 62nd percentile? How should you interpret this? (Source: The College Board) |
Example 4 SAT scores have a mean of 1000 and a standard deviation of 200. a. Suppose a student scores 1400. How many standard deviations are they from the mean? What is their z-score? b. Find the z-score of a student who scores: i. 1700 ii. 850 |
Vocabulary
- experiment - outcome - sample space - tree diagram - event - simple vs compound events | - counting principle - probability - classical vs empirical probability - law of large numbers - complement |
Example 1 Draw a tree diagram for the following experiment: A person is selected from the class, and their sex is determined (male/female). Then a second person is selected, and their sex is determined (m/f). |
Example 2 Consider the following events for the experiment in Example 1. Which outcome(s) do they include? Are they simple or compound events? a. Event A: both people selected are female. b. Event B: at least one person selected is male. |
Example 3 How many outcomes in each experiment? a. Roll a die twice b. Toss a coin three times c. A prospective car buyer can choose between a fixed and a variable interest rate and can also choose a payment period of 36 months, 48 months, or 60 months. d. At a chess tournament, each player will play 5 games, and each game can result in one of three outcomes: win, lose, or tie. e. A telephone number is created by choosing 10 digits. The first and fourth digits are not allowed to be 0. How many different possible telephone numbers are there? |
Example 4 Consider the experiment: You roll a die. Make a list of all possible outcomes, and find the probability of each event: A. You roll a 3. B. You roll an even number. C. You roll a 7. |
Example 5 An experiment is done in which women are stopped on the street and asked whether or not they have played golf at least once. Of 500 women polled, 120 have played golf at least once.
|
Vocabulary
- event - simple event, compound event - probability | - complement - conditional probability (“given”) |
Example 1 Consider the experiment: We roll a die. a. List the sample space (all the outcomes). b. Consider the following events. Which outcome(s) do they include? Are they simple or compound events? i. Event A: We roll the number 4. |
Example 2 In example 1, find the probability of each of the events A, B, and C. |
Example 3 An experiment is done in which women are stopped on the street and asked whether or not they have played golf at least once. Of 500 women polled, 120 have played golf at least once. a. Suppose we stop a woman on the street. If the event E is “she has played golf at least once”, what is P(E) (the probability of event E)? b. Consider the event “she has never played golf”. This is called the complement of event E, written E’. What is the probability P(E’)? |
Example 4 We now consider how probability can change if we are given additional information about a situation. This is idea, called “conditional probability” is absolutely essential -- and tends to give students trouble. I encourage you to study this example very carefully. The women polled in Example 3 were actually asked two questions -- whether they had played golf, and whether their annual income was more or less than $100,000. The combined information is summarized in the table:
Familiarize yourself with the table by answering these questions: 1. Consider event F: “has an income over $100,000”. How many altogether women are in event F? 2. Recall event E: “has played golf”. How many women altogether are in event E? Turn over and continue 3. Are there any women who have income over $100,000 and who also played golf? How many? 4. Are there any women who have income over $100,000 and who have never played golf? How many? 5. What is the probability P(E)? (should be the same answer calculated in Example 3). 6. The absolutely key new idea is to understand phrases like this one: “what is the probability that a woman has played golf, given that the woman has income over $100,000?” The idea is this: suppose that a woman is selected, and we are told that she is one of the ones with income over $100,000. We are no longer considering all 500 women -- we are focussing our attention only on those 90 with income over $100,000. Knowing this, what is the probability that she has played golf? ANSWER: The correct answer to #6 is 0.56, or 5/9. STOP and make sure you know where this answer came from. This kind of probability is called “conditional probability”, and it has it’s own notation -- the previous problem would be written like this: P(E | F) or P( has played golf | income over $100,000) We read “P(E | F)” like this: “the probability of E, given F”, or “the probability that she has played golf, given that she has income over $100,000”. The vertical bar represents the key word given, and the event that comes AFTER the bar is the given information. When you see notation like P(E | F), you should start with the LAST event, the event F -- restrict your attention only to that part of the data that is in event F. 7. What is the probability that a woman has played golf, given that her income is under $100,000? Or, in proper notation, P(played golf | income under $100,000)? 8. Find P(income over $100,000 | has not played golf ). What is the given information? |
Vocabulary
- dependent and independent events - multiplication rule for P(A and B) | - mutually exclusive events - addition rule for P(A or B) |
Example 1 We select a person from the streets of Manhattan. Consider the following events: A = they are female B = they have had their appendix removed C = they are wearing high heels Is there any connection between these events? First consider event A by itself: 1. What is the probability that a randomly selected person is female? Now consider what happens if we have some additional information: 2. What if we are given event C, that the person is wearing high heels – does this affect the probability of event A? 3. What if we are given event B, that the person has had their appendix removed. Does that affect the probability of event A? |
Example 2 In each case, do you think the two events are dependent or independent? a. event A: driving over 85 miles per hour, event B: getting in an accident b. The probability that a knee surgery is successful is 0.85. A surgeon performs a successful knee surgery (event A), and the next day performs another successful knee surgery (event B). c. Tossing a coin and getting Heads (event A), then rolling a die and getting an even number (event B) |
Example 3 a. In a college class, the probability that a randomly selected student is wearing blue jeans is 0.46. The probability that a student is wearing a t-shirt given that they are wearing blue jeans is 0.6. Find the probability that a randomly selected student is wearing jeans and a t-shirt. b. There are 2 blue balls and 3 red balls in a box. If two balls are drawn at random without replacement (that is, we don’t put the balls back in the box), find the probability of drawing first a red ball and then a blue ball. c. The probability that a knee surgery is successful is 0.85. If a surgeon does two knee surgeries in a row, what is the probability that both are successful? d. What is the probability that three knee surgeries in a row are successful? |
Example 4 We roll a die. a. Which two events are mutually exclusive? event A: roll an even number, event B: roll a number greater than 2, event C: roll a 3 or a 5 b. Find P(A), P(B), and P(C). c. Find P(A or C), the probability that either A or C (or both) occurs. Is it simply equal to P(A)+P(B)? d. Find P(A or B). Is it simply equal to P(A)+P(B)? |
Example 5 There is an area of free (but illegal) parking near an inner-city sports arena. The probability that a car parked in this area will be ticketed by police is .35, that the car will be vandalized is .15, and that it will be ticketed and vandalized is .10. Find the probability that a car parked in this area will be ticketed or vandalized. |
Example 6 A blood bank catalogs the types of blood, including positive or negative Rh-factor, given by donors during the last five days. The number of donors who gave each blood type is shown in the table. A donor is selected at random.
1. Find the probability that the donor has type O or type A blood. 2. Find the probability that the donor has type B blood or is Rh-negative. |
Vocabulary
- permutations - factorial - combinations | - nPr - n! - nCr |
Example 1 There are five people running a race, and they will be awarded first, second, third, fourth, and fifth place trophies. How many different ways can the trophies be awarded? |
Example 2 The objective of a 9x9 Sudoku number puzzle is to fill the grid so that each row, each column, and each 3x3 grid contain the digits 1 to 9. How many different ways can the first row of a blank 9x9 Sudoku grid be filled? |
Example 3 There are five people running a race, and the first three will be awarded first, second, and third place trophies. How many different ways can the trophies be awarded? |
Example 4 a. Forty-three race cars started the 2010 Daytona 500. How many ways can the cars finish first, second, and third? b. The board of directors of a company has 12 members. One member is the president, another is the vice president, another is the secretary, and another is the treasurer. How many ways can these positions be assigned? |
Example 5 There are five people, Alice, Bob, Charlie, Dawn and Edgar, in a jury pool. Three of them will be randomly selected to participate in a jury. Make a list of all possible selections. How many possible selections are there? |
Example 6 A state’s department of transportation plans to develop a new section of interstate highway and receives 16 bids for the project. The state plans to hire four of the bidding companies. How many different combinations of four companies can be selected from the 16 bidding companies? |
Example 7 A jury pool consists of 5 men and 7 women. If 4 people are to be randomly selected for a jury, what is the probability a. that all 4 will be women? How many ways are there of choosing 4 people (from the whole group of 12)? How many ways are there of selecting 4 women (out of 7 women)? Now divide to find the probability. b. that 3 will be women and 1 will be a man? How many ways of choosing 3 women? How many ways of choosing 1 man? Multiply these! How many ways of choosing 4 people altogether? |
Example 9 Find the probability of picking five diamonds from a standard deck of playing cards. |
Example 10 A food manufacturer is analyzing a sample of 400 corn kernels for the presence of a toxin. In this sample, three kernels have dangerously high levels of the toxin. If four kernels are randomly selected from the sample, what is the probability that exactly one kernel contains a dangerously high level of the toxin? |
Vocabulary
- random variable - discrete vs continuous variables - discrete probability distribution | - mean variance and standard deviation of a discrete random variable |
Example 1 For each random variable, is it discrete or continuous?
|
Example 2 Is it a probability distribution?
|
Example 3 Is it a probability distribution? Why or why not? A B C x P(x) x P(x) x P(x) 2 0.28 5 0.36 30 0.8 1 0.21 6 0.14 40 -0.4 0 0.43 7 0.44 50 1.2 -1 0.15 8 0.28 9 0.51 |
Example 4 An industrial psychologist administered a personality inventory test for passive-aggressive traits to 150 employees. Each individual was given a score from 1 to 5, where 1 was extremely passive and 5 extremely aggressive. A score of 3 indicated neither trait. The results are shown below. Construct a probability distribution for the random variable x (HINT: what is the probability for each value of x?). Then graph the distribution using a histogram.
|
Example 5 Thirty-five percent of all men living in the States wear a suit to work every day. Consider the experiment in which we select two U.S. men at random, and let the variable x = the number (out of 2) that wear a suit every day. Draw a tree diagram for the experiment and use it to find the probability distribution for x. |
Example 6 The probability distribution for the personality inventory test for passive- aggressive traits discussed in Example 4 is given below.
a. Find the mean score. b. Find the variance and standard deviation of x. |
Example 7 Find the mean and standard deviation of the probability distribution given in the table:
|
Example 8 The expected value of a random variable is simply the mean. It represents what you would expect to happen if the experiment were repeated many many times. At a raffle, 1500 tickets are sold at $2 each for four prizes of $500, $250, $150, and $75. You buy one ticket. If x represents your total gains/losses, find the probability distribution for x.What is the expected value (mean) of x? |
Vocabulary
- binomial experiment - trial, success, failure | - binomial formula - binomial distributions |
Example 1 Microfracture knee surgery has a 75% chance of success on patients with degenerative knees. The surgery is performed on three patients. Find the probability of the surgery being successful on exactly two patients. |
Definition A binomial experiment is an experiment satisfying the following:
Notation: n = number of trials p = probability of success in a single trial q = probability of failure in a single trial ( x = the number of successes in n trials |
Example 2 Is it a binomial experiment?
|
Example 3 A survey indicates that 41% of women in the United States consider reading their favorite leisure-time activity. You randomly select four U.S. women and ask them if reading is their favorite leisure-time activity.
|
Example 4 Sixty percent of households in the United States own a video game console. You randomly select six households and ask them if they own a video game console.
|
Vocabulary
- normal distribution - standard normal distribution | - |
Example 1 Is each random variable discrete or continuous?
For a. we might get a probability distribution like this:
For b., can we make a similar table? |
NORMAL PROBABILITY DISTRIBUTION: When plotted, gives a bell-shaped curve such that:
|
Example 2 1. Which normal curve has a greater mean? 2. Which normal curve has a greater standard deviation?
|
THE STANDARD NORMAL DISTRIBUTION is a normal distribution with a mean of 0 and a standard deviation of 1. |
Example 3 Find the area under the standard normal curve, and write your answer in probability notation:
|
Vocabulary
- z-score | - convert x to z - convert z to x |
Example 1 a. In the standard normal distribution, what is the probability that z is less than 1? b. In the normal distribution with |
Example 2 A survey indicates that people use their cellular phones an average of 1.5 years before buying a new one. The standard deviation is 0.25 year. A cellular phone user is selected at random. Find the probability that the user will use their current phone for less than 1 year before buying a new one. Assume that the variable x is normally distributed. (Adapted from Fonebak) |
Example 3 A survey indicates that for each trip to the supermarket, a shopper spends an average of 45 minutes with a standard deviation of 12 minutes in the store. The lengths of time spent in the store are normally distributed and are represented by the variable x. A shopper enters the store. (a) Find the probability that the shopper will be in the store for each interval of time listed below. (b) Interpret your answer if 200 shoppers enter the store. How many shoppers would you expect to be in the store for each interval of time listed below? a. Between 24 and 54 minutes |
Example 4 In baseball, a batting average is the number of hits divided by the number of at-bats. The batting averages of all major League Baseball players in a recent year can be approximated by a normal distribution, as shown in the diagram below. The mean of the batting averages is 0.262 and the standard deviation is 0.009. (Adapted from ESPN) a. What percent of the players have a batting average of 0.270 or greater? b. If there are 40 players on a roster, how many would you expect to have a batting average of 0.270 or greater? c. How many would you expect to have a batting average below 0.25? |
Example 5 The average midsemester score in this class as reported on the OpenLab last week was 83% with a standard deviation of 13.5%. Assuming the scores are normally distributed, how many students do we expect scored better than a C but worse than an A-? There are 35 students enrolled altogether. HINT: There is an official correspondence between scores and letter grades - where can you find it? |
Vocabulary
- reverse lookup | - convert z to x |
Example 1 a. In the standard normal distribution, find the value of z so that the area to the left of z is 0.3632. b. Find the z-score that has 10.75% of the area to the right. c. What z-score corresponds to the 99th percentile P99? |
Example 2 A veterinarian records the weights of cats treated at a clinic. The weights are normally distributed, with a mean of 9 pounds and a standard deviation of 2 pounds. a. Find the weights x corresponding to z-scores of 1.96, -0.44, and 0. b. What is the most a cat can weigh and still be in the bottom 10% of cats treated at the clinic? |
Example 3 Scores for the California Peace Officer Standards and Training test are normally distributed, with a mean of 50 and a standard deviation of 10. An agency will only hire applicants with scores in the top 10%. What is the lowest score you can earn and still be eligible to be hired by the agency? (Source: State of California) |
Example 4 In a randomly selected sample of women ages 20–34, the mean total cholesterol level is 188 milligrams per deciliter with a standard deviation of 41.3 milligrams per deciliter. Assume the total cholesterol levels are normally distributed. Find the highest total cholesterol level a woman in this 20–34 age group can have and still be in the bottom 1%. (Adapted from National Center for Health Statistics) |
Vocabulary
- continuity correction |
Example 1 50% of all people in America own a credit card.
|
Example 2 In part a of example 1, if x=the number out of 12 who own credit cards, the probability distribution of x looks like this (x is on the first line, P(x) on the second) 0 1 2 3 4 5 6 7 8 9 10 11 12 0.0002 0.0029 0.0161 0.054 0.121 0.1934 0.2256 0.1934 0.1208 0.0537 0.0161 0.0029 0.0002 |
CONTINUITY CORRECTION CHEAT SHEET If we are trying to find the probability that x is... a. between 17 and 35, we subtract .5 from 17 and add .5 to 35. We use: x is between 16.5 and 35.5 b. less than or equal to 100, we add .5 to 100. We use: x is less than 100.5 c. greater than or equal to 86, we subtract .5 from 86. We use: x is greater than 85.5 |
Example 3 Use a continuity correction to convert each of the following binomial intervals to a normal distribution interval.
|
Example 4 Sixty-two percent of adults in the United States have an HDTV in their home. You randomly select 45 adults in the United States and ask them if they have an HDTV in their home. What is the probability that fewer than 20 of them respond yes? (Source: Opinion Research Corporation) |
Example 5 Fifty-eight percent of adults say that they never wear a helmet when riding a bicycle. You randomly select 200 adults in the United States and ask them if they wear a helmet when riding a bicycle. What is the probability that at least 120 adults will say they never wear a helmet when riding a bicycle? (Source: Consumer Reports National Research Center) |
Vocabulary
- sampling distribution of - mean of the sample means |
Example 1 Scenario: How can we determine the mean income of people in the U.S.? Strategy: Take a sample of, for example, 100 people in the U.S., and find the mean of the sample Questions: a. Will the sample mean b. If two different people take two different samples of 100 people, will they get the same answer for the sample mean c. What can we do “get a better answer,” that is, to increase the chances that the sample mean is close to the actual mean? |
Example 2 There are 4 students in a class, and they each took a quiz worth 10 points. Their scores were: Alice 10 Bob 8 Charlie 6 Diane 4 Experiment: choose a person in the class, look at their score.
|
THE CENTRAL LIMIT THEOREM Suppose samples of size n are drawn from a population with mean
Furthermore, the mean
|
Example 3 Cellular phone bills for residents of a city have a mean of $63 and a standard deviation of $11, as shown in the following graph. Random samples of 100 cellular phone bills are drawn from this population and the mean of each sample is determined. a. Find the mean b. What is the probability that the sample mean |
Example 4 Suppose the training heart rates of all 20-year-old athletes are normally distributed, with a mean of 135 beats per minute and standard deviation of 18 beats per minute, as shown in the following graph. Random samples of size 4 are drawn from this population, and the mean of each sample is determined. What is the probability that the sample mean |
Vocabulary
- sample mean - sampling distribution of - mean of the sample means |
Example 1 Cellular phone bills for residents of a city are normally distributed with a mean of $63 and a standard deviation of $11.
|
Example 2 Suppose the training heart rates of 20-year-old athletes are not normally distributed, but have a mean of 135 beats per minute and standard deviation of 18 beats per minute. a. What is the probability that a single athlete’s training heart rate will be greater than 128 bpm? b. For a random sample of 35 what is the probability that the sample mean |
THE CENTRAL LIMIT THEOREM Suppose samples of size n are drawn from a population with mean
Furthermore, the mean
|
Vocabulary
- hypothesis test - claim - null and alternative hypothesis - two types of errors | - level of significance - tails of a test - test statistic - P-value |
Example 1 An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon. The Better Business Bureau wants to know if it should sue the company for false advertising. THINKING QUESTIONS How can we test the claim? If the mean turns out to be more than 50, will we care? Suppose we get a sample mean of xbar = 49.8 mpg. Do we sue? What about xbar = 30 mpg? What about xbar = 45 mpg? What conclusion can we make if xbar = 30 mpg? Is it possible that the company’s claim is still true? What conclusion can we make if xbar = 49.8 mpg? Have we proven that the company’s claim true? |
Example 2 Write the claim in English, and state H0 and Ha in mathematical notation. Which hypothesis represents the claim? 1. A car dealership announces that the mean time for an oil change is less than 15 minutes. 2. A researcher claims that the average daily time spent reading for pleasure by U.S. adults is 24 minutes. 3. An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon. |
TAILS OF A TEST The tails of the test refer to the region in the normal curve corresponding to the alternative hypothesis 1. A left-tailed test occurs when 2. A right-tailed test occurs when 3. A two-tailed test occurs when |
Example 3 A crime is committed and a man is arrested and charged with the crime. Before the trial, what is If we are the jury, what kinds of errors do we need worry about making? Are they equally bad? |
TYPES OF ERRORS A type I error occurs when we reject a true null hypothesis. The level of significance is your maximum allowable probability of making a type I error. It is denoted by A type II error occurs when we fail to reject a false null hypothesis. The probability of a type II error is denoted by , the lowercase Greek letter beta. |
P-VALUES AND THE DECISION RULE The P-value is the area in the tail(s) of the test lying beyond the test statistic. BEWARE: If the test is two-tailed, we have to double the area found in a single tail. DECISION RULE (P-VALUE): 1. If |
Vocabulary
- critical values - test statistic |
Example 1, part I An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon. The Better Business Bureau wants to know if it should sue the company for false advertising. They have asked you to design a test for the company’s claim, and they require a 5% level of significance. |
HYPOTHESIS TESTING FOR LARGE SAMPLES ( Before you take a sample:
After you take a sample:
|
Example 1, part II A sample of 50 hybrid autos are selected from the manufacturer in question and tested. The mean gas mileage for the sample is |
Example 2 The U.S. Department of Agriculture claims that the mean cost of raising a child from birth to age 2 by husband-wife families in the United States is $13,120. A random sample of 500 children (age 2) has a mean cost of $12,925 with a standard deviation of $1745. At a = 0.10, is there enough evidence to reject the claim? (Adapted from U.S. Department of Agriculture Center for Nutrition Policy and Promotion) |
Vocabulary
- t-distribution - degrees of freedom |
Example 1 A used car dealer says that the mean price of a 2008 Honda CR-V is at least $20,500. You suspect this claim is incorrect and find that a random sample of 14 similar vehicles has a mean price of $19,850 and a standard deviation of $1084. Is there enough evidence to reject the dealer’s claim at |
HYPOTHESIS TESTING FOR SMALL SAMPLES ( Before you take a sample:
After you take a sample:
|
Example 2 An industrial company claims that the mean pH level of the water in a nearby river is 6.8. You randomly select 19 water samples and measure the pH of each. The sample mean and standard deviation are 6.7 and 0.24, respectively. Is there enough evidence to reject the company’s claim at |
Example 3 A company that makes cola drinks states that the mean caffeine content per 12-ounce bottle of cola is 40 milligrams. You want to test this claim. During your tests, you find that a random sample of thirty 12-ounce bottles of cola has a mean caffeine content of 39.2 milligrams with a standard deviation of 7.5 milligrams. At |
Vocabulary
- correlation - positive and negative correlation - correlation coefficient |
Example 1 a. For used laptops sold on eBay, age of laptop and price paid b. For companies, the annual amount spent on advertising and the annual profits c. For couples, height of husband and income of wife |
Three kinds of correlation a. positive correlation: as one value increases, the other also increases b. negative correlation: as one value increases, the other decreases (and vice versa) c. no correlation: the values are not related Question: What kind of correlation do you expect for each case in Example 1? |
Correlation Coefficient The correlation coefficient, denoted
Formula: the correlation coefficient is: where |
Example 2 A group of 4 college students were asked how many hours of television they watch each week. Hours of TV: 10 5 7 2 GPA: 2.2 3.2 3.5 3.8
|
Example 3 Monthly incomes and food expenditures for seven households (in hundreds of dollars): Income (x) 35 49 21 39 15 28 25 Food expenses (y) 9 15 7 11 5 8 9 Use the facts that |
Vocabulary
- regression line - slope - y-intercept |
Example 1 A group of 4 college students were asked how many hours of television they watch each week. Hours of TV: 10 5 7 2 GPA: 2.2 3.2 3.5 3.8 Here is a graph of the data: One of your classmates tells you they watch 8 hours of TV a week, and asks you to predict their GPA. Their friend only watches 1.5 hours per week, and also wants you to predict their GPA. |
Defn: A regression line, also called a line of best fit, is the line that “best fits” the data - it is the line that makes the distance from all data points as small as possible. |
Regression Line for Example 1 |
Example 2: The equation for the regression line in Example 1 is: a. Use the equation to make a precise prediction of GPA (y) when Hours of TV is x = 8. b. Make a prediction of GPA when Hours of TV is 1.5 |
THE EQUATION OF THE REGRESSION LINE The equation of the regression line is:
RECALL: m is the slope, b is the y-intercept. |
Example 3: You are visiting Yellowstone National Park and you want to see Old Faithful, the world’s most famous geyser, erupt. Unfortunately, you arrive at the geyser only to find that an eruption has just taken place -- you will have to wait for the next one. A park ranger mentions that that there is a correlation between the duration of each eruption and the time (in minutes) until the next eruption, and provides some data on past eruptions: Duration of eruption: 1.82 4.47 3.88 1.98 2.37 Time until next eruption: 58 86 80 57 61 A tourist says that the eruption that just took place was quite long, about 5 minutes in duration. Find the equation of the regression line and use it to predict how long you will have to wait until the next eruption. |
Vocabulary
- - Goodness-of-fit test | - observed frequency - expected frequency |
Example 1 A major magazine claims that the habits of American coffee-drinkers are distributed as follows: 13% drink 1 cup a week 15% drink 2 cups a week 27% drink 1 cup a day 45% drink 2 cups a day or more To determine whether this claim is accurate, you perform a survey of 1600 randomly selected coffee drinkers. The results of the survey are: 1 cup a week 193 people 2 cups a week 206 people 1 cup a day 462 people 2+ cups a day 739 people At a significance level of 5%, is there enough evidence to reject the magazine’s claim? |
CHI-SQUARED GOODNESS-OF-FIT TEST 1. Identify the null and alternative hypotheses. The null hypothesis is always that the expected distribution (from the claim) matches the observed distribution (from the sample). 2. Find the level of significance 3. Determine the critical value using the 4. Make a table showing p, E, O and p = expected percent for each category (from the claim) E = expected frequency for each category. Formula: O = observed frequency (from the sample). To calculate 5. Verify that every expected frequency E is at least 5. If not, stop - the test fails. 6. Find the test statistic 7. Make a decision to reject, or fail to reject, the null hypothesis. If the test statistic is greater than the critical value, we reject the null hypothesis. Otherwise, we fail to reject. 8. Interpret the decision with regard to the original claim. |
Example 2 According to the M&M website, the mix of colors in their milk chocolate M&Ms is: 24% cyan blue, 20% orange, 16% green, 14% bright yellow, 13% red, 13% brown Using the sample provided and |