Published using Google Docs
Spr13-MAT1272HANDOUTS
Updated automatically every 5 minutes

MAT 1272 Statistics, CityTech, Mr. Reitz, Spr2013

Day 1

Vocabulary

-data, data set
-statistics
-population vs sample

-parameter vs statistic

-qualitative vs quantitative

-sampling

-sampling error

-random sampling

Example 1

Identify the population and the sample

A. A survey of 1000 U.S. adults found that 59% think buying a home is the best investment a family can make.

B.  A study of 33,043 infants in Italy was conducted to find a link between a heart rhythm abnormality and sudden infant death syndrome.

C. A magazine mails a questionnaires to each company in Fortune magazine’s top 100 best companies to work for and receives responses from 85 of them.

Example 2

Qualitative or quantitative?

A. number of people in a family

B. length of a frog’s jump

C. color of a car

D. how many cars sold by a dealer in a day

E. A person’s ethnic background

Example 3

A magazine wants to know the opinions of New York City residents about the city’s response to Hurricane Sandy.  Consider the following methods for choosing a sample.  Are there possible sources of bias?

A.  Contact relief agencies and obtain lists of people affected by the hurricane.  Choose randomly from this list.

B.  Make a list of all street intersections in the city.  Choose 10 intersections at random, send reporters to interview 100 people at each intersection.

C.  Get a list of all phone numbers currently in use in New York City, choose randomly from this list.


Day 2

Vocabulary

-sample size n

-frequency distribution

-frequency f

-histogram

-polygon

-relative frequency

-percentage

-classes

-lower limit and upper limit of a class

-class width

-midpoint

-cumulative frequency

Example 1

During a short period on a weekday afternoon, a department store recorded the age of each of its customers:

20  41  33  25  28  21 39  45  32  36  31  30  29  33  38

Example 2

The following sample data set lists the prices of 30 portable global positioning system (GPS) navigators.  Construct a frequency distribution that has seven classes.  Draw a histogram.

90        130        400        200        350        70        325        250        150        250

275        270        150        130        59        200        160        450        300        130

220        100        200        400        200        250        95        180        170        150

Example 3

The midpoint of a class can be found using the formula:  midpoint = (upper limit - lower limit)/2

Using Example 2 above, find the midpoint of each class.  Draw a polygon, using the midpoints to label the x-axis.

Example 4

The cumulative frequency of a class is the sum (“add up”) of all frequencies of classes up to and including the class.  For example, if the frequency of the first class is 3 and the second class is 5, then the cumulative frequency of the first class is 3, and the second class is 8.

Using Example 2 above, find the cumulative frequency of each class.  Draw a polygon, using the midpoints to label the x-axis.


Day 3

Vocabulary

-stem-and-leaf plot

Example 1

The following are the number of text messages sent last week by the cellular phone users on one floor of a college dormitory.                                

155  159  144  129  105  145  126  116  130  114  122  112  112  142

126  118  118  108  122  121  109  140  126  119  113  117  118  109

109  119  139  139  122    78  133  126  123  145  121  134  124  119  

132  133  124  129  112  126  148  147

Display the data in a stem-and-leaf plot.  What two rows contain the bulk of the data?  What percentage of cell users in the study do those two rows represent?  Based on those two rows, complete the following statement: “Most cell users in the study made between ___ and ___ texts last week.”

Days 4 & 5

Vocabulary/Symbols

Measures of central tendency:

- mean

- median

- mode

-

Measures of Variation / Dispersion

- range

- variance

- standard deviation

-

Example 1

Each semester students fill out an evaluation of their professors, scoring each professor in a number of areas on a scale from 1 (worst)  to 5 (best).  Suppose that a Professor has an overall average score of 4.2.  Would you say that his rating is:

A. Very good                           B. About average                           C. Poor

Example 2

Scores on a recent 10-point quiz in a class of 11 students are below. Find the mean.

        8        10        2        10        10        9        10        0        9        8        10

Example 3

Consider the following two data sets on the ages of all workers in each of two small companies.

Company 1:        47        38        36        40        37        42        38        42

Company 2:        70        39        18        46        27                        

  1. Find the mean, median and mode for Company 1
  2. Find the mean, median and mode for Company 2

Day 6

Vocabulary

- Quartiles Q1, Q2, Q3,

- interquartile range (IQR)

- box-and-whisker plot

- standard score (or z-score)

Example 1

The number of nuclear power plants in the top 15 nuclear power-producing countries in the world are listed. Find the first, second, and third quartiles of the data set.  Find the interquartile range (IQR).
(Source: International Atomic Energy Agency)        

7  18  11  6  59  17  18  54  104  20  31  8  10  15  19

Example 2

Draw a box-and-whisker plot for the nuclear power plant data in Example 1.

Example 3

The graph gives the cumulative frequency distribution for SAT test scores of college-bound students in a recent year.  What test score represents the 62nd percentile? How should you interpret this? (Source: The College Board)

Example 4

SAT scores have a mean of 1000 and a standard deviation of 200.  

a.  Suppose a student scores 1400.  How many standard deviations are they from the mean?  What is their z-score?

b. Find the z-score of a student who scores:           i. 1700                 ii.  850

Day 7

Vocabulary

- experiment

- outcome

- sample space

- tree diagram

- event

- simple vs compound events

- counting principle

- probability

- classical vs empirical probability

- law of large numbers

- complement

Example 1

Draw a tree diagram for the following experiment:

A person is selected from the class, and their sex is determined (male/female).  Then a second person is selected, and their sex is determined (m/f).

Example 2

Consider the following events for the experiment in Example 1.  Which outcome(s) do they include? Are they simple or compound events?

a.  Event A: both people selected are female.       b.  Event B: at least one person selected is male.

Example 3

How many outcomes in each experiment?

a. Roll a die twice

b. Toss a coin three times

c. A prospective car buyer can choose between a fixed and a variable interest rate and can also choose a payment period of 36 months, 48 months, or 60 months.

d. At a chess tournament, each player will play 5 games, and each game can result in one of three outcomes: win, lose, or tie.

e.  A telephone number is created by choosing 10 digits.  The first and fourth digits are not allowed to be 0. How many different possible telephone numbers are there?

Example 4

Consider the experiment: You roll a die.  Make a list of all possible outcomes, and find the probability of each event:

A. You roll a 3.                 B. You roll an even number.                     C. You roll a 7.

Example 5

An experiment is done in which women are stopped on the street and asked whether or not they have played golf at least once.  Of 500 women polled, 120 have played golf at least once.  

  1. Suppose we stop a woman on the street.  If the event E is “she has played golf at least once”, what is P(E) (the probability of event E)?
  2. Consider the event “she has never played golf”.  This is called the complement of event E, written E’.  What is the probability P(E’)?

Day 9

Vocabulary

- event

- simple event, compound event

- probability

- complement

- conditional probability  (“given”)

Example 1

Consider the experiment:  We roll a die.

a. List the sample space (all the outcomes).

b.  Consider the following events.  Which outcome(s) do they include? Are they simple or compound events?

i.  Event A: We roll the number 4.      
ii.  Event B: We roll an even number.
iii. Event C: We roll a number less than 3

Example 2

In example 1, find the probability of each of the events A, B, and C.

Example 3

An experiment is done in which women are stopped on the street and asked whether or not they have played golf at least once.  Of 500 women polled, 120 have played golf at least once.  

a. Suppose we stop a woman on the street.  If the event E is “she has played golf at least once”, what is P(E) (the probability of event E)?

b. Consider the event “she has never played golf”.  This is called the complement of event E, written E’.  What is the probability P(E’)?

Example 4

We now consider how probability can change if we are given additional information about a situation.  This is idea, called “conditional probability”  is absolutely essential -- and tends to give students trouble. I encourage you to study this example very carefully.

The women polled in Example 3 were actually asked two questions -- whether they had played golf, and whether their annual income was more or less than $100,000.  The combined information is summarized in the table:

Played golf

Never played golf

Total

Income over $100,000

50

40

90

Income under $100,000

70

340

410

Total

120

380

500

Familiarize yourself with the table by answering these questions:

1.  Consider event F: “has an income over $100,000”.  How many altogether women are in event F?

2.  Recall event E: “has played golf”.  How many women altogether are in event E?

Turn over and continue

3.  Are there any women who have income over $100,000 and who also played golf? How many?

4.  Are there any women who have income over $100,000 and who have never played golf? How many?

5.  What is the probability P(E)?  (should be the same answer calculated in Example 3).

6.  The absolutely key new idea is to understand phrases like this one:

      “what is the probability that a woman has played golf, given that the woman has income over $100,000?”  

The idea is this:  suppose that a woman is selected, and we are told that she is one of the ones with income over $100,000.  We are no longer considering all 500 women -- we are focussing our attention only on those 90 with income over $100,000.  Knowing this, what is the probability that she has played golf?

ANSWER: The correct answer to #6 is 0.56, or 5/9.  STOP and make sure you know where this answer came from.

This kind of probability is called “conditional probability”, and it has it’s own notation -- the previous problem would be written like this:

P(E | F)    or    P( has played golf | income over $100,000)

We read “P(E | F)” like this: “the probability of E, given F”, or “the probability that she has played golf, given that she has income over $100,000”.  The vertical bar represents the key word given, and the event that comes AFTER the bar is the given information.  When you see notation like P(E | F), you should start with the LAST event, the event F -- restrict your attention only to that part of the data that is in event F.

7. What is the probability that a woman has played golf, given that her income is under $100,000?  Or, in proper notation, P(played golf | income under $100,000)?

8.  Find P(income over $100,000  | has not played golf ).  What is the given information?


Day 10

Vocabulary

- dependent and independent events

- multiplication rule for P(A and B)

- mutually exclusive events

- addition rule for P(A or B)

Example 1

We select a person from the streets of Manhattan.  Consider the following events:

        A = they are female

        B = they have had their appendix removed

        C = they are wearing high heels

Is there any connection between these events? First consider event A by itself:

1.  What is the probability that a randomly selected person is female?  

Now consider what happens if we have some additional information:

2.  What if we are given event C, that the person is wearing high heels – does this affect the probability of event A?

3. What if we are given event B, that the person has had their appendix removed.  Does that affect the probability of event A?

Example 2

In each case, do you think the two events are dependent or independent?

a.  event A: driving over 85 miles per hour, event B: getting in an accident

b.  The probability that a knee surgery is successful is 0.85.  A surgeon performs a successful knee surgery (event A), and the next day performs another successful knee surgery (event B).

c.  Tossing a coin and getting Heads (event A), then rolling a die and getting an even number (event B)

Example 3

a.  In a college class, the probability that a randomly selected student is wearing blue jeans is 0.46.  The probability that a student is wearing a t-shirt given that they are wearing blue jeans is 0.6.  Find the probability that a randomly selected student is wearing jeans and a t-shirt.

b.  There are 2 blue balls and 3 red balls in a box.  If two balls are drawn at random without replacement (that is, we don’t put the balls back in the box), find the probability of drawing first a red ball and then a blue ball.

c.  The probability that a knee surgery is successful is 0.85.  If a surgeon does two knee surgeries in a row, what is the probability that both are successful?  

d. What is the probability that three knee surgeries in a row are successful?

Example 4

We roll a die.

a. Which two events are mutually exclusive?

        event A: roll an even number,      event B: roll a number greater than 2,     event C: roll a 3 or a 5

b. Find P(A), P(B), and P(C).  

c. Find P(A or C), the probability that either A or C (or both) occurs.  Is it simply equal to P(A)+P(B)?

d. Find P(A or B).  Is it simply equal to P(A)+P(B)?

Example 5

There is an area of free (but illegal) parking near an inner-city sports arena. The probability that a car parked in this area will be ticketed by police is .35, that the car will be vandalized is .15, and that it will be ticketed and vandalized is .10. Find the probability that a car parked in this area will be ticketed or vandalized.

Example 6

A blood bank catalogs the types of blood, including positive or negative Rh-factor, given by donors during the last five days. The number of donors who gave each blood type is shown in the table. A donor is selected at random.

O

A

B

AB

Total

Rh Positive

156

139

37

12

344

Rh Negative

28

25

8

4

65

Total

184

164

45

16

409

1. Find the probability that the donor has type O or type A blood.

2. Find the probability that the donor has type B blood or is Rh-negative.


Day 11

Vocabulary

- permutations

- factorial

- combinations

- nPr

- n!

- nCr

Example 1

There are five people running a race, and they will be awarded first, second, third, fourth, and fifth place trophies.  How many different ways can the trophies be awarded?

Example 2

The objective of a 9x9 Sudoku number puzzle is to fill the grid so that each row, each column, and each 3x3 grid contain the digits 1 to 9.  How many different ways can the first row of a blank 9x9 Sudoku grid be filled?

Example 3

There are five people running a race, and the first three will be awarded first, second, and third place trophies.  How many different ways can the trophies be awarded?

Example 4

a.  Forty-three race cars started the 2010 Daytona 500. How many ways can the cars finish first, second, and third?

b.  The board of directors of a company has 12 members. One member is the president, another is the vice president, another is the secretary, and another is the treasurer. How many ways can these positions be assigned?

Example 5

There are five people, Alice, Bob, Charlie, Dawn and Edgar, in a jury pool.  Three of them will be randomly selected to participate in a jury.  Make a list of all possible selections.  How many possible selections are there?

Example 6

A state’s department of transportation plans to develop a new section of interstate highway and receives 16 bids for the project. The state plans to hire four of the bidding companies. How many different combinations of four companies can be selected from the 16 bidding companies?

Example 7

A jury pool consists of 5 men and 7 women.  If 4 people are to be randomly selected for a jury, what is the probability

a.  that all 4 will be women?

How many ways are there of choosing 4 people (from the whole group of 12)?

How many ways are there of selecting 4 women (out of 7 women)?

Now divide to find the probability.

b. that 3 will be women and 1 will be a man?

        How many ways of choosing 3 women?   How many ways of choosing 1 man?  Multiply these!

How many ways of choosing 4 people altogether?

Example 9

Find the probability of picking five diamonds from a standard deck of playing cards.

Example 10

A food manufacturer is analyzing a sample of 400 corn kernels for the presence of a toxin. In this sample, three kernels have dangerously high levels of the toxin. If four kernels are randomly selected from the sample, what is the probability that exactly one kernel contains a dangerously high level of the toxin?


Days 12 & 13

Vocabulary

- random variable

- discrete vs continuous variables

-  discrete probability distribution

- mean variance and standard deviation of a discrete random variable

Example 1

For each random variable, is it discrete or continuous?

  1. The number of cars sold at a dealership during a given month
  2. The time taken to complete an examination
  3. The age of a person, rounded to the nearest year
  4. The number of complaints received at the office of an airline on a given day
  5. The weight of a fish
  6. The price of a house
  7. The number of heads obtained in three tosses of a coin

Example 2

Is it a probability distribution?

Days of rain, x

Probability P(x)

0

0.216

1

0.432

2

0.288

3

0.064

Example 3

Is it a probability distribution?  Why or why not?

          A                                            B                                            C

x        P(x)                        x        P(x)                        x        P(x)

2        0.28                        5        0.36                        30        0.8

1        0.21                        6        0.14                        40        -0.4

0        0.43                        7        0.44                        50        1.2

-1        0.15                        8        0.28                        

                                9        0.51

Example 4

An industrial psychologist administered a personality inventory test for passive-aggressive traits to 150 employees. Each individual was given a score from 1 to 5, where 1 was extremely passive and 5 extremely aggressive. A score of 3 indicated neither trait. The results are shown below.  Construct a probability distribution for the random variable x (HINT: what is the probability for each value of x?). Then graph the distribution using a histogram.

Score, x

Frequency, f

1

24

2

33

3

42

4

30

5

21

Example 5

Thirty-five percent of all men living in the States wear a suit to work every day.  Consider the experiment in which we select two U.S. men at random, and let the variable x = the number (out of 2) that  wear a suit every day.  Draw a tree diagram for the experiment and use it to find the probability distribution for x.

Example 6

The probability distribution for the personality inventory test for passive- aggressive traits discussed in Example 4 is given below.

x

P(x)

1

.16

2

.22

3

.28

4

.20

5

.14

a. Find the mean score.

b. Find the variance and standard deviation of x.

Example 7

Find the mean and standard deviation of the probability distribution given in the table:

x

5

10

15

20

P(x)

.08

.65

.22

.05

Example 8

The expected value of a random variable is simply the mean.  It represents what you would expect to happen if the experiment were repeated many many times.

At a raffle, 1500 tickets are sold at $2 each for four prizes of $500, $250, $150, and $75. You buy one ticket. If x represents your total gains/losses, find the probability distribution for x.What is the expected value (mean) of x?


Day 14

Vocabulary

- binomial experiment

- trial, success, failure

- binomial formula

- binomial distributions

Example 1

Microfracture knee surgery has a 75% chance of success on patients with degenerative knees. The surgery is performed on three patients. Find the probability of the surgery being successful on exactly two patients.  

Definition

A binomial experiment is an experiment satisfying the following:

  1. It consists of some number n of identical repetitions of an experiment.  Each one is called a trial.
  2. Each trial has exactly two outcomes (called success and failure).
  3. The probability of success is p, the probability of failure is q.
  4. The trials are independent.  The outcome of one does not affect another.
  5. The random variable x represents the number of successful trials.

Notation:

n = number of trials

p = probability of success in a single trial

q = probability of failure in a single trial (

x = the number of successes in n trials

Example 2

Is it a binomial experiment?

  1. We flip a coin ten times.  The variable x equals the number of heads.
  2. We select one hundred people on the street and determine if each one is a man or a woman.  The variable x equals the number of women.
  3. We select one hundred people on the street and ask each one the color of their hair.  The variable x equals the number with black hair.
  1. A jar contains 4 red and 4 blue marbles.  We draw three marbles from the jar without replacement.  The variable x equals the number red marbles.

Example 3

A survey indicates that 41% of women in the United States consider reading their favorite leisure-time activity. You randomly select four U.S. women and ask them if reading is their favorite leisure-time activity.

  1. Find the probability that:
  1. exactly two of them respond yes
  2. at least two of them respond yes
  3. at most two respond yes
  1. Create a probability distribution for x = the number that respond yes.
  2. Find the mean for x (the average number that respond yes).  Find the standard deviation for x.

Example 4

Sixty percent of households in the United States own a video game console. You randomly select six households and ask them if they own a video game console.

  1. Find the probability that the number who own a video game console is between 3 and 5..
  2. Find the mean and standard deviation.


Day 16

Vocabulary

- normal distribution

- standard normal distribution

-

Example 1

Is each random variable discrete or continuous?

  1. x = number of pencils in a student's backpack.
  2. y = height of US man.

For a. we might get a probability distribution like this:

x

0

1

2

3

4

P(x)

.20

.40

.20

.15

.05

For b., can we make a similar table?

NORMAL PROBABILITY DISTRIBUTION:  When plotted, gives a bell-shaped curve such that:

  1. The probability that x lies in any interval is always a number between zero and 1.
  2. The total area under the curve is 1.
  3. The curve is symmetric about the mean.
  4. The two tails of the curve extend indefinitely.
  5. Between and (in the center of the curve), the graph curves downward. The graph curves upward to the left of  and to the right of . The points at which the curve changes from curving upward to curving downward are called inflection points.  Finding the inflection points can help you see how large the standard deviation  is.

Example 2

1. Which normal curve has a greater mean?

2. Which normal curve has a greater standard deviation?

 

THE STANDARD NORMAL DISTRIBUTION is a normal distribution with a mean of 0 and a standard deviation of 1.

Example 3

Find the area under the standard normal curve, and write your answer in probability notation:
HINT: For each problem, draw a sketch first, marking z and shading the area you want to find.

  1. to the left of z = -0.32
  2. to the right of z = 0.07
  3. to the right of z = -1
  4. between z = -1.25 and z = 0.76
  5. between z = -2 and z = -0.65


Day 17

Vocabulary

- z-score

- convert x to z

- convert z to x

Example 1

a.  In the standard normal distribution, what is the probability that z is less than 1?

b.  In the normal distribution with  and , what is the probability that x is less than 100?

Example 2

A survey indicates that people use their cellular phones an average of 1.5 years before buying a new one. The standard deviation is 0.25 year. A cellular phone user is selected at random. Find the probability that the user will use their current phone for less than 1 year before buying a new one. Assume that the variable x is normally distributed. (Adapted from Fonebak)

Example 3

A survey indicates that for each trip to the supermarket, a shopper spends an average of 45 minutes with a standard deviation of 12 minutes in the store. The lengths of time spent in the store are normally distributed and are represented by the variable x. A shopper enters the store. (a) Find the probability that the shopper will be in the store for each interval of time listed below. (b) Interpret your answer if 200 shoppers enter the store. How many shoppers would you expect to be in the store for each interval of time listed below?

a. Between 24 and 54 minutes
b. More than 39 minutes

Example 4

In baseball, a batting average is the number of hits divided by the number of at-bats. The batting averages of all major League Baseball players in a recent year can be approximated by a normal distribution, as shown in the diagram below. The mean of the batting averages is 0.262 and the standard deviation is 0.009. (Adapted from ESPN)

a.  What percent of the players have a batting average of 0.270 or greater?

b.  If there are 40 players on a roster,  how many would you expect to have a batting average of 0.270 or greater?

c.  How many would you expect to have a batting average below 0.25?

Example 5

The average midsemester score in this class as reported on the OpenLab last week was 83% with a standard deviation of 13.5%.  Assuming the scores are normally distributed, how many students do we expect scored better than a C but worse than an A-?  There are 35 students enrolled altogether.

HINT: There is an official correspondence between scores and letter grades - where can you find it?


Day 18

Vocabulary

- reverse lookup

- convert z to x

Example 1

a.  In the standard normal distribution, find the value of z so that the area to the left of z is 0.3632.

b.  Find the z-score that has 10.75% of the area to the right.

c.  What z-score corresponds to the 99th percentile P99?
Recall: the 99th percentile means 99% of the total lies below this score.

Example 2

A veterinarian records the weights of cats treated at a clinic. The weights are normally distributed, with a mean of 9 pounds and a standard deviation of 2 pounds.

a.  Find the weights x corresponding to z-scores of 1.96, -0.44, and 0.

b.  What is the most a cat can weigh and still be in the bottom 10% of cats treated at the clinic?

Example 3

Scores for the California Peace Officer Standards and Training test are normally distributed, with a mean of 50 and a standard deviation of 10. An agency will only hire applicants with scores in the top 10%. What is the lowest score you can earn and still be eligible to be hired by the agency? (Source: State of California)

Example 4

In a randomly selected sample of women ages 20–34, the mean total cholesterol level is 188 milligrams per deciliter with a standard deviation of 41.3 milligrams per deciliter. Assume the total cholesterol levels are normally distributed. Find the highest total cholesterol level a woman in this 20–34 age group can have and still be in the bottom 1%. (Adapted from National Center for Health Statistics)


Day 19

Vocabulary

- continuity correction

Example 1

50% of all people in America own a credit card.  

  1. In a random sample of 12 people, what is the probability that 4 to 6 of them own a credit card? (4 decimal places, please)
  2. In a random sample of 150 people, what is the probability that 50 to 80 of them own a credit card?

Example 2

In part a of example 1, if x=the number out of 12 who own credit cards, the probability distribution of x looks like this (x is on the first line, P(x) on the second)

0         1         2         3         4         5         6         7         8         9        10        11        12

0.0002        0.0029        0.0161        0.054        0.121        0.1934        0.2256        0.1934        0.1208        0.0537        0.0161        0.0029        0.0002

CONTINUITY CORRECTION CHEAT SHEET

If we are trying to find the probability that x is...

a. between 17 and 35, we subtract .5 from 17 and add .5 to 35.  We use:  x is between 16.5 and 35.5

b. less than or equal to 100, we add .5 to 100.  We use: x is less than 100.5

c.  greater than or equal to 86, we subtract .5 from 86.  We use: x is greater than 85.5

Example 3

Use a continuity correction to convert each of the following binomial intervals to a normal distribution interval.

  1. The probability of getting between 270 and 310 successes
  2. The probability of getting at least 158 successes
  3. The probability of getting fewer than 63 successes

Example 4

Sixty-two percent of adults in the United States have an HDTV in their home. You randomly select 45 adults in the United States and ask them if they have an HDTV in their home. What is the probability that fewer than 20 of them respond yes? (Source: Opinion Research Corporation)

Example 5

Fifty-eight percent of adults say that they never wear a helmet when riding a bicycle. You randomly select 200 adults in the United States and ask them if they wear a helmet when riding a bicycle. What is the probability that at least 120 adults will say they never wear a helmet when riding a bicycle? (Source: Consumer Reports National Research Center)


Day 20

Vocabulary

- sampling distribution of

- mean of the sample means
- standard deviation of the sample means
 

Example 1

Scenario: How can we determine  the mean  income  of people in the U.S.?  

Strategy: Take a sample of, for example, 100 people in the U.S., and find the mean of the sample .

Questions: a.  Will the sample mean  be equal to the actual mean ?  Explain.

b.  If two different people take two different samples of 100 people, will they get the same answer for the sample mean ? Explain.

c.  What can we do “get a better answer,” that is, to increase the chances that the sample mean is close to the actual mean?

Example 2

There are 4 students in a class, and they each took a quiz worth 10 points.  Their scores were:

Alice        10

Bob        8

Charlie        6

Diane        4

Experiment: choose a person in the class, look at their score.

  1.  What is the mean (average score) ?  What is the standard deviation ?  
    Hint: We learned standard deviation in the second week of class -- look it up!  But for now I’ll just tell you,
  2.  Suppose we select a two-person sample, with replacement.  What is the sample mean ?
  3. Everyone choose their own two-person sample, with replacement, and calculate the sample mean.
  4. What is the average of the ?  This is called the mean of the sample means,  “mu sub x-bar”
    What is the standard deviation of the
    ?  The standard deviation of the sample means, 
  5. Make a probability distribution for , the sample mean (what are the different values of ?  make a frequency count).  This is called the sampling distribution of sample means, or sampling distribution of .


THE CENTRAL LIMIT THEOREM

Suppose samples of size n are drawn from a population with mean  and standard deviation .  Then:

  1.  If the population is normally distributed, then the sampling distribution of  will be normally distributed, for any sample size n.
  2.  If the population is NOT normally distributed, the the sampling distribution of  will still be (approximately) normally distributed as long as the sample size is at least 30  (.

Furthermore, the mean  and standard deviation  of the sample means  are given by:

 and

Example 3

Cellular phone bills for residents of a city have a mean of $63 and a standard deviation of $11, as shown in the following graph. Random samples of 100 cellular phone bills are drawn from this population and the mean of each sample is determined.

a.  Find the mean  and standard deviation  of the sampling distribution of .

b.  What is the probability that the sample mean  will fall between $62 and $65?

Example 4

Suppose the training heart rates of all 20-year-old athletes are normally distributed, with a mean of 135 beats per minute and standard deviation of 18 beats per minute, as shown in the following graph. Random samples of size 4 are drawn from this population, and the mean of each sample is determined.  What is the probability that the sample mean    will be greater than 128 beats per minute?


Day 20

Vocabulary

- sample mean

- sampling distribution of

- mean of the sample means
- standard deviation of the sample means
 

Example 1

Cellular phone bills for residents of a city are normally distributed with a mean of $63 and a standard deviation of $11.

  1. If we choose a cellular customer at random, what is the probability that his bill will be within $5 of the mean (that is, between $61 and $65)?
  2. If we choose 10 customers at random and find the average of their bills   (the mean of the sample, or sample mean), what is the probability that it will be within $5 of the mean?

Example 2

Suppose the training heart rates of 20-year-old athletes are not normally distributed, but have a mean of 135 beats per minute and standard deviation of 18 beats per minute.  

a.  What is the probability that a single athlete’s training heart rate will be greater than 128 bpm?

b.  For a random sample of 35 what is the probability that the sample mean    will be greater than 128 bpm?

THE CENTRAL LIMIT THEOREM

Suppose samples of size n are drawn from a population with mean  and standard deviation .  Then:

  1.  If the population is normally distributed, then the sampling distribution of  will be normally distributed, for any sample size n.
  2.  If the population is NOT normally distributed, the the sampling distribution of  will still be (approximately) normally distributed as long as the sample size is at least 30  (.

Furthermore, the mean  and standard deviation  of the sample means  are given by:

 and

Day 21

Vocabulary

- hypothesis test

- claim

- null and alternative hypothesis

- two types of errors

- level of significance

- tails of a test

- test statistic

- P-value

Example 1

An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon.  The Better Business Bureau wants to know if it should sue the company for false advertising.

THINKING QUESTIONS

How can we test the claim?

If the mean turns out to be more than 50, will we care?

Suppose we get a sample mean of xbar = 49.8 mpg. Do we sue?

        What about xbar = 30 mpg?

        What about xbar = 45 mpg?

What conclusion can we make if xbar = 30 mpg? Is it possible that the company’s claim is still  true?

What conclusion can we make if xbar = 49.8 mpg? Have we proven that the company’s claim true?

Example 2

Write the claim in English, and state H0 and Ha in mathematical notation.  Which hypothesis represents the claim?

1.  A car dealership announces that the mean time for an oil change is less than 15 minutes.

2.  A researcher claims that the average daily time spent reading for pleasure by U.S. adults is 24 minutes.  

3.  An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon.  

TAILS OF A TEST

The tails of the test refer to the region in the normal curve corresponding to the alternative hypothesis  (the region in which we reject ).  There are three types:

1.  A left-tailed test occurs when  has the form  for some number k.

2.  A right-tailed test occurs when  has the form  for some number k.

3.  A two-tailed test occurs when  has the form  for some number k.

Example 3

A crime is committed and a man is arrested and charged with the crime.  

Before the trial, what is ? That is,  do we a) assume the man is innocent, or b) assume the man is guilty?

If we are the jury, what kinds of errors do we need worry about making? Are they equally bad?

TYPES OF ERRORS

A type I error occurs when we reject a true null hypothesis.  The level of significance is your maximum allowable probability of making a type I error. It is denoted by , the lowercase Greek letter alpha.

A type II error occurs when we fail to reject a false null hypothesis.  The probability of a type II error is denoted by , the lowercase Greek letter beta.

P-VALUES AND THE DECISION RULE

The P-value is the area in the tail(s) of the test lying beyond the test statistic.

BEWARE: If the test is two-tailed, we have to double the area found in a single tail.

DECISION RULE (P-VALUE):  1. If , we reject .        2. If , we fail to reject .

Day 22

Vocabulary

- critical values

- test statistic

Example 1, part I

An automobile manufacturer advertises that its new hybrid car has a mean gas mileage of 50 miles per gallon.  The Better Business Bureau wants to know if it should sue the company for false advertising.  They have asked you to design a test for the company’s claim, and they require a 5% level of significance.

HYPOTHESIS TESTING FOR LARGE SAMPLES ()

Before you take a sample:

  1. State the claim. Identify the null hypothesis  and alternative hypothesis .
  2. Specify the level of significance . The level of significance gives the area of the rejection region(s).
  3. Describe the tails of the test.  Sketch the rejection region(s).
  4. Determine the critical value(s).  If the data collected from the sample (the test statistic)  lies outside these values, we will reject the null hypothesis.

After you take a sample:

  1. Find the z-value of the test statistic (from your sample).  , where ,  
  2. Make a decision to reject or fail to reject the null hypothesis.
  3. Interpret the decision in the context of the original claim.

Example 1, part II

A sample of 50 hybrid autos are selected from the manufacturer in question and tested.  The mean gas mileage for the sample is  mpg, and the standard deviation is mpg.

Example 2

The U.S. Department of Agriculture claims that the mean cost of raising a child from birth to age 2 by husband-wife families in the United States is $13,120. A random sample of 500 children (age 2) has a mean cost of $12,925 with a standard deviation of $1745. At a = 0.10, is there enough evidence to reject the claim? (Adapted from U.S. Department of Agriculture Center for Nutrition Policy and Promotion)


Day 23

Vocabulary

- t-distribution

- degrees of freedom

Example 1

A used car dealer says that the mean price of a 2008 Honda CR-V is at least $20,500. You suspect this claim is incorrect and find that a random sample of 14 similar vehicles has a mean price of $19,850 and a standard deviation of $1084. Is there enough evidence to reject the dealer’s claim at ? Assume the population is normally distributed. (Adapted from Kelley Blue Book)

HYPOTHESIS TESTING FOR SMALL SAMPLES ()

Before you take a sample:

  1. State the claim. Identify the null hypothesis  and alternative hypothesis .
  2. Specify the level of significance . The level of significance gives the area of the rejection region(s).
  3. Describe the tails of the test.  Sketch the rejection region(s).
  4. Determine the degrees of freedom .  
  5. Determine the critical value(s) using the t-distribution table.  If the data collected from the sample (the test statistic)  lies outside these values, we will reject the null hypothesis.

After you take a sample:

  1. Find the t-value of the test statistic (from your sample).  , where ,  
  2. Make a decision to reject or fail to reject the null hypothesis.
  3. Interpret the decision in the context of the original claim.

Example 2

An industrial company claims that the mean pH level of the water in a nearby river is 6.8. You randomly select 19 water samples and measure the pH of each. The sample mean and standard deviation are 6.7 and 0.24, respectively. Is there enough evidence to reject the company’s claim at ? Assume the population is normally distributed.

Example 3

A company that makes cola drinks states that the mean caffeine content per 12-ounce bottle of cola is 40 milligrams. You want to test this claim. During your tests, you find that a random sample of thirty 12-ounce bottles of cola has a mean caffeine content of 39.2 milligrams with a standard deviation of 7.5 milligrams. At , can you reject the company’s claim? (Adapted from American Beverage Association)


Day 24

Vocabulary

- correlation

- positive and negative correlation

- correlation coefficient

Example 1

a. For used laptops sold on eBay, age of laptop and price paid

b. For companies, the annual amount spent on advertising and the annual profits

c. For couples, height of husband and income of wife

Three kinds of correlation

a. positive correlation: as one value increases, the other also increases

b. negative correlation: as one value increases, the other decreases (and vice versa)

c. no correlation: the values are not related

Question: What kind of correlation do you expect for each case in Example 1?

Correlation Coefficient

The correlation coefficient, denoted  'rho' for populations and r for samples, is always a number
between -1 and 1.

  • if  is close to 1, there is a positive correlation
  • if  is close to -1, there is a negative correlation
  • if  is close to 0, there is no correlation

Formula: the correlation coefficient is:  

where , ,

Example 2

A group of 4 college students were asked how many hours of television they watch each week.  
This was compared to their GPA.

Hours of TV:        10        5        7        2

GPA:                 2.2        3.2        3.5        3.8

  1. Do you expect hours of TV and GPA to be positively or negatively correlated?
  2. Find the correlation coefficient.
  3. Based on your result, what advice would you give to a student who wants to know if watching TV is related to success in college?

Example 3

Monthly incomes and food expenditures for seven households (in hundreds of dollars):

Income (x)                35        49        21        39        15        28        25

Food expenses (y)        9        15        7        11        5        8        9

Use the facts that  to calculate the correlation coefficient.  Is income positively or negatively correlated to food expenses?


Day 25

Vocabulary

- regression line

- slope

- y-intercept

Example 1

A group of 4 college students were asked how many hours of television they watch each week.  
This was compared to their GPA.

Hours of TV:        10        5        7        2

GPA:                2.2        3.2        3.5        3.8

Here is a graph of the data:

One of your classmates tells you they watch 8 hours of TV a week, and asks you to predict their GPA. Their friend only watches 1.5 hours per week, and also wants you to predict their GPA.

Defn:  A regression line, also called a line of best fit, is the line that “best fits” the data - it is the line that makes the distance from all data points as small as possible.

Regression Line for Example 1

Example 2:

The equation for the regression line in Example 1 is:  

a.  Use the equation to make a precise prediction of GPA (y) when Hours of TV is x = 8.

b.  Make a prediction of GPA when Hours of TV is 1.5

THE EQUATION OF THE REGRESSION LINE

The equation of the regression line is:

 , with  , and

RECALL: m is the slope, b is the y-intercept.  is the average of y-values, and is the average of x-values.  

Example 3:

You are visiting Yellowstone National Park and you want to see Old Faithful, the world’s most famous geyser, erupt.  Unfortunately, you arrive at the geyser only to find that an eruption has just taken place -- you will have to wait for the next one.  A park ranger mentions that that there is a correlation between the duration of each eruption and the time (in minutes) until the next eruption, and provides some data on past eruptions:

Duration of eruption:        1.82        4.47        3.88        1.98        2.37

Time until next eruption:        58        86        80        57        61

A tourist says that the eruption that just took place was quite long, about 5 minutes in duration.  Find the equation of the regression line and use it to predict how long you will have to wait until the next eruption.


Day 25

Vocabulary

-  chi-squared

- Goodness-of-fit test

- observed frequency

- expected frequency

Example 1

A major magazine claims that the habits of American coffee-drinkers are distributed as follows:

13% drink 1 cup a week

15% drink 2 cups a week

27% drink 1 cup a day

45% drink 2 cups a day or more

To determine whether this claim is accurate, you perform a survey of 1600 randomly selected coffee drinkers.  The results of the survey are:

1 cup a week        193 people

2 cups a week        206 people

1 cup a day        462 people

2+ cups a day        739 people

At a significance level of 5%, is there enough evidence to reject the magazine’s claim?

CHI-SQUARED GOODNESS-OF-FIT TEST

1.  Identify the null and alternative hypotheses.  The null hypothesis is always that the expected distribution (from the claim) matches the observed distribution (from the sample).

2.  Find the level of significance , and the degrees of freedom .  WARNING: k is the number of CATEGORIES, not the sample size.

3.  Determine the critical value using the  (chi-squared) table, and sketch the rejection region.  The chi-squared test is ALWAYS right-tailed.

4.  Make a table showing p, E, O and .

        p = expected percent for each category (from the claim)

        E = expected frequency for each category.  Formula:  

        O = observed frequency (from the sample).

        To calculate , first calculate O-E, then square it, then divide by E.

5.  Verify that every expected frequency E is at least 5.  If not, stop - the test fails.

6.  Find the test statistic .

7.  Make a decision to reject, or fail to reject, the null hypothesis.  If the test statistic is greater than the critical value, we reject the null hypothesis.  Otherwise, we fail to reject.

8.  Interpret the decision with regard to the original claim.

Example 2

According to the M&M website, the mix of colors in their milk chocolate M&Ms is:

24% cyan blue, 20% orange, 16% green, 14% bright yellow, 13% red, 13% brown

Using the sample provided and , perform a chi-squared Goodness-of-fit test to test whether the claimed distribution of colors is correct.