1 of 55

Copyright © Cengage Learning. All rights reserved.

4

Statistics

1

2 of 55

Copyright © Cengage Learning. All rights reserved.

4.1

Population, Sample, and Data

2

3 of 55

Objectives

  • Construct a frequency distribution

  • Construct a histogram

  • Construct a pie chart

3

4 of 55

Population, Sample, and Data

The field of statistics can be defined as the science of collecting, organizing, and summarizing data in such a way that valid conclusions and meaningful predictions can be drawn from them.�

The first part of this definition, “collecting, organizing, and summarizing data,” applies to descriptive statistics.��The second part, “drawing valid conclusions and making meaningful predictions,” describes inferential statistics.

4

5 of 55

Population versus Sample

5

6 of 55

Population versus Sample

Who will become the next president of the United States?��During election years, political analysts spend a lot of time and money trying to determine what percent of the vote each candidate will receive. ��However, because there are over 175 million registered voters in the United States, it would be virtually impossible to contact each and every one of them and ask,�“Whom do you plan on voting for?”

6

7 of 55

Population versus Sample

Consequently, analysts select a smaller group of people, determine their intended voting patterns, and project their results onto the entire body of all voters.

Because of time and money constraints, it is very common for researchers to study the characteristics of a small group in order to estimate the characteristics of a larger group.

In this context, the set of all �objects under study is called�the population, and any �subset of the population is�called a sample (see Figure 4.1).

Figure 4.1

Population versus sample.

7

8 of 55

Population versus Sample

When we are studying a large population, we might not be able to collect data from every member of the population, so we collect data from a smaller, more manageable sample.��Once we have collected these data, we can summarize by calculating various descriptive statistics, such as the average value.��Inferential statistics, then, deals with drawing conclusions (hopefully, valid ones!) about the population, based on the descriptive statistics of the sample data.

8

9 of 55

Population versus Sample

Sample data are collected and summarized to help us draw conclusions about the population. A good sample is representative of the population from which it was taken. ��Obviously, if the sample is not representative, the conclusions concerning the population might not be valid. The most difficult aspect of inferential statistics is obtaining a representative sample.��Conclusions are only as reliable as the sampling process and information will usually change from sample to sample.

9

10 of 55

Frequency Distributions

10

11 of 55

Frequency Distributions

The first phase of any statistical study is the collection of data. Each element in a set of data is referred to as a�data point. When data are first collected, the data points might show no apparent patterns or trends.��To summarize the data and detect any trends, we must organize the data. This is the second phase of descriptive statistics.��The most common way to organize raw data is to create a frequency distribution, a table that lists each data point along with the number of times it occurs (its frequency).

11

12 of 55

Frequency Distributions

The composition of a frequency distribution is often easier to see if the frequencies are converted to percents, especially if large amounts of data are being summarized. ��The relative frequency of a data point is the frequency of the data point expressed as a percent of the total number of data points (that is, made relative to the total). The relative frequency of a data point is found by dividing its frequency by the total number of data points in the data set. ��Besides listing the frequency of each data point, a frequency distribution should also contain a column that gives the relative frequencies.

12

13 of 55

Example 1 – Creating a Frequency Distribution: Single Values

While bargaining for their new contract, the employees of 2 Dye 4 Clothing asked their employers to provide daycare service as an employee benefit.��Examining the personnel files of the company’s fifty employees, the management recorded the number of children under six years of age that each employee was caring for.

13

14 of 55

Example 1 – Creating a Frequency Distribution: Single Values

The following results were obtained:

Organize the data by creating a frequency distribution.

Solution:

First, we list each different number in a column, putting them in order from smallest to largest (or vice versa). Then we use tally marks to count the number of times each data point occurs.

cont’d

14

15 of 55

Example 1 – Solution

The frequency of each data point is shown in the third column of Figure 4.2.

cont’d

Figure 4.2

Frequency distribution of data.

15

16 of 55

Example 1 – Solution

To get the relative frequencies, we divide each frequency by 50 (the total number of data points) and change the resulting decimal to a percent, as shown in the fourth column of Figure 4.2.

cont’d

16

17 of 55

Example 1 – Solution

The raw data have now been organized and summarized. �

At this point, we can see that about one-third of the employees have no need for child care (32%), while the remaining two-thirds (68%) have at least one child under 6 years of age who would benefit from company-sponsored daycare.

The most common trend (that is, the data point with the highest relative frequency for the fifty employees) is having one child (36%).

cont’d

17

18 of 55

Grouped Data

18

19 of 55

Grouped Data

When raw data consist of only a few distinct values (for instance, the data in Example 1, which consisted of only the numbers 0, 1, 2, 3, 4, and 5), we can easily organize the data and determine any trends by listing each data point along with its frequency and relative frequency. �

However, when the raw data consist of many nonrepeated data points, listing each one separately does not help us to see any trends the data set might contain.�

In such cases, it is useful to group the data into intervals�or classes and then determine the frequency and relative frequency of each group rather than of each data point.

19

20 of 55

Example 2 – Creating a Frequency Distribution: Grouped Data

Keith Reed is an instructor for an acting class offered through a local arts academy. The class is open to anyone who is at least 16 years old. Forty-two people are enrolled; their ages are as follows:

Organize the data by creating a frequency distribution.

20

21 of 55

Example 2 – Solution

Listing each distinct data point and its frequency might not summarize the data well enough for us to draw conclusions. Instead, we will work with grouped data.�

First, we find the largest and smallest values (62 and 16). Subtracting, we find the range of ages to be�62 16 = 46 years.�

In working with grouped data, it is customary to create between four and eight groups of data points. We arbitrarily choose six groups, the first group beginning at the smallest data point, 16.

21

22 of 55

Example 2 – Solution

To find the beginning of the second group (and hence the end of the first group), divide the range by the number of groups, round off this answer to be consistent with the data, and then add the result to the smallest data point:�

46 ÷ 6 = 7.6666666 … ≈ 8

The beginning of the second group is 16 + 8 = 24, so the first group consists of people from 16 up to (but not including) 24 years of age.

In a similar manner, the second group consists of people from 24 up to (but not including) 32 (24 + 8 = 32) years of age.

cont’d

This is the width of each group.

22

23 of 55

Example 2 – Solution

The remaining groups are formed and the ages tallied in�the same way.

The frequency distribution is shown in Figure 4.3.

cont’d

Figure 4.3

Frequency distribution of grouped data.

23

24 of 55

Example 2 – Solution

Now that the data have been organized, we can observe various trends:�

Ages from 24 to 32 are most common (31% is the highest relative frequency), and ages from 48 to 56 are least common (5% is the lowest). Also, over half the people enrolled (57%) are from 16 to 32 years old.

cont’d

24

25 of 55

Grouped Data

25

26 of 55

Histograms

26

27 of 55

Histograms

When data are grouped in intervals, they can be depicted by a histogram, a bar chart that shows how the data are distributed in each interval.��To construct a histogram, mark off the class limits on a horizontal axis. If each interval has equal width, we draw two vertical axes; the axis on the left exhibits the frequency of an interval, and the axis on the right gives the corresponding relative frequency.��We then draw a rectangle above each interval; the height of the rectangle corresponds to the number of data points�contained in the interval.

27

28 of 55

Histograms

The vertical scale on the right gives the percentage of data contained in each interval. The histogram depicting the distribution of the ages of the people in Keith Reed’s acting class is shown in Figure 4.4.

Figure 4.4

Ages of the people in Keith Reed’s acting class.

28

29 of 55

Histograms

What happens if the intervals do not have equal width? For instance, suppose the ages of the people in Keith Reed’s acting class are those given in the frequency distribution shown in Figure 4.5.

Figure 4.5

Frequency distribution of age.

29

30 of 55

Histograms

With frequency and relative frequency as the vertical scales, the histogram depicting this new distribution is given in Figure 4.6.

Figure 4.6

Why is this histogram misleading?

30

31 of 55

Histograms

Does the histogram in Figure 4.6 give a truthful representation of the distribution? No; the rectangle over the interval from 45 to 65 appears to be larger than the rectangle over the interval from 30 to 45, yet the interval from 45 to 65 contains fewer data than the interval from 30 to 45.��This is misleading; rather than comparing the heights of the rectangles, our eyes naturally compare the areas of the rectangles. Therefore, to make an accurate comparison, the areas of the rectangles must correspond to the relative frequencies of the intervals. This is accomplished by utilizing the density of each interval.

31

32 of 55

Histograms and Relative Frequency Density

32

33 of 55

Histograms and Relative Frequency Density

Density is a ratio. In science, density is used to determine the concentration of weight in a given volume:�Density = weight/volume.��For example, the density of water is 62.4 pounds per cubic foot. In statistics, density is used to determine the concentration of data in a given interval:�Density = (percent of total data)/(size of an interval).��Because relative frequency is a measure of the percentage of data within an interval, we shall calculate the relative frequency density of an interval to determine the concentration of data within the interval.

33

34 of 55

Histograms and Relative Frequency Density

For example, if the interval 20 x < 25 contains eleven out of forty-two data points, then the relative frequency density of the interval is

34

35 of 55

Histograms and Relative Frequency Density

If a histogram is constructed using relative frequency density as the vertical scale, the area of a rectangle will correspond to the relative frequency of the interval, as shown in Figure 4.7.

Figure 4.7

Area equals relative frequency.

35

36 of 55

Histograms and Relative Frequency Density

area = base height

= Δx rfd

=

= f / n

= relative frequency

36

37 of 55

Histograms and Relative Frequency Density

Adding a new column to the frequency distribution given in Figure 4.5, we obtain the relative frequency densities shown in Figure 4.8.

Figure 4.8

Calculating relative frequency density.

37

38 of 55

Histograms and Relative Frequency Density

We now construct a histogram using relative frequency density as the vertical scale. The histogram depicting the distribution of the ages of the people in Keith Reed’s acting�class (using the frequency distribution in Figure 4.8) is shown in Figure 4.9.

Figure 4.9

Ages of the people in Keith Reed’s acting class.

38

39 of 55

Histograms and Relative Frequency Density

Comparing the histograms in Figures 4.6 and 4.9, we see that using relative frequency density as the vertical scale (rather than frequency) gives a more truthful representation of a distribution when the interval widths are unequal.

Figure 4.6

Why is this histogram misleading?

39

40 of 55

Example 3 – Constructing a Histogram: Grouped Data

To study the output of a machine that fills bags with corn chips, a quality control engineer randomly selected and weighed a sample of 200 bags of chips. The frequency distribution in Figure 4.10 summarizes the data. Construct a histogram for the weights of the bags of corn chips.

Figure 4.10

Weights of bags of corn chips.

40

41 of 55

Example 3 – Solution

Because each interval has the same width (Δx = 0.2), we construct a combined frequency and relative frequency histogram. The relative frequencies are given in�Figure 4.11.

Figure 4.11

Relative frequencies.

41

42 of 55

Example 3 – Solution

We now draw coordinate axes with appropriate scales and rectangles (Figure 4.12). Notice the (near) symmetry of the histogram.

cont’d

Figure 4.12

Weights of bags of corn chips.

42

43 of 55

Histograms and Single-Valued Classes

43

44 of 55

Histograms and Single-Valued Classes

The histograms that we have constructed so far have all utilized intervals of grouped data. For instance, the first group of ages in Keith Reed’s acting class was from 16 to 24 years old (16 x < 24).��However, if a set of data consists of only a few distinct values, it may be advantageous to consider each distinct value to be a “class” of data; that is, we utilize single-valued classes of data.

44

45 of 55

Example 4 – Constructing a Histogram: Single Values

A sample of high school seniors was asked, “How many television sets are in your house?” The frequency distribution in Figure 4.13 summarizes the data. Construct a histogram using single-valued classes of data.

Figure 4.13

Frequency distribution.

45

46 of 55

Example 4 – Solution

Rather than using intervals of grouped data, we use a single value to represent each class. Because each class has the same width (Δx = 1), we construct a combined frequency and relative frequency histogram.�

The relative frequencies�are given in Figure 4.14.

Figure 4.14

Calculating relative frequency.

46

47 of 55

Example 4 – Solution

We now draw coordinate axes with appropriate scales and rectangles. In working with single valued classes of data, it is common to write the single value at the midpoint of the base of each rectangle as shown in Figure 4.15.

cont’d

Figure 4.15

Number of television sets.

47

48 of 55

Pie Charts

48

49 of 55

Pie Charts

Many statistical studies involve categorical data—those which is grouped according to some common feature or quality.��One of the easiest ways to summarize categorical data is through the use of a pie chart.��A pie chart shows how various categories of a set of data account for certain proportions of the whole.

49

50 of 55

Pie Charts

Financial incomes and expenditures are invariably shown as pie charts, as in Figure 4.16.

Figure 4.16

How a typical medical dollar is spent.

50

51 of 55

Pie Charts

To draw the “slice” of the pie representing the relative frequency (percentage) of the category, the appropriate central angle must be calculated.��Since a complete circle comprises 360 degrees, we obtain the required angle by multiplying 360° times the relative frequency of the category.

51

52 of 55

Example 5 – Constructing a Pie Chart

What type of academic degree do you hope to earn? The different types of degrees and the number of each type of degree conferred in the United States during the 2010–11 academic year is given in Figure 4.17. Construct a pie chart to summarize the data.

Figure 4.17

Academic degrees conferred, 2010–11.�Source: National Center for Education Statistics.

52

53 of 55

Example 5 – Solution

Find the relative frequency of each category and multiply it by 360° to determine the appropriate central angle.

The necessary calculations are shown in Figure 4.18.

Figure 4.18

Calculating relative frequency and central angles.

cont’d

53

54 of 55

Example 5 – Solution

Now use a protractor to lay out the angles and draw the “slices.” The name of each category can be written directly on the slice, or, if the names are too long, a legend consisting of various shadings may be used. �

Each slice of the pie should contain its relative frequency, expressed as a percent.

cont’d

54

55 of 55

Example 5 – Solution

The whole reason for constructing a pie chart is to convey information visually; pie charts should enable the reader to instantly compare the relative proportions of categorical data. See Figure 4.19.

cont’d

Figure 4.19

Academic degrees conferred, 2010–11.

55