1 of 87

DESCRIPTIVE STATISTICS: A BRIEF INTRODUCTION

BASIC LABORATORY METHODS IN A REGULATED ENVIRONMENT

2 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

3 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

4 of 87

WHY LEARN ABOUT STATISTICS?

  • Statistics provides tools that are used in
    • Quality control
    • Research
    • Measurements
    • Sports

5 of 87

LATER IN THIS COURSE

  • We will use some of these tools to display, interpret, and evaluate our results
    • Ideas
    • Vocabulary
    • A few calculations

6 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

7 of 87

VARIATION

  • There is variation in the natural world
    • People vary
    • Measurements vary
    • Plants vary
    • Weather varies

8 of 87

VARIATION

  • Variation among organisms is the basis of natural selection and evolution

9 of 87

EXAMPLE

  • 100 people take a drug and 75 of them get better
  • 100 people don’t take the drug but 68 get better without it
  • Did the drug help?

10 of 87

VARIABILITY IS THE ISSUE

  • There is variation in response to the illness
  • There is variation in response to the drug
  • So, it’s difficult to figure out if the drug helped

11 of 87

STATISTICS

  • Provides mathematical tools to help arrive at meaningful conclusions in the presence of variability

12 of 87

STATISTICS

  • Might help researchers decide if a drug is helpful or not
  • This is a more advanced application of statistics than we will get into
  • We are interested in “descriptive statistics”

13 of 87

DESCRIPTIVE STATISTICS

  • Descriptive statistics is one area within statistics

  • Provides tools to DESCRIBE, organize and interpret variability in our observations of the natural world

  • We will use tools from descriptive statistics in this course

14 of 87

SOME IMPORTANT VOCABULARY

  • Population:
    • Entire group of events, objects, results, or individuals, all of whom share some unifying characteristic

15 of 87

POPULATIONS

  • Examples:
    • All of a person’s red blood cells
    • All the enzyme molecules in a test tube
    • All the college students in the U.S.

16 of 87

SAMPLE

  • Sample: Portion of the whole population that represents the whole population
  • Example: It is virtually impossible to measure the level of hemoglobin in every cell of a patient
    • Rather, take a sample of the patient’s blood and measure the hemoglobin level

17 of 87

POPULATION AND SAMPLE

18 of 87

MORE ABOUT SAMPLES

  • Representative sample: sample that truly represents the variability in the population -- good sample

19 of 87

TWO VOCABULARY WORDS

  • A sample is random if all members of the population have an equal chance of being drawn

  • A sample is independent if the choice of one member does not influence the choice of another

  • Samples need to be taken randomly and independently in order to be representative

20 of 87

SAMPLING

  • How we take a sample is critical and often complex
  • If sample is not taken correctly, it will not be representative

21 of 87

EXAMPLE

  • How would you sample a field of corn?

22 of 87

EXAMPLE CONT.

  • Did you think about factors that could affect sample
    • Edge of field vs middle
    • Topography (might affect how rain runs off)
    • What is adjacent on each side
    • Other factors?

23 of 87

VARIABLES

  • Variables:
    • Characteristics of a population (or a sample) that can be observed or measured

    • Called variables because they can vary among individuals

24 of 87

VARIABLES

  • Examples:
    • Blood hemoglobin levels

    • Activity of enzymes

    • Test scores of students

25 of 87

VARIABLES

  • A population or sample can have many variables that can be studied
  • Example
    • Same population of six-year-old children can be studied for
      • Height
      • Shoe size
      • Reading level
      • Etc.

26 of 87

DATA

  • Data: Observations of a variable (singular is datum)
    • May or may not be numerical

  • Examples:
    • Heights of all the children in a sample (numerical)
    • Lengths of insects (numerical)
    • Pictures of mouse kidney cells (not numerical)

27 of 87

ALWAYS UNCERTAINTY

  • Even if you take a sample correctly, there is uncertainty when you use a sample to represent the whole population
    • Various samples from the same population are unlikely to be identical

  • So, need to be careful about drawing conclusions about a population, based on a sample – there is always some uncertainty

28 of 87

SAMPLE SIZE

  • If a sample is drawn correctly, then, the larger the sample, the more likely it is to accurately reflect the entire population

  • If it is not done correctly, then a bigger sample may not be any better

  • How does this apply to the corn field?

29 of 87

INFERENTIAL STATISTICS

  • Another branch of statistics

  • Won’t talk about it much

  • Deals with tools to handle the uncertainty of using a sample to represent a population

30 of 87

EXAMPLES

  • In a quality control setting, 15 vials of product from a batch are tested. What is the sample? What is the population?

  • In an experiment, the effect of a carcinogenic compound was tested on 2000 lab rats. What is the sample? What is the population?

  • A clinical study of a new drug was tested on fifty patients. What is the sample? What is the population?

31 of 87

ANSWERS

  • 15 vials, the sample, were tested for QC. The population is all the vials in the batch.

  • The sample is the rats that were tested. The population is probably all lab rats.

  • The sample is the 50 patients tested in the trial. The population is all patients with the same condition.

32 of 87

EXAMPLE PROBLEM

  • An advertisement says that 2 out of 3 doctors recommend Brand X.
    • What is the sample? What is the population?
    • Is the sample representative?
    • Does this statement ensure that Brand X is better than competitors?

33 of 87

ANSWER

Many abuses of statistics relate to poor sampling. The population of interest is all doctors. No way to know what the sample is. The sample could have included only relatives of employees at Brand X headquarters, or only doctors in a certain area. Therefore, the statement does not ensure that most doctors recommend Brand X. It certainly does not ensure that Brand X is best.

34 of 87

DESCRIBING DATA SETS

  • Draw a sample from a population

  • Measure values for a particular variable

  • Result is a data set

35 of 87

DATA SETS

  • Individuals vary, therefore the data set has variation

  • Data without organization is like letters that aren’t arranged into words

36 of 87

DATA SETS

  • Numerical data can be arranged in ways that are meaningful – or that are confusing or deceptive

37 of 87

DESCRIPTIVE STATISTICS

  • Provides tools to organize, summarize, and describe data in meaningful ways

  • Example:
    • Exam scores for a class is the data set
    • What is the variable of interest?
    • Can summarize with the class “average”, what does this tell you?

38 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

39 of 87

AVERAGE OR MEAN

  • A measure that describes a data set, such as the average, is sometimes called a “statistic”

  • Average gives information about the center of the data

  • We will consider mean and average to be synonyms (mathematicians distinguish them)

40 of 87

MEDIAN AND MODE

  • Two other statistics that give information about the center of a set of data

  • Median is the middle value

  • Mode is most frequent value

41 of 87

MEASURES OF CENTRAL TENDENCY

  • Measures that describe the center of a data set are called: Measures of Central Tendency

  • Mean, median, and the mode

42 of 87

HYPOTHETICAL DATA SET

2 5 6 7 8 3 9 3 10 4 7 4 6 11 9

Simplest way to organize them is to put in order:

2 3 3 4 4 5 6 6 7 7 8 9 9 10 11

By inspection they center around 6 or 7

43 of 87

MEAN

  • Add all the numbers together and divide by number of values

2 3 3 4 4 5 6 6 7 7 8 9 9 10 11

What is the mean for this data set?

44 of 87

NOMENCLATURE

  •  

45 of 87

EXAMPLE

  • Data set

2 3 3 4 5 6 7 8 9

    • What is the mode?
    • What is the median?

46 of 87

MEAN OF A POPULATION VERSUS THE MEAN OF A SAMPLE

  •  

47 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

48 of 87

DISPERSION

  • Data sets A and B both have the same average:

A 4 5 5 5 6 6

B 1 2 4 7 8 9

  • But are not the same:
    • A is more clumped around the center of the central value
    • B is more dispersed, or spread out

49 of 87

MEASURES OF DISPERSION

  • Measures of central tendency do not describe how dispersed a data set is
  • Measures of dispersion do; they describe how much the values in a data set vary from one another

50 of 87

MEASURES OF DISPERSION

  • Common measures of dispersion are:
    • Range
    • Variance
    • Standard deviation
    • Coefficient of variation

51 of 87

CALCULATIONS OF DISPERSION

  • Measures of dispersion, like measures of central tendency, are calculated
  • Range is the difference between the lowest and highest values in a data set

52 of 87

EXAMPLE

  • Example data:

2 3 3 4 4 5 6 6 7 7 8 9 9 10 11

  • Range: 11-2 = 9 or, 2 to 11

  • Range is not particularly informative because it is based only on two values from the data set

53 of 87

CALCULATING VARIANCE AND STANDARD DEVIATION

  • Variance and standard deviation measure of the average amount by which each observation varies from the mean

  • Example, data set, lengths of 8 insects:

4 cm 5 cm 6 cm 7 cm

7 cm 7 cm 9 cm 11 cm

54 of 87

DEVIATION

4 cm 5 cm 6 cm 7 cm 7 cm 7 cm 9 cm 11 cm

  • The mean is 7 cm

  • How much do they vary from one another?

  • Intuitively might see how much each point varies from the mean
    • This is called the deviation

55 of 87

CALCULATION OF DEVIATIONS FROM MEAN

4 cm 5 cm 6 cm 7 cm 7 cm 7 cm 9 cm 11 cm

Value-Mean (in cm) Deviation (in cm)

(4-7) - 3

(5-7) - 2

(6-7) - 1

(7-7) 0

(7-7) 0

(7-7) 0

(9-7) +2

(11-7) +4

56 of 87

SUM OF DEVIATIONS

Value-Mean Deviation

(in cm)

(4-7) - 3

(5-7) - 2

(6-7) - 1

(7-7) 0

(7-7) 0

(7-7) 0

(9-7) +2

(11-7) +4

Sum of deviations = 0

57 of 87

SUM OF DEVIATIONS IS ZERO

  • Sum of the deviations from the mean is always zero

  • Therefore, cannot use the average deviation

  • Therefore, mathematicians decided to square each deviation so they will get positive numbers

58 of 87

SUM OF SQUARED DEVIATIONS

Value-Mean Deviation Squared Deviation

(in cm)

(4-7) - 3 9 cm2

(5-7) - 2 4 cm2

(6-7) - 1 1 cm2

(7-7) 0 0

(7-7) 0 0

(7-7) 0 0

(9-7) +2 4 cm2

(11-7) +4 16 cm2

total squared deviation = sum of squares = 34 cm2

59 of 87

VARIANCE

  • Total squared deviation (sum of squares) divided by the number of measurements:

34 cm2 = 4.25 cm2

8

60 of 87

STANDARD DEVIATION (SD)

  •  

61 of 87

VARIANCE OF POPULATION VS SAMPLE

  • Statisticians distinguish between the mean and SD of a population and a sample

  • The variance of a population is called sigma squared, σ2

  • Variance of a sample is S2

62 of 87

SD OF POPULATION VS SAMPLE

  • The standard deviation of a population is called sigma, σ

  • Standard deviation of a sample is S or SD

63 of 87

STANDARD DEVIATION OF A SAMPLE

64 of 87

EXAMPLE PROBLEM

A biotechnology company sells cultures of E. coli. The bacteria are grown in batches that are freeze dried and packaged into vials. Each vial is expected to have 200 mg of bacteria. A QC technician tests a sample of vials from each batch and reports the mean weight and SD.

65 of 87

EXAMPLE CONT.

Batch Q-21 has a mean weight of 200 mg and a SD of 12 mg. Batch P-34 has a mean weight of 200 mg and as SD of 4 mg. Which lot appears to have been packaged in a more controlled fashion?

66 of 87

ANSWER

The SD can be interpreted as an indication of consistency. The SD of the weights of Batch P-34 is lower than of Batch Q-21. Therefore, the weights for vials for Batch P-34 are less dispersed than those for Batch Q-21 and Batch P-34 appears to have been better controlled.

67 of 87

LECTURE OVERVIEW

  • Why learn about statistics?
  • Some basic vocabulary and concepts
  • Measures of central tendency
  • Measures of dispersion
  • Graphical methods, frequency distributions

68 of 87

FREQUENCY DISTRIBUTIONS

  • So far, talked about calculations to describe data sets
  • Now talk about graphical methods

69 of 87

�THE WEIGHTS OF 175 FIELD MICE

(in grams)

19 22 20 24 22 19 27 20 21 22 20 22 24 24 21 25 19 21 20 23 25 22 19 17 20 20 21 25 21 22 27 22 19 22 23 22 25 22 24 23 20 21 22 23 21 24 19 21 22 22 25 22 23 20 23 22 22 26 21 24 23 21 25 20 23 20 21 24 23 18 20 23 21 22 22 25 21 23 22 24 20 21 23 21 19 21 24 20 22 23 20 22 19 22 24 20 25 21 22 22 24 21 22 23 25 21 19 19 21 23 22 22 24 21 23 22 23 28 20 23 26 21 22 24 20 21 23 20 22 23 21 19 20 26 22 20 21 22 23 24 20 21 23 22 24 21 23 22 24 21 22 24 20 22 21 23 26 21 22 23 24 21 23 20 20 21 25 22 20 22 21 21 23 22

70 of 87

FREQUENCY DISTRIBUTION TABLE OF THE WEIGHTS OF FIELD MICE

Weight Frequency

(in grams)

17 1

18 1

19 11

20 25

21 34

22 40

23 27

24 19

25 10

26 4

27 2 28 1

71 of 87

FREQUENCY TABLE

  • Tells us that most mice have weights in the middle of the range, a few are lighter or heavier
  • The word distribution refers to a pattern of variation for a given variable

72 of 87

FREQUENCY DISTRIBUTIONS

  • It is important to be aware of patterns, or distributions, that emerge when data are organized by frequency

  • The frequency distribution can be illustrated as a frequency histogram

73 of 87

FREQUENCY HISTOGRAMS

  • X-axis is units of measurement, in this example, weight in grams

  • Y-axis is the frequency of a particular value

  • For example, 11 mice weighed 19 g

  • The values for these 11 mice are illustrated as a bar

74 of 87

FREQUENCY HISTOGRAMS

  • Note that when the mouse data were collected, a mouse recorded as 19 grams actually weighed between 18.5 g and 19.4 g.

  • Therefore, the bar spans an interval of 1 gram

75 of 87

FIRST FOUR BARS

WEIGHTS IN GRAMS

17 18 19 20

F

R

E

Q

U

E

N

C

Y

76 of 87

CONSTRUCTING A FREQUENCY HISTOGRAM

  • Divide the range of the data into intervals
  • It is simplest to make each interval (class) the same width
  • No set rule as to how many intervals to have
  • For example, length data might be 1-9 cm, 10-19 cm, 20-29 cm and so on

77 of 87

FREQUENCY HISTOGRAMS

  • Count the number of observations that are in each interval
  • Make a frequency table with each interval and the frequency of values in that interval
  • Label the axes of a graph with the intervals on the X-axis and the frequency on the Y-axis

78 of 87

FREQUENCY HISTOGRAMS

  • Draw in bars where the height of a bar corresponds to the frequency of the value

  • Center the bars above the midpoint of the class interval

  • For example, if the interval is 0-9 cm, then the bar should be centered at 4.5 cm

79 of 87

EXAMPLE PROBLEM

Multiple Choice:

What frequency distribution is illustrated in this histogram?

a. All the mice are of the same weight.

b. There are the same number of mice in each weight class.

c. Neither of the above.

80 of 87

ANSWER

b is correct. The histogram illustrates a situation where there are four mice in each weight class.

81 of 87

NORMAL FREQUENCY DISTRIBUTION

  • If weights of very many lab mice were measured, would likely have a frequency distribution that looks like a bell shape, also called the “normal distribution

82 of 87

NORMAL DISTRIBUTION

WEIGHT

F

R

E

Q

U

E

N

C

Y

83 of 87

NORMAL DISTRIBTION

  • Very important

  • Examples:
    • Heights of humans
    • Measure same thing over and over, measurements will have this distribution
    • In this course we talk a lot about measurements

84 of 87

CALCULATIONS AND GRAPHICAL METHODS

  • Related

  • The center of the peak of a normal curve is the mean, the median and the mode

  • Values are evenly spread out on either side of that high point

85 of 87

CALCULATIONS AND GRAPHICAL METHODS

  • The width of the normal curve is related to the SD

  • The more dispersed the data, the higher the SD and the wider the normal curve

  • Exact relationship is in text, not go into it this semester

86 of 87

ASSIGNMENT

  • To be sure you understand these calculations and ideas, practice on problems 1-16 pages 359-362 in your textbook

  • The remaining problems in the chapter are optional for the purposes of this course

87 of 87

TO DELVE DEEPER INTO THE TOPICS IN THIS LECTURE

  • Chapter 14 in Basic Laboratory Methods for Biotechnology: Textbook and Laboratory Reference, 3rd Edition has more explanation of these statistical methods and a number of problems to help master basic statistical tools.