1 of 26

Exploratory Data Analysis�And�Statistical Thinking

2 of 26

Numerical Summaries

  • Mean: mean(a) = (2+3+4+5+9) / 5 = 4.6

  • Median: 4

  • Variance: var(a) = ( 2.6^2 + 1.6^2 + 0.6^2 + 0.4^2 + 4.4^2 ) / (5-1) = 7.3

  • Standard deviation: sd(a) = sqrt(7.3) = 2.7

a = c(2, 3, 4, 5, 9)

Unbiasedness

3 of 26

Data Visualization Principles�From Karl Broman https://kbroman.org/AdvData/19_datavis_notes.pdf

  • Be accurate and clear.
  • Let the data speak.
  • Science not sales.
  • In tables, avoid too many digits. Don’t drop ending 0’s.

  • Most examples presented in this lecture were taken from Karl Broman and Rafael Irizarry.

4 of 26

Show the data

3D barplot is unnecessary. Humans are bad at seeing in 3D dimensions.

5 of 26

Boxplot with data points

Median (50th percentile)

25th percentile

75th percentile

Minimum excluding outliers

Maximum excluding outliers

6 of 26

Piecharts

Can you rank the percentage of the browser? Can you see the difference from 2000 to 2015?

Humans are not good at quantifying and comparing angles.

7 of 26

Barplots are better

Humans are good at visualizing and comparing length.

8 of 26

Mistaken data labels

Source: Introduction to GPT5 https://www.youtube.com/watch?v=0Uu_VJeVVfo

9 of 26

Know when to include 0

  • By avoiding 0, relatively small differences can be made to look much bigger than they actually are.

10 of 26

Don’t sort alphabetically

11 of 26

Consider logs

12 of 26

Use color to compare groups

13 of 26

Use color to compare groups

Avoid using red-green contrast to make it friendly to people who are color blind.

14 of 26

Avoid too many digits and align the numbers

15 of 26

Association/Correlation is not Causation

  • Perhaps the most important concept to learn for statistical thinking.

16 of 26

What is causation?

  • Taking ibuprofen reduces pain levels

  • Causation is stronger than correlation. It means change in one thing causes another thing to change. Causation implies correlation, but not vice versa.

17 of 26

Confounding

X

Y

C

Confounder: Change in C causes change in X and Y

Exposure/Treatment

Outcome

Sometimes we know confounder and can adjust for it, but often we don’t know.

Identify confounder requires statistical thinking and domain knowledge.

18 of 26

Simpson’s Paradox�X and Y negatively correlated

19 of 26

Simpson’s Paradox�X and Y negatively correlated�Z negatively correlated with X; X and Y are actually positively correlated

20 of 26

1973 UC Berkley Graduate Admission

It appears that there is negative correlation between gender and admission rate.

Stratified by department, more women applied to departments with lower admission rate. Within each department, there is not much correlation between gender and admission rate.

21 of 26

Clinical Trials

Patients who take Drug A have lower mortality than those who don’t.

Can we say Drug A is effective in reducing mortality?

Sicker patients may avoid Drug A.

Healthier patients may be more likely to take Drug A.

22 of 26

Randomized Clinical Trials (RCT)

Randomization: Randomly assign each individual to treatment/control group.

Sicker and healthier patients have the same probability of being assigned Drug A.

Overall, randomization ensures that the distribution of sicker and healthier patients is expected to be similar between the treatment and control groups.

This applies to other confounders as well.

23 of 26

Other Notable Confounding Examples in Biology��Genetic Association Between a Variant and Type 2 Diabetes�500 cases and 500 controls

Population

Cases

Controls

Frequency of a genetic variant X

A (East Asian)

300

100

60%

B (European)

200

400

10%

Overall, variant X appears much more common in cases (300*0.6 + 200*0.1) / 500 = 40%

than controls (100*0.6 + 400*0.1) / 500 = 20%

Naively, you might conclude: “SNP X increases diabetes risk.”

Here, ancestry is a confounder because:

Population A has higher SNP X frequency.

Population A also has higher T2D prevalence (maybe due to diet, lifestyle, or other genetic factors).

24 of 26

Other Notable Confounding Examples in Biology�Batch Effects

  • Are Human and mouse gene expression data more similar across tissues or across species?

25 of 26

In a DNA sequencing machine, there are different lanes. Human samples and mouse samples are confounded by lanes.

26 of 26

What is a correct study design?

  • Distributing human and mouse tissue randomly across lanes.