Exploratory Data Analysis�And�Statistical Thinking
Numerical Summaries
a = c(2, 3, 4, 5, 9)
Unbiasedness
Data Visualization Principles�From Karl Broman https://kbroman.org/AdvData/19_datavis_notes.pdf
�
Show the data
3D barplot is unnecessary. Humans are bad at seeing in 3D dimensions.
Boxplot with data points
Median (50th percentile)
25th percentile
75th percentile
Minimum excluding outliers
Maximum excluding outliers
Piecharts
Can you rank the percentage of the browser? Can you see the difference from 2000 to 2015?
Humans are not good at quantifying and comparing angles.
Barplots are better
Humans are good at visualizing and comparing length.
Mistaken data labels
Source: Introduction to GPT5 https://www.youtube.com/watch?v=0Uu_VJeVVfo
Know when to include 0
Don’t sort alphabetically
Consider logs
Use color to compare groups
Use color to compare groups
Avoid using red-green contrast to make it friendly to people who are color blind.
Avoid too many digits and align the numbers
Association/Correlation is not Causation
What is causation?
Confounding
X
Y
C
Confounder: Change in C causes change in X and Y
Exposure/Treatment
Outcome
Sometimes we know confounder and can adjust for it, but often we don’t know.
Identify confounder requires statistical thinking and domain knowledge.
Simpson’s Paradox�X and Y negatively correlated
Simpson’s Paradox�X and Y negatively correlated�Z negatively correlated with X; X and Y are actually positively correlated
1973 UC Berkley Graduate Admission
It appears that there is negative correlation between gender and admission rate.
Stratified by department, more women applied to departments with lower admission rate. Within each department, there is not much correlation between gender and admission rate.
Clinical Trials
Patients who take Drug A have lower mortality than those who don’t.
Can we say Drug A is effective in reducing mortality?
Sicker patients may avoid Drug A.
Healthier patients may be more likely to take Drug A.
Randomized Clinical Trials (RCT)
Randomization: Randomly assign each individual to treatment/control group.
Sicker and healthier patients have the same probability of being assigned Drug A.
Overall, randomization ensures that the distribution of sicker and healthier patients is expected to be similar between the treatment and control groups.
This applies to other confounders as well.
Other Notable Confounding Examples in Biology��Genetic Association Between a Variant and Type 2 Diabetes�500 cases and 500 controls
Population | Cases | Controls | Frequency of a genetic variant X |
A (East Asian) | 300 | 100 | 60% |
B (European) | 200 | 400 | 10% |
Overall, variant X appears much more common in cases (300*0.6 + 200*0.1) / 500 = 40%
than controls (100*0.6 + 400*0.1) / 500 = 20%
Naively, you might conclude: “SNP X increases diabetes risk.”
Here, ancestry is a confounder because:
Population A has higher SNP X frequency.
Population A also has higher T2D prevalence (maybe due to diet, lifestyle, or other genetic factors).
Other Notable Confounding Examples in Biology�Batch Effects
In a DNA sequencing machine, there are different lanes. Human samples and mouse samples are confounded by lanes.
What is a correct study design?