Week 4: Exploratory Data Analysis
Introduction to Data Visualization
W4995.003 Spring 2025
00 Quiz
01 Sparks Presentation
02 Revisit: C ⊂ O ⊂ Q
03 Monthly Budget Redesigns
04 What is EDA?
05 EDA Examples: Air Pollution, Stop & Frisk
06 EDA Assignment: Group critique
Quiz
5 min
Closed book
Slide via Jeff Heer
02
Review: Categorical, Ordinal, & Quantitative
Data Types: C ⊂ O ⊂ Q
C: Categorical
Operations: =, ≠
Categories are of equal importance, or “equidistant”
O: Ordered
Operations: =, ≠, <, >
Items of equal importance, or “equidistant”
Q-Interval (location of zero arbitrary)
Operations: =, ≠, <, >, -
Can measure distances or spans, only delta (i.e. intervals) may be compared
Q-Ratio (zero fixed)
Operations: =, ≠, <, >, -, %
Can measure ratios or proportions e.g. Length, Mass, Temp, counts and amounts
Slide via Jeff Heer
William Playfair (1786)
Image via Wikipedia
Inventor of line charts, bar charts, and pie charts.
British pounds
X-axis: year (Q)
Y-axis: currency (Q)
Color: imports/exports (C)
From the Pudding
Rectangle Area: laugh share (Q)
Rectangle Position: category (C)
Color Hue: category (double-encoded) (C)
Map of the Market (Wattenberg 2000)
Rectangle Area: market cap (Q)
Rectangle Position: market sector (C)
Color Hue: loss vs. gain (C)
Color Value: magnitude of loss or gain (Q)
Exercise: data type of zip code?
Ben Fry, Zipdecode (1999)
Zip code trivia (and why it is categorical)
Via wikipedia
Note: it’s difficult for a single O/Q value to span 2D
Via wikipedia
One option:
The Hilbert Curve
Compare & Contrast
Left: John Hopkins dashboard, right: Bloomberg
Compare & Contrast
NYTimes
Compare & Contrast
NYTimes
XKCD/NYTimes
03
Monthly Budgets
Stream graphs
A variation on the stacked area chart
03
What is Exploratory Data Analysis?
Design
Computer Graphics
HCI
Psychology
Statistics
Cartography
dataviz
EDA is...
A style of data analysis that employs graphical & statistical techniques to:
EDA always precedes formal (confirmatory) data analysis.
“The greatest value of a picture is when it forces us to notice what we never expected to see.”
— John Tukey
Anscombe’s Quartet
Anscombe’s Quartet (1973)
Mean & Variance
uX = 9.0, σX = 11
uY = 7.5, σY = 4.125
Linear Regression
Y = 3 + 0.5X
R^2 = 0.67
Tukey 1977
Based on insights developed at Bell Labs in ’60s
Introduced new techniques for visualizing and summarizing data:
Boxplot 5-Number Summary
Boxplot 5-Number Summary
Note: Max/Min in Boxplots
https://www.leansigmacorporation.com/box-plot-with-minitab/
Counterintuitive: less space means more density
https://www.leansigmacorporation.com/box-plot-with-minitab/
EDA differs from classical analysis
Exploratory data analysis is sometimes compared
to detective work: it is the process of gathering evidence.
Confirmatory data analysis is comparable to a
court trial: it is the process of evaluating evidence.
We use graphics in data analysis to...
Use visual pattern detection to guide analysis
Identifying trends
Constant
Periodic
Linear
Exponential
Identifying trends
Identifying trends
Simple distributions
Uniform
Normal
Exponential
Graphics via Petra Isenberg
Simple distributions
Log-Normal
Skewed right (tail on right)
Skewed left
Graphics via Petra Isenberg
Use EDA to help you guess appropriate models
https://blog.cloudera.com/blog/2015/12/common-probability-distributions-the-data-scientists-crib-sheet/
What we do when we do EDA...
Next steps, iterative:
“The goal of data exploration is to generate many promising leads that you can later explore in more depth.”
R for Data Science, Wickham
Characteristics of exploratory graphs
04
EDA Example: EPA Air Pollution
Via Roger Peng, John Hopkins Biostatistics
Dataset: Air Pollution in the US
Overview of Process
Dataset: Air Pollution in the US
pm25 fips region long lat
9.771 01003 east -87.75 30.59
9.994 01027 east -85.84 33.27
10.689 01033 east -87.73 34.73
11.337 01049 east -85.80 34.46
12.120 01055 east -86.03 34.02
10.828 01069 east -85.35 31.19
Via Roger Peng, EPA 2008-2010 https://aqs.epa.gov/aqsweb/documents/data_mart_welcome.html
Boxplot: PM2.5
Boxplot with overlaid feature
Histogram: PM2.5
Histogram with different bin size
Second variable: Categorical
Small multiples: PM2.5 x Region
Small multiples: PM2.5 x Region
the levels in eastern counties are on average higher than the levels in western counties.
Multiple histograms: PM2.5 x Region
Scatterplot: PM2.5 x Latitude
South
North
Scatterplot: PM2.5 x Latitude
highest levels of PM2.5 tend to be in the middle region of the country.
~South
~North
Scatterplot: PM2.5 x Latitude x Region
East: black
West: red
Easier if separate: PM2.5 x Latitude x Region
04
EDA Example: NYPD Stop & Frisk
“Stop, Question, and Frisk” by Bloomberg’s NYPD
“Stop, Question, and Frisk” by Bloomberg’s NYPD
“Stop, Question, and Frisk” by Bloomberg’s NYPD
Hypotheses & questions
Data Source
Hypothesis 1
The number of people stopped, questioned, and frisked in New York City declined after Floyd v. City of New York on Oct. 31, 2013.
Overall, looks like significant decline in number of stops
Plotting by day, we see there might be misformatted or missing data.
Missing data?
Zooming in we see:
Number of stops seems to fluctuate significantly but regularly.
This could be due to weekend/week-
day or reporting patterns.
Overall, stop and frisks declined after the Floyd decision.
However, there was already a downward trend from early 2012 in the number of stops.
Further research is needed to understand how the court orders were implemented.
Lawsuit
Decision
Hypothesis 2
After adopting the stricter “probable cause” standard, are police making a) fewer and b) more effective stops?
Most stops don’t result in an arrest.
In fact only about 1 out of 10 results in an arrest.
Over 2013-2015, with decline in number of stops, number of arrests have also declined.
By 2015, percentage of stops that resulted in arrests were more than double that of 2013.
Hypothesis 3
With the change in procedure, is racial disparity also narrowing?
There’s a lot of messy data in the field.
Using valid values from data spec sheet, we can filter out invalid values.
Already we see disproportionate representation.
When splitting into into small multiples by year, overall decrease in number of stops masks any changes we can see. What we’re interested in comparing is part-of-whole.
Abela, Advanced Presentations by Design, 2013 redrawn by Berinato in Good Charts
There may be a very slight trend towards less racial disparity.
A slightly clearer view of the same, since very slight differences in slope are hard to perceive.
Summary: Exploratory Data Analysis
Summary: Exploratory Data Analysis
Additional EDA Resources
https://itl.nist.gov/div898/handbook/eda/eda.htm
http://greenteapress.com/thinkstats2/html/thinkstats2002.html
https://www.youtube.com/watch?v=pgSLSYLNEq0
06
EDA: Brainstorming session
Form pairs with someone not in your group
3 min each
Describe the topics you’re examining and the data set you’ve found, then share your initial charts
5 min
Work independently to come up with 3 hypotheses for the other group’s topic. Suggest at least one additional variable that could be interesting.
Form pairs with someone not in your group
5 min
Compare your hypotheses with your partner’s and choose which seem most interesting/fruitful for further investigation
06
Tableau Concepts
Tableau/Polaris Contribution
Taxonomy of vis types
Stolte, et. al. Polaris, ACM 2008.
Tableau UI chart types
Valid chart types for the selected dimensions & measures
Tableau Fields: “Scale” of Fields
Nominal
Ordinal
Q-Ratio
Q-Interval
Discrete
Continuous
Continuous
Discrete
Table Algebra to specify configurations
Operands are the database fields
Each operand interpreted as a set {...}
Continuous and Discrete fields treated differently
Three operators
concatenation (+)
cross product (x)
nest (/)
Via Jeff Heer.
x
/
+
/
+
x
Via Jeff Heer.
/
+
x
Via Jeff Heer.
GROUP BY category, region, segment
Relational Data Model
Example: U.S. Census Data
People Count: # of people in group
Year: 1850–2000 (every decade)
Age: 0–90+
Sex: Male, Female
Martial Status: married,
never married,
widowed,
divorced
Example via Jeffrey Heer.
“Roll-Up” and “Drill-Down”
Examine population by year and age?
Roll-up the data along the desired dimensions
SELECT year, age, SUM(people)
FROM census
GROUP BY year, age
Via Jeff Heer.
Dimensions
Measure
Drill-Down
see the breakdown by marital status?
Drill-down into additional dimensions
SELECT year, age, marst, SUM(people) FROM census
GROUP BY year, age, marst
Via Jeff Heer.
Data Cube Concept
Via Jeff Heer.
Via Jeff Heer.
Via Jeff Heer.
Questions?
Next Week…
Topics Next Week
Checklist For Next Week