CMSC/DATA 320
Experimental Design
Introduction to Data Science
Instructor
Fardina Fathmiul Alam
fardina@umd.edu
Lecture 02
Establishing Causal Relationships
A classic cartoon highlighting the importance of proper control, measurement, and the potential pitfalls in experimental setups.
Source: hawaii.edu/fishlab/NearsideFrame.htm
2
Today's Objectives
Goal & Focus
Today, we'll cover the basics of experimental design, including how to plan and conduct experiments. The goal is to help you design and analyze experiments more effectively by identifying variables, hypotheses, and confounding factors.
Syllabus & Key Topics
Companion Textbook Reading
Chapter 3: ffalam.github.io/CMSC320TextBook/chapter3/Chapter_3_0.html
3
Experimental Design in Data Science?
The process of planning, conducting, and analyzing experiments to test hypotheses and gather meaningful data for data-driven decisions.
Data science fundamentally involves making decisions based on data.
4
Asking the Right Questions
before solving a Data Science problem is a great start! And be specific!
How many visitors did the website 'X' receive last week?
Why did website traffic drop last weekend?
What will the weather be like for the next 10 days?
Does offering free shipping increase the number of purchases?
Which courses to offer to maximize enrollment next semester?
Optimization Criteria
Objective or goal function we want to maximize or minimize.
Methodology Note: Causal Experimental Design
We use general experimental design to collect reliable data, but we strictly need causal experimental design (for causal and prescriptive problems) when testing the exact effect of real-world interventions (e.g., A/B testing).
Descriptive
Diagnostic
Prescriptive
Causal
Predictive
Different questions lead to different models, data needs, and evaluation criteria.
5
Topics
01. Introduction & Variables Next Topic
02. Hypothesis
Formulating and testing a testable statement.
03. Confounder Variables
04. Dealing with Confounders
05. Methods for Collecting Data
06. Bias & Blinding
6
Example: Online Retail
The Scenario
You are a data scientist testing whether changing the color of the "Buy Now" button on your website affects the Click-Through Rate (CTR).
Question to Consider
How can we set up an experiment to collect data in this case?
Buy It Now
What is Your Problem Definition?
Find which version (Option A default or Option B red) is more likely to maximize the CTR.
What is your Optimization Criteria? What we want to maximize?
Select the button option that leads to a Higher CTR (percentage of users who click on the button after seeing it).
7
Experimental Setup
How can we set up an experiment to collect data in this case?
Control Group
Original Website
Views existing button color. Experiences no changes; serves as the baseline.
Treatment Group
Modified Website
Sees different color for "Buy Now" button. Experiences the active change.
Buy It Now
Track click-through rates (CTR) for both groups over a specified testing period.
Evaluate performance differences between Control and Treatment CTRs post-experiment.
Determine if the button color change statistically influenced the final
outcome.
Data Size / Sample: Number of website visitors
Experimental Design
Identifying Key Variables
Data Size / Sample: Number of website visitors
Control Group: Views original blue button. Experiences no active changes (baseline).
Treatment Group:Views different red button color. Experiences the test change.
Goal: Comparing these groups isolates and identifies the true effect of the Independent Variable on the Dependent Variable.
Buy It Now
Manipulated
Independent Variable
Button Color
The factor we actively change (e.g., Blue vs. Red "Buy It Now") to observe its effect.
Measured Outcome
Dependent Variable
Click-Through Rate
The outcome metric measured to evaluate if changing the color made a difference.
Ques: What are the variables here?
Summary: Variables, Population, and Groups
9
Once the research problem is defined, identify the key variables of interest, establish your study groups, and define the target population.
1. Variables of Interest
Independent Variable (IV)
The variable that is manipulated or changed by the researcher to observe its direct effects.
Examples: Drug dosage, new algorithm, marketing strategy.
Dependent Variable (DV)
The outcome being measured, which is expected to change in direct response to the IV.
Examples: Patient recovery rate, sales revenue, user engagement.
2. Study Groups
Treatment Group
Receives the specific treatment or active intervention (the IV is applied directly to this group).
Control Group
Does not receive the intervention. Serves as the baseline for exact comparison.
Why compare both?
Comparing these groups isolates and identifies the true effect of the IV on the DV.
3. Population & Sample
Defining the Scope:
Always clearly specify the broader population or the specific subset sample that your scientific research will focus on.
A well-defined sample is essential for drawing accurate, generalizable conclusions.
10
Topics
01. Introduction & Variables
02. Hypothesis Next Topic
Formulating and testing a testable statement.
03. Confounder Variables
04. Dealing with Confounders
05. Methods for Collecting Data
06. Bias & Blinding
11
Come up with a Hypothesis
A hypothesis is a testable statement you want to evaluate.
"If X is true, then Y should happen."
What is a hypothesis?
Describes the relationship between variables and outcomes
How do we test it?
Run an experiment or make observations
12
Brainstorming Time
Hypothesis: "If the amount of average study time (___ variable?) is increased, then average exam scores (___ variable?) will also increase."
01. Question 1
What is the optimization goal/criteria here? How are independent and dependent variables related?
02. Question 2
While independent variables affect dependent variables, multiple independent variables (can / can not) influence each other.
Good experimental design aims to minimize correlated variables.
13
Topics
01. Introduction & Variables
02. Hypothesis
Formulating and testing a testable statement.
03. Confounder Variables Next Topic
04. Dealing with Confounders
05. Methods for Collecting Data
06. Bias & Blinding
14
Confounders (Before Data Collection)
Confounder: An external variable that affects the DV and distorts the IV → DV relationship if not controlled.
Why it matters: Can lead to incorrect conclusions. Not the study focus, but must be managed.
Examples
01. Exercise Scenario
"If the exercise duration is extended, then the average calories burned will also increase."
02. Literacy Scenario
"As books read increases, average literacy also increases."
Confounder: metabolic rate
Confounders: age, socioeconomic status etc.
More Example: Experimental Design Flow : Polling
01. Problem Formulation
Imagine you're a data scientist tasked with predicting the outcome of a political election using a dataset of voter preferences.
Goal: Design an experiment that accurately represents the entire population's voting behavior.
02. The Challenge
How do we know which candidate is ahead in a complex, multi-demographic environment?
Key Question: Can we create the IDEAL POLLING?
The Solution: Eliminate Confounding Variables
Minimize factors like Sample Bias, geographic representation, Population Proportion Bias, and Demographic Mismatch to make the collected data as representative and accurate as possible.
16
Topics
01. Introduction & Variables
02. Hypothesis
Formulating and testing a testable statement.
03. Confounder Variables
04. Dealing with Confounders Next Topic
05. Methods for Collecting Data
06. Bias & Blinding
17
Methodology
Some Ways to Deal with Confounder
Design Stage (Before Data Collection)
01. Randomization
Randomly assign units to groups to balance confounders on average.
• Random Sampling
Select representative subset to reduce bias.
• Stratified
Group by characteristic, then randomize.
02. Restriction
Limit the study to only one level of a confounder.
Key Benefit
By keeping the confounding factor completely constant, it is prevented from varying and introducing bias into results.
03. Matching
Pair or group units with highly similar confounder values.
Application
Enables direct comparisons across treatment and control groups by aligning corresponding subjects.
04. Replication
Repeat the entire experiment under similar conditions.
Outcome
Verifies consistency and substantially reduces the overall influence of any uncontrolled confounders.
Analysis Stage (After Data Collection)
Regression / multivariable models, Statistical adjustment
Example: Control Confounder Variable in an Experiment
If the amount of study time ( independent variable) is increased, then exam scores ( dependent variable) will also increase.
18
DESIGN IDEA 01
Stratified Randomization
Treatment
Treatment
Stratify students by prior knowledge (high/medium/low):
DESIGN IDEA 02
Block Design (Matched Pair)
Pair participants by prior knowledge:
19
Try by yourself: control the effect of “age”
“As books read increases, avg. literacy also increases.”
Goal: Measure the age of each individual to see and control the confounding effects of age on literacy.
Experimental Design: How to design?
Method 01
Random Sampling
Select a completely representative subset of the population to reduce overall age bias.
Method 02
Stratified Randomization
Group (stratify) participants by age groups first, then randomize treatment within each group.
Method 03
Block Design (Match Pair)
Pair participants of identical or highly similar ages, assigning one to control and one to treatment.
20
Topics
01. Introduction & Variables
02. Hypothesis
Formulating and testing a testable statement.
03. Confounder Variables
04. Dealing with Confounders
05. Methods for Collecting Data Next Topic
06. Bias & Blinding
21
Methods for Collecting Data
If pre-existing datasets are not available
Method A
Observational Studies
Observe and record data (variables) without intervening or manipulating variables.
E.g. Observing animal behavior in a natural habitat without any external influence.
Method B
Surveys
Collect information through structured, carefully designed questionnaires or interviews.
E.g. Conducting a survey to gather opinions on a political issue.
Method C
Experiments
We actively change or manipulate something to see what happens and measure the response.
Key: Allows researchers to establish clear cause-and-effect relationships.
Method D
Simulations
Create artificial scenarios to model real-world situations for safe or efficient data collection.
E.g. Using a computer simulation to study traffic patterns in a city.
A. Observational Studies
Cross-sectional studies: data collected at one time point
Example: survey people’s exercise habits today
Retrospective (case-control) studies: look back at past exposure
Example: compare smoking history of lung cancer patients vs. non-patients
Prospective (longitudinal/cohort) studies: follow a group (cohort) over time
Example: track smokers and non-smokers for 10 years
23
B. Surveys
A specific type of observational study
Method Overview
Collect data using structured questions or questionnaires to study specific populations.
Real-world Example: Surveying university students about their weekly study habits and exam-related stress levels.
The Survey Design Process
24
D. Simulation
Use a computer or mathematical model to mimic real-world systems
When to Use
Highly useful when real-world experiments are too costly, risky, or entirely impractical.
What-If Scenarios
Allows safe, repeated testing of diverse operational theories under strictly controlled settings.
Practical Example
Simulate city traffic flow to study road congestion without changing actual physical infrastructure.
25
Topics
01. Introduction & Variables
02. Hypothesis
Formulating and testing a testable statement.
03. Confounder Variables
04. Dealing with Confounders
05. Methods for Collecting Data
06. Bias & Blinding Next Topic
26
Placebo Effect
Improvement occurs due to belief in treatment, not the treatment itself
Key Characteristics
• Can bias results significantly if not controlled in studies.
• Highly common across medical, psychological, and behavioral research.
Example: Patients feel physically better after receiving a simple sugar pill they believe is active medicine.
Minimizing Bias through Blinding Designs
Design 01
Single-Blind
Participants are unaware of their group assignment, but researchers know who receives the treatment.
Design 02
Double-Blind
Neither the participants nor the researchers running the trial know who receives treatment vs. placebo.
Key Takeaway
Why Blinding Matters
Using a placebo helps keep participants completely unaware of their group, directly reducing experimental bias.
27
Key Takeaways
The Fundamental Rule of Data Collection
Your data must representative of the population you want to study.
Keep in mind that
It is almost impossible to be certain that your experiment has completely removed all forms of bias. It is necessary to consider possible sources of bias and highlight them in your analysis. Ideally, future experiments would improve upon your method by iteratively eliminating those sources of bias.
28
Quick Class Task
Identify which method for collecting data (observational study, experiment, simulation, or survey) is best in each of the following situations and explain your answer.
01
Earthquake Impact
The effect of a severe earthquake would have on the Salt Lake Valley.
02
Catalog Coupon Effectiveness
Whether or not a certain coupon attached to the outside of a catalog makes recipients more likely to order products.
03
Smoking & Health
Whether or not smoking has an effect on coronary heart disease.
04
Household Income
Determining the average household income of homes in Salt Lake City.