1
Applied Data Analysis (CS401)
Maria Brbic / Robert West
Lecture 5
Regression for disentangling data
09 Oct 2024
Linear regression as you know it
2
Scalar product (a.k.a. dot product) of 2 vectors
Example with one predictor
3
X
y
𝛽1: intercept
𝛽2: slope
y ≈ 𝛽1 + 𝛽2X
3
Linear regression as you know it
4
?
Optimality criterion: least squares
5
Use cases of regression
6
Regression as comparison of�average outcomes
7
Example with one binary predictor Xi
yi = 𝛽1 + 𝛽2Xi + 𝜖i .
kid_score = 78 + 12 · mom_hs + error
8
0
1
mom_hs
kid_score
20 60 100 140
No
Yes
mean kid_score for moms who didn’t finish high school: 78
mean kid_score for moms who finished high school: 78 + 12 = 90
One binary predictor Xi:�Interpretation of fitted parameters 𝛽
yi = 𝛽1 + 𝛽2Xi + 𝜖i .
9
One binary predictor Xi:�Interpretation of fitted parameters 𝛽
yi = 𝛽1 + 𝛽2Xi + 𝜖i .
10
So why not just compute the two means separately and then compare them?
What a mean monkey!
Example with one continuous predictor Xi
yi = 𝛽1 + 𝛽2Xi + 𝜖i .
kid_score = 26 + 0.6 · mom_iq + error
11
mom_iq
kid_score
0 50 100
estimated (hypothetical) mean kid_score for moms with IQ = 0: 26
estimated mean kid_score for moms with IQ = 100: 26 + 0.6 · 100 = 86
0 50 100 150
One continuous predictor Xi:�Interpretation of fitted parameters 𝛽
yi = 𝛽1 + 𝛽2Xi + 𝜖i .
12
Example with multiple predictors
yi = 𝛽1 + 𝛽2Xi2 + 𝛽3Xi3 + 𝜖i .
kid_score = 26 + 6 · mom_hs + 0.6 · mom_iq + error
No
Yes
13
Example with multiple predictors
kid_score = 26 + 6 · mom_hs + 0.6 · mom_iq + error
mom_iq
kid_score
20 60 100 140
80 100 120 140
kids of moms who didn’t finish high school:
intercept = 26
slope = 0.6
kids of moms who finished high school:
intercept = 26 + 6 = 32
slope = 0.6
14
Example with interaction of predictors
yi = 𝛽1 + 𝛽2Xi2 + 𝛽3Xi3 + 𝛽4Xi2Xi3 + 𝜖i .
kid_score = −11 + 51 · mom_hs + 1.1 · mom_iq − 0.5 · mom_hs · mom_iq + error
No
Yes
15
Example with interaction of predictors
kid_score = −11 + 51 · mom_hs + 1.1 · mom_iq − 0.5 · mom_hs · mom_iq + error
mom_iq
kid_score
20 60 100 140
80 100 120 140
kids of moms who didn’t finish high school:
intercept = −11
slope = 1.1
kids of moms who finished high school:
intercept = −11 + 51 = 40
slope = 1.1 − 0.5 = 0.6
16
So why not just compute the two means separately and then compare them?
17
So why not just compute the two means separately and then compare them?
avg kid_score 90 | avg kid_score 90 |
avg kid_score 78 | avg kid_score 78 |
Mom finished high school
Mom�didn’t finish high school
Mom drives Mercedes
Mom doesn’t drive Mercedes
990 women | 10 women |
10 women | 990 women |
Mom finished high school
Mom�didn’t finish high school
Mom drives Mercedes
Mom doesn’t drive Mercedes
18
19
THINK FOR A MINUTE:
What is the mean outcome for Mercedes-driving moms vs. for non-Mercedes-driving moms?�Compare the two means! What does the comparison tell you about the link between Mercedes-driving and kid_score?
(Feel free to discuss with your neighbor.)
mean kid_score 90 | mean kid_score 90 |
mean kid_score 78 | mean kid_score 78 |
Mom finished high school
Mom�didn’t finish high school
Mercedes
990 women | 10 women |
10 women | 990 women |
Mom finished high school
Mom�didn’t finish high school
No Mercedes
Mercedes
No Mercedes
Aha!
20
Course eval (“indicative feedback”) open until �Sun 13 Oct 13th �Go to https://isa.epfl.ch now!
21
Quantifying uncertainty
22
Quantifying uncertainty
p-value: probability of estimating such an extreme coefficient if the true coefficient were zero�(= null hypothesis)
Aha!
23
Residuals and R2
Variance of�outcomes y
24
residual
Residuals and R2
Variance of�outcomes y
Aha!
25
Coefficient of determination: R2
R2 = 0.147
R2 = 0.865
26
Coefficient of determination: R2
27
Coefficient of determination: R2
R2 = 0.67 everywhere!
28
Assumptions made in regression modeling
29
Assumptions for regression modeling
30
Assumptions for regression modeling (2)
��But very flexible: we require linearity in predictors (not necessarily in raw inputs); predictors can be arbitrary functions of raw inputs, e.g.,�- logarithms, polynomials, reciprocals, … �- interactions (i.e., products) of multiple inputs�- discretization of raw inputs, coded as indicator variables
31
Assumptions for regression modeling (3)
less important�in practice
32
Transformations of predictors and outcomes
33
Transformations of predictors
34
Mean-centering of predictors
-100 -50 0 50
mean kid_score for moms with mean IQ: 86
0 50 100
mom_iq
kid_score
0 50 100
(hypothetical) mean kid_score for moms with IQ = 0: 26
0 50 100 150
35
After mean-centering of predictors, …
… you have a convenient interpretation of coefficients 𝛽j of main predictors (i.e., non-interaction predictors):
36
Standardization via z-scores
37
Logarithmic outcomes
38
Logarithmic outcomes: Interpreting coefficients
39
Going beyond linear regression for comparing means
40
Beyond linear regression:�generalized linear models
41
Beyond comparing means; or, A taste of causality: “Difference in differences”
42
Beyond comparing means; or, A taste of causality: “Difference in differences” (2)
43
a
b
c
d
Beyond comparing means; or, A taste of causality: “Difference in differences” (2)
44
a
b
c
d
What a treat!
A bonanza of causality:
Next lecture!
45
$#1t!, my banana is non-linear…
Summary
46
Feedback
47
Give us feedback on this lecture here: https://go.epfl.ch/ada2024-lec5-feedback
Credits
48
Bonus: Logarithmic outcomes and predictors
Interpretation of coefficient of logarithmic predictor:
49