Maria Glymour
Department of Epidemiology & Biostatistics
University of California, San Francisco
Introduction to Using Directed Acyclic Graphs in Dementia Research
Organization
Why bother learning to use DAGs?
Inferring Causation From Association
Statistical association between two variables X and Y may be due to:
How we can use this
Organization
Causal Directed Acyclic Graphs
X
Y
U
A
A
Y
U
X
Show your assumptions about the causal relationships among X, Y, and possible covariates in a causal diagram:
A
Y
B
X
E
Terminology
Colliders vs Non-Colliders
B
C
A
Colliders: common effects
Non-Colliders:
common causes (=confounders)
Or mediators
B
C
A
B
C
A
D-separation
Places where the above doesn’t hold: (1) if there are perfectly offsetting balance between two paths or (2) the variables violate the consistency assumption, e.g., when we think a variable has a specific value (e.g.,”amyloid positive” or “high BMI”) there are actually different versions of that variable, and those different versions have different causes and effects (eg based on measuring amyloid fibrils vs plaques or central vs peripheral adiposity).
D-separation
Places where the above doesn’t hold: (1) if there are perfectly offsetting balance between two paths or (2) the variables violate the consistency assumption, e.g., when we think a variable has a specific value (e.g.,”amyloid positive” or “high BMI”) there are actually different versions of that variable, and those different versions have different causes and effects (eg based on measuring amyloid fibrils vs plaques or central vs peripheral adiposity).
In small print because you shouldn’t worry about it now. But I hope you/we come back to this someday.
D-separation: intuition
Example Causal Diagrams
X
Y
U
A
X
Y
B
E
A
X
Y
A3
A1
A2
A
Y
B
X
E
3 Approaches to Demonstrating Causation
X
Y
U
X
Y
U
Z
X
Y
U
M1
M2
Identifying causal effects: the Back Door Criterion
X
Y
Z
DAGs Help With Research Design
Organization
Confounding
Education
Dementia
Depression
Confounding
Education
Income
Depression
Dementia
Unreliable Measures
Depression
Dementia
CESD1
ε1
Unreliable measures of a confounder
Education
Income
Depression
Dementia
High-school completion
Causal Structure for an RCT, a quasi-experiment, an instrumental variable
Education
Dementia
Random Assignment/Instrument
Exposure
Randomized Trial: Randomly Assign Therapy to Reduce Depression
Education
Depression
Dementia
Random Assignment
Anti-Dep Drug
Randomized Trial: Randomly Assign Therapy to Reduce Depression
Education
Depression
Dementia
Random Assignment
Anti-Dep Drug
Things that go wrong with trials
Education
Depression
Dementia
Random Assignment
Anti-Dep Drug
Loss to Follow-Up in Trials
Participation in outcome assessment
Dementia Incidence
Randomization to intervention
Depression
Physical
activity
Loss to Follow-Up in Trials
Participation in outcome assessment
Dementia Incidence
Randomization to intervention
Cognition Declined
Physical
activity
Selection/Survival Bias as a Threat to Internal Validity in Observational Studies
Survival to age 65
Dementia
Education
Some gene
Selection Bias
Survival to age 65
Dementia
Education
Some gene
Here, we assume education has no effect on dementia.
Would it be statistically associated with dementia among people in your cohort of 65+ year olds?
Selection Bias Compromises Internal Validity
Heart Failure
Mortality
Obesity (pre-HF)
Some gene
Obesity (post-HF)
Representing Direct Effects and Challenges in Estimating Direct Effects
M
Y
X
U
- Conventional approaches to estimating direct effects condition/adjust for the mediator (M), but if there is a confounder of the M-Y association, then X and Y would be associated conditional on M even if there was no indirect effect.
M
Y
X
U
Why would you condition on a collider?
Some “colliders” are not optional:
Some objections to DAGs
Some objections to DAGs
Some objections to DAGs
Quizlet
end
Can you quantify the bias from collider stratification?
Unreliable Measures
Depression
Mortality
CESD1
ε1
Unemployment
Unreliable Measures in Analyses of Change
X
C1
Change in C1
Y1
ε1
U
Y2- Y1
DAGs for missing data
Final quizlet: Estrogen Therapy and Bone Mineral Density in African-American and Caucasian Women
Controlling for body size and composition, the authors examined the association between estrogen therapy and bone mineral density in older African-American and Caucasian women. In 1992–1998, 443 African-American and 989 Caucasian women aged 45–87 years were assessed for medication use, laboratory variables, behavioral characteristics, and bone mineral density. The mean age was 61.3 (95% confidence interval: 60.3, 62.3) years in African Americans and 71.0 (95% confidence interval: 70.4, 71.7) years in Caucasians (P < 0.001). All measures of body size and composition were significantly greater in the African-American women compared with Caucasian women (P < 0.001). As expected, African Americans had significantly higher bone mineral density at all 4 sites independent of age, weight, body composition, estrogen use, and lifestyle factors. Although Caucasians were significantly more likely to currently use estrogen (48.9% vs. 33.9%; P < 0.001), African Americans not using estrogen had significantly higher bone mineral density at all sites except the spine than Caucasians who were using estrogen. Regression models including age and lean mass explained the most variation in bone mineral density (R2 range = 0.13–0.37). Results suggest that higher levels of bone mineral density in African-American women were not due to estrogen use.
Getting Rich on Collider Bias
Getting Rich on Collider Bias
Door with the Prize
Door You Choose First
Door Monty Hall Opens
Organization
Demonstrating Causation
X
Y
U
X
Y
U
Z
X
Y
U
M1
M2
Random assignment🡪X🡪Y
Intuition for IV for health researchers: RCT
Intuition for IV
Instrument Assumptions
Randomization is an instrument for the influence of X on Y if:
Examples of Instrumental Variables
X
Y
U
Z
Examples of Instrumental Variables
X
Y
U
Z
Caveats to IV
X
Y
U
Z
IV versus Covariate Adjustment Approaches
IV
COVARIATE ADJUSTMENT
IV versus Covariate Adjustment Approaches
IV
COVARIATE ADJUSTMENT
In different settings, one or the other set of assumptions may be more plausible. Often evidence is most convincing if combined across designs, because assumptions required are so different.
Limitations and controversies
Pitfalls due to consistency violations: it may be a big deal in dementia research
Structures of Aβ monomer, fibril and oligomers
Guo-fang Chen1, * et al. Acta Pharmacol Sin 2017;
38: 1205–1235; doi: 10.1038/aps.2017.28
Counterfactuals do not always satisfy intuitions about “causation”
Counterfactuals do not always satisfy intuitions about “causation”
Counterfactuals do not always satisfy intuitions about “causation”
Often unclear how to draw the DAG: Difference in difference
Year * Group
Outcome
Eligible Group
Year of Policy Change
E(Y)=b0+b1*Group+b2*Year+b3*Group*Year
Year * Group
Outcome
Eligible Group
Year of Policy Change
M
U1
U2
Conditional on group and year of change, there are no confounders of the association between the interaction and the outcome.
Often unclear how to draw the DAG: Change scores vs change
X
C1
Change in C1
Y1
ε1
Y2- Y1
X
Change in Y
Y1
U
Y2
If you draw the DAG as on the left, adjusting for Y1 induces collider bias and a spurious association between X and change score.
If you draw the DAG as on the right, adjusting for Y1 is necessary to block the direct effect of Y1 on Y2.
I chose the DAG on the left because I consider change in C a biological phenomenon, indirectly measured with test scores. Changing the test scores would not change your brain, but changing your brain would change your test scores.
Causal discovery
Conclusions
Summary
In recent decades, there has been an outpouring of methods explicitly focusing on supporting causal inferences from observational (non-randomized) data. Much of this work is based on a counterfactual framework for causation and relies on Directed Acyclic Graphs (DAGs) as a core tool to represent causal hypotheses. In this talk, I will introduce the counterfactual account of causation, illustrate the distinction between counterfactual definitions of causal parameters and statistical comparisons used to estimate those parameters. I will introduce DAGs and the d-separation rule allowing us to link causal structures represented in DAGs to statistical associations. Many common biases encountered in epidemiologic studies are fruitfully represented in DAGs, and we will review a few examples. Finally, time permitting, we will introduce the use of instrumental variables, linking to a causal framework. Illustrative examples will be drawn from my research on dementia and stroke.
64
64
Conditioning on a Collider
If two variables are statistically independent, but have a common effect, then, within levels of this effect, they will be statistically dependent.
Really.
Usually.
65
65
A collider anecdote
Some tall people are fast, and some are slow.
Some short people are fast, and some are slow.
Knowing that somebody in the general population is short does not give you information about whether they are fast or slow.
NBA ball players must be either very tall, or very fast.
If you know an NBA ball player is short… what do you know about his speed?
66
66
A collider anecdote
Some tall people are fast, and some are slow.
Some short people are fast, and some are slow.
Knowing that somebody in the general population is short does not give you information about whether they are fast or slow.
NBA ball players must be either very tall, or very fast.
If you know an NBA ball player is short… what do you know about his speed?
67
NBA player
Speed
Height
67
A collider anecdote
I throw a party, and I only invite people who are either very rich or very funny.
You come to my party (you are very funny) and get stuck talking to the most boring person you have ever met.
Is he rich?
68
68
A collider anecdote
I throw a party, and I only invite people who are either very rich or very funny.
You come to my party (you are very funny) and get stuck talking to the most boring person you have ever met.
Is he rich?
69
Party invitation
Funniness
Wealth
69
A Collider Simulation: Please try this at home.
Make up some rules following the structure of the DAG, such that X and Y are independent causes of Z, then assess association between X and Y after conditioning on Z:
70
70
X and Y are independent; X and Z are positively associated
71
Scatter X , Y
Scatter X , Z
Independent
Positively Associated
71
X is inversely associated with Y, conditional on Z
72
Scatter X , residual (Y|Z)
Inversely Associated
72
Quizlet
73
X4
X2
X3
X1
X5
X4
X2
X3
X1
X5
If you know that one of these three causal structures generated the data, and you have perfect measures of all 5 variables, can you tell which one is correct?
X1
X4
X2
X3
X1
X5
U
73
Quizlet
74
X4
X2
X3
X1
X5
X4
X2
X3
X1
X5
X1, X2, and X3 predict X5, but are independent of X5 conditional on X4.
X4
X2
X3
X1
X5
U
X1, X2, and X3 are independent of X5, unless you condition on X4, when they become associated with X5
X1 and X3 are independent of X5, unless you condition on X2 or X4, when they become associated with X5
74
Organization
75
75