1 of 53

Jiyong Park

Bryan School of Business and Economics

University of North Carolina at Greensboro

jiyong.park@uncg.edu

Session 18. Synthetic Control / Causal Discovery

1

Korea Summer Session on Causal Inference 2021

Korea Summer Session on Causal Inference 2021

Session Website: https://sites.google.com/view/causal-inference2021

Module 2. Machine Learning for Causal Inference

Toolkit for Causal Inference

: Synthetic Control / Causal Discovery

2 of 53

Session 18. Synthetic Control / Causal Discovery

2

Korea Summer Session on Causal Inference 2021

Korea Summer Session on Causal Inference 2021

Synthetic Control

: Mimicking the Counterfactual

3 of 53

Session 18. Synthetic Control / Causal Discovery

3

Korea Summer Session on Causal Inference 2021

Counterfactual Revisited

  • From the potential outcomes framework, causation is defined as the difference in potential outcomes after the treatment.
    • “What if the treatment was not applied?”
    • Causal effect = (Actual outcome for treated if treated) – (Potential outcome for treated if not treated)

Counterfactual

Counterfactual

Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

3

2

1

1

3

0

1

4

0

1

4 of 53

Session 18. Synthetic Control / Causal Discovery

4

Korea Summer Session on Causal Inference 2021

Counterfactual Revisited

  • From the potential outcomes framework, causation is defined as the difference in potential outcomes after the treatment.
    • “What if the treatment was not applied?”
    • Causal effect = (Actual outcome for treated if treated) – (Potential outcome for treated if not treated)

Counterfactual

Counterfactual

Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

3

1

ATET = 1

2

1

1

1

3

0

2

1

ATEC = 1

4

0

2

1

Ignorability ≡ Exchangeability ≡ Unconfoundedness ≡ Exogeneity

ATE

How? (PO) Ceteris Paribus

(SCM) Backdoor/Frontdoor Criterion

5 of 53

Session 18. Synthetic Control / Causal Discovery

5

Korea Summer Session on Causal Inference 2021

Recall the Research Design

The treatment is clearly defined, which allows to distinguish the treatment and control groups.

There are observations before and after the treatment.

Treatment Assignment w/o Randomization

Subjects select into the treatment.

Subjects are assigned for the treatment.

The treatment is assigned by an external, unexpected shock/event.

The treatment is assigned by an arbitrary threshold/cutoff.

No

Instrumental Variable

Yes

No

Self-Selection

Exogenous Shock

Discontinuity

Quasi-Experiment

The primary purpose is causal inference, and random assignment is feasible.

Matching

There are sufficient (matchable) observations and/or no information on the functional form between the treatment and the outcome.

Regression

Selection on Observables

Yes

DID + (Matching)

RD

Yes

No

Yes

No

Causal Diagram

Note that this might depend on the research context.

Yes

Randomized Controlled Trial

No

There are variables that predict the treatment, but not relate to the error term.

6 of 53

Session 18. Synthetic Control / Causal Discovery

6

Korea Summer Session on Causal Inference 2021

Importance of Panel Data for Causal Inference

  • Panel data structure allows us to observe the outcomes of the treatment group in the absence of treatment.

Treatment Group

After Period

Actual Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

0

0

1

1

1

1

3

2

1

0

0

0

1

1

1

1

3

0

0

0

1

0

1

0

1

4

0

0

0

0

0

1

0

1

7 of 53

Session 18. Synthetic Control / Causal Discovery

7

Korea Summer Session on Causal Inference 2021

Importance of Panel Data for Causal Inference

  • Panel data structure allows us to observe the outcomes of the treatment group in the absence of treatment.

Treatment Group

After Period

Actual Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

0

0

1

ATET = 1

1

1

1

3

1 + 0.5

2

1

0

0

0

1

1

1

1

0 + 0.5

3

0

0

0

1

0

1

0

1

4

0

0

0

0

0

1

0

1

Average increase by 0.5

Difference-in-Differences (DID)

Parallel trends assumption

For ATET, strict ignorabiility assumption is not necessary.

8 of 53

Session 18. Synthetic Control / Causal Discovery

8

Korea Summer Session on Causal Inference 2021

Basic Idea of Synthetic Controls

  • A combination of untreated units often provides a more appropriate counterfactual for the treated units.

Treatment Group

After Period

Actual Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

0

0

1

1

1

1

3

2

1

0

0

0

1

1

1

1

3

0

0

0

1

0

1

0

1

4

0

0

0

0

0

1

0

1

(2) Predicting the counterfactual (a.k.a. synthetic control)

Synthetic Controls (SC)

 

 

(1) Training a prediction model

Even parallel trends assumption may not be necessary.

9 of 53

Session 18. Synthetic Control / Causal Discovery

9

Korea Summer Session on Causal Inference 2021

Basic Idea of Synthetic Controls

  • A combination of untreated units often provides a more appropriate counterfactual for the treated units.

“The synthetic control approach developed by Abadie, Diamond, and Hainmueller (2010, 2014) and Abadie and Gardeazabal (2003) is arguably the most important innovation in the policy evaluation literature in the last 15 years. This method builds on difference-in-differences estimation, but uses systematically more attractive comparisons.” (Athey and Imbens 2017, p. 9)

Athey, S. and Imbens, G.W., 2017. The state of applied econometrics: Causality and policy evaluation. Journal of Economic Perspectives31(2), pp.3-32.

10 of 53

Session 18. Synthetic Control / Causal Discovery

10

Korea Summer Session on Causal Inference 2021

Case Study (1) Impact of California Anti-Tobacco Legislation

  • Example: California tobacco control program (Proposition 99) (Abadie et al. 2010)

Abadie, A., Diamond, A. and Hainmueller, J., 2010. Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association105(490), pp.493-505.

The treated unit (California) is not comparable to untreated units (other states), but their trends become parallel since the late 1970s.

11 of 53

Session 18. Synthetic Control / Causal Discovery

11

Korea Summer Session on Causal Inference 2021

Case Study (1) Impact of California Anti-Tobacco Legislation

  • Example: California tobacco control program (Proposition 99) (Abadie et al. 2010)

Abadie, A., Diamond, A. and Hainmueller, J., 2010. Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association105(490), pp.493-505.

12 of 53

Session 18. Synthetic Control / Causal Discovery

12

Korea Summer Session on Causal Inference 2021

Case Study (1) Impact of California Anti-Tobacco Legislation

  • Example: California tobacco control program (Proposition 99) (Abadie et al. 2010)

Abadie, A., Diamond, A. and Hainmueller, J., 2010. Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association105(490), pp.493-505.

13 of 53

Session 18. Synthetic Control / Causal Discovery

13

Korea Summer Session on Causal Inference 2021

Case Study (1) Impact of California Anti-Tobacco Legislation

  • Example: California tobacco control program (Proposition 99) (Abadie et al. 2010)

Abadie, A., Diamond, A. and Hainmueller, J., 2010. Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association105(490), pp.493-505.

Arkhangelsky, D., Athey, S., Hirshberg, D.A., Imbens, G.W. and Wager, S., 2019. Synthetic difference in differences (No. w25532). National Bureau of Economic Research.

14 of 53

Session 18. Synthetic Control / Causal Discovery

14

Korea Summer Session on Causal Inference 2021

Case Study (2) Impact of Reunification on West Germany

  • We are never able to observe the counterfactual on the reunification, but we can constitute a synthetic control using similar countries that best resembles the economic trajectory of West Germany.

Abadie, A., 2021. Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature59(2), pp.391-425.

Transparency of the counterfactual is one of the most attractive features of the synthetic control.

Donor Pool

15 of 53

Session 18. Synthetic Control / Causal Discovery

15

Korea Summer Session on Causal Inference 2021

How to Construct the Synthetic Control

  • Original method (Abadie and Gardeazabal 2003; Abadie et al. 2010)
    • Instead of predicting the outcome directly, it aims to choose weights of control units to minimize the difference in the pre-intervention values of predictors of the outcome (with a constraint that weights are nonnegative and sum is one).

Abadie, A. and Gardeazabal, J., 2003. The economic costs of conflict: A case study of the Basque Country. American Economic Review, 93(1), pp.113-132.

Abadie, A., Diamond, A. and Hainmueller, J., 2010. Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association105(490), pp.493-505.

Predictor for the treated unit

Weighted predictor for the untreated units

There could be multiple predictors that contribute differently to the synthetic control.

Treatment effect for the treated unit after intervention

Counterfactual computed using the synthetic control (weighted control units that best resemble the treated unit)

16 of 53

Session 18. Synthetic Control / Causal Discovery

16

Korea Summer Session on Causal Inference 2021

How to Construct the Synthetic Control

  • There are many methods to minimize the difference in the outcome of interest, or predictors of the outcome, between the treated unit and a combination of the untreated units in the pre-treatment period.

Doudchenko, N. and Imbens, G.W., 2016. Balancing, regression, difference-in-differences and synthetic control methods: A synthesis (No. w22791). National Bureau of Economic Research.

17 of 53

Session 18. Synthetic Control / Causal Discovery

17

Korea Summer Session on Causal Inference 2021

How to Construct the Synthetic Control

Basically, the synthetic control approach is the prediction problem.

“Like for the lasso, the goal of synthetic controls is out-of-sample prediction” (Abadie 2021, p. 408)

Abadie, A., 2021. Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature59(2), pp.391-425.

18 of 53

Session 18. Synthetic Control / Causal Discovery

18

Korea Summer Session on Causal Inference 2021

Sensitivity Tests for Synthetic Controls

  • Sensitivity tests to the choice of predictors / the choice of the donor pool
  • Placebo tests for backdating / untreated units

Abadie, A., 2021. Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature59(2), pp.391-425.

Amjad, M., Shah, D. and Shen, D., 2018. Robust synthetic control. Journal of Machine Learning Research, 19(1), pp.802-852.

Case Study (2) Impact of Reunification on West Germany

Case Study (1) Impact of California Anti-Tobacco Legislation

19 of 53

Session 18. Synthetic Control / Causal Discovery

19

Korea Summer Session on Causal Inference 2021

Sensitivity Tests for Synthetic Controls

  • A train-test split approach can also be applied to the synthetic control.
    • Train – Test – Treat – Compare (TTTC) process

Varian, H.R., 2016. Causal inference in economics and marketing. Proceedings of the National Academy of Sciences113(27), pp.7310-7315.

Pre-treatment period

20 of 53

Session 18. Synthetic Control / Causal Discovery

20

Korea Summer Session on Causal Inference 2021

What if There is No Control Group?

  • In the case of no control group, time-series forecasting models can be used to predict the counterfactual.
    • The time-series approach works well for modellable changes (e.g., seasonal effects, long-term trends, short-term fluctuations), though it is challenging to account for structural changes by other factors.
    • It is also called an interrupted time-series, or quasi-experimental time-series analysis.

Treatment Group

After Period

Actual Treatment

Potential Outcomes

Causal Effect

Subject i

1

1

0

0

1

1

1

1

3

2

1

0

0

0

1

1

1

1

Time-Series Approach

time-series forecasting

21 of 53

Session 18. Synthetic Control / Causal Discovery

21

Korea Summer Session on Causal Inference 2021

What if There is No Control Group?

  • In the case of no control group, time-series forecasting models can be used to predict the counterfactual.
    • The time-series approach works well for modellable changes (e.g., seasonal effects, long-term trends, short-term fluctuations), though it is challenging to account for structural changes by other factors.
    • It is also called an interrupted time-series, or quasi-experimental time series analysis.

Brodersen, K.H., Gallusser, F., Koehler, J., Remy, N. and Scott, S.L., 2015. Inferring Causal Impact Using Bayesian Structural Time-Series Models. The Annals of Applied Statistics, 9(1), pp.247-274.

22 of 53

Session 18. Synthetic Control / Causal Discovery

22

Korea Summer Session on Causal Inference 2021

Case Study (3) Uber’s Application of Synthetic Control

  • Why are A/B tests not feasible in some cases?

23 of 53

Session 18. Synthetic Control / Causal Discovery

23

Korea Summer Session on Causal Inference 2021

Case Study (3) Uber’s Application of Synthetic Control

  • First alternative to A/B test: Difference-in-differences across cities

24 of 53

Session 18. Synthetic Control / Causal Discovery

24

Korea Summer Session on Causal Inference 2021

Case Study (3) Uber’s Application of Synthetic Control

  • Second alternative to A/B test: Synthetic control across cities

25 of 53

Session 18. Synthetic Control / Causal Discovery

25

Korea Summer Session on Causal Inference 2021

Case Study (4) Causal Analysis of GS25 vs CU

  • Causal inference and causal mindset are important for data-driven, evidence-based reasoning.
    • Think of the following questions for yourself, and the preliminary answers will be provided later on Slack.

(1) GS25 를 둘러싼 이슈는 인과추론 문제인가?

(2) 인과적인 효과를 어떻게 정의하고 측정할 수 있을까?

(3) GS25 와 CU 를 비교하는 건 합당할까?

만약 아니라면, 대안은 무엇인가?

(4) 분석 결과를 어떻게 신뢰할 수 있을까?

26 of 53

Session 18. Synthetic Control / Causal Discovery

26

Korea Summer Session on Causal Inference 2021

Requirements for Synthetic Control

  • Contextual Requirements
    • Size of the Effect and Volatility of the Outcome
    • Availability of a Comparison Group
    • No Anticipation
    • No Interference
    • Convex Hull Condition
    • Time Horizon

Abadie, A., 2021. Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature59(2), pp.391-425.

  • Data Requirements
    • Aggregate Data on Predictors and Outcomes
    • Sufficient Pre-intervention Information
    • Sufficient Post-intervention Information

27 of 53

Session 18. Synthetic Control / Causal Discovery

27

Korea Summer Session on Causal Inference 2021

Recommended Reading and Watching

Abadie, A., 2021. Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature59(2), pp.391-425.

Pioneer of synthetic control approach

Guest Talk by Alberto Abadie - Synthetic Controls

(https://www.youtube.com/watch?v=nKzNp-qpE-I)

28 of 53

Session 18. Synthetic Control / Causal Discovery

28

Korea Summer Session on Causal Inference 2021

Korea Summer Session on Causal Inference 2021

Causal Discovery

: Identifying Causal Relationships from Data

29 of 53

Session 18. Synthetic Control / Causal Discovery

29

Korea Summer Session on Causal Inference 2021

Toward Knowledge Discovery

30 of 53

Session 18. Synthetic Control / Causal Discovery

30

Korea Summer Session on Causal Inference 2021

Toward Knowledge Discovery

Theory → Evidence (Data)

Evidence (Data) → Theory

31 of 53

Session 18. Synthetic Control / Causal Discovery

31

Korea Summer Session on Causal Inference 2021

Data Generation Process and Causal Discovery

  • Data Generation Process: Causal Graph → Data
  • Causal Discovery: Data → Causal Graph

Ma, S. and Statnikov, A., 2017. Methods for computational causal discovery in biomedicine. Behaviormetrika44(1), pp.165-191.

Causal Effect Identification and Estimation

32 of 53

Session 18. Synthetic Control / Causal Discovery

32

Korea Summer Session on Causal Inference 2021

Overall Structure of Causal Discovery

(2) What is the Markov equivalence class?

(1) What assumptions are required?

+ Acyclicity for DAG

(3) How to learn causal structures?

(4) How to test conditional independence?

33 of 53

Session 18. Synthetic Control / Causal Discovery

33

Korea Summer Session on Causal Inference 2021

Causal Markov and Faithfulness Assumptions

  • Causal Markov Assumption: A node is dependent only on its descendants in the graph. In other words, a node is independent on other variables, conditional on its causes.
  • Faithfulness Assumption: Nodes that are causally connected in a particular way in the graph are probabilistically dependent.

 

 

 

 

Causal Markov Assumption

Causal Markov Assumption + Faithfulness Assumption

34 of 53

Session 18. Synthetic Control / Causal Discovery

34

Korea Summer Session on Causal Inference 2021

Violation of Faithfulness Assumption

Source: Brady Neal’s lecture notes

  • Distinct causal paths that have opposite effects could cancel out each other.

35 of 53

Session 18. Synthetic Control / Causal Discovery

35

Korea Summer Session on Causal Inference 2021

Recall the Conditional (In-)Dependence (Association)

    • As they are, mediators and confounders do establish an association between nodes, but colliders do not.

Mediator (Chain)

Confounder (Fork)

Collider (Immorality)

information flow

information flow

information flow

X

M

Y

X

C

Y

X

Z

Y

X and Y are d-connected.

X and Y are d-connected.

X and Y are d-separated.

36 of 53

Session 18. Synthetic Control / Causal Discovery

36

Korea Summer Session on Causal Inference 2021

Recall the Conditional (In-)Dependence (Association)

    • After conditioning on them (controlling for, or blocking), mediators and confounders does not establish an association between nodes, but colliders do.

Mediator (Chain)

Confounder (Fork)

Collider (Immorality)

M

information flow

C

information flow

Z

information flow

X

Y

X

Y

X

Y

X and Y are d-separated.

To estimate the direct causal effect of X on Y, mediators should be blocked.

X and Y are d-separated.

To estimate the direct or indirect causal effect of X on Y, confounders should be blocked.

X and Y are d-connected.

To estimate the direct or indirect causal effect of X on Y, colliders should NOT be blocked.

37 of 53

Session 18. Synthetic Control / Causal Discovery

37

Korea Summer Session on Causal Inference 2021

Markov Equivalence Class

  • Markov Equivalence Class: A set of DAGs that encode the same set of conditional independencies.

Eberhardt, F., 2016. Introduction to the foundations of causal discovery. International Journal of Data Science and Analytics2(3), pp.81-91.

The “V” structures (colliders, immorality) play a critical role as it has only one structure for the same class.

38 of 53

Session 18. Synthetic Control / Causal Discovery

38

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms

  • Causal discovery algorithms can be classified into two types: (i) constraint-based and (ii) score-based.

Constraint-based algorithms are based on conditional independence constraints.

Score-based algorithms generate a number of candidate causal graphs, assign a score to each, and select a final graph based on the scores.

PC Algorithm

(Peter Spirtes and Clark Glymour)

FCI Algorithm

(Fast Causal Inference)

Assuming no unobserved confounders

Assuming unobserved confounders

GES Algorithm

(Greedy Equivalence Search)

39 of 53

Session 18. Synthetic Control / Causal Discovery

39

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms (1) PC Algorithm

  • Step 1. Start with a complete undirected graph.
  • Step 2. Eliminate edges between variables that are unconditionally independent.
  • Step 3. For each pair of variables having an edge between them, eliminate the edge if they are independent, conditional on a subset of variables with edges to them (increasing the size of subsets 1 to n).
  • Step 4. Identify a “V” structure (collider, immorality) and orient edges.
  • Step 5. Orient the remaining edges not to be a collider (i.e., orientation propagation).

Ground Truth

Skeleton

40 of 53

Session 18. Synthetic Control / Causal Discovery

40

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms (2) FCI Algorithm

  • FCI algorithm is similar to PC algorithm, but it further assumes that there could be an unmeasured confounder between nodes, except the “Y” structures.

Note that causal discovery algorithms do not necessarily provide complete causal information

PC algorithm

FCI algorithm

41 of 53

Session 18. Synthetic Control / Causal Discovery

41

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms (2) FCI Algorithm

  • FCI algorithm is similar to PC algorithm, but it further assumes that there could be an unmeasured confounder between nodes, except the “Y” structures.

Ground Truth

Unmeasured confounder

Graph after removing conditional independence

Graph after orienting the “V” structures

Can be an arrow head or tail

When will it become an arrow tail (i.e., causal effect of X on Y)?

42 of 53

Session 18. Synthetic Control / Causal Discovery

42

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms (2) FCI Algorithm

  • FCI algorithm is similar to PC algorithm, but it further assumes that there could be an unmeasured confounder between nodes, except the “Y” structures.

Ground Truth

Unmeasured confounder

Graph after removing conditional independence

Graph after orienting the “V” structures

A

B

A

B

A

B

If there is an unmeasured confounder between X and Y, A (or B) and Y cannot be independent conditional on X.

43 of 53

Session 18. Synthetic Control / Causal Discovery

43

Korea Summer Session on Causal Inference 2021

Causal Discovery Algorithms (3) GES Algorithm

  • Step 1. Start with an empty graph containing no edges.
  • Step 2. Greedily add edges (dependencies) one at a time in the orientation that maximize some fit score, such as Bayesian Information Score (BIC) (the lower, the better fit).
  • Step 3. Map the resulting model to the corresponding Markov equivalence class.
  • Step 4. Continue Steps 2 and 3 until the score can no longer be improved.
  • Step 5. Remove edges one at a time as long as it maximizes the score (e.g., decreases the BIC).
  • Step 6. Continue Step 5 until no further edges can be removed.

44 of 53

Session 18. Synthetic Control / Causal Discovery

44

Korea Summer Session on Causal Inference 2021

Summary of Causal Discovery Algorithms

LiNGAM: Linear, non-gaussian, acyclic model

PNL: post-non-linear causal model

ANM: non-linear additive noise model

Glymour, C., Zhang, K. and Spirtes, P., 2019. Review of causal discovery methods based on graphical models. Frontiers in Genetics10, p.524.

FCM (functional causal model)

45 of 53

Session 18. Synthetic Control / Causal Discovery

45

Korea Summer Session on Causal Inference 2021

Recommended Reading

Glymour, C., Zhang, K. and Spirtes, P., 2019. Review of causal discovery methods based on graphical models. Frontiers in Genetics10, p.524.

46 of 53

Session 18. Synthetic Control / Causal Discovery

46

Korea Summer Session on Causal Inference 2021

Conditional Independence Tests

  • Conditional independence tests depend on the distributions of variables in Bayesian networks (causal graphs).

1) Discrete Bayesian networks (categorical variables)

2) Discrete Bayesian networks (ordered factors)

3) Gaussian Bayesian networks (continuous normal variables)

4) Non-Gaussian Bayesian networks (continuous variables)

47 of 53

Session 18. Synthetic Control / Causal Discovery

47

Korea Summer Session on Causal Inference 2021

Practical Guidance for Causal Discovery

Practical causal analysis is not a matter of pressing a few buttons. There are multiple algorithms available, many of them are poorly tested, some of them are poor implementations of good algorithms, some of them are just plain poor algorithms, all of them have choices of parameters, and all of them have conditions

on the data distributions and other assumptions under which they will be informative rather than misleading.” (Glymour et al. 2019, p. 11)

Glymour, C., Zhang, K. and Spirtes, P., 2019. Review of causal discovery methods based on graphical models. Frontiers in Genetics10, p.524.

48 of 53

Session 18. Synthetic Control / Causal Discovery

48

Korea Summer Session on Causal Inference 2021

Practical Guidance for Causal Discovery

  • Causal discovery algorithms work asymptotically (i.e., with a large volume of data).
  • Distributions of the variables play a critical role in conditional independence tests.
  • Domain knowledge may help the causal discovery.

Shen, X., Ma, S., Vemuri, P. and Simon, G., 2020. Challenges and opportunities with causal discovery algorithms: application to Alzheimer’s pathophysiology. Scientific Reports10(1), pp.1-12.

Without background knowledge

With trivial background knowledge

(demographic variables cannot be caused by others)

49 of 53

Session 18. Synthetic Control / Causal Discovery

49

Korea Summer Session on Causal Inference 2021

Wrap-Up

  • The synthetic control method aims to constitute a combination of the untreated units (donor pool) that resembles the outcome of the treated unit in the absence of the treatment.
    • From the potential outcomes framework, the synthetic controls share the spirit with the difference-in-differences.
    • However, the synthetic control approach is basically the prediction problem, where ML can be in play.

  • Causal discovery aims to identify causal relationships (causal graph) from large quantities of data through computational methods.
    • Causal discovery algorithms can be classified into two types: (i) constraint-based and (ii) score-based.
    • Given that it depends on conditional independence, the distributions of the variables need to be carefully considered.
    • For effective causal discovery, domain knowledge may be helpful.

50 of 53

Session 18. Synthetic Control / Causal Discovery

50

Korea Summer Session on Causal Inference 2021

Bridging the Social Science and Computer Science

  • One of the most cited resistance among social scientists against the structural causal models stems from the inability to validate causal structures that represent phenomena of interest.

“In general it is easy to come up with arguments for the presence of links: as anyone who has attended an empirical economics seminar knows, the difficult part is coming up with an argument for the absence of such effects that convinces the audience.” (Imbens 2020, p. 1140)

Imbens, G.W., 2020. Potential outcome and directed acyclic graph approaches to causality: Relevance for empirical practice in economics. Journal of Economic Literature58(4), pp.1129-79.

Can the causal discovery be a remedy for this concern?

(maybe not as of now, but will be the case in the near future)

51 of 53

Session 18. Synthetic Control / Causal Discovery

51

Korea Summer Session on Causal Inference 2021

Korea Summer Session on Causal Inference 2021

Final Remarks

52 of 53

Session 18. Synthetic Control / Causal Discovery

52

Korea Summer Session on Causal Inference 2021

Korea Summer Session on Causal Inference

53 of 53

End of Document

Session 18. Synthetic Control / Causal Discovery

53

Korea Summer Session on Causal Inference 2021