1 of 74

Causal inference and the design of impact evaluations in international development and humanitarian settings

Dr. Wael Moussa

Scientist, FHI 360

16/10/2019

1

2 of 74

Research Design

  • The design stage is usually the first step in any research

  • This is where you create your plan of attack and address the following questions:
    • What is the causal relationship of interest?
    • What is the ideal scenario to test this relationship?
    • How do you intend on identifying the hypothesized relationship? (identification strategy)
    • What analytic tools are you going to use?

2

3 of 74

What is the causal relationship?

  • Descriptive and correlational research provides a lot of value, informs researchers and policy makers of the current state of the world

  • However, being able to inform questions about cause and effect are more valuable in that they are actionable
    • i.e. a correlation is good to know, a causal relationship is better

3

4 of 74

Causal Inference

  • Identifying a causal relationship helps make predictions about the consequences of …
    • changes in certain circumstances
    • shocks to the system
    • implementation of policies

  • Enables construction of an alternative…a counterfactual
    • A scenario in which an actual “event” never happened, or vice versa
    • Compare with actuality

  • Allows you to measure the impact of this “event” in isolation of anything else happening

  • But, it is not always a simple task

4

5 of 74

Causal Inference

  •  

5

6 of 74

Causal Inference

  • In other words, we are constructing a counterfactual observation
    • A counterfactual is something of a “what if” scenario
    • What would have happened to the outcome of interest if the Treatment was never administered

  • e.g. You are interested in examining the impacts of harboring a large population of refugees in your country,
    • The counterfactual would be the state of your country in a world where the refugees were able to remain in their homes

6

7 of 74

The ideal experiment

  • What ideal circumstances will allow me to capture the causal effects of X on Y?

  • Example
    • Imagine it’s 1986 and you were offered to buy stock in this unknown company called Microsoft (they did something with computers)
    • You decided to stuff your money under your mattress
    • What was the effect of that decision? How would you evaluate it?
      • Get a time machine and change your decision in 1986
      • Compare present life (mattress) to present life (Microsoft)

7

8 of 74

The ideal experiment

  • Obviously, the ideal experiment is not possible
    • You cannot observe an individual/subject experiencing two different states of the world, simultaneously

  • Still, it is worth thinking about these solutions
    • Helps visualize and formulate a solution to your research problem
    • Helps formulate the causal question more precisely
    • Helps get a better understanding of the appropriate comparison
    • The mechanism of the experiment allows you to think of which factors to manipulate and which to hold constant

  • Sometimes, thinking along these lines can lead to very creative research

8

9 of 74

Example

  • Social scientists have always been interested in topics of race and sex
    • This has spurred a vast literature on discrimination “are certain people treated differently because they are perceived to be black or white/female or male?”

  • Bertrand and Mullainathan (2004) wanted to investigate whether discrimination truly existed in the labor market
    • Constructed a number of CV/resumes and sent out fake job applications to a number of employers in Boston and Chicago
    • The experiment randomized the assignment of African-American sounding names versus white sounding names to the same resumes

9

10 of 74

Example (cont’d)

  • The title of the paper is “Are Emily and Greg more employable than Lakisha and Jamal?”

  • There answer was yes
    • White names were more likely to receive callbacks than African-American names with similar resumes
    • Callbacks were more responsive to resume quality for white names than for African-American ones
    • The findings were consistent across occupation, industry, and firm size

10

11 of 74

Constructing the counterfactual

  •  

11

12 of 74

Constructing the counterfactual

  •  

12

13 of 74

Constructing the counterfactual

  • The two groups of individuals/subjects should be statistically indistinguishable from each other, but differ only in their receipt of the treatment

  • IF we can prove that the two groups are similar except for their treatment status
  • THEN any observed differences in the outcome can be confidently attributed to the treatment

13

14 of 74

Constructing the counterfactual

  •  

14

15 of 74

Constructing the counterfactual

  • What makes a good constructed counterfactual?
  • The two groups should, on average, be the same
    • We are simply comparing the overall distributions of the two groups

  1. The two groups should react to the same treatment in the same way
    • You want to be sure that the counterfactual you created is accurate
    • The control group would equally benefit from the treatment as the treatment group

  • The two groups must not be exposed to other treatments differentially
    • This means you want to make sure there are no confounders in your study
    • e.g. we are studying the impact of bonus cash on the purchase of candy, but the treatment group were also provided with more trips to the candy store

15

16 of 74

Constructing the counterfactual

  •  

16

bias

17 of 74

Constructing the counterfactual

    • Is it possible that even after ensuring the two groups look similar, the groups may still be different (unobservable factors are critical)

    • Let’s imagine you have an on-the-job training that requires employees to stay an hour after work
      • Enrollment is voluntary
      • Assume the intervention group and the non-intervention group are similar in terms of demographics, SES, health, etc. but are they really similar though?
      • It could be that employees who volunteered have more “drive” than those who did not want to volunteer
      • Differences in their outputs are now a combination of the treatment and differences in their level of “drive”

17

18 of 74

What is your identification strategy?

There are three broad categories that most causal research designs fall into

  1. Experimental: think randomized controlled trials (RCTs)
  2. Natural experiment: the experiment is induced naturally, outside of the researcher’s control
  3. Quasi-experimental: no random assignment, but you create an artificial experiment with ex-post data

  • If none of the above are possible, then do an observational study
    • Identify the target population, measure their outcomes, and changes in their observable attributes
    • Regression modeling (OLS, Fixed effects, HLM, SEM, etc.)
    • Provide descriptive and correlational analyses, but no causality!

18

19 of 74

Experiments

  • The researcher is in control of assigning the treatment (best assignment is random assignment)—“randomized” and also in control of the environment—“controlled

  • Makes it much easier to isolate the impact of the treatment from any outside factors

  • Example: Negative Income Tax 1960s and 1970s (New Jersey)
    • The research study wanted to investigate the effects of social welfare
    • Government subsidy if household income falls below a threshold
    • The social experiment randomized the threshold for different families
    • Found that families receiving higher subsidies were substituting for earned income

19

20 of 74

Natural Experiments

  • The researcher is not in control of assigning the treatment

  • An experiment-like environment formed as a natural occurrence

  • Even though it is not an artificial experiment, it still holds the same properties as one

  • Example: Angrist and Evans (1998)
    • Wanted to study the relationship between children and parents’ labor supply
    • The used the sex of the first two births as a natural experiment (parents with same-sex siblings are more likely to have a third child)
    • An additional child causes the labor supply of parents to decrease, more so for the mother than the father
    • Family income also declines

20

21 of 74

Quasi-experiments

  • What is done is done; the researcher shows up after the fact and still needs to measure the causal effect of the treatment
  • Treatments are non-randomly assigned, but there are ways around this to get at causality
  • Constructing an artificial sample that resembles an experiment
  • Example: Lechner (1999)
    • Effect of vocational training on earnings and unemployment in post-unification East Germany
    • Training was not assigned randomly—training was during business hours and was very demanding
    • People self-selected into the training—individual were more educated, more likely to already be working in scientific fields
    • Matched individuals who received training to similar individuals how did not
    • Effects on earnings were positive

  • There are drawbacks though
    • A lot of times, you cannot control for unobservable traits that may influence the outcome

21

22 of 74

Methods

22

23 of 74

Linear Regression

  • Linear regression is the most popular method in empirical research
  • Intuitive interpretation of often complex statistical relationships
  • More robust than simple difference in means, cross-tabulations, and correlations
  • It is a flexible tool in that it enables you to form conditional expectations
    • Allows you to make ceteris paribus statements with your data (all else being equal)
    • e.g. all else being equal, a 10% increase in neighborhood robberies lowers housing values by 2%

23

24 of 74

Conditional Expectation Function

  •  

24

25 of 74

CEF Example

25

 

26 of 74

CEF and Linear Regression

  • The essence of any research is to uncover the CEF

  • We can use linear regression to approximate the CEF or provide us with insight into what the CEF looks like

  • Unlike the example in the previous slide, linear regression will draw a straight line that minimizes the sum of squared errors

26

27 of 74

CEF and Linear Regression Example

27

28 of 74

Linear Regression

  •  

28

29 of 74

Linear regression and Causality

  •  

29

30 of 74

Fixed Effects – Panel Data

  •  

30

31 of 74

Fixed Effects – Panel Data

31

Without Fixed Effects

32 of 74

Fixed Effects – Panel Data

32

With Fixed Effects

33 of 74

Fixed Effects – Panel Data

  •  

33

34 of 74

Fixed Effects – Panel Data

  • Example:
    • Refugee households receive cash transfers in 2018 but did not receive anything in 2017
    • If I follow the same households between 2017 and 2018 (longitudinal sample)
    • Then I can measure the changes in their outcomes when they received the cash transfer

    • Problem—are the changes in nutritional outcomes strictly due to the receipt of the cash transfer?
    • What if households received a training that increased their knowledge of nutrition between 2017 and 2018?

  • We still need a valid counterfactual!

34

35 of 74

How do I make sure I have causality?

  •  

35

Observed

characteristics

Unobserved

characteristics

36 of 74

Randomized Controlled Trials (RCTs)

  •  

36

0

0

 

37 of 74

Randomized Controlled Trials (RCTs)

  • Sometimes randomization at the person level is not always possible
    • You can randomize at a higher cluster level (Cluster-Randomized Trials)
    • Example: Teacher professional development intervention can only be assigned at the school level—although the purpose is to improve student learning
    • There is no consequence to your design in terms of producing causal effects
    • However, because the assignment is at the school level, you must cluster your standard errors at the school level as well!

  • Remember, the higher the level of clustering, the less power you have
    • Counter that by getting more clusters (more schools, districts, villages, etc.)

37

38 of 74

Non-Random Assignment

  • Can I still estimate the causal relationship when the treatment is assigned non-randomly?

  • Yes, you can!
    • Difference-in-differences
    • Regression discontinuity
    • Propensity score matching
    • Other (not within the scope of this seminar)

  • However, certain conditions apply

38

39 of 74

Difference-in-Differences

  • The difference-in-differences (DD) is another approach that enables us to estimate the causal effect from a treatment with possibly non-random assignment

  • The DD approach is a way to estimate a counterfactual that relies on timing and group differences that can lead to causal inference

  • IF you can satisfy these conditions:
    • Group composition is stable over time
    • Parallel trends in pre-treatment outcomes for the treatment and control groups
    • Common shocks for the treatment and control groups in the post-treatment period

39

40 of 74

Difference-in-Differences

  • The DD approach has a very intuitive appeal
    • DD does not require baseline balance in the outcomes, i.e. does not require randomization to produce causal effects

    • First, it computes the change in the outcome (y) between the pre- and post-treatment period for the treatment group – this is the first difference

    • Second, it computes the change in the outcome between the pre- and post-treatment period for the control group – this is the second difference (the counterfactual)

    • The final step takes the difference between the two differences to estimate the treatment effect

40

41 of 74

Difference-in-Differences

  • The basic structure of a difference-in-differences (diff-in-diff for short) includes 2 groups and 2 time periods
    • Treatment group: group that received a treatment or intervention
    • Control group: group that did not receive any treatment or intervention

    • Pre-treatment period: the time-period prior to anyone receiving the treatment
    • Post-treatment period: the time-period after the treatment group received the treatment/intervention

  • Again, this is just the basic structure, the DD framework can be generalized to multiple treatments and staggered treatment periods

41

42 of 74

Difference-in-Differences

42

Pre-Treatment

Post-Treatment

Treatment Group

A – No Treatment

B – Treatment

Control Group

C – No Treatment

D – No Treatment

Diff-in-diff = (BA) – (DC)

43 of 74

Difference-in-Differences

  •  

43

44 of 74

Difference-in-Differences

  •  

44

45 of 74

Difference-in-Differences

  •  

45

46 of 74

Difference-in-Differences: Assumptions

  • Common shocks in the post-treatment period
    • This is equivalent to making sure there are no different interventions or events coinciding with the treatment for the treatment and control groups
    • Example:
      • Treatment and control group receive no cash transfers in period 1
      • Treatment occurs in period 2, but the treatment group also receives a “personal finance” training in period 2
      • Control group experiences no changes in period 2

  • Composition of treatment and control group should be stable from pre- to post-treatment periods

  • Parallel trends of outcomes for treatment and control groups in the pre-treatment period

46

47 of 74

Difference-in-Differences

Conclusion: Control group is appropriate

47

Post-Treatment

Treatment

Group

Treatment Effect

Treatment

Time

Pre-treatment trends

are parallel

Counterfactual

Common shocks

48 of 74

Difference-in-Differences

Conclusion: Control group is inappropriate

48

Post-Treatment

Correct Treatment Effect

Bias

Not Parallel

Estimated Counterfactual

Estimated Treatment Effect

Correct Counterfactual

49 of 74

Diff-in-Diff Example

  • Fallah, Krafft, Wahba (2018)—Economic Research Forum
  • Evaluation of the impact of the influx of Syrian refugees on employment in Jordan

  • Strategy:
    • Treatment group = regions of Jordan with high % of Syrian refugees
    • Control group = regions with stable and low & of Syrian refugees
    • Treatment onset = 2010; data from 2004 – 2016
    • Retrospective data using recall labor market information to determine trends

  • Finding: No significant pre-post differences for T group relative to C group

49

50 of 74

Diff-in-Diff in Program Impact Evaluations

  • For most, if not all, international development programs, getting enough data to determine a pre-treatment trend is not possible

  • If you can randomize the treatment, you can use the DD framework to show that there is balance along observables
  • Here, you can safely assume the pre-trends are parallel (since it’s unobservable)

  • Running a DD equation will produce similar results as a simple difference in means in the post-period

50

51 of 74

Regression Discontinuity

  •  

51

52 of 74

Regression Discontinuity

  • Things to keep in mind with an RD is that the treatment is not random; it IS determined by another variable

  • The RD design compares the outcomes for people who lie just below and just above the discontinuity

  • The idea is that these two groups of people are not that different, and the only thing that separates them is their location relative to the cutoff
    • If this is true, then the assignment of the treatment around the cutoff is as good as random

  • As such, the treatment effect is local to the area around the discontinuity (not generalizable beyond cutoff) 🡪 Produces a Local Average Treatment Effect (LATE)

52

53 of 74

Regression Discontinuity

  • For an illustrative example, let’s take a look at kindergarten eligibility in NYC

  • A child is eligible to enroll in kindergarten if his/her 5th birthday falls on or before Dec 31 of that year (otherwise, wait one more year)

  • Couple this with the compulsory attendance law in NYC, which states that a student may choose to drop out of school on his/her 17th birthday

53

54 of 74

Regression Discontinuity

  • These rules create a discontinuity in the duration of compulsory schooling before a student can choose whether or not to remain in school
    • Children whose birthday is in December start when they are just under 5 years old
    • Children whose birthday is in January start when they are just under 6 years old
    • This means that the December group has to stay in school for 12 years (17-5=12) before they can drop out
    • The January group has to stay in school for 11 years before they can drop out (17-6=11)

54

55 of 74

Regression Discontinuity

55

Treatment = +1 year of schooling

Area of interest

56 of 74

Regression Discontinuity – Condition 1

56

 

test for difference in density

57 of 74

Regression Discontinuity – Condition 2

  •  

57

58 of 74

Regression Discontinuity

  •  

58

Local Average Treatment Effect

59 of 74

Regression Discontinuity

  •  

59

This is now condition 3

60 of 74

Regression Discontinuity

  •  

60

61 of 74

Regression Discontinuity

  •  

61

62 of 74

Fuzzy RDD – Example

62

 

 

63 of 74

63

Fuzzy RDD – Example

 

64 of 74

Propensity Score Matching

  • This is a non-experimental technique whose purpose is to mimic that of an actual randomized experiment in the absence of randomization

  • The basic setup is as follows:
    • The treatment is assigned non-randomly
    • Create a synthetic control group from non-treated sample to estimate the counterfactual—call it the matched control group
    • Here, individuals who received the treatment are matched to untreated individuals based on their probability of receiving the treatment—call it the propensity score
    • The treatment effect is the difference in expected outcomes, conditional on the propensity score

64

65 of 74

Propensity Score Matching

  • How do we construct the matched control group?

  • There are several matching algorithms we can use to create a control group that looks like the treatment group
    • Nearest neighbor matching
    • Caliper matching
    • Inverse probability weights
    • More

  • All these techniques, however, require the estimation of a propensity score

65

66 of 74

Propensity Score Matching

  •  

66

67 of 74

Propensity Score Matching

  •  

67

68 of 74

Propensity Score Matching

  •  

68

69 of 74

Propensity Score Matching

  • Steps to propensity score matching:
  • Estimate the probability of receiving the treatment as a function of observed characteristics
  • Compute predicted probabilities—this is the propensity score
    • This computes the probability that any given observation would have received the treatment based only on its observed attributes
  • Use the propensity scores to find observations that match the treated group but never received the treatment
  • Check that the sample is balanced between treatment and matched group
  • Run the regression of interest to your study using only the treated and matched sample, or using propensity scores as weights
    • This step varies depending on the matching algorithm you decide to choose

69

70 of 74

Propensity Score Matching

  • Example: we want to estimate the effect of a reading intervention on reading scores

  • We estimate the probability of receiving the treatment as a function of students’ observed characteristics

  • The graph plots the propensity score distribution for treated and untreated children

  • You can tell that the two groups are different

70

71 of 74

Propensity Score Matching

  • Example: we want to estimate the effect of a reading intervention on reading scores

  • We estimate the probability of receiving the treatment as a function of students’ observed characteristics

  • The graph plots the propensity score distribution for treated and untreated children

  • After matching, they now look much more similar!

71

72 of 74

Propensity Score Matching

  •  

72

73 of 74

Propensity Score Matching

  • Example: Evaluation of a women’s empowerment intervention in Metn district in Lebanon (OXFAM, 2015)

  • Treatment included awareness raising and legal consultation for women

  • Comparison group was drawn from 4 different regions of Lebanon with similar religious composition
    • Metn, Deir El Ahmar, Merjeyoun, and Zahle

73

74 of 74

Thank you!

Feel free to direct your questions to

Wael Moussa: wmoussa@fhi360.org

74