1 of 55

Metadata

  • Title: High Frequency Checks & Back Checks
  • Purpose: Introduce and practice IPA’s user-written commands for monitoring incoming data quality
  • Learning Goals and Key Takeaways:
  • - The role of data quality in IPA’s Minimum Must-Dos and IPA’s user-written commands that achieve those standards
  • - Understand format of IPA’s High Frequency Checks and Back Checks commands
  • - Experience setting up and interpreting both tools to incorporate in data flow
  • Date Created:
  • Created by: Christopher Boyer
  • Last Edited on: 10/05/2020
  • Last Edited by: Rosemarie Sandino

2 of 55

High Frequency Checks & Back Checks

Rosemarie Sandino

Technical Products Coordinator

RST, USA 2021

3 of 55

Data Flow – Importing

Raw Survey Data

(.csv)

Master .do

Import .do

Merge .do

Cleaning .do

Prepped Survey Data

(.dta)

  1. Drop unnecessary variables
  2. Generate new variables
  3. Rename variables
  4. Label variables
  5. Format variables
  6. Only do the cleaning necessary to do checks!

Maintain untouched raw data files!

4 of 55

Data Flow – Monitoring

Enumerator Performance

Prepped Survey Data

(.dta)

High Frequency Checks (HFC)

Back Checks (Audits)

Daily

Checks

Weekly Checks

Spot Checks

-bcstats-

Daily

Report

Weekly Report

BC Diffs

Logical Errors/Discrepancies

Threats to Research Quality

5 of 55

High-Frequency Checks

A high-frequency check is a check of some element of the data collection process, completed on a regular basis as new data comes in. They provide information about:

      • The quality of the data
      • Threats to research validity
      • Enumerator performance
      • The quality of the survey
      • The data flow (are there systemic flaws?)
    • Run checks on your data as often as you can
      • Consider running some checks every day, some every few days, some at the beginning of data collection, etc.

High Frequency Checks & Back Checks

6 of 55

Reasons for Data Quality Issues

    • Enumerator
      • doesn’t ask question clearly
      • doesn’t probe or make respondent feel comfortable answering the question
      • rushes through survey or skips questions
      • accidentally types the wrong thing or chooses the wrong answer
    • Respondent
      • doesn’t understand the question
      • doesn’t feel comfortable answering the question
      • is tired of answering questions and starts answering randomly
    • Survey
      • is confusing
      • choice lists don’t include the answer the respondent wants to choose
      • is too long and cumbersome
      • programming error leads to irrelevant question or typo

High Frequency Checks & Back Checks

7 of 55

Acting on Data Quality Issues

Consider the reason behind the data quality issue to address it

    • What could have caused the issue?
    • Can it be fixed? How can it be fixed?
    • Who should fix it? What should the response be?
    • What is the process for responding?

Remember that no data is perfect

    • Prioritize issues you can see and address
    • Create a plan for addressing each type of check

High Frequency Checks & Back Checks

8 of 55

Data Quality Action Plans

    • Discuss with your PI in advance what is most important to monitor
    • Create a schedule for running HFCs and sending reports to PIs
    • Write out a plan for action steps in advance and how you will address each check you run
      • Run check on key variables for outliers.
      • For outliers past a certain threshold (that you have created with your PI), give field managers Excel file of flagged values
      • Field managers call enumerators with flagged values and ask them to confirm
      • Make replacements in replacements file or mark them as “okay” so they don’t show up again

High Frequency Checks & Back Checks

9 of 55

HFC Template Workflow

HFC Outputs

Distribute among Field and Research Teams

  • HFC Replacements
  • Re-training
  • Reminders
  • Incentives
  • Praise
  • Survey programming changes

10 of 55

Data Collection

Back checks (MMD for IPA)

    • Re-ask within 1-3 days
    • Think about types and variance
    • Back check > 10% of surveys
    • Stratify random assignment by surveyor
    • Quickly compare and take action using bcstats
    • Backcheck forms should be included in IRB submissions

High Frequency Checks & Back Checks

11 of 55

bcstats to the Rescue!!

  • What is it?
  • bcstats is an IPA user-written Stata command to compare back check data to survey data and summarize results.
  • Syntax:

bcstats, surveydata(filename) bcdata(filename) ///

id(varlist) [options]

12 of 55

Linking the two

Why we do back checks in the first place:

  • Identify survey falsifications (type 1)
  • Evaluate enumerator performance (type 2)
  • Identify enumerator effects or other sources of bias (HFCs/type3)
  • Evaluate the reliability of our measures (type 3)

13 of 55

Data Flow

Run master do file to:

    • Import data
    • Merge and/or Prep Data
      • Creates: prepped dta file
    • Run High-Frequency Checks
      • Creates: HFC outputs, tracking data, enumerator performance
    • Make replacements
      • Creates: adjusted dta file
    • Run back check analysis
      • Optional: prep back check data as well

14 of 55

Data Management System

  • Consolidate high frequency checks, tracking, and back checks into the same system.
  • Create your folder structure/data flow with readme files
  • Input and output sheets in Excel that allow you to turn on/off checks, set key variables for checks, formatted summaries

15 of 55

Relative Working Directories

  • Relative working directories reference your existing working directory
  • ../
    • Move back one folder

16 of 55

Data Flow

17 of 55

Minimum Checks

Survey Logic Checks

1. Check that all interviews are completed

2. Check that there are no duplicate observations

3. Check that all surveys have consent

4. Check that certain variables have no missing values

5. Check that follow up record data match original

6. Check for skip pattern or survey logic violations

7. Check that no variable has all missing values

8. Check for hard/soft constraint violations

9. Check specify other for new categories or recodes

10. Check that date values fall within survey range

11. Check for outliers in unconstrained variables

18 of 55

Minimum Checks

Enumerator Checks

1. Check the percentage of “don’t know” and “refusal”

values for each variable by enumerator

2. Check the percentage of missing values for each

variable by enumerator.

3. Check the percentage of survey refusals by

enumerator

4. Check the number of surveys per day by

enumerator

5. Check average interview duration by enumerator

6. Check that duration of consent and other key

variables by enumerator

7. Check for systematic differences in responses by

enumerator

19 of 55

Minimum Checks

Research Quality Checks

1. Survey progress towards recruitment goals.

2. Summary of key research variables.

 3. Two way summaries of research variables by

demographic/geographic characteristics.

 4. Refusal/not found rates by treatment status.

 5. Maps/GIS

20 of 55

Additional Checks

Research Quality Checks

1. Text audit summaries

2. Field comments summaries

 3. Progress reports

 4. Variable statistics by enumerator

21 of 55

Data Import Workflow

1. Download .csv using SurveyCTO Sync

3. Stata Data Set

2. SurveyCTO do-file

22 of 55

HFC Template Workflow

1. HFC Inputs

2. Run Simple Do-file on Data

3. HFC Output Lists

23 of 55

HFC Template Workflow

1. HFC Outputs

3. HFC Replacements

2. Distribute among Field and Research Teams

  • Re-training
  • Reminders
  • Incentives
  • Praise
  • Survey programming changes

24 of 55

HFC Template Workflow

1. HFC Inputs

3. HFC Output Lists

4. HFC Replacements

2. Run Simple Do-file on Data

25 of 55

File Overview: Input

HFC Input

26 of 55

File Overview: Output

HFC Output

27 of 55

File Overview: Replacement

  • A new instructions sheet that explains how to fill out replacements
  • Option to drop, replace, or “okay” observations
  • Automatically run in DMS

28 of 55

Dashboards

29 of 55

Excel Output - 1. Differences File

The basic idea is to harmonize Excel output with the HFCs

30 of 55

Excel Output - 1. Differences File

The first section identifies the id, survey enum, and back checker

31 of 55

Excel Output - 1. Differences File

The second section identifies the type (1, 2, or 3) and the variable name and label (question text)

32 of 55

Excel Output - 1. Differences File

The third section shows the survey value and the corresponding back check value. Users can also specify additional variables to display (e.g. date)

33 of 55

Excel Output - 2. Error rates by variable

The second tab shows the error rates by variable (i.e. the number of differences divided by the total number of responses) sorted from highest to lowest rate of errors.

34 of 55

Excel Output - 3. Error rates by enumerator

The third tab shows the error rates by enumerator sorted from highest to lowest.

35 of 55

Excel Output - 4. Error rates by back checker

The fourth tab shows the error rates by back checker sorted from highest to lowest.

36 of 55

Excel Output - 5. Statistical test of stability

The fifth tab shows the results of a two-sample mean comparison test of stability for continuous type 2 and type 3 variables.

A p-value below 0.05 suggests that back check data are significantly different than survey data (which could mean improper surveying is skewing the results)

37 of 55

Excel Output - 6. Statistical test of stability

The sixth tab shows the results of a two-sample mean comparison test of stability for binary type 2 and type 3 variables.

38 of 55

Excel Output - 7. Reliability of survey measures

The seventh tab shows the simple response variance and the reliability rato.

39 of 55

  • High Frequency Checks Manual is on the Global Help Desk!

40 of 55

  • Back Check Manual is on the Global Help Desk!

41 of 55

HFCs Exercise (60 min.)

If you already have ipacheck and bcstats installed:

Open Stata, set your working directory, and type the code:

ipacheck new, exercise

Open exercise_instructions.pdf and begin!

If you have not installed the commands, go to installation instructions on website

42 of 55

Thank you

poverty-action.org

povertyactionlab.org

Thank you

poverty-action.org

43 of 55

Credits

  • This presentation was created by Christopher Boyer and references the user-written commands ipacheck (created by Christopher Boyer) and bcstats (created by Matthew White).

44 of 55

References

  • Addae, Kai. 2017. “Back Check Manual.” Innovations for Poverty Action Global Help Desk. Last modified December 19, 2017. https://povertyaction.force.com/support/s/article/back-check-manual.
  • Innovations for Poverty Action. 2017. “high-frequency-checks”. Github repository. https://github.com/PovertyAction/high-frequency-checks.
  • Innovations for Poverty Action. 2017. “bcstats”. Github repository. https://github.com/PovertyAction/bcstats.
  • Sandino, Rosemarie. 2018. “High Frequency Checks Manual.” Innovations for Poverty Action Global Help Desk. Last modified June 25, 2018. https://povertyaction.force.com/support/s/article/hfcs-manual.

45 of 55

Measurement Theory

var(ε)

Let’s be honest with ourselves when answering a question, respondents are likely drawing their answer at random from some distribution of responses about the true value.

46 of 55

Measurement Theory

var(ε)

For example, if you ask me multiple times “how many times have you washed your hands in the last day” I’m likely to give a slightly different answer depending on my memory, stress level, etc.

47 of 55

Measurement Theory

var(ε)

Across an entire survey this variability in interview responses will tend to increase the variance of an indicator… and possibly decrease our power for detecting significant effects.

var(y)

yij = tij + εi

48 of 55

Measurement Theory

var(ε)

Under certain assumptions (i.e. the survey and back check responses are independent), it can be shown that variance of the measurement error is equivalent to the simple response variance (SRV).

SRV = var(yi2- yi1) = var(εi)

49 of 55

Measurement Theory

The SRV, by itself, is not a very useful tool for assessing the reliability of different questions. However, it can be used to calculate the reliability ratio

SRV = var(yi2- yi1)

Reliability Ratio = 1 - SRV / var(yi)

50 of 55

Measurement Theory

A reliability ratio of greater than 0.8 is considered “high”. The UN recommends a reliability ratio of at least 0.5 in all important survey outcomes.

51 of 55

Measurement Theory

However, to measure the reliability ratio effectively we must take care to ensure that the conditions of the interview in the survey and the back check are as close to the same as possible.

52 of 55

Measurement Theory

A note of caution: there is another source of error, response bias, which is a systematic difference in successive responses.

bias

53 of 55

Measurement Theory

For example, suppose I am very conscious of what FEMALE interviewers might think of me if I reported my true hand washing habits, so then I systematically over report hand washing when the interviewer is female.

bias

H0: μ1 = μ2

H1: μ1 ≠ μ2

54 of 55

Measurement Theory

Unfortunately, the only way to identify this type of error is to either (1) directly measure the “true value” or (2) change the conditions under which the back check is done.

bias

H0: μ1 = μ2

H1: μ1 ≠ μ2

55 of 55

Summary

  • The reliability ratio and SRV give an idea of how much variability is due to instability in your instrument, but estimating them requires conditions of survey and back check to match.

  • On the other hand, stability checks can give you a sense of potential response bias, but estimating them requires knowing the true value or changing the conditions of the back check.