1 of 16

Tracking Data Quality in �Online and Laboratory Pools

Liman Wang, Weishan Zhang, Chenyu Wang, Donald Lyons, Aastha Mittal, Rachel Tran, Don A. Moore, Juliana Schroeder

2 of 16

Changing Data Sources

  • In the past 15 years, we've become dependent on online pools

  • Pros:
    • inexpensive
    • fast
    • (somewhat) diverse

  • Cons:
    • participant professionalization
    • distraction / multitasking
    • participant fraud

3 of 16

Measuring Data Quality

  • Attention:
    • What's a reasonable level of attention?
    • How should we measure attention?
  • Comprehension:
    • Participants understand the intended question
  • Effort:
    • How hard are participants trying
  • Common sense/general knowledge:
    • Participants have basic background info
  • Authenticity:
    • Humans, not AI

4 of 16

Our Study

  • 2021-2024 (+ 2026 now running but not included in analysis yet)
    • CloudResearch MTurk
    • Prolific
    • Xlab online (UC Berkeley’s Experimental Social Science Lab online pool)
    • CDR online (U Chicago’s Center for Decision Research Lab online pool)
    • Credamo (online pool in China; starting in 2023)
    • Community pools: Xlab’s online (2021, 2022); CDR’s in-person (Mindworks, 2024, 2026)
    • Xlab in-person and CloudResearch Connect in 2026 only
  • Aim for 200 per pool each year

https://osf.io/knvad/

5 of 16

Results – Attention Checks

Self-reported attention.Overall, how closely do you usually pay attention to the instructions and tasks when you complete surveys on this platform? Please be completely honest; you won’t be penalized for your response.” ��(1 = Not at all closely; 5 = Extremely closely).

Check embedded in instructions. “In the following survey, you will be asked questions that are related to how you approach a multitude of situations. For these questions please answer them as truthfully as possible. If you are in doubt between some of the answers, choose the answer that was your first choice. To ensure the validity of our results, it is important that you read any information that you are provided carefully. You will be asked questions regarding the information that you are given. For example, if you are reading this, please ignore this question and select: Italy. What country are you in?” �(1 = Italy ; 2 = Spain; 3 = United States; 4 = South Africa)

Check embedded in item. “How well do the following statements describe your personality? I see myself as someone who

...select ‘strongly agree’ for this item.”

Common sense question. “Normally, how often is the U.S. Presidential Election?”

(1 = Annually; 2 = Every 2 years; 3 = Every 4 years; 4 = Every decade; 5 = There is not a strict rule on this.)

6 of 16

Results - Reading Comprehension

"Sally and Molly are best friends. They live together in an apartment and have many friends over. Sally's best friend is Tom. He loves to ski and read. Molly's best friend is Samantha. She loves to travel with Molly. Their next destination is going to be Greece. Using the text above, who does each of the pronouns relate to?"

  • he, she, they, their

7 of 16

Results – Effort on Writing Tasks

Writing quality. �“Please describe in detail the plot of the last movie that you watched. Please write at least four sentences. The more detail you can provide, the better.”

“Please describe the best survey you have taken part in on this platform. Please write at least three sentences.”

Coding:

  • 0 = gibberish
  • 1 = didn't pass, but made sense
  • 2 = just pass the requirement
  • 3 = detailed

8 of 16

Results – Survey Experience

9 of 16

Results – Participant Demographics

10 of 16

Correlations

Additional Measures

  • Basic arithmetic problems
  • Self-report multitasking: “How often do you do other things at the same time as completing surveys?”

11 of 16

In summary

  • There are multiple dimensions of data quality

  • Different types of quality for different types of pools
    • Online pools excel in answering attention checks correctly
    • Lab pools are better at effort/thought (e.g., writing/comprehension quality)

  • Relative stability in quality across years��….but the AI arms race is just starting in earnest

12 of 16

Thank you!

Don Moore, dm@berkeley.edu

Materials, data, preregistrations: https://osf.io/knvad/

Project sponsor:

UC Berkeley’s Experimental Social Science Laboratory

13 of 16

Bot Check 1 — Typing Behavior

“Please describe in detail the plot of the last movie that you watched. Please write at least four sentences. The more detail you can provide, the better. (Please describe your own experience and do not copy or paste from online sources.)”

— We embed JavaScript in Qualtrics to capture typing behavior:

  • Inter-keystroke intervals
  • Distributional properties of typing intercals (mean, median, SD)
  • Copy-and-paste events (e.g., length and timing of pasted text)

14 of 16

Bot Check 2 — Hidden Instructions

15 of 16

Bot Check 3 — Visual Illusion

16 of 16

Bot Check 4 — Trajectory Prediction