1 of 13

Evidence of Fraud in an Influential Field Experiment About Dishonesty

Uri, Joe, & Leif and anonymous authors

August 17, 2021

Data Colada

Presented by Chen Rozenshtein

2 of 13

The Original Paper, 2012

Sign first, lie less

A three-study paper in PNAS reported that asking people to sign a pledge of honesty before they fill out a form (top of the page) makes them more honest than signing after (bottom of the page).

The idea was influential: taught in classrooms, cited widely, and adopted by real organizations moving the signature line to the top of their forms.

Study 1

Lab - self-reported task

Study 2

Lab - self-reported task

Study 3

Field experiment at an auto insurer

Study 3 is where this story goes

3 of 13

The Insurance Field Experiment

Customers reported the current odometer reading for up to four insured cars. Half signed the honesty pledge at the top of the form, half at the bottom.

13,488

customers

+2,400

more miles reported

by the sign-at-top group

+10.3%

higher mileage - the

headline “honesty” effect

4

cars per policy,

max

  • “miles driven” was never reported directly. It was computed as Time-2 reading minus a much earlier Time-1 reading on file at the company.

4 of 13

The First Crack, 2020

A replication that failed - and an oddity

A 2020 follow-up (Kristal, Whillans & the five original authors) ran six studies that all failed to replicate the lab effect. Re-analyzing Study 3, they found something stranger.

“the randomization failed (or may have even failed to occur as instructed).”

the 2020 authors

~15,000 mi

Baseline gap between conditions

BEFORE random assignment

~2,400 mi

Analyzed difference

AFTER random assignment

The decisive move: to be transparent, they posted all the raw data on OSF. An anonymous team downloaded it - and that posted file is where the fabrication was caught.

5 of 13

The Investigation

Data Colada, 2021

Uri Simonsohn, Joe Simmons and Leif Nelson, with an anonymous team, walked through four independent red flags.

1

Impossible distribution

Miles driven are uniform 0-50,000 and stop dead at 50k

2

No rounding at Time 2

Humans round odometer readings; this data never does

3

Impossible “twins”

Every row has a near-duplicate in a second font

4

No rounding in Cambria

The duplicated rows also lost all rounding

6 of 13

Anomaly 1

A distribution that can't be real

Real driving is lopsided: many people drive a moderate amount, a few drive a lot. Here, every value from 0 to 50,000 miles is equally likely - statistically indistinguishable from a uniform random draw (p = .84) - and nothing exceeds 50,000.

hard wall

at 50,000

Likely recipe

Add a uniform random number, capped at 50,000, onto each baseline mileage.

=RANDBETWEEN(0,50000)

Same flat shape appears for all four cars (all p > .78).

7 of 13

Anomaly 2

People tend to round. The data didn't.

When people report mileage by hand, many give round numbers. The earlier Time-1 readings on file show exactly that. The experimental Time-2 readings show none of it - as if a machine produced them.

24.5%

of Time-1 readings end in zero - normal human rounding

10.3%

at Time 2 - flat, like a random number generator

8 of 13

Two Fonts

The baseline column for Car #1 is printed in two different fonts - exactly half the rows in Calibri, half in Cambria. The forensic reading: the data started in Calibri, then was duplicated in Cambria with a small random number added to disguise the copy.

6,744 rows in Calibri and 6,744 in Cambria - an exact 50/50 split that a real dataset would probably never produce.

9 of 13

Two Fonts

The baseline column for Car #1 is printed in two different fonts - exactly half the rows in Calibri, half in Cambria. The forensic reading: the data started in Calibri, then was duplicated in Cambria with a small random number added to disguise the copy.

id

font

baseline_car1

update_car1

4

Cambria

23,912

59,136

5

Calibri

16,862

59,292

6

Calibri

147,738

167,895

7

Calibri

18,780

49,811

9

Cambria

28,993

63,707

Color added to mark the font of each baseline value.

6,744 rows in Calibri and 6,744 in Cambria - an exact 50/50 split that a real dataset would probably never produce.

10 of 13

Anomaly 3

Impossibly identical “driving twins”

Every Calibri customer has a Cambria match whose mileage, on all four cars, is just slightly higher - always by less than 1,000 miles. Across the whole file, no exceptions.

Recipe for the twins: duplicate each row, then add =RANDBETWEEN(0,1000) to every baseline mileage.

car 1

car 2

car 3

car 4

Calibri row

49,675

17,709

27,357

64,428

Cambria twin

50,350

18,421

27,714

64,784

difference

+675

+712

+357

+356

One of 22 four-car pairs - the pattern holds for two-car and one-car customers too.

11 of 13

So, what is fabricated?

Time-1 data

Real rows were duplicated, then nudged with a random number to hide the copy.

Time-2 data

The entire experimental column was invented with a random generator, capped at 50,000.

Who fabricated it? The data can't say.

Only one author was ever in contact with the insurer that ran the study. That leaves three logical possibilities:

That author

The posted file's metadata lists Dan Ariely as its creator and Nina Mazar as the last to modify it; the post draws no conclusion about who is responsible.

Someone at the insurance company

Someone in that author's lab

12 of 13

Authors’ Responses

DA

Dan Ariely

Led the study

Partnered with the insurer and was the only author in contact with it. Says the company collected, entered and merged the data before sending it to him; he did not suspect or test it. Duke's research-integrity office reviewed the matter.

NM

Nina Mazar

Received the file

Says she was not involved in conducting the field study and has no knowledge of who fabricated the data. Notes that the 2020 randomization concern is what led the team to post the data publicly - which made the discovery possible.

MB

Max Bazerman

Pushed to retract

Recalls flagging implausible Study-3 data as early as 2011 but being reassured. In 2020 he argued for retraction; only he and Shu were in favor, so it didn't happen. Wishes he had acted independently.

FG

Francesca Gino

Not involved in Study 3

Says she had no role in running or analyzing the field study and no suspicion at publication. Thanks the investigators and regrets not taking a stronger stance for retraction in 2020.

Lead author Lisa Shu did not post a public response. Summaries here paraphrase the four authors' posted statements.

13 of 13

Thanks!

Any questions?

Chen Rozenshtein

chenrozenstein@gmail.com