1 of 6

Popper

01

OUR PROBLEM → EVERY TEAM’S PROBLEM

We got the fix in minutes.

Proving it took longer.

US

We became the

review bottleneck.

EVERY TEAM

Coding agents make this

everyone’s bottleneck.

TRUST IS NOW THE SCARCE RESOURCE

2 of 6

Popper

02

THE FALSE PASS

A passing check can still prove nothing.

One test. Two revisions. The result looked reassuring.

BEFORE THE FIX

PASS

AFTER THE FIX

PASS

≠ PROOF

If it passed before the change, it did not prove the change.

3 of 6

Popper

03

WHAT POPPER DOES

Every PR claim has to earn trust.

Popper turns a pull request’s promise into an adversarial experiment.

01 / ATTACK

Fireworks

Extract the claim.

Generate tests to falsify it.

02 / EXECUTE

Daytona

Run each test on the

before and after code.

03 / COMPARE

CodeRabbit

Set execution evidence beside

an independent review opinion.

Braintrust traces the evidence

CopilotKit lets reviewers interrogate it

A human decides

4 of 6

Popper

04

LIVE DEMO

Watch the claim meet the evidence.

90-SECOND DEMO

01

Load a recorded PR run

02

See evidence disagree

03

Let the human decide

5 of 6

Popper

05

WHAT WE LEARNED

We had to verify the verifier.

The hard part was preserving the meaning of evidence.

PASS ≠ PROOF

Pass before + after

That test is inconclusive.

TIMEOUT ≠ FAILURE

Sandbox unavailable

That is missing evidence.

OPINION ≠ EVIDENCE

Review says “looks good”

Only execution can prove behavior.

WHEN PROOF IS MISSING, THE GATE STAYS CLOSED

6 of 6

Popper

06

THE PROMISE

AI makes

the claim.

Popper makes it

earn trust.

THE HUMAN STILL DECIDES