Popper
01
OUR PROBLEM → EVERY TEAM’S PROBLEM
We got the fix in minutes.
Proving it took longer.
US
We became the
review bottleneck.
EVERY TEAM
Coding agents make this
everyone’s bottleneck.
TRUST IS NOW THE SCARCE RESOURCE
Popper
02
THE FALSE PASS
A passing check can still prove nothing.
One test. Two revisions. The result looked reassuring.
BEFORE THE FIX
PASS
AFTER THE FIX
PASS
≠ PROOF
If it passed before the change, it did not prove the change.
Popper
03
WHAT POPPER DOES
Every PR claim has to earn trust.
Popper turns a pull request’s promise into an adversarial experiment.
01 / ATTACK
Fireworks
Extract the claim.
Generate tests to falsify it.
02 / EXECUTE
Daytona
Run each test on the
before and after code.
03 / COMPARE
CodeRabbit
Set execution evidence beside
an independent review opinion.
Braintrust traces the evidence
CopilotKit lets reviewers interrogate it
A human decides
Popper
04
LIVE DEMO
Watch the claim meet the evidence.
90-SECOND DEMO
01
Load a recorded PR run
02
See evidence disagree
03
Let the human decide
Popper
05
WHAT WE LEARNED
We had to verify the verifier.
The hard part was preserving the meaning of evidence.
PASS ≠ PROOF
Pass before + after
That test is inconclusive.
TIMEOUT ≠ FAILURE
Sandbox unavailable
That is missing evidence.
OPINION ≠ EVIDENCE
Review says “looks good”
Only execution can prove behavior.
WHEN PROOF IS MISSING, THE GATE STAYS CLOSED
Popper
06
THE PROMISE
AI makes
the claim.
Popper makes it
earn trust.
THE HUMAN STILL DECIDES