Large Language Monkeys: Scaling Inference-Time Compute with Repeated Sampling
Brad Brown*, Jordan Juravsky*, Ryan Ehrlich*, Ronald Clark, Quoc Le, Chris Ré, Azalia Mirhoseini
Why Care About Scaling Inference-Time Compute?
1. Unexploited Axis for Scaling
2. Concentrated Axis for Scaling
- Spin ‘til you win on the problem you care about.
Why Care About Scaling Inference-Time Compute?
1. Unexploited Axis for Scaling
2. Concentrated Axis for Scaling
- Spin ‘til you win on the problem you care about.
We Did the Simplest Possible Thing
How far can we get with a for loop?
Main Findings:
Repeated sampling:
Main Findings:
Repeated sampling:
Main Findings:
Repeated sampling:
Main Findings:
Repeated sampling:
How Do We Evaluate Repeated Sampling?
Coverage: As the number of samples increases, what fraction of problems can we solve using any sample that was generated?
Repeated Sampling is Effective Across Tasks
Repeated Sampling is Effective Across Tasks
pass@250 = 56% on SWE-bench Lite using DeepSeek-Coder-V2-Instruct
current single-attempt SOTA is 43%
Repeated Sampling can be Cost-Effective
Repeated Sampling is Effective Across Models
Main Findings:
Repeated sampling:
Scaling Laws For Repeated Sampling
Scaling Laws For Repeated Sampling
Main Findings:
Repeated sampling:
Verification Difficulty Varies by Task
Common Verification Methods Don’t Scale
Verification Requires Finding Needles in the Haystack
Future/Ongoing Work
Thanks!
Training Scaling Laws Work Well!
Repeated Sampling can be Cost-Effective