1 of 24

Large Language Monkeys: Scaling Inference-Time Compute with Repeated Sampling

Brad Brown*, Jordan Juravsky*, Ryan Ehrlich*, Ronald Clark, Quoc Le, Chris Ré, Azalia Mirhoseini

2 of 24

Why Care About Scaling Inference-Time Compute?

1. Unexploited Axis for Scaling

  • Training = $billions, inference OOMs less

2. Concentrated Axis for Scaling

- Spin ‘til you win on the problem you care about.

3 of 24

Why Care About Scaling Inference-Time Compute?

1. Unexploited Axis for Scaling

  • Training = $billions, inference OOMs less

2. Concentrated Axis for Scaling

- Spin ‘til you win on the problem you care about.

4 of 24

We Did the Simplest Possible Thing

How far can we get with a for loop?

5 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification.

6 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification.

7 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification.

8 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification.

9 of 24

How Do We Evaluate Repeated Sampling?

Coverage: As the number of samples increases, what fraction of problems can we solve using any sample that was generated?

10 of 24

Repeated Sampling is Effective Across Tasks

11 of 24

Repeated Sampling is Effective Across Tasks

pass@250 = 56% on SWE-bench Lite using DeepSeek-Coder-V2-Instruct

current single-attempt SOTA is 43%

12 of 24

Repeated Sampling can be Cost-Effective

13 of 24

Repeated Sampling is Effective Across Models

14 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification to exploit in practice.

15 of 24

Scaling Laws For Repeated Sampling

16 of 24

Scaling Laws For Repeated Sampling

17 of 24

Main Findings:

Repeated sampling:

  1. Is effective across multiple tasks and models.

  • Can often yield predictable benefits (i.e. scaling laws).

  • Requires scalable verification.

18 of 24

Verification Difficulty Varies by Task

  • In some settings, like code, formal proofs, etc., we have access to tools for automatically verifying solutions.

  • In other settings, not so much. We’re forced to resort to methods like:
    • Majority voting.
    • Reward models.

19 of 24

Common Verification Methods Don’t Scale

20 of 24

Verification Requires Finding Needles in the Haystack

21 of 24

Future/Ongoing Work

  • (In Progress) Building AI coders from the ground up that scale inference-time compute.
    • Amortizing expensive ops (e.g. reading every file) across many samples.
    • Let models write their own tests – another place to scale compute!

  • Better algorithms than repeated sampling.
    • Combining sequential and parallel scaling.
    • Removing isolation between parallel branches.

  • Designing scalable verification.

22 of 24

Thanks!

23 of 24

Training Scaling Laws Work Well!

24 of 24

Repeated Sampling can be Cost-Effective