1 of 70

Why Has Predicting Downstream Capabilities of Foundation Models with Scale Remained Elusive?

Sanmi Koyejo

Hailey Schoelkopf

Brando Miranda

Gabriel Mukobi

Stella Biderman

Herbie Bradley

Varun Madan

Adam Ibrahim

Rylan Schaeffer

2 of 70

Background: Predictability in Large Neural Networks

3 of 70

Background: Predictability in Large Neural Networks

4 of 70

Background: Surprise in Large Language Models

5 of 70

Background: Surprise in Large Language Models

6 of 70

Background: Aggregate downstream capabilities are predictable but only beyond a certain scale

7 of 70

Background: Specific downstream capabilities are predictable but only beyond a certain scale

8 of 70

Why has predicting specific downstream capabilities with scale remained elusive?

9 of 70

Many reasons, but we contribute a new one!

10 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

11 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

  • Reason: Downstream performance

12 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

  • Reason: Downstream performance is computed from negative log likelihoods

13 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

  • Reason: Downstream performance is computed from negative log likelihoods via a sequence of transformations

14 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

  • Reason: Downstream performance is computed from negative log likelihoods via a sequence of transformations that progressively deteriorate the statistical relationship

15 of 70

Many reasons, but we contribute a new one!

  • Note: The reason is specific to multiple-choice question-answering (MCQA) benchmarks

  • Reason: Downstream performance is computed from negative log likelihoods via a sequence of transformations that progressively deteriorate the statistical relationship between performance and scale

16 of 70

Approach: MCQA Sample Correlations

17 of 70

How MCQA performance is computed (in equations)

18 of 70

How MCQA performance is computed (in pictures)

19 of 70

How MCQA performance is computed (in pictures)

20 of 70

How MCQA performance is computed (in pictures)

21 of 70

How MCQA performance is computed (in pictures)

22 of 70

How MCQA performance is computed (in pictures)

23 of 70

How MCQA performance is computed (in pictures)

Claim: The predictability of performance deteriorates with each subsequent transformation

24 of 70

Approach: Correlations Between Scores & Compute

0

0

0

0

0

0

0

0

1

1

0

0

0

1

1

1

1

0

0

1

0

1

1

1

1

1

1

1

Models

Samples

Scores on Benchmark + Metric

25 of 70

Approach: Correlations Between Scores & Compute

0

0

0

0

0

0

0

0

1

1

0

0

0

1

1

1

1

0

0

1

0

1

1

1

1

1

1

1

Models

Samples

7

8

9

10

Models

Scores on Benchmark + Metric

log10�FLOP

26 of 70

Approach: Correlations Between Scores & Compute

0

0

0

0

0

0

0

0

1

1

0

0

0

1

1

1

1

0

0

1

0

1

1

1

1

1

1

1

Models

Samples

7

8

9

10

Models

log10�FLOP

Scores on Benchmark + Metric

Correlations

27 of 70

Approach: Correlations Between Scores & Compute

0

0

0

0

0

0

0

0

1

1

0

0

0

1

1

1

1

0

0

1

0

1

1

1

1

1

1

1

Models

Samples

7

8

9

10

Models

Scores on Benchmark + Metric

Correlations

log10�FLOP

28 of 70

Idealized Score-Compute Correlation Distributions

29 of 70

Idealized Score-Compute Correlation Distributions

30 of 70

Idealized Score-Compute Correlation Distributions

31 of 70

Idealized Score-Compute Correlation Distributions

32 of 70

Methodology: Model Families and NLP Benchmarks

Model Families

  1. Pythia
  2. Cerebras-GPT
  3. OLMo
  4. INCITE
  5. LLM360

Model Details (at a glance)

  • All “base” models i.e. no finetuning
  • 70M to 13B parameters
  • Up to 2.4T tokens

NLP Benchmarks

  1. ARC Easy
  2. ARC Challenge
  3. HellaSwag
  4. MathQA
  5. MCTACO
  6. MMLU (57 subsets separately)
  7. OpenbookQA
  8. PIQA
  9. RACE
  10. SciQ
  11. SIQA
  12. WinoGrande
  13. XWinoGrad En

33 of 70

Results

34 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

35 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

36 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

37 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

38 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

PDF

39 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

PDF

Survival Function

40 of 70

Ex: ARC-Challenge and log pVocab(Correct Choice)

PDF

1 - ECDF

Survival Function

41 of 70

Score-compute correlations decrease under the sequence of transformations to compute Accuracy

42 of 70

Score-compute correlations decrease under the sequence of transformations to compute Accuracy

43 of 70

Score-compute correlations decrease under the sequence of transformations to compute Accuracy

44 of 70

Score-compute correlations decrease under the sequence of transformations to compute Accuracy

45 of 70

This is very qualitative…

Can we be more quantitative?

46 of 70

Statistic of score-compute correlation distributions reveals decrease of correlation values

Statistic:

Mean

47 of 70

Statistic of score-compute correlation distributions reveals decrease of correlation values

Statistic:

Median

48 of 70

Statistic of score-compute correlation distributions reveals decrease of correlation values

Statistic:

AUC of

Survival Function

49 of 70

What causes the decorrelation�between scores and compute?

50 of 70

What causes the decorrelation of scores and compute?

Answer: Probability masses on incorrect choices

51 of 70

What causes the decorrelation of scores and compute?

Answer: Probability masses on incorrect choices

52 of 70

What causes the decorrelation of scores and compute?

Answer: Probability masses on incorrect choices

53 of 70

What causes the decorrelation of scores and compute?

Answer: Probability masses on incorrect choices

54 of 70

What causes the decorrelation of scores and compute?

Answer: Probability masses on incorrect choices

55 of 70

How does mass on incorrect choices change with compute?

56 of 70

How does mass on incorrect choices change with compute?

57 of 70

How does mass on incorrect choices change with compute?

58 of 70

How does mass on incorrect choices change with compute?

59 of 70

How does mass on incorrect choices change with compute?

60 of 70

How does mass on incorrect choices change with compute?

61 of 70

(Preliminary) Scaling Behavior of Incorrect Choices

62 of 70

(Preliminary) Scaling Behavior of Incorrect Choices

63 of 70

(Preliminary) Scaling Behavior of Incorrect Choices

64 of 70

(Preliminary) Scaling Behavior of Incorrect Choices

65 of 70

(Preliminary) Scaling Behavior of Incorrect Choices

66 of 70

How does mass on incorrect choices change with compute?

Answers:

  1. Probability mass increases on incorrect answers with increasing compute

  • As best as we can tell, there is no “tipping point” where probability mass on incorrect choices begins decreasing

67 of 70

Conclusions

68 of 70

Metrics play a vital role in the predictability of LLM performance with scale.

  • Preferred (often discontinuous) metrics systematically degrade predictability
  • Partial credit and probability score metrics are (more) predictable

69 of 70

Criticism: “Metrics we care about are discontinuous.”

  • I think this misses the point of the work
    • We are most interested in predictability
    • We also argue that “partial-credit” metrics are more useful for science and decision-making
  • Perhaps, though there are often construct validity gaps
    • Multiple-choice + Accuracy is the most common metric in benchmarks
    • Real-world use cases are rarely well captured by this. “Does passing the bar imply usefulness for law?”

Criticism: “Of course, metrics affect results; this is trivial!”

This does not explain *your setting of interest*

  • Agreed, we are working on it. Join us!
  • Towards a science of scaling-predictable and human-interpretable evaluation!

70 of 70

Next Steps

  1. Why does mass increase on incorrect choices with increasing compute?
    1. Disentangle contributions of 3 potential causes:
      1. Models place more mass on syntactically valid strings
      2. Models learn to choose one of the available options
      3. Incorrect choices are semantically close to the correct answer
  2. Scaling laws for per-sample incorrect choices
    • 4 different methods - how well does each work?
  3. Gadre’s Finding: Explained by more data->lower variance, or more data->match the pretraining distribution?
  4. Science of Scaling-Predictable and Human-Interpretable Evals
    • Move beyond multiple-choice question-answering benchmarks

Thank you!