Why Has Predicting Downstream Capabilities of Foundation Models with Scale Remained Elusive?
Sanmi Koyejo
Hailey Schoelkopf
Brando Miranda
Gabriel Mukobi
Stella Biderman
Herbie Bradley
Varun Madan
Adam Ibrahim
Rylan Schaeffer
Background: Predictability in Large Neural Networks
Background: Predictability in Large Neural Networks
Background: Surprise in Large Language Models
Background: Surprise in Large Language Models
Background: Aggregate downstream capabilities are predictable but only beyond a certain scale
Background: Specific downstream capabilities are predictable but only beyond a certain scale
Why has predicting specific downstream capabilities with scale remained elusive?
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Many reasons, but we contribute a new one!
Approach: MCQA Sample Correlations
How MCQA performance is computed (in equations)
How MCQA performance is computed (in pictures)
How MCQA performance is computed (in pictures)
How MCQA performance is computed (in pictures)
How MCQA performance is computed (in pictures)
How MCQA performance is computed (in pictures)
How MCQA performance is computed (in pictures)
Claim: The predictability of performance deteriorates with each subsequent transformation
Approach: Correlations Between Scores & Compute
0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 1 | 1 | 0 | 0 | 0 | 1 |
1 | 1 | 1 | 0 | 0 | 1 | 0 |
1 | 1 | 1 | 1 | 1 | 1 | 1 |
Models
Samples
Scores on Benchmark + Metric
Approach: Correlations Between Scores & Compute
0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 1 | 1 | 0 | 0 | 0 | 1 |
1 | 1 | 1 | 0 | 0 | 1 | 0 |
1 | 1 | 1 | 1 | 1 | 1 | 1 |
Models
Samples
7 |
8 |
9 |
10 |
Models
Scores on Benchmark + Metric
log10�FLOP
Approach: Correlations Between Scores & Compute
0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 1 | 1 | 0 | 0 | 0 | 1 |
1 | 1 | 1 | 0 | 0 | 1 | 0 |
1 | 1 | 1 | 1 | 1 | 1 | 1 |
Models
Samples
7 |
8 |
9 |
10 |
Models
log10�FLOP
Scores on Benchmark + Metric
|
Correlations
Approach: Correlations Between Scores & Compute
0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 | 1 | 1 | 0 | 0 | 0 | 1 |
1 | 1 | 1 | 0 | 0 | 1 | 0 |
1 | 1 | 1 | 1 | 1 | 1 | 1 |
Models
Samples
7 |
8 |
9 |
10 |
Models
Scores on Benchmark + Metric
| | | | | | |
Correlations
log10�FLOP
Idealized Score-Compute Correlation Distributions
Idealized Score-Compute Correlation Distributions
Idealized Score-Compute Correlation Distributions
Idealized Score-Compute Correlation Distributions
Methodology: Model Families and NLP Benchmarks
Model Families
Model Details (at a glance)
NLP Benchmarks
Results
Ex: ARC-Challenge and log pVocab(Correct Choice)
Ex: ARC-Challenge and log pVocab(Correct Choice)
Ex: ARC-Challenge and log pVocab(Correct Choice)
Ex: ARC-Challenge and log pVocab(Correct Choice)
Ex: ARC-Challenge and log pVocab(Correct Choice)
Ex: ARC-Challenge and log pVocab(Correct Choice)
Survival Function
Ex: ARC-Challenge and log pVocab(Correct Choice)
1 - ECDF
Survival Function
Score-compute correlations decrease under the sequence of transformations to compute Accuracy
Score-compute correlations decrease under the sequence of transformations to compute Accuracy
Score-compute correlations decrease under the sequence of transformations to compute Accuracy
Score-compute correlations decrease under the sequence of transformations to compute Accuracy
This is very qualitative…
Can we be more quantitative?
Statistic of score-compute correlation distributions reveals decrease of correlation values
Statistic:
Mean
Statistic of score-compute correlation distributions reveals decrease of correlation values
Statistic:
Median
Statistic of score-compute correlation distributions reveals decrease of correlation values
Statistic:
AUC of
Survival Function
What causes the decorrelation�between scores and compute?
What causes the decorrelation of scores and compute?
Answer: Probability masses on incorrect choices
What causes the decorrelation of scores and compute?
Answer: Probability masses on incorrect choices
What causes the decorrelation of scores and compute?
Answer: Probability masses on incorrect choices
What causes the decorrelation of scores and compute?
Answer: Probability masses on incorrect choices
What causes the decorrelation of scores and compute?
Answer: Probability masses on incorrect choices
How does mass on incorrect choices change with compute?
How does mass on incorrect choices change with compute?
How does mass on incorrect choices change with compute?
How does mass on incorrect choices change with compute?
How does mass on incorrect choices change with compute?
How does mass on incorrect choices change with compute?
(Preliminary) Scaling Behavior of Incorrect Choices
(Preliminary) Scaling Behavior of Incorrect Choices
(Preliminary) Scaling Behavior of Incorrect Choices
(Preliminary) Scaling Behavior of Incorrect Choices
(Preliminary) Scaling Behavior of Incorrect Choices
How does mass on incorrect choices change with compute?
Answers:
Conclusions
Metrics play a vital role in the predictability of LLM performance with scale.
Criticism: “Metrics we care about are discontinuous.”
Criticism: “Of course, metrics affect results; this is trivial!”
This does not explain *your setting of interest*
Next Steps
Thank you!