Lecture Notes: Beyond Benchmarks: Building a Science of AI Measurement
Presenter: Samni Koyejo (Keynote)
Focus: Building trustworthy AI through validity-centered evaluation, cost-efficient tools, and critical awareness of synthetic data risks.
Notes compiled by an LLM + Vukosi Marivate
1. Introduction
Context & Motivation
- Benchmarks—from early character recognition to modern evaluations like GPQA—have driven AI progress. Yet the community often anchors grand claims (e.g., “general reasoning”) on narrow benchmark results, risking misinterpretation.
- Samni Koyejo advocates for a validity-centered evaluation framework: benchmarks must clearly support the specific claims they’re used to justify.
Validity Framework: “Measurement to Meaning”
Paper: Measurement to Meaning: A Validity-Centered Framework for AI Evaluation (Salaudeen et al., May 2025) (arXiv)
- Highlights that validity isn’t a property of benchmarks themselves, but of how well evidence supports intended claims.
- Proposes five critical validity dimensions:
- Content Validity – Does the dataset encompass all relevant scenarios?
- Criterion Validity – Does performance align with existing trusted measurements?
- Construct Validity – Are we actually measuring the intended abstract trait (e.g., reasoning), not just memorization or pattern matching?
- External Validity – Will results generalize across contexts, domains, distributions?
- Consequential Validity – What real-world effects does the claim have?
- Case Example: Excelling on GPQA (graduate-level biology/chemistry/physics MCQs) may support criterion validity for answering questions—but not construct validity for general reasoning.
Probabilistic & Efficient Evaluation with IRT
2. Making AI Evaluation More Efficient & Reliable
Problem with Standard Practices
- Traditional evaluation (simple accuracy averages) is expensive, noisy due to varied question difficulty, and often non-adaptive.
Solution: Applying Item Response Theory (IRT)
- Paper: Reliable and Efficient Amortized Model-Based Evaluation (Truong et al., Mar 2025) details this approach (covered in a Stanford news article) (news.stanford.edu).
- Uses IRT to separately model model capability and item difficulty, allowing more nuanced performance assessment.
- Amortized Calibration: Models question difficulty from content embeddings—reduces need for extensive labeling.
- Adaptive Question Generation: Synthesizes new items tailored to fill gaps in capability assessment.
- Outcome: Evaluation becomes cost-effective and scalable, enabling dynamic and continuous assessment capability.
Synthetic Data & Model Collapse
3. Understanding Model Collapse & Synthetic Data Risks
Key Concern
- Recursively training models on synthetic outputs (without real data) can degrade performance over generations—a phenomenon known as model collapse.
Evidence & Insights
- Paper: Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data (Gerstgrasser et al., Apr 2024) (dlab.berkeley.edu, arXiv)
- Confirms that if real data is replaced entirely by synthetic output, model collapse occurs.
- Demonstrates that if synthetic data is accumulated alongside real data, model performance remains stable—even as the real data proportion diminishes.
- Follow-Up Work: Collapse or Thrive? Perils and Promises of Synthetic Data in a Self‑Generating World (Kazdan et al., Oct 2024 / ICLR 2025) supports these findings across multiple modeling tasks and use cases (arXiv).
- Media & Context: A CACM news article also echoes Koyejo’s point: real-world data typically accumulates, so catastrophic collapse is unlikely—and mixing synthetic with real data slows degradation (cacm.acm.org).
- Summary: While model collapse is theoretically sound, realistic data workflows (adding rather than replacing real data) mitigate it.
4. Open Questions & Challenges
- Capability vs. Impact: Capability measurement research is not well-aligned with policy and societal impact considerations.
- Benchmark Formats: Heavy reliance on multiple-choice tests fails to reflect model behavior in realistic or interactive settings.
- Designing AI-Native Tests: We may need evaluation tools specifically designed for AI—not borrowed from human assessment paradigms.
- Institutional Oversight: There's limited governance or accountability around how claims are derived from benchmarks.
Take Action
For Researchers | For Industry | For Policy & Civil Society |
Avoid vague claims—explicitly state what your benchmark does and doesn’t measure. | Invest in IRT-based or adaptive evaluation systems, beyond static leaderboards. | Demand transparency on the validity evidence behind benchmark-based claims. |
Report structured evidence across the five validity dimensions. | Use amortized calibration and adaptive synthesis to evaluate more efficiently. | Encourage development of AI-tailored evaluation methodologies. |
Diversify evaluation formats—open-ended, interactive, multi-step tasks. | Refrain from exaggerated capability marketing; ground statements in measured performance. | Support frameworks and policies that tie claims to evidence and impact considerations. |