Lecture Notes: Beyond Benchmarks: Building a Science of AI Measurement

Presenter: Samni Koyejo (Keynote)
Focus: Building trustworthy AI through validity-centered evaluation, cost-efficient tools, and critical awareness of synthetic data risks.

Notes compiled by an LLM + Vukosi Marivate

1. Introduction

Context & Motivation

  • Benchmarks—from early character recognition to modern evaluations like GPQA—have driven AI progress. Yet the community often anchors grand claims (e.g., “general reasoning”) on narrow benchmark results, risking misinterpretation.
  • Samni Koyejo advocates for a validity-centered evaluation framework: benchmarks must clearly support the specific claims they’re used to justify.

Validity Framework: “Measurement to Meaning”

Paper: Measurement to Meaning: A Validity-Centered Framework for AI Evaluation (Salaudeen et al., May 2025) (arXiv)

  • Highlights that validity isn’t a property of benchmarks themselves, but of how well evidence supports intended claims.

  • Proposes five critical validity dimensions:

  1. Content Validity – Does the dataset encompass all relevant scenarios?
  2. Criterion Validity – Does performance align with existing trusted measurements?
  3. Construct Validity – Are we actually measuring the intended abstract trait (e.g., reasoning), not just memorization or pattern matching?
  4. External Validity – Will results generalize across contexts, domains, distributions?
  5. Consequential Validity – What real-world effects does the claim have?

  • Case Example: Excelling on GPQA (graduate-level biology/chemistry/physics MCQs) may support criterion validity for answering questions—but not construct validity for general reasoning.


Probabilistic & Efficient Evaluation with IRT

2. Making AI Evaluation More Efficient & Reliable

Problem with Standard Practices

  • Traditional evaluation (simple accuracy averages) is expensive, noisy due to varied question difficulty, and often non-adaptive.

Solution: Applying Item Response Theory (IRT)

  • Paper: Reliable and Efficient Amortized Model-Based Evaluation (Truong et al., Mar 2025) details this approach (covered in a Stanford news article) (news.stanford.edu).

  • Key Contributions:

  • Uses IRT to separately model model capability and item difficulty, allowing more nuanced performance assessment.
  • Amortized Calibration: Models question difficulty from content embeddings—reduces need for extensive labeling.
  • Adaptive Question Generation: Synthesizes new items tailored to fill gaps in capability assessment.

  • Outcome: Evaluation becomes cost-effective and scalable, enabling dynamic and continuous assessment capability.

Synthetic Data & Model Collapse

3. Understanding Model Collapse & Synthetic Data Risks

Key Concern

  • Recursively training models on synthetic outputs (without real data) can degrade performance over generations—a phenomenon known as model collapse.

Evidence & Insights

  • Paper: Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data (Gerstgrasser et al., Apr 2024) (dlab.berkeley.edu, arXiv)

  • Confirms that if real data is replaced entirely by synthetic output, model collapse occurs.
  • Demonstrates that if synthetic data is accumulated alongside real data, model performance remains stable—even as the real data proportion diminishes.

  • Follow-Up Work: Collapse or Thrive? Perils and Promises of Synthetic Data in a Self‑Generating World (Kazdan et al., Oct 2024 / ICLR 2025) supports these findings across multiple modeling tasks and use cases (arXiv).

  • Media & Context: A CACM news article also echoes Koyejo’s point: real-world data typically accumulates, so catastrophic collapse is unlikely—and mixing synthetic with real data slows degradation (cacm.acm.org).

  • Summary: While model collapse is theoretically sound, realistic data workflows (adding rather than replacing real data) mitigate it.

4. Open Questions & Challenges

  1. Capability vs. Impact: Capability measurement research is not well-aligned with policy and societal impact considerations.
  2. Benchmark Formats: Heavy reliance on multiple-choice tests fails to reflect model behavior in realistic or interactive settings.
  3. Designing AI-Native Tests: We may need evaluation tools specifically designed for AI—not borrowed from human assessment paradigms.
  4. Institutional Oversight: There's limited governance or accountability around how claims are derived from benchmarks.

Take Action

For Researchers

For Industry

For Policy & Civil Society

Avoid vague claims—explicitly state what your benchmark does and doesn’t measure.

Invest in IRT-based or adaptive evaluation systems, beyond static leaderboards.

Demand transparency on the validity evidence behind benchmark-based claims.

Report structured evidence across the five validity dimensions.

Use amortized calibration and adaptive synthesis to evaluate more efficiently.

Encourage development of AI-tailored evaluation methodologies.

Diversify evaluation formats—open-ended, interactive, multi-step tasks.

Refrain from exaggerated capability marketing; ground statements in measured performance.

Support frameworks and policies that tie claims to evidence and impact considerations.