1 of 14

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, Julian Nyarko

Stanford Law School, Snorkel AI, Harvey

2 of 14

Is Legal Practice More Like Science Or Art?

  • Rules-based, ground truth is objective
  • Body of settled knowledge
  • “Does this reaction produce the expected compound?”
  • “Does this contract clause trigger the specified payment obligation under the formula?”

  • Judgment-based, ground truth is subjective
  • Holistic comparisons are common
  • “Is this painting any good, and what about it makes it so?”
  • “Is this predictive legal memo any good, and what about it makes it so?”

1 2 3 4 5 6 7 8 9 10

Natural Science

Fine Art

3 of 14

Motivation

Evaluation methods help improve legal AI models

But it’s unclear which of the 2 commonly used methods is better

Train/improve model

Pick evaluation method

Collect expert data using method

  • Legal AI models are systematically tested using real tasks
  • Expert feedback is solicited with an “evaluation method”

Memo

—- – -- –

—-- - — –

– - - – —-

Memo 2

- —- – —

- - - —- -

– - – - - –

10/15 points

“Memo 1 is stronger”

Rubric-Based Scoring

Comparative Judgment

Memo 1

— - - –

- — - —-

—-- — —

4 of 14

51 Expert Lawyers Across Various Practice Areas Made This Evaluation Study Possible

The typical contributing lawyer had 10 years of experience

Concentrated in litigation and/or transactions

  • Lawyers from an AmLaw 100 firm, an AmLaw 200 firm, and Snorkel AI completed evaluations
  • >20% were partners/counsel and nearly 40% were senior associates
  • Practice areas most relevant to commercial work
  • 71% experienced in litigation, and 63% in transactions
  • Other areas: regulatory, employment, tax, and IP
  • Advantage of comparative judgments over rubrics held across experience level

Years of Practice

Frequency

8+

63%

12+

33%

20+

14%

5 of 14

The Lawyers Performed More Than 3000 Evaluations On 30 Tasks From Harvey’s BigLaw Bench

  • All tasks had a rubric; some had attached PDF documents
  • Tasks were broadly split across Litigation and Transactions

Example Transactional Task

6 of 14

Ground Truth Quality Levels Were Constructed Using Prompt-Controlled Variation

  • Claude Opus 4.6 given 6�dimensions:�
    • Analytical depth
    • Precision
    • Completeness
    • Reasoning clarity
    • Judgment
    • Nuance
  • 3 quality levels, 50 versions to account for sampling variability
  • 30 x 50 = 1500 generations per quality level

7 of 14

Quality Levels Were Validated By LLM-as-a-Judge And A Practicing BigLaw Partner

  • Observed values correct ~70% of the time (noisy, but not too noisy)�
  • Stratified sample verified by BigLaw partner�
  • Creates limitations to external validity, but no accepted standard to do this

Ground-truth quality generations (Opus 4.6)

Observed quality measurements (GPT-5.4)

8 of 14

50 Versions Of Each Quality Level Were Generated To Reduce Sampling Variability

50

50

50

50

Excellent

Good

Intermediate

Lawyer sees

Rubric

per task

Rubric Evaluations

Preference Evaluations

Quality Level Samples

Lawyer sees

per task

9 of 14

Key Finding

Legal AI work product should be evaluated by asking experienced lawyers which answer they would trust, not just by tallying rubric points.

For model development, preference data from side-by-side comparison may offer better supervision signal than rubric as it captures tacit, holistic quality that rubrics, fixed to predefined criteria, tend to miss.

Rubrics still matter for explaining strengths and weaknesses, but the data suggest they are less reliable as the sole measure of overall legal quality.

Memo

—- – -- –

—-- - — –

– - - – —-

Memo 2

- —- – —

- - - —- -

– - – - - –

10/15 points

“Memo 1 is stronger”

Rubric-Based Scoring

Comparative Judgment

Memo 1

— - - –

- — - —-

—-- — —

10 of 14

We Measure Quality Recovery Two Ways: For Benchmarking And For Individual Judgments

How well the pooled signal recovers the whole ordering

How well a single signal recovers a two-item ordering

BT on prefs

Correlate with truth, aggregate on task

Highly sample size dependent

Correct pick?

Translate rubrics into preferences

Win rate isn’t at task-level

11 of 14

The Advantage For Comparative Judgments Seems To Persist Across Experience Level; Plus CJs Are Twice As Fast

The effect seems to persist across experience level

And comparative judgments take half the time of rubric scoring

CJs better for this lawyer

Rubrics better for this lawyer

Each dot is a lawyer

12 of 14

Comparative Judgment Was Better At Identifying The Higher-Quality Legal AI Work Product Than Rubric Scoring

Comparative judgments put work products in the right quality order

While rubrics do so only�very rarely

  • CJs correctly order better > worse 67% of the time, and recover the full ordering with correlation 0.908
  • Rubrics order better > worse 54% of the time and recover the full ordering with correlation 0.150

*Differences significant under hierarchical bootstrap at at the lawyer, then lawyer-task level

13 of 14

These Results Have Implications For Benchmarkers, Firms, And Labs

How should we evaluate AI models/tools in knowledge work domains?

How can we build on this evaluation to inform research and legal practice?

“Output from Model 1 is stronger”

Output from Model 1

- — - —-

—-- — —

Output from Model 2

- —- – —

- - - —- -

Aggregate judgments from many lawyers

Pick better model

  • Beyond choosing tools, the dataset can make explicit the tacit and holistic markers of quality that experienced lawyers recognize but rarely articulate

Aggregate judgments from LLM autograders

  • Our research suggests that LLM autograders perform similarly to human experts, and could be used as a cost-effective alternative for evaluating models at scale

14 of 14

Questions?