JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, Julian Nyarko
Stanford Law School, Snorkel AI, Harvey
Is Legal Practice More Like Science Or Art?
1 2 3 4 5 6 7 8 9 10
Natural Science
Fine Art
Motivation
Evaluation methods help improve legal AI models
But it’s unclear which of the 2 commonly used methods is better
Train/improve model
Pick evaluation method
Collect expert data using method
Memo
—- – -- –
—-- - — –
– - - – —-
Memo 2
- —- – —
- - - —- -
– - – - - –
10/15 points
“Memo 1 is stronger”
Rubric-Based Scoring
Comparative Judgment
Memo 1
— - - –
- — - —-
—-- — —
51 Expert Lawyers Across Various Practice Areas Made This Evaluation Study Possible
The typical contributing lawyer had 10 years of experience
Concentrated in litigation and/or transactions
Years of Practice | Frequency |
8+ | 63% |
12+ | 33% |
20+ | 14% |
The Lawyers Performed More Than 3000 Evaluations On 30 Tasks From Harvey’s BigLaw Bench
Example Transactional Task
Ground Truth Quality Levels Were Constructed Using Prompt-Controlled Variation
Quality Levels Were Validated By LLM-as-a-Judge And A Practicing BigLaw Partner
Ground-truth quality generations (Opus 4.6)
Observed quality measurements (GPT-5.4)
50 Versions Of Each Quality Level Were Generated To Reduce Sampling Variability
50
50
50
50
Excellent
Good
Intermediate
Lawyer sees
Rubric
per task
Rubric Evaluations
Preference Evaluations
Quality Level Samples
Lawyer sees
per task
Key Finding
Legal AI work product should be evaluated by asking experienced lawyers which answer they would trust, not just by tallying rubric points.
For model development, preference data from side-by-side comparison may offer better supervision signal than rubric as it captures tacit, holistic quality that rubrics, fixed to predefined criteria, tend to miss.
Rubrics still matter for explaining strengths and weaknesses, but the data suggest they are less reliable as the sole measure of overall legal quality.
Memo
—- – -- –
—-- - — –
– - - – —-
Memo 2
- —- – —
- - - —- -
– - – - - –
10/15 points
“Memo 1 is stronger”
Rubric-Based Scoring
Comparative Judgment
Memo 1
— - - –
- — - —-
—-- — —
We Measure Quality Recovery Two Ways: For Benchmarking And For Individual Judgments
How well the pooled signal recovers the whole ordering
How well a single signal recovers a two-item ordering
BT on prefs
Correlate with truth, aggregate on task
Highly sample size dependent
Correct pick?
Translate rubrics into preferences
Win rate isn’t at task-level
The Advantage For Comparative Judgments Seems To Persist Across Experience Level; Plus CJs Are Twice As Fast
The effect seems to persist across experience level
And comparative judgments take half the time of rubric scoring
CJs better for this lawyer
Rubrics better for this lawyer
Each dot is a lawyer
Comparative Judgment Was Better At Identifying The Higher-Quality Legal AI Work Product Than Rubric Scoring
Comparative judgments put work products in the right quality order
While rubrics do so only�very rarely
*Differences significant under hierarchical bootstrap at at the lawyer, then lawyer-task level
These Results Have Implications For Benchmarkers, Firms, And Labs
How should we evaluate AI models/tools in knowledge work domains?
How can we build on this evaluation to inform research and legal practice?
“Output from Model 1 is stronger”
Output from Model 1
- — - —-
—-- — —
Output from Model 2
- —- – —
- - - —- -
Aggregate judgments from many lawyers
Pick better model
Aggregate judgments from LLM autograders
Questions?