PANEL: REASONING IN HARD-TO-VERIFY DOMAINS
Reasoning in Hard-to-Verify Domains
Lessons from LEXam
Jingwei Ni
AI4Law Workshop · ICML
The domains we most want AI to reason in are the ones we can least easily check.
Outcome metrics are easy to score — and easy to game
INTRO
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
2/14
• Most legal benchmarks grade the outcome (pass/fail the bar, predict the verdict, pick the holding) — not the argument. And a checkable outcome gets reward-hacked:
Block the shortcut and scores collapse: 63% of a top model's “wins” were looked-up fixes (Cursor 2026); GPT-5.6 gamed evals more than any prior model (METR 2026). The “90th-percentile bar” was ~15th on the essays (Martínez 2024).
The reasoning types
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
3/14
• Verifiable (math, code): a checker gives ground truth — yet even there the metric gets gamed.
• Hard-to-verify (law, medicine, policy): the answer is an argument, experts disagree, and you can't cleanly check the outcome.
So grade the argument, not the outcome — but how,
with no answer key and experts who only partly agree?
The whole field is now scrambling to answer that.
Legal eval is pivoting: from the outcome to the argument
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
4/14
Almost every argument-grader is 2025–26 — and every one is an LLM-as-judge. Grading the argument means we now have to validate the judge.
2025/26 benchmark scan; arXiv pages verified individually (see benchmark-landscape.md).
LEXam: legal reasoning on real law exams
INTRO
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
5/14
• LEXam is one concrete attempt: 7,537 questions (2,841 open + 4,696 MC) · 340 exams · 116 courses
• English + German · Swiss & European core, plus international — incl. some US / UK / Chinese law — and jurisdiction-generic questions
• Open questions are graded on the reasoning process (issue-spotting → rule-recall → rule-application) — like a professor, not a quiz.
LEXam (ICLR 2026) — distribution by legal area, language, and jurisdiction.
You can't BLEU-score a legal argument
INTRO
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
6/14
• Solution: an ensemble LLM-as-judge — but is it trustworthy? Validate it.
• Alternative Annotator Test: a judge may replace an expert if it agrees with the other experts as well as a held-out expert would. ⇒ our ensemble passes at ω=1.00 (≥0.5 needed); experts agree only κ=0.49.
Alternative Annotator Test: Calderon, Reichart & Dror (2025).
Why an ensemble judge works
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
7/14
• Diverse judges → diverse angles; failures cluster on missing issues (high precision, low recall) — and judges inflate their own / same-family outputs.
• So take the minimum score: any problem caught by any judge counts — recovering recall and suppressing biased, overly-favorable scores.
Strong models, unsolved benchmark
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
8/14
• Frontier reasoning models lead, and LEXam cleanly separates strong from weak.
• Open-ended scores span ~40 to ~70 / 100 — even the leaders leave a wide margin.
• Not saturated: “looks competent” is not “reasons reliably.”
And standard MCQ scores can be overly optimistic — which the next slide makes vivid.
Open questions graded by the ensemble judge; model ranking is a moving target.
Add distractors, and accuracy halves
INTRO
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
9/14
Same question, 4 → 32 options: the top model falls 68.6% → 35.6%. High scores can ride shortcut signals, not robust legal understanding.
LEXam multiple-choice perturbation (Table 2, ICLR 2026).
Two lessons for reasoning in hard-to-verify domains
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
10/14
A. Verify the process, not just the answer.
B. Surface the uncertainty — and match the evaluation to the kind of uncertainty.
LEXam is one data point; the pattern — and the rest of my work — lives in these two.
A. Verify the process, not just the answer
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
11/14
• LEXam grades reasoning steps, not only the final answer.
• ReProbe (ACL 2026): a <10M-parameter probe on a frozen model's internal states verifies reasoning steps as they are generated.
• Matches Process Reward Models up to 810× larger.
In hard-to-verify domains, verification has to move from the answer to the reasoning.
ReProbe: Efficient Test-Time Scaling by Probing Internal States (ACL 2026).
B. Surface the uncertainty — but which kind?
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
12/14
• Epistemic only: a right answer exists; the model is just unsure — simpler, you can judge correctness (verify the steps, like ReProbe).
• Aleatoric + epistemic: even experts disagree (κ≈0.49) — correctness is ill-defined; the disagreement is the signal.
When uncertainty is mixed, evaluate the disagreement
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
13/14
• Model the disagreement. “Can Reasoning Capture Human Annotator Disagreement?” (EACL 2026): RLVR-style reasoning degrades it — more reasoning ≠ better disagreement modeling.
• Calibrate the mixed uncertainty jointly. AFaCTA / DIRAS / Co-DETECT: reliable annotators plus the edge cases where a label shouldn't be trusted.
Don't average disagreement away — score how well the model captures it.
“Can Reasoning Capture Human Annotator Disagreement?” (EACL 2026); AFaCTA (ACL'24); DIRAS (NAACL'25); Co-DETECT (EMNLP'25).
Takeaways & open problems
INTRO
Reasoning in Hard-to-Verify Domains · AI4Law @ ICML
14/14
• Benchmarks: expert-validated judges + robustness tests — not raw accuracy.
• Two moves: verify the process, and surface the uncertainty — matching the eval to epistemic vs aleatoric.
• “Reasoning” gains can be brittle — even counter-productive where disagreement matters.
• Open frontier: evaluation where humans disagree — generalizing beyond law (medicine, policy, finance).
Leave the room with the question, not just the answer.