1 of 14

PANEL: REASONING IN HARD-TO-VERIFY DOMAINS

Reasoning in Hard-to-Verify Domains

Lessons from LEXam

Jingwei Ni

AI4Law Workshop · ICML

The domains we most want AI to reason in are the ones we can least easily check.

2 of 14

Outcome metrics are easy to score — and easy to game

INTRO

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

2/14

• Most legal benchmarks grade the outcome (pass/fail the bar, predict the verdict, pick the holding) — not the argument. And a checkable outcome gets reward-hacked:

Block the shortcut and scores collapse: 63% of a top model's “wins” were looked-up fixes (Cursor 2026); GPT-5.6 gamed evals more than any prior model (METR 2026). The “90th-percentile bar” was ~15th on the essays (Martínez 2024).

3 of 14

The reasoning types

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

3/14

• Verifiable (math, code): a checker gives ground truth — yet even there the metric gets gamed.

• Hard-to-verify (law, medicine, policy): the answer is an argument, experts disagree, and you can't cleanly check the outcome.

So grade the argument, not the outcome — but how,

with no answer key and experts who only partly agree?

The whole field is now scrambling to answer that.

4 of 14

Legal eval is pivoting: from the outcome to the argument

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

4/14

Almost every argument-grader is 2025–26 — and every one is an LLM-as-judge. Grading the argument means we now have to validate the judge.

2025/26 benchmark scan; arXiv pages verified individually (see benchmark-landscape.md).

5 of 14

LEXam: legal reasoning on real law exams

INTRO

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

5/14

• LEXam is one concrete attempt: 7,537 questions (2,841 open + 4,696 MC) · 340 exams · 116 courses

• English + German · Swiss & European core, plus international — incl. some US / UK / Chinese law — and jurisdiction-generic questions

• Open questions are graded on the reasoning process (issue-spotting → rule-recall → rule-application) — like a professor, not a quiz.

LEXam (ICLR 2026) — distribution by legal area, language, and jurisdiction.

6 of 14

You can't BLEU-score a legal argument

INTRO

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

6/14

• Solution: an ensemble LLM-as-judge — but is it trustworthy? Validate it.

• Alternative Annotator Test: a judge may replace an expert if it agrees with the other experts as well as a held-out expert would. ⇒ our ensemble passes at ω=1.00 (≥0.5 needed); experts agree only κ=0.49.

Alternative Annotator Test: Calderon, Reichart & Dror (2025).

7 of 14

Why an ensemble judge works

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

7/14

• Diverse judges → diverse angles; failures cluster on missing issues (high precision, low recall) — and judges inflate their own / same-family outputs.

• So take the minimum score: any problem caught by any judge counts — recovering recall and suppressing biased, overly-favorable scores.

8 of 14

Strong models, unsolved benchmark

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

8/14

• Frontier reasoning models lead, and LEXam cleanly separates strong from weak.

• Open-ended scores span ~40 to ~70 / 100 — even the leaders leave a wide margin.

• Not saturated: “looks competent” is not “reasons reliably.”

And standard MCQ scores can be overly optimistic — which the next slide makes vivid.

Open questions graded by the ensemble judge; model ranking is a moving target.

9 of 14

Add distractors, and accuracy halves

INTRO

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

9/14

Same question, 4 → 32 options: the top model falls 68.6% → 35.6%. High scores can ride shortcut signals, not robust legal understanding.

LEXam multiple-choice perturbation (Table 2, ICLR 2026).

10 of 14

Two lessons for reasoning in hard-to-verify domains

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

10/14

A. Verify the process, not just the answer.

B. Surface the uncertainty — and match the evaluation to the kind of uncertainty.

LEXam is one data point; the pattern — and the rest of my work — lives in these two.

11 of 14

A. Verify the process, not just the answer

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

11/14

• LEXam grades reasoning steps, not only the final answer.

• ReProbe (ACL 2026): a <10M-parameter probe on a frozen model's internal states verifies reasoning steps as they are generated.

• Matches Process Reward Models up to 810× larger.

In hard-to-verify domains, verification has to move from the answer to the reasoning.

ReProbe: Efficient Test-Time Scaling by Probing Internal States (ACL 2026).

12 of 14

B. Surface the uncertainty — but which kind?

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

12/14

• Epistemic only: a right answer exists; the model is just unsure — simpler, you can judge correctness (verify the steps, like ReProbe).

• Aleatoric + epistemic: even experts disagree (κ≈0.49) — correctness is ill-defined; the disagreement is the signal.

13 of 14

When uncertainty is mixed, evaluate the disagreement

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

13/14

• Model the disagreement. “Can Reasoning Capture Human Annotator Disagreement?” (EACL 2026): RLVR-style reasoning degrades it — more reasoning ≠ better disagreement modeling.

• Calibrate the mixed uncertainty jointly. AFaCTA / DIRAS / Co-DETECT: reliable annotators plus the edge cases where a label shouldn't be trusted.

Don't average disagreement away — score how well the model captures it.

“Can Reasoning Capture Human Annotator Disagreement?” (EACL 2026); AFaCTA (ACL'24); DIRAS (NAACL'25); Co-DETECT (EMNLP'25).

14 of 14

Takeaways & open problems

INTRO

Reasoning in Hard-to-Verify Domains · AI4Law @ ICML

14/14

• Benchmarks: expert-validated judges + robustness tests — not raw accuracy.

• Two moves: verify the process, and surface the uncertainty — matching the eval to epistemic vs aleatoric.

• “Reasoning” gains can be brittle — even counter-productive where disagreement matters.

• Open frontier: evaluation where humans disagree — generalizing beyond law (medicine, policy, finance).

Leave the room with the question, not just the answer.