1 of 23

Improving Reliability�in Legal Artificial Intelligence

Harry Surden

Associate Director: Stanford University CodeX Center

Hatfield Chaired Professor of Law and Technology University of Colorado

LLMxLaw Conference, University of Cambridge

27 June 2026

2 of 23

About

Harry Surden

Associate Director Stanford University CodeX Center for Legal Informatics

Professor of Law University of Colorado (Hatfield Chair of Law and Technology)

Faculty Director Silicon Flatirons Center Artificial Intelligence Initiative

Academic Research: Artificial Intelligence and Law

Computable Contracts

Applied Legal Informatics

Background in Computer Science and Law:

Software Engineer for Cisco Systems and Bloomberg LP prior to law

3 of 23

Overview: LLMs and Law

  • Where We Have Been

  • Where We Are Today

  • Harnesses and Frameworks to Improve Reliability

  • Where We Are Likely Going

4 of 23

Where We

Have Been:

The LLM Revolution

(2023 – 2026)

5 of 23

Taking Stock: Three Years In (2023-2026)

  • AI Model Capabilities
    • Models shifted from simple text generation to reasoning, coding, and tool use
    • Agentic workflows can now plan and execute multi-step tasks
    • Harnesses and Frameworks

  • Performance
    • Major improvements in coding, math reasoning, and scientific problem solving
    • Many early benchmarks saturated; harder reasoning benchmarks emerged
    • Rough rule of thumb: ~10% - 20% improvement per year, compounding, on many knowledge tasks

  • Ecosystem
    • Open-source models rapidly caught up with frontier models
    • The frontier/open gap narrowed to roughly 3-6 months

6 of 23

Taking Stock: LLM Revolution (2023-2026)

Specialized Legal LLM / AI Tools

  • Legal research + citations
    • Lexis+ AI / Protégé, Westlaw CoCounsel, vLex Vincent AI

  • Drafting, summarization, document analysis
    • Briefs, memos, contracts, discovery, deposition summaries

  • Firm-specific AI platforms
    • Harvey, Legora, Spellbook, Clearbrief

  • Frontier AI Generalist Systems
    • OpenAI ChatGPT, Anthropic Claude, Google Gemini

7 of 23

Where We Are Today

8 of 23

AI Systems are Remarkably Effective In Law

Strengths of Modern AI Systems

  • They can now perform useful work in:
    • Basic and medium level legal question-answering
    • Document summarization
    • First-draft generation
    • Contract and document analysis
    • Information retrieval
    • Issue spotting
    • Comparison across documents
    • Explanation of complex legal materials

9 of 23

Limits and Reliability Risks

  • Factual errors
    • Hallucinated authorities, quotes, procedural facts
    • Missing, stale, or non-controlling law
    • Overstated confidence

  • Judgment errors
    • “Hard cases” can cause issues
    • Wrong holding scope, Wrong interpretation or analogy
    • Missing factual, procedural, or experiential nuance
    • Bad advice despite plausible form

  • Workflow / reliance errors
    • System lacks key facts
    • Human reviewer cannot easily verify the answer
    • Polished output creates false confidence

10 of 23

Embarrassing and Costly Errors

1,300+

documented cases of AI-fabricated citations

11 of 23

Harnesses and Frameworks to Improve Reliability

12 of 23

What Is a "Harness” for AI Reliability?

Improved Legal Reliability Will Come From

Building Systems or "harnesses" around models:

Legal knowledge bases, verification layers, multi-modal fusion, evaluation datasets, workflow constraints, and human review

13 of 23

What Is a "Harness” for AI Reliability?

  • A harness is the "scaffolding" around the model
      • The systems, tools, checks, and workflows that sit between the user and the LLM.

  • It is a framework that works to improve reliability by controlling input and output:
    • Inputs - screens the query, facts, jurisdiction, and task
    • Sources - retrieves from trusted legal knowledge bases
    • Memory, Skills, and Standards workflows, reusable skills, organizational knowledge, best practices
    • Process - decomposes tasks, uses tools, applies templates
    • Verification - checks citations, quotes, authorities, and reasoning
    • Review - routes high-risk outputs to human lawyers and other AI systems
    • Evaluation - logs failures and tests the system over time

  • General Harness Examples:
      • Claude Code
      • OpenAI Codex

14 of 23

Why are harnesses important for reliability?

  • Legal AI errors can occur at every stage of the workflow.

    • Reliability depends on whether the system can:
      • understand the user’s actual legal task
      • identify the relevant jurisdiction, facts, and procedural posture
      • retrieve the right legal materials
      • ground conclusions in trusted sources
      • check citations, quotations, and authorities
      • flag uncertainty and contrary authority
      • route high-risk outputs to human legal review for judgment or analysis
      • provide systematic methods for detecting errors or failures
      • learn from failures over time

    • A harness helps turn raw AI model output into a controlled legal workflow.

15 of 23

Input Verification

  • Examine the query before answering
    • Jurisdiction and governing law
    • Time period and currency of authority

    • Source type: case, statute, contract, policy, record
    • User goal: research, drafting, negotiation, strategy
    • Sensitive or privileged information

  • What would make this unsafe to rely on?

16 of 23

Grounded Results

  • Make the basis visible
    • Retrieval based upon reliable sources
    • Show the source in addition to answer
    • Separate quoted authority from model inference

  • Mark gaps, uncertainty, and conflicts
    • Keep cited propositions close to the source text
    • Prefer 'I do not know' to unsupported fluency
    • UI Designed to Help Humans catch Errors

17 of 23

Output Verification

  • Have multiple models check the work
    • Model Fusion: multiple models - accept overlap, investigate disagreement

    • Adversarial collaboration: a second system searches for flaws
      • Citation and quote checks
      • Rule/scope checks: does the authority support this proposition?
      • Multi-model comparison for high-stakes answers

18 of 23

Internal Evaluations

  • Evaluations or “Evals”
    • A suite of controlled tests with known correct and incorrect answers
    • Used to test the reliability of a system over time
    • Crucial for scoring and evaluating the reliability of any system
    • Should be judged both by human lawyers and other AI systems

  • Build a repository of mistakes and accurate outputs
    • Gold examples: representative matters with known-good answers
    • Failure bank: hallucinations, bad holdings, bad advice
    • Regression tests: rerun whenever the model, prompts, retrieval, or UI changes

  • Compare to humans: measure relative reliability
    • Humans make mistakes too

19 of 23

Legal Reliability Improved

  • Reduce Errors
    • Better models
    • Better data and retrievals
    • More reliable harnesses

  • Detect Errors
    • Systematic means of detecting errors
    • Detecting Errors from Multiple angles
      • Multiple models
      • Human Review
    • UI Considerations

  • Risk Allocation
    • Extra attention for high-risk workflows
    • Audit trails
    • “Haste makes waste”

20 of 23

Where We Are Going

21 of 23

Where Are We Going (2026 - 2029)

  • Where Are We Going? The Next 1-3 Years
    • Legal AI will become more like a reliability-engineered legal system.

  • Likely Developments
    • More capable reasoning models
      • better at multi-step legal analysis, planning, structured problem solving.
      • continued ~10% - 20% annual improvement, compounding

    • More tool-using systems
      • models connected to legal databases, document repositories
      • citation checkers, calendars, matter files, etc

    • More grounded answers
      • legal conclusions increasingly tied to retrieved statutes,
      • cases, regulations, contracts, internal work product

22 of 23

Where Are We Going (2026 - 2029)

  • Likely Developments
    • More agentic workflows
      • systems that break tasks into steps: retrieve, analyze, draft, critique, verify, revise

    • More verification layers
      • automatic checks for citations, quotations, authority status, jurisdiction, dates, and internal consistency

    • More internal evaluations
      • firms and legal departments building gold-standard examples,
      • known-failure datasets, and regression tests

23 of 23

  • Questions?

  • Contact Info
    • Professor Harry Surden
      • Associate Director Stanford CodeX Center
      • Professor of Law, University of Colorado
      • hsurden@colorado.edu

Questions