1 of 23

Improving Reliability�in Legal Artificial Intelligence�

Harry Surden

Associate Director: Stanford University CodeX Center

Hatfield Chaired Professor of Law and Technology University of Colorado

​

LLMxLaw Conference, University of Cambridge

27 June 2026

2 of 23

About

Harry Surden

​

Associate Director Stanford University CodeX Center for Legal Informatics

Professor of Law University of Colorado (Hatfield Chair of Law and Technology)

Faculty Director Silicon Flatirons Center Artificial Intelligence Initiative

​

Academic Research: Artificial Intelligence and Law

Computable Contracts

Applied Legal Informatics

​

Background in Computer Science and Law:

Software Engineer for Cisco Systems and Bloomberg LP prior to law

​

​

​

​

​

3 of 23

Overview: LLMs and Law

​

  • Where We Have Been

​

  • Where We Are Today

​

  • Harnesses and Frameworks to Improve Reliability

​

  • Where We Are Likely Going

4 of 23

Where We

Have Been:

The LLM Revolution

(2023 – 2026)

5 of 23

Taking Stock: Three Years In (2023-2026)

​

  • AI Model Capabilities
    • Models shifted from simple text generation to reasoning, coding, and tool use
    • Agentic workflows can now plan and execute multi-step tasks
    • Harnesses and Frameworks

​

  • Performance
    • Major improvements in coding, math reasoning, and scientific problem solving
    • Many early benchmarks saturated; harder reasoning benchmarks emerged
    • Rough rule of thumb: ~10% - 20% improvement per year, compounding, on many knowledge tasks

​

  • Ecosystem
    • Open-source models rapidly caught up with frontier models
    • The frontier/open gap narrowed to roughly 3-6 months

6 of 23

Taking Stock: LLM Revolution (2023-2026)

​

Specialized Legal LLM / AI Tools

  • Legal research + citations
    • Lexis+ AI / Protégé, Westlaw CoCounsel, vLex Vincent AI

​

  • Drafting, summarization, document analysis
    • Briefs, memos, contracts, discovery, deposition summaries

​

  • Firm-specific AI platforms
    • Harvey, Legora, Spellbook, Clearbrief

​

  • Frontier AI Generalist Systems
    • OpenAI ChatGPT, Anthropic Claude, Google Gemini

​

7 of 23

Where We Are Today

8 of 23

AI Systems are Remarkably Effective In Law

​

Strengths of Modern AI Systems

  • They can now perform useful work in:
    • Basic and medium level legal question-answering
    • Document summarization
    • First-draft generation
    • Contract and document analysis
    • Information retrieval
    • Issue spotting
    • Comparison across documents
    • Explanation of complex legal materials

​

​

9 of 23

Limits and Reliability Risks

​

  • Factual errors
    • Hallucinated authorities, quotes, procedural facts
    • Missing, stale, or non-controlling law
    • Overstated confidence

​

  • Judgment errors
    • “Hard cases” can cause issues
    • Wrong holding scope, Wrong interpretation or analogy
    • Missing factual, procedural, or experiential nuance
    • Bad advice despite plausible form

​

  • Workflow / reliance errors
    • System lacks key facts
    • Human reviewer cannot easily verify the answer
    • Polished output creates false confidence

10 of 23

Embarrassing and Costly Errors

1,300+

documented cases of AI-fabricated citations

11 of 23

Harnesses and Frameworks to Improve Reliability

12 of 23

What Is a "Harness” for AI Reliability?

​

Improved Legal Reliability Will Come From

Building Systems or "harnesses" around models:

​

Legal knowledge bases, verification layers, multi-modal fusion, evaluation datasets, workflow constraints, and human review

​

13 of 23

What Is a "Harness” for AI Reliability?

  • A harness is the "scaffolding" around the model
      • The systems, tools, checks, and workflows that sit between the user and the LLM.

​

  • It is a framework that works to improve reliability by controlling input and output:
    • Inputs - screens the query, facts, jurisdiction, and task
    • Sources - retrieves from trusted legal knowledge bases
    • Memory, Skills, and Standards – workflows, reusable skills, organizational knowledge, best practices
    • Process - decomposes tasks, uses tools, applies templates
    • Verification - checks citations, quotes, authorities, and reasoning
    • Review - routes high-risk outputs to human lawyers and other AI systems
    • Evaluation - logs failures and tests the system over time

​

  • General Harness Examples:
      • Claude Code
      • OpenAI Codex

​

14 of 23

Why are harnesses important for reliability?

  • Legal AI errors can occur at every stage of the workflow.

​

    • Reliability depends on whether the system can:
      • understand the user’s actual legal task
      • identify the relevant jurisdiction, facts, and procedural posture
      • retrieve the right legal materials
      • ground conclusions in trusted sources
      • check citations, quotations, and authorities
      • flag uncertainty and contrary authority
      • route high-risk outputs to human legal review for judgment or analysis
      • provide systematic methods for detecting errors or failures
      • learn from failures over time

​

    • A harness helps turn raw AI model output into a controlled legal workflow.

15 of 23

Input Verification

  • Examine the query before answering
    • Jurisdiction and governing law
    • Time period and currency of authority

​

    • Source type: case, statute, contract, policy, record
    • User goal: research, drafting, negotiation, strategy
    • Sensitive or privileged information

​

  • What would make this unsafe to rely on?

16 of 23

Grounded Results

  • Make the basis visible
    • Retrieval based upon reliable sources
    • Show the source in addition to answer
    • Separate quoted authority from model inference

​

  • Mark gaps, uncertainty, and conflicts
    • Keep cited propositions close to the source text
    • Prefer 'I do not know' to unsupported fluency
    • UI Designed to Help Humans catch Errors

17 of 23

Output Verification

  • Have multiple models check the work
    • Model Fusion: multiple models - accept overlap, investigate disagreement

​

    • Adversarial collaboration: a second system searches for flaws
      • Citation and quote checks
      • Rule/scope checks: does the authority support this proposition?
      • Multi-model comparison for high-stakes answers

18 of 23

Internal Evaluations

  • Evaluations or “Evals”
    • A suite of controlled tests with known correct and incorrect answers
    • Used to test the reliability of a system over time
    • Crucial for scoring and evaluating the reliability of any system
    • Should be judged both by human lawyers and other AI systems

​

  • Build a repository of mistakes and accurate outputs
    • Gold examples: representative matters with known-good answers
    • Failure bank: hallucinations, bad holdings, bad advice
    • Regression tests: rerun whenever the model, prompts, retrieval, or UI changes

​

  • Compare to humans: measure relative reliability
    • Humans make mistakes too

19 of 23

Legal Reliability Improved

  • Reduce Errors
    • Better models
    • Better data and retrievals
    • More reliable harnesses

​

  • Detect Errors
    • Systematic means of detecting errors
    • Detecting Errors from Multiple angles
      • Multiple models
      • Human Review
    • UI Considerations

​

  • Risk Allocation
    • Extra attention for high-risk workflows
    • Audit trails
    • “Haste makes waste”

20 of 23

Where We Are Going

21 of 23

Where Are We Going (2026 - 2029)

  • Where Are We Going? The Next 1-3 Years
    • Legal AI will become more like a reliability-engineered legal system.

​

  • Likely Developments
    • More capable reasoning models
      • better at multi-step legal analysis, planning, structured problem solving.
      • continued ~10% - 20% annual improvement, compounding

​

    • More tool-using systems
      • models connected to legal databases, document repositories
      • citation checkers, calendars, matter files, etc

​

    • More grounded answers
      • legal conclusions increasingly tied to retrieved statutes,
      • cases, regulations, contracts, internal work product

22 of 23

Where Are We Going (2026 - 2029)

  • Likely Developments
    • More agentic workflows
      • systems that break tasks into steps: retrieve, analyze, draft, critique, verify, revise

​

    • More verification layers
      • automatic checks for citations, quotations, authority status, jurisdiction, dates, and internal consistency

​

    • More internal evaluations
      • firms and legal departments building gold-standard examples,
      • known-failure datasets, and regression tests

23 of 23

​

  • Questions?

​

  • Contact Info
    • Professor Harry Surden
      • Associate Director Stanford CodeX Center
      • Professor of Law, University of Colorado
      • hsurden@colorado.edu

​

​

Questions