1 of 23

AIDA �DESIGN & GOVERNANCE�

ARTIFICIAL INTELLIGENCE DRAFTING ASSISTANT

2 of 23

HOW AIDA HELPS

  • Consistent application of laws and rules
  • Reduces administrative burden
    • Assists with grammar
    • Includes legal citations
  • Produces draft based on examiner’s decision, the facts and applicable laws
  • Indicates disagreement, with escalation

3 of 23

GUARDRAILS

  • Human enters findings of fact
    • AIDA does not rely on transcript analysis
  • Human enters the decision
    • AIDA does not make a recommendation
    • AIDA will produce a flag if it disagrees with the examiner’s decision, which is then escalated to the SAE for review
  • Governance
    • Training, including ethical use of AI
    • Feedback informs improvements
    • LLM selection
    • Mitigating automation bias

4 of 23

TRANSCRIPTION BIAS*

  • Performance is unreliable
    • Research show significant variability based on audio quality, speaker demographics, accent, and gender.
  • Overall Accuracy
    • Best case:  Whisper 1.75% Word Error Rate (WER), or 0.5% when ignoring minor formatting differences
    • Typical performance: 10-12% WER for native English speakers under standard conditions
  • Key Error Type
    • Hallucinations: Low risk but LLMs invent phrases approximately 1% of the time

* Research gathered by Zarak Khan, New Jersey Innovation Authority

5 of 23

TRANSCRIPTION DISPARITIES*

  • Race and Ethnicity
    • Legacy systems showed significant bias (35% WER for Black speakers vs. 19% for White speakers)
    • Modern models like Whisper have narrowed but not eliminated this gap
  • Accent and Language Background
    • North American English: Most accurate
    • British/Australian English: Moderately accurate
    • Non-native speakers: Less accurate, with Vietnamese and Thai accents showing highest error rates
  • Gender
    • Female speakers: ~11.7% WER
    • Male speakers: ~8.0% WER
    • Indicates measurable gender-based performance gap

* Research gathered by Zarak Khan, New Jersey Innovation Authority

6 of 23

LLM SELECTION

  • Quality
    • This refers to the accuracy, effectiveness, and overall performance of the LLM in completing the desired tasks.
  • Latency
    • Measures the time it takes for the LLM to process and respond to a query. Lower latency is crucial for real-time applications.
  • Cost
    • The pricing models of LLM providers can vary significantly. Consider the upfront costs, ongoing subscription fees, and any usage-based charges.
  • Context size
    • This defines the amount of preceding text the LLM can analyze to understand the context of a query. Al larger window can lead to better comprehension but may also impact latency.

The LLM landscape is rapidly evolving. Providers are continually improving their models, leading to shifts in quality, latency, and cost over time.

7 of 23

AUTOMATION BIAS

  • Uncritical Acceptance:
    • Users may accept AI recommendations without question, even where contradict training, believing outputs are fully accurate
    • Reality is algorithms have flaws, complexity of the situation may not have been contemplated (multiple vs. single issue)
  • Not new:
    • Observed in fields like aviation, finance and healthcare – industry with prevalent systems
  • Risk:
    • Increases in environments where decision has high impact

8 of 23

MITIGATING AUTOMATION BIAS

  • Human judgment paramount
  • AI training, focused on AI limitations and the importance of critical thinking
  • AI as an intern
  • Question. It. Always.

9 of 23

MANAGING AUTOMATION BIAS

  • Design principles that help user engage critically with the AI tool (flags, review, resources)
  • Error detection
  • Feedback mechanisms – continued improvement loops
  • Build in human oversight
    • Regular monitoring/audit/adjustment
    • Authority to override
    • Regular retraining
    • Diversity in review

10 of 23

DECISION FLAG

  • If AIDA does not feel it can draft a decision based on the facts as applied to legal authority it will issue a flag.
  • That case will then be escalated to SAEs to review and make a final determination.

11 of 23

AE enters findings of fact and their decision into AIDA

Decision

Findings of Fact

If AIDA cannot draft the letter based on the AE’s decision, it will produce an SAE review alert

AIDA produces

a draft decision letter

AE reviews and edits

decision letter

AE submits final decision to Salesforce for

mailing to parties

AE conducts hearing

AIDA Workflow

12 of 23

EVALUATION & FEEDBACK

  • Providing feedback is essential to the health and development of AIDA
  • A mandatory starred rating system has been integrated into the tool
  • Examiners are required to rate each output on a scale of 1 to 5.
  • A feedback box is also used to capture your explanations of the issues.
    • If AIDA gets it wrong it’s important to explain what was incorrect about the decision

13 of 23

FEEDBACK RATINGS

14 of 23

ETHICAL USE AND GOVERNANCE OF AI

  • Importance of ethical AI development and use
  • Key ethical considerations:
    • Bias and fairness
    • Transparency and explainability
    • Privacy and data protection
    • Accountability and responsibility
  • Organization policies & guidelines

15 of 23

16 of 23

17 of 23

18 of 23

19 of 23

20 of 23

21 of 23

22 of 23

SUMMARY

AIDA…

  • Human centered
  • Uses Gen-AI to help draft decisions
  • Requires examiner input and validation to function
  • Requires oversite
  • Supports does not replace

AIDA is not…

  • Making the decisions
  • Perfect
    • Limited issues
    • Needs care & correction
    • Needs human in the loop
  • The source of truth for any legal interpretations or decisions
  • Collecting or exposing personal information

23 of 23

THANK YOU!