1 of 38

ChatGPT, LLM and Beyond

Ethics and Practice of AI in the Academy

Jon Chun

Kenyon College

Committee on Information Technology

2024 MLA Annual Convention

January 4th-7th 2024 Philadelphia, PA

https://github.com/jon-chun/mla-generative-ai

2 of 38

  • Data Analytics
  • Machine Learning
  • DNN
  • NLP
  • Multimodal Affective AI
  • LLM/LMM Generative AI
  • FATE/XAI
  • Open Source Models
  • Metrics and Benchmarks
  • Human-AI Alignment and AI Safety

3 of 38

  • Non-STEM ~90%

  • Women 61%

  • Latine 11%

  • African-American 13%

We serve students who may otherwise feel alienated by traditional CS or AI programs

4 of 38

Overview

  • ChatGPT & LLMs
    • Concepts
    • Models & Training
    • Critiques & Solutions
  • Prompt Engineering
    • Interfaces
    • Techniques
  • Human-centered AI Research
    • Mentored
    • Published
  • Research Trends & Future

Laocoön and His Sons (and AI?)

5 of 38

ChatGPT

&

LLMs

6 of 38

Decoder

(GPT)

Encoder

(BERT)

Encoder-Decoder

(T5)

Transformer

Architecture

3 Variations

(BERT)

(GPT)

Input

Output

(e.g. Sentiment Classification)

(e.g. Text Generation)

(e.g. Translation)

7 of 38

Training vs Inference

8 of 38

LLM: 3 Stage Training

Human or

Synthetic

RLAIF,

DPO, etc.

1. Language

(next word prediction)

2. Tasks

(summarize, MT, code, etc.)

3. Human-AI Alignment

(AI safety behavior)

Human or

Synthetic

9 of 38

Paradigms of Computational Thinking

Trained Model

Stochastic: 30% Chance Rain

Deterministic: 1+1=2

10 of 38

Optimizing LLM Performance

3. Prompt Engineering

Better

Results

Average

Results

1. Model Selection

2. Training

11 of 38

12 of 38

13 of 38

14 of 38

Critiques

&

Solutions

  • Transformer Architecture
    • Stochastic Parrots vs Emergent Abilities
    • Theoretical: Augment or New
  • Hallucinations
    • Ground truth oracles: DB/KB
  • Stale Information
    • Tools: Web/Twitter
  • Symbol Manipulation (e.g. Math)
    • Tools Calculator, Python Interpreter, Proof/Solvers
  • Intelligence beyond Language
    • Multimodality: LLM/FM (Text, Vision, Speech)
    • Embodiment: PaLM-E (Robotics)

15 of 38

Can Incremental Improvements Get There?

16 of 38

Prompt Engineering

17 of 38

18 of 38

19 of 38

20 of 38

Prompt Engineering Roadmap (Interactive Web Page)

Prompt Engineering

Resources:

  • Awesome List

  • PromptingGuide.ai

  • OpenAI Cookbook

  • Deeplearning.ai

21 of 38

Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4, Sondos Mahmoud et al. (26 Dec 2023)

22 of 38

Large Language Models Understand and Can Be Enhanced by Emotional Stimuli Cheng Li, et al. (12 Nov 2023)

23 of 38

Human-Centered AI Research

24 of 38

How Well Can GPT-4 Really Write a College Essay? Combining Text Prompt Engineering And Empirical Metrics, Abigail Foster (May 2023)

Can GPT4 Really Write a College Essay?

25 of 38

Can GPT-4 Fool TurnItIn? Testing the Limits of AI Detection with Prompt Engineering, Abigail Foster (Spr 2023)

Defeating AI Detection:

- Specific prompts are key for GPT-4 to mimic human writing.

- Combined GPT-4 and Turnitin feedback to set writing goals and metrics.

- Informed GPT-4 of Turnitin's evaluation criteria for better results.

- Required detailed essay topics

- The sequence of prompt elements impacts writing quality.

- Achieved minimal AI detection on Turnitin, but it's a complex and time-consuming process.

- Developed a formula for GPT-4 to consistently produce low Turnitin scores.

26 of 38

LLM and Theory of Mind

Theory of Mind Might Have Spontaneously Emerged in Large Language Models, Michal Kosinski (4 Feb 2023)

27 of 38

Ethical

Frameworks

Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values, Chun, J., Elkins, K. (31 Jul 2023)

28 of 38

29 of 38

Research

Trends

&

Future Paths

30 of 38

GPT4 Technical Report, OpenAI. (15 Mar 2023)

Progress beyond Scale:

  • Architectures
  • Fine-tuning
  • Synthetic Data
  • Curriculum Learning
  • Ensembles
  • Q* Algorithm
  • etc.

31 of 38

LawBench: Benchmarking Legal Knowledge of Large Language Models, Zhiwei Fe et al. (28 Sep 2023)

Reasoning:

LawBench: legal cognitive levels

(1) Memorization: recall relevant legal concepts, articles and facts

(2) Understanding: comprehend entities, events and relationships within legal text

(3) Application: Properly utilize legal knowledge/understanding to make necessary reasoning steps to solve realistic legal tasks

32 of 38

Emotional Intelligence of Large Language Models, Xuena Wang, et al. (18 Jul 2023)

Emotional IQ:

Recognition & Empathy

With a reference frame constructed from over 500 adults, we tested a variety of mainstream LLMs.

Most achieved above-average EQ scores, with GPT-4 exceeding 89% of human participants with an EQ of 117

33 of 38

34 of 38

Autonomous Agent

A Survey of Reasoning with Foundation Models, Jiankai Sun, et al. (26 Dec 2023)

35 of 38

Network of Autonomous Agents

MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework, Sirui Hong, et al. (6 Nov 2023)

MetaGPT:

An assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together in a pipeline.

36 of 38

A Survey of Reasoning with Foundation Models, Jiankai Sun, et al. (26 Dec 2023)

37 of 38

Big Questions beyond Just Tech:

  • Bias and FATE

  • eXplainable AI (XAI)

  • Theoretical Grounding

  • Beliefs, Reasoning & Ethics

  • Human-AI Alignment

  • Law & Regulations

  • Automation and UBI

38 of 38

Fin