ChatGPT, LLM and Beyond
Ethics and Practice of AI in the Academy
Jon Chun
Kenyon College
Committee on Information Technology
2024 MLA Annual Convention
January 4th-7th 2024 Philadelphia, PA
https://github.com/jon-chun/mla-generative-ai
We serve students who may otherwise feel alienated by traditional CS or AI programs
Overview
Laocoön and His Sons (and AI?)
ChatGPT
&
LLMs
Decoder
(GPT)
Encoder
(BERT)
Encoder-Decoder
(T5)
Transformer
Architecture
3 Variations
(BERT)
(GPT)
Input
Output
(e.g. Sentiment Classification)
(e.g. Text Generation)
(e.g. Translation)
Training vs Inference
LLM: 3 Stage Training
Human or
Synthetic
RLAIF,
DPO, etc.
1. Language
(next word prediction)
2. Tasks
(summarize, MT, code, etc.)
3. Human-AI Alignment
(AI safety behavior)
Human or
Synthetic
Paradigms of Computational Thinking
Trained Model
Stochastic: 30% Chance Rain
Deterministic: 1+1=2
Optimizing LLM Performance
3. Prompt Engineering
Better
Results
Average
Results
1. Model Selection
2. Training
Critiques
&
Solutions
Can Incremental Improvements Get There?
Prompt Engineering
Prompt Engineering Roadmap (Interactive Web Page)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4, Sondos Mahmoud et al. (26 Dec 2023)
Large Language Models Understand and Can Be Enhanced by Emotional Stimuli Cheng Li, et al. (12 Nov 2023)
Human-Centered AI Research
How Well Can GPT-4 Really Write a College Essay? Combining Text Prompt Engineering And Empirical Metrics, Abigail Foster (May 2023)
Can GPT4 Really Write a College Essay?
Can GPT-4 Fool TurnItIn? Testing the Limits of AI Detection with Prompt Engineering, Abigail Foster (Spr 2023)
Defeating AI Detection:
- Specific prompts are key for GPT-4 to mimic human writing.
- Combined GPT-4 and Turnitin feedback to set writing goals and metrics.
- Informed GPT-4 of Turnitin's evaluation criteria for better results.
- Required detailed essay topics
- The sequence of prompt elements impacts writing quality.
- Achieved minimal AI detection on Turnitin, but it's a complex and time-consuming process.
- Developed a formula for GPT-4 to consistently produce low Turnitin scores.
LLM and Theory of Mind
Theory of Mind Might Have Spontaneously Emerged in Large Language Models, Michal Kosinski (4 Feb 2023)
Ethical
Frameworks
Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values, Chun, J., Elkins, K. (31 Jul 2023)
Research
Trends
&
Future Paths
GPT4 Technical Report, OpenAI. (15 Mar 2023)
Progress beyond Scale:
LawBench: Benchmarking Legal Knowledge of Large Language Models, Zhiwei Fe et al. (28 Sep 2023)
Reasoning:
LawBench: legal cognitive levels
(1) Memorization: recall relevant legal concepts, articles and facts
(2) Understanding: comprehend entities, events and relationships within legal text
(3) Application: Properly utilize legal knowledge/understanding to make necessary reasoning steps to solve realistic legal tasks
Emotional Intelligence of Large Language Models, Xuena Wang, et al. (18 Jul 2023)
Emotional IQ:
Recognition & Empathy
With a reference frame constructed from over 500 adults, we tested a variety of mainstream LLMs.
Most achieved above-average EQ scores, with GPT-4 exceeding 89% of human participants with an EQ of 117
Autonomous Agent
A Survey of Reasoning with Foundation Models, Jiankai Sun, et al. (26 Dec 2023)
Network of Autonomous Agents
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework, Sirui Hong, et al. (6 Nov 2023)
MetaGPT:
An assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together in a pipeline.
A Survey of Reasoning with Foundation Models, Jiankai Sun, et al. (26 Dec 2023)
Big Questions beyond Just Tech:
Fin