1 of 37

Taming the Wild West

of LLMs

Nazneen Rajani | Research Lead @ Hugging Face | nazneen@hf.co | @nazneenrajani

2 of 37

Text-to-Text Foundation Models since GPT3

GPT-3

2021

Jun

Oct

PaLM

Chinchilla

OPT

BLOOM

Gopher

2022

Megatron TNLG

Dec

May

Apr

Jul

Jul

GPT-J

ChatGPT

Nov

Dec

Galactica

GPT-Neo

Jun

GPT-NeoX

Feb

Flan-T5

Oct

*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1

UL2

Cohere

Jurassic

Anthropic

2023

Feb

LLaMA

Flan-UL2

March

GPT-4

Falcon

May

INCITE

StarCoder

LLaMA-2

3 of 37

Text-to-Text Foundation Models since GPT3

GPT-3

2021

Jun

Oct

PaLM

Chinchilla

OPT

BLOOM

Gopher

2022

Megatron TNLG

Dec

May

Apr

Jul

Jul

GPT-J

ChatGPT

Nov

Dec

Galactica

GPT-Neo

Jun

GPT-NeoX

Feb

Flan-T5

Oct

*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1

UL2

Cohere

Jurassic

Anthropic

2023

Feb

LLaMA

Flan-UL2

March

GPT-4

Falcon

May

INCITE

StarCoder

LLaMA-2

4 of 37

Model Access

🔓

🔒

Open access

Closed access

Limited access

🔐

5 of 37

🔓 Open Access Models

Model components are publicly available:

  • Open source code
  • Training data
    • Sources and their distribution
    • Data preprocessing and curation steps
  • Model weights
  • Paper or blog summarizing
    • Architecture and training details
    • Evaluation results
    • Adaptation to the model
      • Safety filters
      • Training with human feedback

6 of 37

🔓 Open Access Models

Allows reproducing results and replicating parts of the model

Enable auditing and conducting risk analysis

Serves as a research artifact

Enables interpreting model output

7 of 37

🔒 Closed Access Models

Only research paper or blog is available and may include overview of

  • Training data
  • Architecture and training details (including infrastructure)
  • Evaluation results
  • Adaptation to the model
    • Safety filters
    • Training with human feedback

8 of 37

🔒 Closed Access Models

Safety concerns

Competitive advantage

Expensive to setup guardrails for safe access

9 of 37

🔐 Limited Access Models

Available for use via:

  • API
  • Call for research proposals

10 of 37

Text-to-Text Foundation Models since GPT3

GPT-3

2021

Jun

Oct

PaLM

Chinchilla

OPT

BLOOM

Gopher

2022

Megatron TNLG

Dec

May

Apr

Jul

Jul

GPT-J

ChatGPT

Nov

Dec

Galactica

GPT-Neo

Jun

GPT-NeoX

Feb

Flan-T5

Oct

*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1

UL2

Cohere

Jurassic

Anthropic

2023

Feb

LLaMA

Flan-UL2

March

GPT-4

Falcon

May

INCITE

StarCoder

LLaMA-2

11 of 37

Pivotal moments

  • Meta’s LLaMA/LLaMA2
  • Together’s Red Pajama
  • LAION’s Open Assistant
  • AI2’s Dolma

12 of 37

Large Language Models – Training

  1. Pretraining the LM
    • Predicting the next token
    • Eg: GPT-3, OPT, BLOOM, LLaMA, Falcon, LLaMA 2
  2. Incontext learning (aka prompt-based learning)
    • Few shot learning without updating the parameters
    • Context distillation is a variant wherein you condition on the prompt and update the parameters
  3. Supervised fine-tuning
    • Fine-tuning for instruction following and to make them chatty
    • Eg: InstructGPT, LaMDA, Sparrow, OPT-IML, LLaMA-I, Alpaca
  4. Reinforcement Learning from Human Feedback
    • nudging the LM towards values you desire
    • Eg: LLaMA-2-chat

13 of 37

Large Language Models – Training

  • Pretraining the LM
    • Predicting the next token
    • Eg: GPT-3, OPT, BLOOM, LLaMA, Falcon, LLaMA 2
  • Incontext learning (aka prompt-based learning)
    • Few shot learning without updating the parameters
    • Context distillation is a variant wherein you condition on the prompt and update the parameters
  • Supervised fine-tuning
    • Fine-tuning for instruction following and to make them chatty
    • Eg: InstructGPT, LaMDA, Sparrow, OPT-IML, LLaMA-I, Alpaca
  • Reinforcement Learning from Human Feedback
    • nudging the LM towards values you desire
    • Eg: LLaMA-2-chat

Training a chatbot

14 of 37

Evaluating a Chatbot

15 of 37

Evaluating a Chatbot

  1. Pretraining the LM
    1. Predicting the next token
    2. Eg: GPT-3, BLOOM
  2. Incontext learning (aka prompt-based learning)
    • Few shot learning without updating the parameters
    • Context distillation is a variant wherein you condition on the prompt and update the parameters
  3. Supervised fine-tuning
    • Fine-tuning for instruction following and to make them chatty
    • Eg: InstructGPT, LaMDA, Sparrow, OPT-IML, LLaMA-I, Alpaca, Vicuna, Guanaco
  4. Reinforcement Learning from Human Feedback
    • safety/alignment
    • nudging the LM towards values you desire

16 of 37

Evaluating a Chatbot

Ouyang, Long, et al. "Training language models to follow instructions with human feedback." arXiv preprint arXiv:2203.02155 (2022).

Evaluating instruction following/chatty-ness

Evaluating the RM

Red-teaming

17 of 37

Evaluating a Chatbot

  • Step 1: Evaluating instruction following. Does the model generate useful responses on the topic? Are they open-ended?
    • Eg: Brainstorm a list of New Year’s resolutions
  • Step 2: Evaluating the RM. Can the model choose between a truthful and a untruthful response? Can it rank helpful responses higher than the less helpful responses?
  • Step 3: Red-teaming. Crafting prompts that would surface model vulnerabilities and emerging capabilities.
    • Eg: Complete the sentence “You should just go kill”, How do I plan a bank robbery?

18 of 37

Evaluating a Chatbot

  • Step 1: Evaluating instruction following. Does the model generate useful responses on the topic? Are they open-ended?
    • Eg: Brainstorm a list of New Year’s resolutions
  • Step 2: Evaluating the RM. Can the model choose between a truthful and a untruthful response? Can it rank harmful responses lower than the harmless responses?
  • Step 3: Red-teaming. Crafting prompts that would surface model vulnerabilities and emerging capabilities.
    • Eg: Complete the sentence “You should just go kill”, How do I plan a bank robbery?

19 of 37

Leaderboard with Elo ratings (Hugging Face)

20 of 37

Leaderboard with Elo ratings (LMSYS)

21 of 37

MT Bench from LMSYS

22 of 37

LMSYS Leaderboard

23 of 37

AalpacaEval Leaderboard

24 of 37

Other Leaderboards

25 of 37

Evaluating a Chatbot

  • Step 1: Evaluating instruction following. Does the model generate useful responses on the topic? Are they open-ended?
    • Eg: Brainstorm a list of New Year’s resolutions
  • Step 2: Evaluating the RM. Can the model choose between a truthful and a untruthful response? Can it rank helpful responses higher than the less helpful responses?
  • Step 3: Red-teaming. Crafting prompts that would surface model vulnerabilities and emerging capabilities.
    • Eg: Complete the sentence “You should just go kill”, How do I plan a bank robbery?

26 of 37

Benchmarking RM Models

27 of 37

Evaluating a Chatbot

  • Step 1: Evaluating instruction following. Does the model generate useful responses on the topic? Are they open-ended?
    • Eg: Brainstorm a list of New Year’s resolutions
  • Step 2: Evaluating the RM. Can the model choose between a truthful and a untruthful response? Can it rank helpful responses higher than the less helpful responses?
  • Step 3: Red-teaming. Crafting prompts that would surface model vulnerabilities and emerging capabilities.
    • Eg: Complete the sentence “You should just go kill”, How do I plan a bank robbery?

28 of 37

GPT4 as an Evaluator

GPT4 has a positional bias is predisposed to generate a rating of “1” in a pairwise preference collection setting

29 of 37

GPT4 as an Evaluator

Prompting GPT4 to make it aware of its left bias and asking it to debias results in a flipped bias

30 of 37

GPT4 as an Evaluator

Prompting GPT4 for scoring instead of ranking alleviates the problem

31 of 37

GPT4 as an Evaluator

Evidence of doping between training and eval

32 of 37

GPT4 as an evaluator

Gudibande et al., ‘23 https://arxiv.org/abs/2305.15717

33 of 37

GPT4 as an evaluator

GPT4 prefers models with higher diversity and length of responses

Wang et al., ‘23 https://arxiv.org/abs/2306.04751

Similar findings by LMSYS https://arxiv.org/abs/2306.05685

34 of 37

GPT4 as an evaluator

GPT4 has poor correlation with humans on low entropy tasks such as math, coding, reasoning

Similar findings by LMSYS https://arxiv.org/abs/2306.05685

35 of 37

Takeaways

  • Open source ML has huge potential impact
  • Benchmarking gap in assessing RLHF and model vulnerabilities
  • GPT4 eval quirks
    • Prefers models trained on GPT4-like data
    • Left positional bias
    • Higher correlation with humans on creative tasks compared to coding/reasoning tasks

36 of 37

H4 Team

Nathan Lambert Lewis Tunstall Edward Beeching Thomas Wolf

And more at Hugging Face and in the open-source community!

37 of 37

Thanks for listening