Taming the Wild West
of LLMs
Nazneen Rajani | Research Lead @ Hugging Face | nazneen@hf.co | @nazneenrajani
Text-to-Text Foundation Models since GPT3
GPT-3
2021
Jun
Oct
PaLM
Chinchilla
OPT
BLOOM
Gopher
2022
Megatron TNLG
Dec
May
Apr
Jul
Jul
GPT-J
ChatGPT
Nov
Dec
Galactica
GPT-Neo
Jun
GPT-NeoX
Feb
Flan-T5
Oct
*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1
UL2
Cohere
Jurassic
Anthropic
2023
Feb
LLaMA
Flan-UL2
March
GPT-4
Falcon
May
INCITE
StarCoder
LLaMA-2
Text-to-Text Foundation Models since GPT3
GPT-3
2021
Jun
Oct
PaLM
Chinchilla
OPT
BLOOM
Gopher
2022
Megatron TNLG
Dec
May
Apr
Jul
Jul
GPT-J
ChatGPT
Nov
Dec
Galactica
GPT-Neo
Jun
GPT-NeoX
Feb
Flan-T5
Oct
*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1
UL2
Cohere
Jurassic
Anthropic
2023
Feb
LLaMA
Flan-UL2
March
GPT-4
Falcon
May
INCITE
StarCoder
LLaMA-2
Model Access
🔓
🔒
Open access
Closed access
Limited access
🔐
🔓 Open Access Models
Model components are publicly available:
🔓 Open Access Models
Allows reproducing results and replicating parts of the model
Enable auditing and conducting risk analysis
Serves as a research artifact
Enables interpreting model output
🔒 Closed Access Models
Only research paper or blog is available and may include overview of
🔒 Closed Access Models
Safety concerns
Competitive advantage
Expensive to setup guardrails for safe access
🔐 Limited Access Models
Available for use via:
Text-to-Text Foundation Models since GPT3
GPT-3
2021
Jun
Oct
PaLM
Chinchilla
OPT
BLOOM
Gopher
2022
Megatron TNLG
Dec
May
Apr
Jul
Jul
GPT-J
ChatGPT
Nov
Dec
Galactica
GPT-Neo
Jun
GPT-NeoX
Feb
Flan-T5
Oct
*only LLMs with >1B parameters & EN as the main training language are shown. Comprehensive list: https://crfm.stanford.edu/helm/v1.0/?models=1
UL2
Cohere
Jurassic
Anthropic
2023
Feb
LLaMA
Flan-UL2
March
GPT-4
Falcon
May
INCITE
StarCoder
LLaMA-2
Pivotal moments
Large Language Models – Training
Large Language Models – Training
Training a chatbot
Evaluating a Chatbot
Evaluating a Chatbot
Evaluating a Chatbot
Ouyang, Long, et al. "Training language models to follow instructions with human feedback." arXiv preprint arXiv:2203.02155 (2022).
Evaluating instruction following/chatty-ness
Evaluating the RM
Red-teaming
Evaluating a Chatbot
Evaluating a Chatbot
Leaderboard with Elo ratings (Hugging Face)
Leaderboard with Elo ratings (LMSYS)
MT Bench from LMSYS
LMSYS Leaderboard
AalpacaEval Leaderboard
Other Leaderboards
Evaluating a Chatbot
Benchmarking RM Models
Evaluating a Chatbot
GPT4 as an Evaluator
GPT4 has a positional bias is predisposed to generate a rating of “1” in a pairwise preference collection setting
GPT4 as an Evaluator
Prompting GPT4 to make it aware of its left bias and asking it to debias results in a flipped bias
GPT4 as an Evaluator
Prompting GPT4 for scoring instead of ranking alleviates the problem
GPT4 as an Evaluator
Evidence of doping between training and eval
GPT4 as an evaluator
Gudibande et al., ‘23 https://arxiv.org/abs/2305.15717
GPT4 as an evaluator
GPT4 prefers models with higher diversity and length of responses
Wang et al., ‘23 https://arxiv.org/abs/2306.04751
Similar findings by LMSYS https://arxiv.org/abs/2306.05685
GPT4 as an evaluator
GPT4 has poor correlation with humans on low entropy tasks such as math, coding, reasoning
Similar findings by LMSYS https://arxiv.org/abs/2306.05685
Takeaways
H4 Team
Nathan Lambert Lewis Tunstall Edward Beeching Thomas Wolf
And more at Hugging Face and in the open-source community!
Thanks for listening