1 of 50

DEEP LEARNING JP

[DL Papers]

LLMの評価について

Keno Harada, D1, the University of Tokyo

http://deeplearning.jp/

2 of 50

2

3 of 50

よく見るleaderboard, 何を評価している?

3

4 of 50

日本語ベンチマークJGLUE

4

5 of 50

実際に解いてみよう

https://docs.google.com/forms/d/e/1FAIpQLSc2YtM_ItuZwqmIHmKdSW88cFf_Z_w2Myom68mhZKowChrbcA/viewform 

自由記述の問題はChatGPT使わずに答えてください🙏

5

6 of 50

LLMの評価について

  • Perplexity
  • Downstreamタスクでの性能
    • よく使われるNLPベンチマークの理解
    • 分類問題, 生成問題
  • LLMの安全性・信頼性
  • 評価の難しさ

6

7 of 50

7

8 of 50

Perplexity 次単語予測の性能

8

9 of 50

学習中眺めるもの

9

10 of 50

10

11 of 50

11

12 of 50

12

13 of 50

13

14 of 50

14

15 of 50

15

16 of 50

16

17 of 50

17

18 of 50

18

19 of 50

19

20 of 50

20

21 of 50

21

22 of 50

Can we include learned components in our evaluation metrics?

22

23 of 50

How do we evaluate open-ended text generation?

23

24 of 50

24

25 of 50

25

26 of 50

26

27 of 50

実際に出たヤバイ例は元論文Table17へ

27

28 of 50

28

29 of 50

29

30 of 50

30

31 of 50

31

32 of 50

DEEP LEARNING JP

[DL Papers]

SCIBENCH: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Keno Harada, D1, the University of Tokyo

http://deeplearning.jp/

33 of 50

新しいベンチマーク SCIBENCHの提案

  • ScienceQAやGSM8Kはgrade-levelで簡単
  • MATHhあhigh-school levelと言いつつ四則演算+alphaくらいなのでreasoning ability測れない
  • AGIEvalやCEvalは選択肢問題がメイン
  • オンラインの教材を作っているためleakが起きている可能性あり
  • CoT promptとか外部ツールの利用提案されているけどまだ課題ありそう
  • よく使われる教科書から695問作成(Open), 中間試験などから作った問題104問(Close)
    • 自由記述
    • 物理, 熱力学, 古典力学, 量子化学, 物理化学, 微積分, 統計, 微分方程式

33

34 of 50

SCIBENCHでの評価と発見

  • Gpt-3.5-turboとgpt-4で評価
    • CoT prompting
    • Prompting to use external use

34

35 of 50

35

36 of 50

DEEP LEARNING JP

[DL Papers]

Judging LLM-as-a-judge with MT-Bench and Chatbot Arena

Keno Harada, D1, the University of Tokyo

http://deeplearning.jp/

37 of 50

MT-BenchとChatbot Arenaの提案, LLMによる評価の分析

  • MT-bench: 自由記述でchat botのの複数回の会話のやり取りとinstruction-following能力を測る
    • 80の質問と3000の投票, 30,000の会話データ
    • Writing, roleplay, extraction, reasoning, math, coding, STEM, humanities/social science
  • Chatbot Arena: 2つのchatbotsの出力を人間が比べて評価
  • LLM-as-a-judgeの分析
    • Position bias
    • Verbosity bias
    • Self-enhancement bias
    • Limited reasoning ability
    • GPT-4の判断は人間の判断と80%くらいmatch

37

38 of 50

MT-Benchの例

38

39 of 50

既存のベンチマーク

  • Core-knowledge benchmarks
    • MMLU, HellaSwag, ARC, Wino-Grande, HumanEval, GSM-8K, AGIEval
  • Instruction-following benchmarks
    • Flan, Self-instruct, NaturalInstructions, Super-NaturalInstructions
  • Conversational benchmarks
    • CoQA, MMDialog, OpenAssistant

39

40 of 50

LLM-as-a-Judge

  • 人間にいちいち評価してもらうの大変
  • 自動的に評価しようと思っても模範回答的なものがないと大変
    • 模範解答があるとROUGEやBLEUで測れる

40

41 of 50

Pairwise comparison

41

42 of 50

Single answer grading

42

43 of 50

Reference-guided grading

43

44 of 50

Position bias

  • GPT-3.5とVicuna-13Bのモデルの回答において、テキストの順番上先にあるものを良いと評価する
    • 人間にも見られる現象

44

45 of 50

Verbosity bias

  • 長い出力を好む傾向
    • 同じ内容の箇条書きを増やして検証
    • GPT-4は他のモデルに比べて影響は小さい

45

46 of 50

Self-enhancement bias

  • モデル自身の回答を好むbias

46

47 of 50

Limited capability in grading math and reasoning questions

  • 問題解ける能力はあるのに、間違った記述に引っ張られて判定を間違える

47

48 of 50

対応策

  • Swapping positions
    • 位置を入れ替えてどちらも同じ解答が選ばれたら勝利とする
  • Few-shot judge
  • Chain of thought and reference guide judge
  • Fine-tuning a judge model

48

49 of 50

High agreement between GPT-4 and humans

  • 人間同士のvoteの一致率が81%, 人間とGPT-4のvoteの一致率が85%

49

50 of 50

RLHF: Reward Modeling

  • Metaのtest setでも他のベンチマークでも他のモデルを凌駕
    • GPT-4に「どっちの文章が良いか選んで」というプロンプトで判断させたら他のモデルよりもMetaのtest setで良い性能
  • めっちゃ良い、というような違いが分かりやすいほど正答率も上がる
  • モデルサイズが大きくなればなるほど良いし、データも増えれば正答率上がる
    • InstructGPTの時は6Bを採用、175Bだと不安定になったという報告が

50