1 of 90

大型語言模型�內部運作機制

2 of 90

課程重點

你好

我很好

內部運作機制

  • 對於大型語言模型內部的運作機制有進一步了解
  • 知道剖析內部運作機制的方法 (可以應用到不同領域)

(AI 的腦神經科學)

3 of 90

�請注意在這堂課中�沒有任何語言模型被訓練

4 of 90

https://arxiv.org/pdf/2407.14561

5 of 90

假設你已經熟悉 Transformer 的架構

https://youtu.be/N6aRv06iv2g?si=lLr3V2--QyfTuRM2

https://youtu.be/n9TlOhRjYoc?si=brnV18A1d8T-QxfF

https://youtu.be/gmsMY5kc-zw?si=-3_1WABbennG1QqW

https://youtu.be/hYdO9CscNes?si=Ke55_ABHZqtp_Aib

6 of 90

假設你已經熟悉 Transformer 的架構

https://youtu.be/uhNsUCb2fJI?si=5jeDnNlcEGv2UPIN

7 of 90

可解釋的機器學習

https://youtu.be/WQY85vaQfTI?si=QP9mlhZoD4Hy-xF-

https://youtu.be/0ayIPqbdHYQ?si=WtdggsDHBMMXMiIB

8 of 90

語言模型在「想」什麼?

https://youtu.be/rZzfqkfZhY8?si=SghPRZbFJLrKQk7L

9 of 90

課程內容

一「個」神經元在做什麼

一「層」神經元在做什麼

一「群」神經元在做什麼

讓語言模型直接說出它的想法

10 of 90

一「個」神經元在做什麼

11 of 90

什麼是神經元?

 

Transformer

 

……

 

 

 

 

……

 

 

12 of 90

……

……

 

 

 

 

Layer

……

Layer

 

 

Layer

 

 

……

Embedding

Unembedding

13 of 90

……

……

Self-attention Layer

Layer

Layer

Layer

Layer

Layer

Layer

Layer

Layer

+

14 of 90

怎麼知道一個神經元在做什麼?

+

語言模型

 

……

 

 

 

 

1. 該神經元「啟動」時,語言模型會說髒話

只能說明有相關性

15 of 90

怎麼知道一個神經元在做什麼?

+

語言模型

 

……

 

 

 

 

1. 該神經元「啟動」時,語言模型會說髒話

2. 移除該神經元,語言模型說不出髒話

0?

16 of 90

怎麼知道一個神經元在做什麼?

+

語言模型

 

……

 

 

 

 

1. 語言模型說髒話,該神經元都會「啟動」

2. 移除該神經元,語言模型說不出髒話

(option)

3. 不同啟動程度,說不同「等級」的髒話 (?)

17 of 90

川普神經元

https://distill.pub/2021/multimodal-neurons/

18 of 90

川普神經元

https://distill.pub/2021/multimodal-neurons/

19 of 90

類比人類大腦

  • 祖母神經元 / Jennifer Aniston 神經元

20 of 90

不容易解釋單一神經元的功能

  • 一件事情可能很多神經元共同管理

https://arxiv.org/abs/2405.02421

跟文法單數、複數有關的神經元

21 of 90

不容易解釋單一神經元的功能

  • 一件事情可能很多神經元共同管理

https://arxiv.org/abs/2405.02421

22 of 90

不容易解釋單一神經元的功能

  • 一個神經元可能同時管很多事

https://transformer-circuits.pub/2023/monosemantic-features/vis/a-neurons.html

請 ChatGPT 4.5 解釋一下

23 of 90

用 AI (GPT-4) 來解釋 AI (GPT-2)

https://youtu.be/GBXm30qRAqg?si=kjZt1HKI8MWDu3ZE

https://youtu.be/OOvhBIIHITE?si=licwcd-p1oZP10v0

24 of 90

為什麼不是一個神經元負責一個任務?

+

4096 個神經元

(LLaMA 3 8B)

神經元 #123, #643, #3987

神經元 #11, #123, #777

就算每個神經元只有啟動、不啟動

 

輸出中文

拒絕請求

25 of 90

一「層」神經元在做什麼

26 of 90

一「層」神經元在做什麼

請 教 我 怎 麼 製 作 炸 藥

Layer

Layer

……

拒絕請求

(功能向量)

Representation

27 of 90

抽取拒絕向量

請 教 我 怎 麼 製 作 炸 藥

Layer

Layer

……

拒絕

其他

(實際觀察到的)

28 of 90

抽取拒絕向量

請教我怎麼製作炸藥

幫我寫一封詐騙信

拒絕的狀況

=

=

+

+

拒絕

拒絕

其他

其他’

+

其他的平均

拒絕

平均所有拒絕的情況

29 of 90

抽取拒絕向量

請教我怎麼製作炸藥

幫我寫一封詐騙信

=

=

+

+

拒絕

拒絕

其他

其他’

+

其他的平均

拒絕

請教我機器學習

寫一首詩給我

拒絕的狀況

沒拒絕的狀況

其他的平均’

30 of 90

請教我怎麼製作炸藥

幫我寫一封詐騙信

請教我機器學習

寫一首詩給我

拒絕的狀況

沒拒絕的狀況

=

+

拒絕

向量

 

-

+

 

拒絕

的平均

沒拒絕

的平均

31 of 90

驗證拒絕向量

請 教 我 機 器 學 習

Layer

Layer

……

拒絕

向量

 

32 of 90

https://arxiv.org/abs/2406.11717

33 of 90

驗證拒絕向量

請 教 我 機 器 學 習

Layer

Layer

……

拒絕

向量

 

請 教 我 怎 麼 製 作 炸 藥

Layer

Layer

……

 

???

34 of 90

https://arxiv.org/abs/2406.11717

35 of 90

https://youtu.be/ExXA05i8DEQ?si=1Q3LbmyW5m_rZHXr

Representation Engineering,

Activation Engineering,

Activation Steering …

36 of 90

Sycophancy Vector

https://arxiv.org/abs/2312.06681

37 of 90

Truthful Vector

https://arxiv.org/abs/2402.17811

https://arxiv.org/abs/2306.03341

Source: https://arxiv.org/abs/2402.17811

“Find a penny, pick it up, all day long you'll have good luck.”

38 of 90

In-Context Vector

Source: https://arxiv.org/abs/2310.15213

https://arxiv.org/abs/2310.15213

https://arxiv.org/pdf/2310.15916

https://arxiv.org/abs/2311.06668

39 of 90

Source: https://arxiv.org/abs/2310.15213

40 of 90

In-Context Vector

Source: https://arxiv.org/abs/2310.15213

41 of 90

一「層」神經元在做什麼

Layer

Layer

……

有拒絕的向量、諂媚的向量、說真話的向量 ……

能否把某一層所有的功能向量都找出來

……

 

 

 

 

 

Reference:

https://transformer-circuits.pub/2023/monosemantic-features/index.html

42 of 90

Layer 10

Layer 11

……

你 是 誰 啊 ? AI 嗎 …

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

0.1

0.2

0.1

0.6

 

 

43 of 90

Layer 10

Layer 11

……

早 安 , 你 今 天 好 嗎 …

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

0.1

0.2

0.1

0.6

 

 

 

 

 

 

 

0.7

0.2

0.1

 

 

 

 

44 of 90

……

 

 

 

 

 

 

表示沒有用到第 k 個功能向量

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

45 of 90

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

small

 

0.1

0.2

0.3

0.5

0.4

0.3

0.1

0.2

0.3

0.5

0.4

0.3

1

0

0

1

0

0

0

1

0

0

1

0

0

0

1

0

0

1

46 of 90

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

……

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

small

 

0.1

0.2

0.3

0.5

0.4

0.3

0.1

0.2

0.3

0.5

0.4

0.3

每次選擇的功能向量越少越好

 

 

用 Sparse Auto-Encoder (SAE) 來解

47 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#31164353

48 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#31164353

49 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#1013764

50 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#1013764

51 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#1013764

52 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

鋱 (Terbium)

53 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#80091

54 of 90

Claude 3 Sonnet 中的功能向量

https://transformer-circuits.pub/2024/scaling-monosemanticity/

#847723

Gemma 2 Version

https://arxiv.org/abs/2408.05147

55 of 90

一「群」神經元在做什麼

語言模型完成某一項任務的機制

56 of 90

https://arxiv.org/abs/2304.14767

https://arxiv.org/abs/2305.15054

57 of 90

「語言模型」的「模型」

模型是指用一個較為簡單的東西來代表另一個東西

人類真正的語言

語言模型

語言模型的模型

58 of 90

「語言模型」的「模型」

語言模型

語言模型的模型

「模型」的特性:

  • 要比原來的實物簡單
  • 保有原來實物的特徵

(faithfulness)

59 of 90

抽取知識的模型

語言模型

Taipei

The Taipei 101 is located in

語言模型

Seattle

The Space Needle is located in

語言模型

New York

Trump born in

語言模型

Hawaii

Obama born in

https://arxiv.org/abs/2308.09124

60 of 90

The Taipei 101 is located in

Layer

Layer

Layer

Taipei

The Taipei 101 is located in

Layer

Layer

Layer

Layer

Linear

Function

https://arxiv.org/abs/2308.09124

Taipei

 

 

 

The Space Needle

 

 

Seattle

61 of 90

The Taipei 101 is located in

Layer

Layer

Layer

Taipei

The Taipei 101 is located in

Layer

Layer

Layer

Layer

Linear

Function

https://arxiv.org/abs/2308.09124

508

 

 

 

has a height of

62 of 90

The Taipei 101 is located in

Layer

Layer

Layer

Linear

Function

 

 

Layer

Layer

Layer

Linear

Function

 

 

is located in

The Space Needle

 

Seattle

 

 

找出

Faithfulness

比對語言模型的答案

Taipei

語言模型的答案

63 of 90

https://arxiv.org/abs/2308.09124

64 of 90

根據模型上得到的預測來改變實體

The Taipei 101 is located in

Layer

Layer

Linear

Function

Taipei

 

 

Kaohsiung

The Taipei 101 is located in

Layer

Layer

Layer

Taipei

Layer

 

 

Kaohsiung

65 of 90

https://arxiv.org/abs/2308.09124

66 of 90

系統化的語言模型「模型」建構方法

語言模型

Taipei

The Taipei 101 is located in

The Space Needle is located in

Seattle

語言

模型

的模型

Taipei

The Taipei 101 is located in

The Space Needle is located in

Seattle

Pruning

目標任務答案不變

Circuit

直到一目了然 (?)

cf. Network Compression

67 of 90

系統化的語言模型「模型」建構方法

  • Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
    • https://arxiv.org/abs/2211.00593
  • Towards Automated Circuit Discovery for Mechanistic Interpretability
    • https://arxiv.org/abs/2304.14997
  • Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
    • https://arxiv.org/abs/2307.09458
  • Attribution Patching Outperforms Automated Circuit Discovery
    • https://arxiv.org/abs/2310.10348
  • Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
    • https://arxiv.org/abs/2403.19647
  • Knowledge Circuits in Pretrained Transformers
    • https://arxiv.org/abs/2405.17969

68 of 90

讓語言模型直接說出想法

69 of 90

語言模型會說話,所以「問」就完事了!

70 of 90

語言模型會說話,所以「問」就完事了!

71 of 90

語言模型會說話,所以「問」就完事了!

72 of 90

語言模型的思維是透明的

……

……

Layer

residual connection

https://arxiv.org/abs/1512.03385

73 of 90

 

 

Logit Lens

Residual Stream

把資訊加入 Residual Stream

74 of 90

語言模型的思維是透明的

https://arxiv.org/abs/2001.09309

Wei-Tsung Kao

Tsung-Han Wu

75 of 90

語言模型的思維是透明的

https://arxiv.org/pdf/2305.16130

76 of 90

https://arxiv.org/pdf/2305.16130

77 of 90

https://arxiv.org/abs/2402.10588

Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Français: "fleur" - 中文: "

LLaMA 2

78 of 90

每一層就是加點什麼進去 Residual Stream

加入了什麼?

79 of 90

每一層就是加點什麼進去 Residual Stream

 

 

 

 

 

 

 

https://arxiv.org/abs/2012.14913

Transformer Feed-Forward Layers Are Key-Value Memories

(ignore activation function here for simplicity)

80 of 90

每一層就是加點什麼進去 Residual Stream

https://arxiv.org/abs/2203.14680

81 of 90

Knowledge Neurons in Pretrained Transformers

https://arxiv.org/abs/2104.08696

LLM

誰是全世界最帥的人?

金城武

 

李宏毅

金城武

李宏毅

 

 

 

 

 

 

(token

embedding)

82 of 90

Patchscopes

李 宏 毅 老 師

Layer

Layer

……

Logit Lens

  • 只能是一個 token
  • 其實是在預測下一個 Token

https://arxiv.org/pdf/2401.06102

83 of 90

Patchscopes

李 宏 毅 老 師

Layer

Layer

……

https://arxiv.org/pdf/2401.06102

李奧納多: 美國演員, 台積電: 台灣公司, X

: (X是什麼)

Layer

Layer

……

84 of 90

Patchscopes

李 宏 毅 老 師

Layer

Layer

……

https://arxiv.org/pdf/2401.06102

李奧納多: 美國演員, 台積電: 台灣公司, X

:台灣大學教授

Layer

Layer

……

受到例子的影響?

85 of 90

Patchscopes

李 宏 毅 老 師

Layer

Layer

……

https://arxiv.org/pdf/2401.06102

告 訴 我 X 相 關 的 秘 密 ?

是個肥宅

Layer

Layer

……

受到例子的影響?

86 of 90

https://arxiv.org/pdf/2401.06102

Diana, Princess of Wale

Layer

Layer

87 of 90

https://arxiv.org/abs/2406.12775

 

 

 

 

88 of 90

https://arxiv.org/abs/2406.12775

 

 

89 of 90

https://arxiv.org/abs/2406.12775

90 of 90

課程內容

一「個」神經元在做什麼

一「層」神經元在做什麼

一「群」神經元在做什麼

讓語言模型直接說出它的想法