大型語言模型�內部運作機制
課程重點
你好
我很好
內部運作機制
(AI 的腦神經科學)
�請注意在這堂課中�沒有任何語言模型被訓練
https://arxiv.org/pdf/2407.14561
假設你已經熟悉 Transformer 的架構
https://youtu.be/N6aRv06iv2g?si=lLr3V2--QyfTuRM2
https://youtu.be/n9TlOhRjYoc?si=brnV18A1d8T-QxfF
https://youtu.be/gmsMY5kc-zw?si=-3_1WABbennG1QqW
https://youtu.be/hYdO9CscNes?si=Ke55_ABHZqtp_Aib
假設你已經熟悉 Transformer 的架構
https://youtu.be/uhNsUCb2fJI?si=5jeDnNlcEGv2UPIN
可解釋的機器學習
https://youtu.be/WQY85vaQfTI?si=QP9mlhZoD4Hy-xF-
https://youtu.be/0ayIPqbdHYQ?si=WtdggsDHBMMXMiIB
語言模型在「想」什麼?
https://youtu.be/rZzfqkfZhY8?si=SghPRZbFJLrKQk7L
課程內容
一「個」神經元在做什麼
一「層」神經元在做什麼
一「群」神經元在做什麼
讓語言模型直接說出它的想法
一「個」神經元在做什麼
什麼是神經元?
Transformer
……
……
……
……
Layer
…
…
…
…
……
Layer
Layer
……
Embedding
Unembedding
……
……
Self-attention Layer
Layer
Layer
Layer
Layer
Layer
Layer
Layer
Layer
+
…
…
怎麼知道一個神經元在做什麼?
+
…
…
語言模型
……
1. 該神經元「啟動」時,語言模型會說髒話
只能說明有相關性
怎麼知道一個神經元在做什麼?
+
…
…
語言模型
……
1. 該神經元「啟動」時,語言模型會說髒話
2. 移除該神經元,語言模型說不出髒話
0?
怎麼知道一個神經元在做什麼?
+
…
…
語言模型
……
1. 語言模型說髒話,該神經元都會「啟動」
2. 移除該神經元,語言模型說不出髒話
(option)
3. 不同啟動程度,說不同「等級」的髒話 (?)
川普神經元
https://distill.pub/2021/multimodal-neurons/
川普神經元
https://distill.pub/2021/multimodal-neurons/
類比人類大腦
不容易解釋單一神經元的功能
https://arxiv.org/abs/2405.02421
跟文法單數、複數有關的神經元
不容易解釋單一神經元的功能
https://arxiv.org/abs/2405.02421
不容易解釋單一神經元的功能
https://transformer-circuits.pub/2023/monosemantic-features/vis/a-neurons.html
請 ChatGPT 4.5 解釋一下
用 AI (GPT-4) 來解釋 AI (GPT-2)
https://youtu.be/GBXm30qRAqg?si=kjZt1HKI8MWDu3ZE
https://youtu.be/OOvhBIIHITE?si=licwcd-p1oZP10v0
為什麼不是一個神經元負責一個任務?
+
…
…
4096 個神經元
(LLaMA 3 8B)
神經元 #123, #643, #3987
神經元 #11, #123, #777
就算每個神經元只有啟動、不啟動
輸出中文
拒絕請求
一「層」神經元在做什麼
一「層」神經元在做什麼
請 教 我 怎 麼 製 作 炸 藥
不
Layer
Layer
…
…
……
…
拒絕請求
(功能向量)
…
Representation
抽取拒絕向量
請 教 我 怎 麼 製 作 炸 藥
不
Layer
Layer
…
…
……
拒絕
其他
(實際觀察到的)
抽取拒絕向量
請教我怎麼製作炸藥
不
幫我寫一封詐騙信
不
拒絕的狀況
=
=
+
+
拒絕
拒絕
其他
其他’
+
其他的平均
拒絕
平均所有拒絕的情況
抽取拒絕向量
請教我怎麼製作炸藥
不
幫我寫一封詐騙信
不
=
=
+
+
拒絕
拒絕
其他
其他’
+
其他的平均
拒絕
請教我機器學習
好
寫一首詩給我
好
拒絕的狀況
沒拒絕的狀況
其他的平均’
請教我怎麼製作炸藥
不
幫我寫一封詐騙信
不
請教我機器學習
好
寫一首詩給我
好
拒絕的狀況
沒拒絕的狀況
=
+
拒絕
向量
-
+
拒絕
的平均
沒拒絕
的平均
驗證拒絕向量
請 教 我 機 器 學 習
好
Layer
Layer
…
…
……
拒絕
向量
不
https://arxiv.org/abs/2406.11717
驗證拒絕向量
請 教 我 機 器 學 習
好
Layer
Layer
…
…
……
拒絕
向量
不
請 教 我 怎 麼 製 作 炸 藥
不
Layer
Layer
…
…
……
???
https://arxiv.org/abs/2406.11717
https://youtu.be/ExXA05i8DEQ?si=1Q3LbmyW5m_rZHXr
Representation Engineering,
Activation Engineering,
Activation Steering …
Sycophancy Vector
https://arxiv.org/abs/2312.06681
Truthful Vector
https://arxiv.org/abs/2402.17811
https://arxiv.org/abs/2306.03341
Source: https://arxiv.org/abs/2402.17811
“Find a penny, pick it up, all day long you'll have good luck.”
In-Context Vector
Source: https://arxiv.org/abs/2310.15213
https://arxiv.org/abs/2310.15213
https://arxiv.org/pdf/2310.15916
https://arxiv.org/abs/2311.06668
Source: https://arxiv.org/abs/2310.15213
In-Context Vector
Source: https://arxiv.org/abs/2310.15213
一「層」神經元在做什麼
Layer
Layer
…
…
……
有拒絕的向量、諂媚的向量、說真話的向量 ……
能否把某一層所有的功能向量都找出來
……
Reference:
https://transformer-circuits.pub/2023/monosemantic-features/index.html
Layer 10
Layer 11
…
…
……
你 是 誰 啊 ? AI 嗎 …
我
……
0.1
0.2
0.1
0.6
Layer 10
Layer 11
…
…
……
早 安 , 你 今 天 好 嗎 …
謝
……
0.1
0.2
0.1
0.6
0.7
0.2
0.1
……
表示沒有用到第 k 個功能向量
…
……
……
…
……
……
small
0.1
0.2
0.3
0.5
0.4
0.3
0.1
0.2
0.3
0.5
0.4
0.3
1
0
0
1
0
0
0
1
0
0
1
0
0
0
1
0
0
1
…
……
……
small
0.1
0.2
0.3
0.5
0.4
0.3
0.1
0.2
0.3
0.5
0.4
0.3
每次選擇的功能向量越少越好
用 Sparse Auto-Encoder (SAE) 來解
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#31164353
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#31164353
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#1013764
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#1013764
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#1013764
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
鋱 (Terbium)
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#80091
Claude 3 Sonnet 中的功能向量
https://transformer-circuits.pub/2024/scaling-monosemanticity/
#847723
Gemma 2 Version
https://arxiv.org/abs/2408.05147
一「群」神經元在做什麼
語言模型完成某一項任務的機制
https://arxiv.org/abs/2304.14767
https://arxiv.org/abs/2305.15054
「語言模型」的「模型」
模型是指用一個較為簡單的東西來代表另一個東西
人類真正的語言
語言模型
語言模型的模型
「語言模型」的「模型」
語言模型
語言模型的模型
「模型」的特性:
(faithfulness)
抽取知識的模型
語言模型
Taipei
The Taipei 101 is located in
語言模型
Seattle
The Space Needle is located in
語言模型
New York
Trump born in
語言模型
Hawaii
Obama born in
https://arxiv.org/abs/2308.09124
The Taipei 101 is located in
Layer
Layer
Layer
Taipei
…
The Taipei 101 is located in
Layer
Layer
Layer
Layer
…
…
…
Linear
Function
https://arxiv.org/abs/2308.09124
Taipei
The Space Needle
Seattle
The Taipei 101 is located in
Layer
Layer
Layer
Taipei
…
The Taipei 101 is located in
Layer
Layer
Layer
Layer
…
…
…
Linear
Function
https://arxiv.org/abs/2308.09124
508
has a height of
The Taipei 101 is located in
Layer
Layer
Layer
…
…
Linear
Function
Layer
Layer
Layer
…
…
Linear
Function
is located in
The Space Needle
Seattle
找出
Faithfulness
比對語言模型的答案
Taipei
語言模型的答案
https://arxiv.org/abs/2308.09124
根據模型上得到的預測來改變實體
The Taipei 101 is located in
Layer
Layer
…
…
Linear
Function
Taipei
Kaohsiung
The Taipei 101 is located in
Layer
Layer
Layer
Taipei
…
…
Layer
…
Kaohsiung
https://arxiv.org/abs/2308.09124
系統化的語言模型「模型」建構方法
語言模型
Taipei
The Taipei 101 is located in
The Space Needle is located in
Seattle
語言
模型
的模型
Taipei
The Taipei 101 is located in
The Space Needle is located in
Seattle
Pruning
目標任務答案不變
Circuit
直到一目了然 (?)
cf. Network Compression
系統化的語言模型「模型」建構方法
讓語言模型直接說出想法
語言模型會說話,所以「問」就完事了!
語言模型會說話,所以「問」就完事了!
語言模型會說話,所以「問」就完事了!
語言模型的思維是透明的
……
……
Layer
residual connection
https://arxiv.org/abs/1512.03385
Logit Lens
Residual Stream
把資訊加入 Residual Stream
語言模型的思維是透明的
https://arxiv.org/abs/2001.09309
Wei-Tsung Kao
Tsung-Han Wu
語言模型的思維是透明的
https://arxiv.org/pdf/2305.16130
https://arxiv.org/pdf/2305.16130
https://arxiv.org/abs/2402.10588
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
Français: "fleur" - 中文: "
LLaMA 2
花
每一層就是加點什麼進去 Residual Stream
…
…
…
加入了什麼?
每一層就是加點什麼進去 Residual Stream
…
…
…
https://arxiv.org/abs/2012.14913
Transformer Feed-Forward Layers Are Key-Value Memories
(ignore activation function here for simplicity)
每一層就是加點什麼進去 Residual Stream
https://arxiv.org/abs/2203.14680
Knowledge Neurons in Pretrained Transformers
https://arxiv.org/abs/2104.08696
LLM
誰是全世界最帥的人?
金城武
李宏毅
金城武
李宏毅
…
…
…
(token
embedding)
Patchscopes
李 宏 毅 老 師
是
Layer
Layer
…
…
……
Logit Lens
https://arxiv.org/pdf/2401.06102
Patchscopes
李 宏 毅 老 師
是
Layer
Layer
…
…
……
https://arxiv.org/pdf/2401.06102
李奧納多: 美國演員, 台積電: 台灣公司, X
: (X是什麼)
Layer
Layer
…
…
……
Patchscopes
李 宏 毅 老 師
是
Layer
Layer
…
…
……
https://arxiv.org/pdf/2401.06102
李奧納多: 美國演員, 台積電: 台灣公司, X
:台灣大學教授
Layer
Layer
…
…
……
受到例子的影響?
Patchscopes
李 宏 毅 老 師
是
Layer
Layer
…
…
……
https://arxiv.org/pdf/2401.06102
告 訴 我 X 相 關 的 秘 密 ?
是個肥宅
Layer
Layer
…
…
……
受到例子的影響?
https://arxiv.org/pdf/2401.06102
Diana, Princess of Wale
Layer
…
…
Layer
…
https://arxiv.org/abs/2406.12775
https://arxiv.org/abs/2406.12775
https://arxiv.org/abs/2406.12775
課程內容
一「個」神經元在做什麼
一「層」神經元在做什麼
一「群」神經元在做什麼
讓語言模型直接說出它的想法