Large Language Models
Lecture 14
Alignment of Language Models:
Reward Maximization
Krishnendu Ghosh
Why Instruction Tuning not enough?
Taxonomy of Alignment methods
Reinforcement Learning
Reinforcement Learning
Reinforcement Learning
Reinforcement Learning
LLM as a Policy
LLM as a Policy
LLM as a Policy
LLM as a Policy
Who/What is the Reward Model?
LLM as a Reward Model
Architecture of the Reward Model
Bradley-Terry Preference Model I
Bradley-Terry Preference Model II
MLE for BT models
An intuitive view
Where does the Data come from?
Publicly Available Preference Data
Constitutional AI for Preferences
Reward Maximization Objective
Why care about closeness to πref?
Formulating Objective:
Reward Maximization
Formulating Objective:
Closeness to πref
Combining Objective
Takeaways & what next?
Regularized Reward Maximization
Objective
Regularized Reward
How to maximize REINFORCE?
Computing the derivative
The log-derivative trick
Monte Carlo approximation
Expanding the gradient
Implementing REINFORCE
Problems with REINFORCE
REINFORCE with Q functions
Q-function & Value function
Q-function to Advantage function
REINFORCE with advantage functions
Implementing the Value function
Learning the Value function
Vanilla Policy Gradient
Problems
REINFORCE with importance weights
Proximal Policy Optimization
PPO-CLIP
PPO-CLIP with +ve advantage
PPO-CLIP with -ve advantage
The PPO-CLIP algorithm
Things to remember
Policy Gradient/PPO for
LLM Alignment
Direct Preference Optimization on preferences
The non-parametric case
Optimal Policy & Reward (π∗, r∗)
Parametric Policy & Reward (πθ, rθ)
Training the reward function
The DPO objective
Interpreting the objective
Interpreting β
PPO vs DPO
Why is DPO biased?
Why is DPO biased?
Why is DPO biased?
Why is DPO biased?
Deal with OOD Bias in DPO?
Performance Comparison:
Offline vs Online DPO
Main Takeaways
Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/