1 of 69

Large Language Models

Lecture 14

Alignment of Language Models:

Reward Maximization

Krishnendu Ghosh

2 of 69

Why Instruction Tuning not enough?

3 of 69

Taxonomy of Alignment methods

4 of 69

Reinforcement Learning

5 of 69

Reinforcement Learning

6 of 69

Reinforcement Learning

7 of 69

Reinforcement Learning

8 of 69

LLM as a Policy

9 of 69

LLM as a Policy

10 of 69

LLM as a Policy

11 of 69

LLM as a Policy

12 of 69

Who/What is the Reward Model?

13 of 69

LLM as a Reward Model

14 of 69

Architecture of the Reward Model

15 of 69

Bradley-Terry Preference Model I

16 of 69

Bradley-Terry Preference Model II

17 of 69

MLE for BT models

18 of 69

An intuitive view

19 of 69

Where does the Data come from?

20 of 69

Publicly Available Preference Data

21 of 69

Constitutional AI for Preferences

22 of 69

Reward Maximization Objective

23 of 69

Why care about closeness to πref?

24 of 69

Formulating Objective:

Reward Maximization

25 of 69

Formulating Objective:

Closeness to πref

26 of 69

Combining Objective

27 of 69

Takeaways & what next?

28 of 69

Regularized Reward Maximization

Objective

29 of 69

Regularized Reward

30 of 69

How to maximize REINFORCE?

31 of 69

Computing the derivative

32 of 69

The log-derivative trick

33 of 69

Monte Carlo approximation

34 of 69

Expanding the gradient

35 of 69

Implementing REINFORCE

36 of 69

Problems with REINFORCE

37 of 69

REINFORCE with Q functions

38 of 69

Q-function & Value function

39 of 69

Q-function to Advantage function

40 of 69

REINFORCE with advantage functions

41 of 69

Implementing the Value function

42 of 69

Learning the Value function

43 of 69

Vanilla Policy Gradient

44 of 69

Problems

45 of 69

REINFORCE with importance weights

46 of 69

Proximal Policy Optimization

47 of 69

PPO-CLIP

48 of 69

PPO-CLIP with +ve advantage

49 of 69

PPO-CLIP with -ve advantage

50 of 69

The PPO-CLIP algorithm

51 of 69

Things to remember

52 of 69

Policy Gradient/PPO for

LLM Alignment

53 of 69

Direct Preference Optimization on preferences

54 of 69

The non-parametric case

55 of 69

Optimal Policy & Reward (π, r)

56 of 69

Parametric Policy & Reward (πθ, rθ)

57 of 69

Training the reward function

58 of 69

The DPO objective

59 of 69

Interpreting the objective

60 of 69

Interpreting β

61 of 69

PPO vs DPO

62 of 69

Why is DPO biased?

63 of 69

Why is DPO biased?

64 of 69

Why is DPO biased?

65 of 69

Why is DPO biased?

66 of 69

Deal with OOD Bias in DPO?

67 of 69

Performance Comparison:

Offline vs Online DPO

68 of 69

Main Takeaways

69 of 69

Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/