1 of 43

Lecture 5��The policy gradient methods

1

Instructor: Ercan Atam

Institute for Data Science & Artificial Intelligence

Course: DSAI 642- Advanced Reinforcement Learning

2 of 43

2

List of contents for this lecture

  • Policy gradient

  • Reinforce algorithm

3 of 43

3

Relevant readings for this lecture

  • Chapter 9 of Shiyu Zhao, “Mathematical Foundations of Reinforcement learning”, Springer, 2025.

(Most slides at the beginning are modified/improved versions from the presentation slides of Shiyu Zhao)

  • Chapter 2 of Laura Graesser and Wah Loon Keng, “Foundations of Deep Reinforcement Learning:

Theory and Practice in Python”, Addison-Wesley Professional, 2019.

  • Chapter 13 of Richard S. Sutton and Andrew G. Barto, “Reinforcement Learning: An Introduction”,

Second Edition, MIT Press, Cambridge, MA, 2018.

  • Chapter 8 of Nimish Sanghi, “Deep Reinforcement Learning with Python”, 2nd Edition, Apress, 2024.

4 of 43

4

Table-based representation of policies

In Table-based methods, policies are represented by tables: the action probabilities of all states are stored in a table.

Table: A tabular representation of a policy. There are nine states and five actions for each state.

5 of 43

5

Disadvantages of table-based representation

6 of 43

6

Function representation of policies (1)

The idea of function approximation can be applied not only to represent state/action values, but also to represent policies.

Figure: Function representations of policies. The functions may have different structures.

(a)

(b)

7 of 43

7

Function representation of policies (2)

8 of 43

8

Function representation of policies (3)

ANN

ANN

9 of 43

9

Basic idea of the policy gradient method

10 of 43

10

Metrics to define optimal policies: metric 1-average state value (1)

11 of 43

11

Metrics to define optimal policies: metric 1-average state value (2)

12 of 43

12

Metrics to define optimal policies: metric 1-average state value (3)

13 of 43

13

Metrics to define optimal policies: metric 1-average state value (4)

14 of 43

14

Metrics to define optimal policies: metric 2-average reward (1)

15 of 43

15

Metrics to define optimal policies: metric 2-average reward (2)

16 of 43

16

Metrics to define optimal policies: metric 2-average reward (3)

17 of 43

17

Summary of two metrics

18 of 43

18

Remarks about two metrics

19 of 43

19

Relationship between the two metrics

20 of 43

20

Gradients of metrics

21 of 43

21

A unified expression for gradients of metrics (1)

22 of 43

22

A unified expression for gradients of metrics (2)

23 of 43

23

A unified expression for gradients of metrics (3)

24 of 43

24

A unified expression for gradients of metrics (4)

25 of 43

25

A unified expression for gradients of metrics (5)

26 of 43

26

Gradient-ascent algorithm (1)

27 of 43

27

Gradient-ascent algorithm (2)

28 of 43

28

No need for “Markov property” in policy gradient approaches

  • It is worth noting an important observation regarding the Markov property.

  • The derivation of policy gradient does not explicitly rely on the Markov assumption.

  • Since the Bellman equations have not been used at any point in the derivation, policy gradient methods can, in principle, be applied even in non-Markovian environments.

29 of 43

29

REINFORCE

30 of 43

30

Improving REINFORCE (1)

  • Our formulation of the REINFORCE algorithm estimates the policy gradient using Monte Carlo sampling with a single trajectory.

  • This is an unbiased estimate of the policy gradient, but one disadvantage of this approach is that it has a high variance (because the returns can vary significantly from trajectory to trajectory).

  • Factors for significant return variations:

1. Actions have some randomness because they are sampled from a probability distribution.

2. The starting state may vary per episode.

3. The environment transition function may be stochastic.

  • Next, we will introduce a baseline to reduce the variance of the estimate.

31 of 43

31

Improving REINFORCE (2)

32 of 43

32

How to do sampling in REINFORCE?

33 of 43

33

How to interpret REINFOCE? (1)

34 of 43

34

How to interpret REINFOCE? (2)

35 of 43

35

How to interpret REINFOCE? (3)

36 of 43

36

+s, -s of policy-based methods (1)

+s:

    • Better convergence: Policy-based methods often converge more reliably, as they directly optimize the policy rather than relying on the two-step procedure used in value-based methods.

    • Effective in high-dimensional and continuous action spaces: Policy-based methods can handle high dimensional (more than one control input) and continuous action spaces.

    • Adaptability to complex environments: Policy-based methods adapt well to dynamic settings, enabling effective decisions even when the optimal policy changes over time due to environment change.

    • Learn stochastic policies: Policy-based methods naturally support stochastic policies, promoting a better balance between exploration and exploitation.

37 of 43

37

+s, -s of policy-based methods (2)

-s:

    • Sample inefficiency and computational overload: Policy-based methods require many more episodes to learn optimal behavior since they are on-policy (cannot use old data) and has a low learning rate (to prevent fast policy updates).

    • High variance in gradient estimates: Vanilla policy-based methods do not maintain value functions, causing policy evaluation to be inefficient. Therefore, evaluation must be carried out over high number of episodes to keep the variance of policy estimate within a reasonable variance.
      • To mitigate the variance issue, they require variance-reduction techniques (e.g., baselines, advantage functions).

    • Sensitivity to hyperparameters: For example, performance can be highly sensitive to learning rate and other architecture parameters.

38 of 43

38

Appendix 1: Intuition behind the update rule of vanilla policy gradient (1)

39 of 43

39

Appendix 1: Intuition behind the update rule of vanilla policy gradient (2)

  • We could summarize the whole explanation by saying that policy optimization is all about trial and error.

  • You roll out multiple trajectories.

  • The probability of all the actions along the trajectory is increased for those trajectories that are good.

  • For the trajectories that are bad, the probability of all the actions along those bad trajectories are reduced.

40 of 43

40

Appendix 2:Comparing what policy gradient is doing with maximum likelihood model building

  • In Equation (a) , we are just increasing the probability of the actions to increase the overall probability of trajectories that were observed.

  • In Equation (a), we are not making any differentiation for the good vs bad trajectories.

  • In the case of policy gradients in Equation (b) , we are doing something similar to MLE, but with the addition that:
    • We are weighing the log probability gradients with the return of the trajectory so that the probability of good trajectories is increased,
    • And the probability of bad trajectories is decreased.

41 of 43

41

Appendix 3: Derivation of gradient for MLE objective (1)

42 of 43

42

Appendix 3: Derivation of gradient for MLE objective (2)

43 of 43

References �(utilized for preparation of lecture notes or Matlab code)

  • Laura Graesser and Wah Loon Keng, “Foundations of Deep Reinforcement Learning: Theory and Practice in Python”, Addison-Wesley Professional, 2019.
  • Shiyu Zhao, “Mathematical Foundations of Reinforcement learning”, Springer, 2025
  • Richard S. Sutton and Andrew G. Barto, “Reinforcement Learning: An Introduction”, Second Edition, MIT Press, Cambridge, MA, 2018.
  • Nimish Sanghi, “Deep Reinforcement Learning with Python”, 2nd Edition, Apress, 2024.

43