Lecture 5��The policy gradient methods
1
Instructor: Ercan Atam
Institute for Data Science & Artificial Intelligence
Course: DSAI 642- Advanced Reinforcement Learning
2
List of contents for this lecture
3
Relevant readings for this lecture
(Most slides at the beginning are modified/improved versions from the presentation slides of Shiyu Zhao)
Theory and Practice in Python”, Addison-Wesley Professional, 2019.
Second Edition, MIT Press, Cambridge, MA, 2018.
4
Table-based representation of policies
In Table-based methods, policies are represented by tables: the action probabilities of all states are stored in a table.
Table: A tabular representation of a policy. There are nine states and five actions for each state.
5
Disadvantages of table-based representation
6
Function representation of policies (1)
The idea of function approximation can be applied not only to represent state/action values, but also to represent policies.
Figure: Function representations of policies. The functions may have different structures.
(a)
(b)
7
Function representation of policies (2)
8
Function representation of policies (3)
ANN
ANN
9
Basic idea of the policy gradient method
10
Metrics to define optimal policies: metric 1-average state value (1)
11
Metrics to define optimal policies: metric 1-average state value (2)
12
Metrics to define optimal policies: metric 1-average state value (3)
13
Metrics to define optimal policies: metric 1-average state value (4)
14
Metrics to define optimal policies: metric 2-average reward (1)
15
Metrics to define optimal policies: metric 2-average reward (2)
16
Metrics to define optimal policies: metric 2-average reward (3)
17
Summary of two metrics
18
Remarks about two metrics
19
Relationship between the two metrics
20
Gradients of metrics
21
A unified expression for gradients of metrics (1)
22
A unified expression for gradients of metrics (2)
23
A unified expression for gradients of metrics (3)
24
A unified expression for gradients of metrics (4)
25
A unified expression for gradients of metrics (5)
26
Gradient-ascent algorithm (1)
27
Gradient-ascent algorithm (2)
28
No need for “Markov property” in policy gradient approaches
29
REINFORCE
30
Improving REINFORCE (1)
1. Actions have some randomness because they are sampled from a probability distribution.
2. The starting state may vary per episode.
3. The environment transition function may be stochastic.
31
Improving REINFORCE (2)
32
How to do sampling in REINFORCE?
33
How to interpret REINFOCE? (1)
34
How to interpret REINFOCE? (2)
35
How to interpret REINFOCE? (3)
36
+s, -s of policy-based methods (1)
+s:
37
+s, -s of policy-based methods (2)
-s:
38
Appendix 1: Intuition behind the update rule of vanilla policy gradient (1)
39
Appendix 1: Intuition behind the update rule of vanilla policy gradient (2)
40
Appendix 2:Comparing what policy gradient is doing with maximum likelihood model building
41
Appendix 3: Derivation of gradient for MLE objective (1)
42
Appendix 3: Derivation of gradient for MLE objective (2)
References �(utilized for preparation of lecture notes or Matlab code)
43