Lecture 8.� Actor-Critic Algorithm�
Sookyung Kim�
1
Recap: Policy Gradient
This is reward from time-step t to end �because of the causality
2
Taxonomy of RL algorithm
Actor-Critic
3
Improving the policy gradient: �Lowering Variance
t
4
Baseline Trick:�Lowering Variance
b
=
Advantage function
5
State & State-action Value Function
6
Value Function Fitting
7
Policy Evaluation
8
Monte Carlo evaluation with function approximation
9
Can we do better?
Bootstrapping
Might incorrect
Low variance (its not from single sample, but output from NN), �high bias (NN estimation can be bias)
10
Policy Evaluation Example
11
From Evaluation to Actor Critic
12
Putting Discount Factors
13
Actor-critic algorithm (with discount)
14