1 of 14

Lecture 8.� Actor-Critic Algorithm�

Sookyung Kim

1

2 of 14

Recap: Policy Gradient

This is reward from time-step t to end �because of the causality

2

3 of 14

Taxonomy of RL algorithm

Actor-Critic

3

4 of 14

Improving the policy gradient: �Lowering Variance

t

4

5 of 14

Baseline Trick:�Lowering Variance

b

 

=

Advantage function

5

6 of 14

State & State-action Value Function

 

6

7 of 14

Value Function Fitting

 

 

 

 

7

8 of 14

Policy Evaluation

8

9 of 14

Monte Carlo evaluation with function approximation

9

10 of 14

Can we do better?

Bootstrapping

Might incorrect

Low variance (its not from single sample, but output from NN), �high bias (NN estimation can be bias)

10

11 of 14

Policy Evaluation Example

11

12 of 14

From Evaluation to Actor Critic

12

13 of 14

Putting Discount Factors

 

13

14 of 14

Actor-critic algorithm (with discount)

14