1 of 30

Part 2 - Going Beyond Behaviour Cloning with Off-Policy Reinforcement Learning

Adam Jelley

University of Edinburgh

2 of 30

Improving on Behaviour Cloning

  • What if we want to improve on the performance of the agent that collected the offline data?

    • Off-Policy Reinforcement Learning?

3 of 30

Review of Off-Policy RL

  • Off-policy RL algorithms learn from data collected from a policy which may be different to current policy (usually stored from earlier policy in replay buffer)
  • E.g. Q-learning

4 of 30

Review of Off-Policy RL - DDPG

  • Q-learning generalised to continuous actions -> Actor-Critic
    • Actor learns to maximise Q function
    • Critic learns Q function for actor policy
    • Example of generalised policy iteration

5 of 30

Review of Off-Policy RL - TD3

  • Effectively DDPG with “tricks”:
    • Clipped double Q-learning (“Twin”): Two Q networks to reduce overestimation of Q-values
    • Delayed policy updates (“Delayed”): Update value functions more frequently than policy
    • Target policy smoothing: Add noise to target actions to avoid local optima

6 of 30

Review of Off-Policy RL - SAC

  • Adopted the TD3 “tricks” (twin networks and delayed actor updates, but no noise required as policy smoothing already captured by max entropy objective)

  • Both SAC and TD3 are regularly used for continuous domains for their sample efficiency (due to being off-policy)

7 of 30

Improving on Behaviour Cloning with Off-Policy RL

  • Naively applying off-policy algorithms to offline RL (treating the offline data as the replay buffer) usually leads to policy collapse
  • As actor and critic optimise, go further out of distribution of training data
  • No additional interactions to provide feedback on extrapolation error
  • Values become noisy and policy ends up getting similar returns to random…

How can we handle the inevitable distribution shift when optimising return?

  • Early approaches introduce regularisation towards given dataset:
    • Policy Regularisation (regularise actions) e.g. TD3+BC
    • Critic Regularisation (regularise values) e.g. CQL

8 of 30

9 of 30

A Minimalistic Approach to Offline RL - Introduction

10 of 30

A Minimalistic Approach to Offline RL - Theory

11 of 30

A Minimalistic Approach to Offline RL - Experiments

12 of 30

A Minimalistic Approach to Offline RL - Experiments

13 of 30

14 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Introduction

Use uncertainty not regularisation!

15 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Uncertainty Penalisation with Q-Ensemble

16 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Uncertainty Penalisation with Q-Ensemble

17 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Uncertainty Penalisation with Q-Ensemble

18 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Uncertainty Penalisation with Q-Ensemble

19 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Ensemble Gradient Diversification

20 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Ensemble Gradient Diversification

21 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Ensemble Gradient Diversification

22 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Ensemble Gradient Diversification

23 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

24 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

25 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

26 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

27 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

28 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Experiments

29 of 30

Uncertainty Based Offline-RL with Diversified Q-Ensemble - Conclusion

30 of 30

Critique and Further Research Directions

  • Core idea makes a lot of sense
    • Use uncertainty rather than regularisation to control distribution shift

  • Implementation can be improved
    • Building on SAC means underlying approach is well-developed
    • Ensemble similarity loss term seems quite empirical - how can we get a better measure of uncertainty?
    • With a more complete posterior Q distribution, how do we best utilise this uncertainty in both the actor and the critic?

  • Despite shortcomings, approach works very well (stable, efficient and performant), at least on standard benchmarks

  • Potential for this approach to be applied in the online/fine-tuning setting too, using the offline data to solve the “cold-start” problem