1 of 44

1

Federated Methods to speed up Reinforcement Learning

Sajad Khodadadian

Georgia Institute of Technology

2 of 44

2

Reinforcement Learning

3 of 44

3

Reinforcement Learning

My research: Theoretical Foundations of Reinforcement Learning

This talk: Federated Reinforcement Learning

4 of 44

4

Kendall et. al.: “Learning to Drive in a Day”, 2018

Reinforcement Learning

RL is data intensive!

factordaily.com

Multiple data collecting agents

5 of 44

5

Reinforcement Learning

 

shengpu-tang.me

Privacy matters!

6 of 44

6

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

7 of 44

7

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

8 of 44

8

Background on MDP Theory

 

 

Agent

Environment

 

 

9 of 44

9

 

Background on MDP Theory

 

 

 

Discount factor

Reward function

 

Initial state and action

 

policy

10 of 44

10

Background on MDP Theory

 

 

 

 

 

 

 

 

 

  • TD-learning

 

 

 

 

Markovian noise

11 of 44

11

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

12 of 44

12

Vanilla Distributed Reinforcement Learning

12

Agent 1

 

 

Central Agent

Agent j

Local Agents

 

 

 

 

  • Do not want to send data
  • Privacy concerns

13 of 44

13

  • Do not want to send data
  • Privacy concerns

Federated Reinforcement Learning

Agent 1

 

Local observations & policy

not shared with central agent

 

Agent j

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

14 of 44

14

Federated Reinforcement Learning

 

1Shen, et al. "Asynchronous advantage actor critic: Non-asymptotic analysis and linear speedup." arXiv preprint arXiv:2012.15511 (2020).

Conjecture

Yes! We are the first to show this.

Open problem: is there linear speedup in federated RL?

15 of 44

15

 

 

 

 

 

 

 

 

 

 

16 of 44

16

 

 

 

Convergence

Bias

Convergence Variance

17 of 44

17

 

 

 

 

Convergence

Bias

Convergence Variance

Higher order

 

 

 

 

 

 

18 of 44

18

Federated TD-learning

Agent 1

 

 

Agent j

 

 

 

 

 

 

 

 

 

 

 

19 of 44

19

Linear Speedup in Federated TD-Learning

 

 

 

 

 

 

 

 

 

 

20 of 44

20

  • Federated Supervised Learning:

  • Federated Reinforcement Learning:

Linear Speedup in Federated Learning

    • Linear speedup is possible [Khaled, et. al. ‘20], [Spiridonoff, et. al. ‘21], [….]
    • Key ingredient in these results: The noise is i.i.d.

    • [Wai ’20] [Zeng, et. al., ‘20]
    • They have linear penalty
    • Key challenge: Markov noise.
    • Linear speedup in federated TD-learning under i.i.d. noise assumption [Shen, et. al., ‘20]

21 of 44

21

Stochastic Approximation

 

 

TD-learning, off-policy

TD-learning, on-policy:

Contraction

22 of 44

22

Stochastic Approximation

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Extend to Federated!

23 of 44

23

Federated Stochastic Approximation (FedSAM)

Agent 1

 

 

Agent j

 

 

 

 

 

 

 

 

 

 

 

24 of 44

24

Federated Stochastic Approximation (FedSAM)

 

 

 

 

25 of 44

25

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

26 of 44

26

Proof Outline

 

  • i.i.d. noise
  • Markovian noise
  • Markovian noise

 

 

27 of 44

27

Proof Outline

 

variance

28 of 44

28

Proof Outline

 

variance

Moreau envelop

Refined analysis

29 of 44

29

Multiple Agents, Synchronous, i.i.d. Noise

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

30 of 44

30

Multiple Agents, Synchronous, Markovian Noise

 

 

?

 

 

 

 

 

 

 

 

 

 

 

 

?

?

?

 

Our refined analysis:

 

Simple bound

 

31 of 44

31

Multiple Agents, Synchronous, Markovian Noise

 

 

 

 

 

 

 

 

 

 

 

 

Mixing time

 

32 of 44

32

Multiple Agents, Asynchronous, Markovian Noise

 

Consensus error due to local updates

 

 

 

33 of 44

33

Key Takeaways

  1. Federated Reinforcement Learning

  • Key ideas in the proof

 

  • Lyapunov argument
  • Refined analysis based on Markov chain mixing
  • Analysis of synchronization error separately

34 of 44

34

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

35 of 44

35

My other works

 

 

 

Focus of this talk

Focus of my other work

36 of 44

36

My other works

Actor-Critic

Critic

Actor

 

 

 

  • Finite Sample Analysis of Off-policy Natural Actor-Critic Algorithm[KCM, International Conference on Machine Learning ‘20]

  • Finite Sample Analysis of Off-policy Natural Actor-Critic Algorithm With Linear Function Approximation [CKM, IEEE Control Syst. Lett. ‘22] [CKM , CDC ‘22]

 

suboptimal

37 of 44

37

My other works

Actor-Critic

Critic

Actor

 

  • On the Linear (and Super-Linear) Convergence of Natural Policy Gradient Algorithm [KJVM, Syst. Control. Lett. ‘22], [KJVM, CDC ‘21]

38 of 44

38

My other works

Actor-Critic

Critic

Actor

  • Finite Sample Analysis of Two-Time-Scale Natural Actor-Critic Algorithm [KDMJ, IEEE Trans. Automat. Contr. ‘22]

Two-Time-Scale Actor-Critic

39 of 44

39

My other works

                • I also have two pieces of work on Information Theoretic Fairness in Machine Learning

40 of 44

40

Outline

        • Background on Reinforcement Learning
        • Federated Reinforcement Learning

        • My other works
        • Future research directions
                  • Proof sketch

41 of 44

41

Multi Time-Scale Stochastic Approximation

Safe Reinforcement Learning

Markov games

Multi-agent Reinforcement Learning

42 of 44

42

Reinforcement Learning for Operations Research

43 of 44

43

Fairness in Reinforcement Learning

Fairness

44 of 44

44

Multi-Agent Systems

Reinforcement Learning

Fairness

Stochastic Processes

Control Theory

Information Theory

Optimization

Dynamic Programming