Comparative Analysis of Deep Reinforcement Learning for Trajectory Estimation of Autonomous Vehicles In Highways
Prepared by:
Md Adnan Faisal Hossain
ID: 1606063
Supervised by:
Dr. S.M. Mahbubur Rahman
EEE, BUET
Motivation for Autonomous Driving
2
Highway Autonomous Driving
A Review of Motion Planning for Highway Autonomous Driving - Laurène et al
3
Why use Reinforcement Learning?
4
Reinforcement Learning at a glance
5
Policy (π)
Reinforcement Learning Algorithm
AGENT
ENVIRONMENT
State (s)
Action (a)
Reward (Rt)
Policy
Update
Mathematical Foundations of Reinforcement Learning
6
Reinforcement Learning is a method to obtain an optimal policy that generates an optimal state-action-value function from a given Markov Decision Process (MDP)
MDP – (s, A, {Psa}, γ, R)
S : set of states
A : set of actions
Psa(s`) : state transition probability
γ: Discount factor
Rt : Reward from current action in current state
Gt : Total discounted reward (payoff)
Vπ(s) : Value function
V*(s) : Optimal Value function
Qπ(s, a) : Action-value (Q) function
π(s) : RL Policy
π*(s) : Optimal RL Policy
Expected Future Reward
Basic Actor-Critic Reinforcement Learning Algorithm
7
For each episode:
Repeat for each time step:
Actor
Policy
Network
Critic
Policy
Network
Current State (s)
Action values (a)
Q(s, a)
Actor
Target
Network
Critic
Target
Network
Next State (s’)
Action values (a’)
Q’(s, a)
Description of the work
8
Contribution to the literature:
Techniques incorporated to improve performance:
Reinforcement Learning Algorithms
“Continuous control with deep reinforcement learning” - Timothy et al
“Proximal Policy Optimization Algorithms” – John et al
“Trust Region Policy Optimization” – John et al
“Asynchronous Methods for Deep Reinforcement Learning” – Volodymyr et al
9
All the algorithms are implemented in an actor-critic style DRL model
Deep Reinforcement Learning model setup
10
Typical end-to-end Deep RL model setup for autonomous driving –
Deep RL model setup in the proposed literature –
Deep Reinforcement Learning model setup
"Optimal trajectories for time-critical street scenarios using discretized terminal manifolds" - Werling et al
11
Trajectories τ
s
Ego vehicle
s0
d
s
d
s
d
Action space
T: Set of spatiotemporal trajectories called lattice
T = {τ1 , τ2 ……. τm}
Trajectory τ = [ s(t) , d(t) ] , τ ϵ R2
Policy Network
12
(15, 2, 30)
Number of filters
Actor
(1,2)
Ego
(14,2) obstacles
(128, 128)
Front
Image
(128, 128)
Rear
Image
(128, 128, 3, 30)
(128, 128, 3, 30)
(𝛼, 𝛽)
State from past 30 time steps
Conv 2D
(32, 4, 2)
Conv 2D
(64, 4, 1)
Conv 2D
(64, 3, 1)
Global Average pooling
Conv 3D
(64, 4, 2)
Conv 3D
(32, 4, 1)
Global Average pooling + reshape
Conv 3D
(64, 4, 2)
Conv 3D
(32, 4, 1)
Global Average pooling + reshape
+
+
FC (64)
FC (32)
FC (64)
FC (32)
Critic
TD error
Output
(3,1)
tf
df
vf
Filter size
Stride
(1, 1, 128)
Q(s, a)
Input States
Continuous Actions
Reward Function
"Integrating Deep Reinforcement Learning with Model-based Path Planners for Automated Driving" - Ekim et al
13
, if collision or off-road
, otherwise
, speed gain > 5 km/h
, otherwise
Rt : Reward obtained for performing action a in state s
rc : Penalty for collision
rv : reward for driving above minimum speed
rlc : lane change reward/penalty
rd : reward for safety
Training Process
"Congested Traffic States in Empirical Observations and Microscopic Simulations" - Martin et al
14
Training steps:
Safety Net:
TC : Calculated time to collision
THB : Threshold for hard-braking
TB : Threshold for braking
as: Safety action
Simulation Environment
"CARLA: An Open Urban Driving Simulator" - Alexey et al
15
CARLA Environment setup:
1
Ego
2
3
4
5
6
7
8
9
10
11
12
13
14
Sample Simulations
16
17
Quantitative Analysis
"Quantitative Evaluation of Autonomous Driving in CARLA" - Shang et al
18
Performance Metrics:
TTC(s)
time(s)
TTCmin
Quantitative Analysis
19
Results:
| TRPO | DDPG | A2C | PPO2 |
Speed | 85 | 67 | 73 | 53 |
Safety | 39 | 62 | 51 | 51 |
Comfort | 48 | 79 | 53 | 21 |
Average | 57 | 69 | 59 | 42 |
Collision rate (%) | 13 | 4 | 9 | 87 |
Mean Reward Curve (Training Curve):
20
TRPO
DDPG
PPO2
A2C
Conclusion and Future Work
21
Conclusion:
Future work: