1 of 47

Lecture 4��DQN and Its Extensions

1

Instructor: Ercan Atam

Institute for Data Science & Artificial Intelligence

Course: DSAI 642- Advanced Reinforcement Learning

2 of 47

2

List of contents for this lecture

  • DQN: Deep Q-Networks

  • DQN with experience replay

  • Double DQN

  • Duelling DQN

  • Prioritized experience replay

3 of 47

3

Relevant readings for this lecture

  • Chapter 8 of Shiyu Zhao, “Mathematical Foundations of Reinforcement learning”, Springer, 2025.

(Some slides at the beginning are modified/improved versions from the presentation slides of Shiyu Zhao)

  • Chapter 2 of Laura Graesser and Wah Loon Keng, “Foundations of Deep Reinforcement Learning:

Theory and Practice in Python”, Addison-Wesley Professional, 2019.

  • Chapter 13 of Richard S. Sutton and Andrew G. Barto, “Reinforcement Learning: An Introduction”,

Second Edition, MIT Press, Cambridge, MA, 2018.

  • Chapter 7 of Nimish Sanghi, “Deep Reinforcement Learning with Python”, Second Edition, APress,

2024.

(Some slides at the end are modified/improved versions from the presentation slides of Alina Vereshchaka)

4 of 47

4

A convention on notation

5 of 47

5

Deep Q-Learning (1)

6 of 47

6

Deep Q-Learning (2)

7 of 47

7

Deep Q-Learning (3)

8 of 47

8

Deep Q-Learning (4)

9 of 47

9

Deep Q-Learning (5)

10 of 47

10

Deep Q-Learning (6)

11 of 47

11

Deep Q-Learning: Two networks (1)

12 of 47

12

Deep Q-Learning: Two networks (2)

13 of 47

13

Deep Q-Learning: Experience replay (1)

14 of 47

14

Deep Q-Learning: Experience replay (2)

15 of 47

15

Deep Q-Learning: Experience replay (3)

16 of 47

16

Deep Q-Learning: Experience replay (4)

17 of 47

17

Polyak averaging in DQN (1)

Polyak averaging (also called soft target updates) is a technique used in DQN to improve training stability

and convergence by gradually updating the target network parameters, rather than updating them periodically

—which can sometimes cause abrupt changes.

18 of 47

18

Polyak averaging in DQN (2)

 

Why is it useful?

  • Smoother target changes, leading to more stable Q-value estimates.
  • Reduced risk of divergence due to abrupt parameter shifts.
  • Better performance in complex or noisy environments.
  • Faster convergence in many cases due to more consistent learning signals.

Advantages over periodic (hard) updates:

19 of 47

19

DQN Algorithm

“m”: for the main network

“tr”: for the target network

20 of 47

20

Double DQN (DDQN) (1)

Remember “Double Q-Learning”:

21 of 47

21

Double DQN (DDQN) (2)

22 of 47

22

Double DQN (DDQN) (3)

23 of 47

23

Summary: DQN versus Double DQN

24 of 47

24

Double DQN Algorithm

25 of 47

25

Duelling DQN (1)

Definition: Advantage Function

 

26 of 47

26

Duelling DQN (2)

27 of 47

27

Duelling DQN (3)

Example:

28 of 47

28

Duelling DQN (4)

Example (continue):

29 of 47

29

Duelling DQN (5)

30 of 47

30

Duelling DQN (6)

31 of 47

31

Example 1 for understanding duelling DQN motivation (1)

32 of 47

32

Example 1 for understanding duelling DQN motivation (2)

33 of 47

33

Example 2 for understanding duelling DQN motivation

34 of 47

34

Why can duelling DQN be more sample efficient?

35 of 47

35

Improvements provided by the duelling DQN

Fig: Improvements of the dueling architecture over a baseline single network.

36 of 47

36

Duelling DQN: summary

37 of 47

37

Prioritized Experience Replay (PER)

38 of 47

38

PER: TD Error

39 of 47

39

PER: How to get priorities? (1)

40 of 47

40

PER: How to get priorities? (2)

41 of 47

41

PER: Adjustment through use of importance sampling (1)

42 of 47

42

PER: Adjustment through use of importance sampling (2)

43 of 47

43

PER: Adjustment through use of importance sampling (3)

44 of 47

44

PER: Adjustment through use of importance sampling (4)

45 of 47

45

Double DQN with PER

46 of 47

46

PER: Summary

47 of 47

References �(utilized for preparation of lecture notes or Matlab code)

  • Shiyu Zhao, “Mathematical Foundations of Reinforcement learning”, Springer, 2025.
  • Laura Graesser and Wah Loon Keng, “Foundations of Deep Reinforcement Learning: Theory and Practice in Python”, Addison-Wesley Professional, 2019.
  • Richard S. Sutton and Andrew G. Barto, “Reinforcement Learning: An Introduction”, Second Edition, MIT Press, Cambridge, MA, 2018.
  • https://cse.buffalo.edu/~avereshc/rl_fall19/lecture_14_1_Dueling_DQN_and_PER.pdf

47