1 of 23

NEU/PSY 330

Reinforcement learning

Q

0225/2019

2 of 23

Rescorla Wagner learning rule

3 of 23

RW learning rule

An equivalent expression:

What does it mean?

4 of 23

RW learning rule

An equivalent expression:

What does it mean?

5 of 23

RW learning rule

An equivalent expression:

What does it mean?

6 of 23

RW learning rule

An equivalent expression:

What does it mean?

decides the step size of this movement

7 of 23

Q

What’s the limitation of the RW rule?

What’s the key idea of TD?

8 of 23

TD

9 of 23

TD, logic

If my previous expectation and my updated expectation are different …

… there is something to be learned

Q:

  • why is this reasonable?
  • what makes the updated expectation more accurate?

My previous expectation

My updated expectation

10 of 23

TD, intuition

An equivalent expression:

“Weight space”

decides the step size of this movement

11 of 23

Today, we will replicate …

12 of 23

dopamine signals RPE

Schultz 1992

13 of 23

dopamine signals RPE

Schultz 1992

14 of 23

dopamine signals RPE

Schultz 1992

Montague et al 1996

15 of 23

dopamine signals RPE

Trial 1 ?

Trial 30 ?

Trial 50 ?

Schultz 1992

16 of 23

dopamine signals RPE

Trial 1 no prediction, reward occurs

Trial 30 R predicted, no reward

Trial 50 R predicted, reward occurs

Schultz 1992

17 of 23

The training data, single trial

Trial w/o reward

Trial w/ reward

18 of 23

The training data, full experiment

19 of 23

Dopamine response, full experiment

20 of 23

The training data, single trial

Later trials…

Early trials...

21 of 23

The training data, full experiment

22 of 23

Extinction of response to the sensory cue

23 of 23

The exercises... let’ take a look together!

Question?