NEU/PSY 330
Reinforcement learning
Q
0225/2019
Rescorla Wagner learning rule
RW learning rule
An equivalent expression:
What does it mean?
RW learning rule
An equivalent expression:
What does it mean?
RW learning rule
An equivalent expression:
What does it mean?
RW learning rule
An equivalent expression:
What does it mean?
decides the step size of this movement
Q
What’s the limitation of the RW rule?
What’s the key idea of TD?
TD
TD, logic
If my previous expectation and my updated expectation are different …
… there is something to be learned
Q:
My previous expectation
My updated expectation
TD, intuition
An equivalent expression:
“Weight space”
decides the step size of this movement
Today, we will replicate …
dopamine signals RPE
Schultz 1992
dopamine signals RPE
Schultz 1992
dopamine signals RPE
Schultz 1992
Montague et al 1996
dopamine signals RPE
Trial 1 ?
Trial 30 ?
Trial 50 ?
Schultz 1992
dopamine signals RPE
Trial 1 no prediction, reward occurs
Trial 30 R predicted, no reward
Trial 50 R predicted, reward occurs
Schultz 1992
The training data, single trial
Trial w/o reward
Trial w/ reward
The training data, full experiment
Dopamine response, full experiment
The training data, single trial
Later trials…
Early trials...
The training data, full experiment
Extinction of response to the sensory cue
The exercises... let’ take a look together!
Question?