Curiosity-driven Exploration�by Self-supervision Prediction�Deepak Pathak, Pukit Agrawal, Alexei Efros, Trevor Darrell�ICML 2017
Resources:�https://archive.org/details/Redwood_Center_2017_10_11_Deepak_Pathak
LLNL-PRES-705839
1
Motivation
LLNL-PRES-705839
2
Not so intelligent AI
Labeling Training Set
Test in Testset
Cat
LLNL-PRES-705839
3
Life-long learning in human
Passive
Is imitation learning the route to humanoid robot?, Schaal, 1999
Active
The Scientist in a crib, Gopnik, Metzoff &Kuhl,1999
machines
LLNL-PRES-705839
4
Reinforcement Learning �
Extrinsic reward�(Extrinsic of agent )
LLNL-PRES-705839
5
Reinforcement Learning �
Typically, RL requires very dense reward
Mnih et al, Nature 2015
Jaderberg et al, ICLR 2017
Extrinsic reward
LLNL-PRES-705839
6
Rewards are sparse in real world …
Now
In real world, reward can delivered by days, months or years!
30 years later
“curiosity”
LLNL-PRES-705839
7
Intrinsic Motivation/Curiosity
LLNL-PRES-705839
8
Model
LLNL-PRES-705839
9
How can we design “curiosity”
Prediction Uncertainty
Visitation Counts
LLNL-PRES-705839
10
On ‘Mental Model’ of the prediction …
If the organism carries a small scale model of external reality and its own possible actions within its head, it is able to try out various alternatives, conclude which is the best of them, react to future situations before they arise, utilize the knowledge of the past in dealing with present and the future, and in every way react in much fuller, safer and more competent manner to emergencies which face it. ��- Kenneth Craik, 1943 chapter 5, page 61
LLNL-PRES-705839
11
Design of curiosity
Action
Observation
LLNL-PRES-705839
12
Design of curiosity
Action
Observation
LLNL-PRES-705839
13
Design of curiosity
Action
Observation
LLNL-PRES-705839
14
Design of curiosity
Action
Observation
LLNL-PRES-705839
15
Design of curiosity
Action
Observation
Do nothing
LLNL-PRES-705839
16
Design of curiosity
LLNL-PRES-705839
17
Design of curiosity
LLNL-PRES-705839
18
Overview
Train a Model
Predict consequences of the action
Bad prediction 🡪 higher curiosity
LLNL-PRES-705839
19
Training RL with external reward
External Reward
Policy Network
Curiosity Reward (intrinsic)�
+
LLNL-PRES-705839
20
Intrinsic curiosity module (ICM)
Forward�Model
Curiosity in pixel-space is hard
[Schimidhuber 2001]
LLNL-PRES-705839
21
LLNL-PRES-705839
22
LLNL-PRES-705839
23
Intrinsic curiosity module (ICM)
Forward�Model
Curiosity in pixel-space is hard
LLNL-PRES-705839
24
Intrinsic curiosity module (ICM)
Forward�Model
LLNL-PRES-705839
25
Experiments
LLNL-PRES-705839
26
Self-supervised Curiosity
LLNL-PRES-705839
27
No external reward, only curiosity
LLNL-PRES-705839
28
No external reward, only curiosity
LLNL-PRES-705839
29
Do these skills generalize?�: Main problem of RL🡪 not generalizable
Trained on level 1
Tested on level 2
Trained on level 1
Tested on level 3
Curriculum learning
(teacher’s forcing)
LLNL-PRES-705839
30
Evaluate curiosity + extrinsic reward�
VizDoom Game�Note: Agent does not have access to Map
LLNL-PRES-705839
31
Evaluate curiosity + extrinsic reward�
“Sparse”
Very Sparse
LLNL-PRES-705839
32
Is it robust?
LLNL-PRES-705839
33
Is it robust?
Robust to irrelevant part of the observation.
LLNL-PRES-705839
34
Demo video
LLNL-PRES-705839
35
Summary
LLNL-PRES-705839
36
LLNL-PRES-705839
37