1 of 37

Curiosity-driven Exploration�by Self-supervision Prediction�Deepak Pathak, Pukit Agrawal, Alexei Efros, Trevor Darrell�ICML 2017

Resources:�https://archive.org/details/Redwood_Center_2017_10_11_Deepak_Pathak

LLNL-PRES-705839

1

2 of 37

Motivation

LLNL-PRES-705839

2

3 of 37

Not so intelligent AI

  • Task-specific
  • Not adaptive
  • Expensive

Labeling Training Set

Test in Testset

Cat

LLNL-PRES-705839

3

4 of 37

Life-long learning in human

Passive

Is imitation learning the route to humanoid robot?, Schaal, 1999

Active

The Scientist in a crib, Gopnik, Metzoff &Kuhl,1999

machines

LLNL-PRES-705839

4

5 of 37

Reinforcement Learning

  • Reinforcement Learning:�Input: Given environment which provides numerical reward signal, and agent which act inside of that environment �Output: Let agent learn how to take actions(policy) in order to maximize reward.

  • Goal: Learn how to take actions in order to maximize reward
  • Design RL: Objective, State, Action, Reward

Extrinsic reward�(Extrinsic of agent )

LLNL-PRES-705839

5

6 of 37

Reinforcement Learning

Typically, RL requires very dense reward

Mnih et al, Nature 2015

Jaderberg et al, ICLR 2017

Extrinsic reward

LLNL-PRES-705839

6

7 of 37

Rewards are sparse in real world …

Now

In real world, reward can delivered by days, months or years!

30 years later

“curiosity”

LLNL-PRES-705839

7

8 of 37

Intrinsic Motivation/Curiosity

LLNL-PRES-705839

8

9 of 37

Model

LLNL-PRES-705839

9

10 of 37

How can we design “curiosity”

  • Let the agent Incentives to go to area where it didn’t visit before.
  • Problem: No generalization
  • Building model of prediction
  • Potentially generalize to novel scenarios

Prediction Uncertainty

Visitation Counts

LLNL-PRES-705839

10

11 of 37

On ‘Mental Model’ of the prediction …

If the organism carries a small scale model of external reality and its own possible actions within its head, it is able to try out various alternatives, conclude which is the best of them, react to future situations before they arise, utilize the knowledge of the past in dealing with present and the future, and in every way react in much fuller, safer and more competent manner to emergencies which face it. ��- Kenneth Craik, 1943 chapter 5, page 61

LLNL-PRES-705839

11

12 of 37

Design of curiosity

Action

Observation

LLNL-PRES-705839

12

13 of 37

Design of curiosity

Action

Observation

LLNL-PRES-705839

13

14 of 37

Design of curiosity

Action

Observation

LLNL-PRES-705839

14

15 of 37

Design of curiosity

Action

Observation

LLNL-PRES-705839

15

16 of 37

Design of curiosity

Action

Observation

Do nothing

LLNL-PRES-705839

16

17 of 37

Design of curiosity

 

 

 

 

LLNL-PRES-705839

17

18 of 37

Design of curiosity

 

 

 

 

 

LLNL-PRES-705839

18

19 of 37

Overview

Train a Model

Predict consequences of the action

Bad prediction 🡪 higher curiosity

LLNL-PRES-705839

19

20 of 37

Training RL with external reward

External Reward

 

 

Policy Network

Curiosity Reward (intrinsic)�

+

 

 

LLNL-PRES-705839

20

21 of 37

Intrinsic curiosity module (ICM)

 

 

 

 

Forward�Model

 

 

Curiosity in pixel-space is hard

[Schimidhuber 2001]

LLNL-PRES-705839

21

22 of 37

LLNL-PRES-705839

22

23 of 37

LLNL-PRES-705839

23

24 of 37

Intrinsic curiosity module (ICM)

  • Encode relevant parts
    • That the agent can affect
    • That can affect the agent

 

 

 

 

Forward�Model

 

 

 

 

Curiosity in pixel-space is hard

LLNL-PRES-705839

24

25 of 37

Intrinsic curiosity module (ICM)

 

 

 

 

Forward�Model

 

 

 

 

 

 

 

LLNL-PRES-705839

25

26 of 37

Experiments

LLNL-PRES-705839

26

27 of 37

Self-supervised Curiosity

  1. Does it work when no external reward, only curiosity?
  2. Do learned skills generalize?
  3. Evaluate curiosity + extrinsic reward
  4. Is it robust?

LLNL-PRES-705839

27

28 of 37

No external reward, only curiosity

  • Mario:�Learning to play with no reward
    • At the start of training, agent repeat similar actions
    • After some training
      • Emergent behavior
        • Jumping enemies, pipes and pits
        • Killing enemies
        • Staying alive�:Skills learned in order to become curious (unpredictable)
      • Learn to cross over 30% of level 1

LLNL-PRES-705839

28

29 of 37

No external reward, only curiosity

  • VizDoom (3D navigation):� Coverage during exploration
    • Visitation of curious (ICM) > visitation of random exploration

LLNL-PRES-705839

29

30 of 37

Do these skills generalize?�: Main problem of RL🡪 not generalizable

Trained on level 1

Tested on level 2

Trained on level 1

Tested on level 3

Curriculum learning

(teacher’s forcing)

LLNL-PRES-705839

30

31 of 37

Evaluate curiosity + extrinsic reward�

VizDoom Game�Note: Agent does not have access to Map

LLNL-PRES-705839

31

32 of 37

Evaluate curiosity + extrinsic reward�

“Sparse”

Very Sparse

LLNL-PRES-705839

32

33 of 37

Is it robust?

LLNL-PRES-705839

33

34 of 37

Is it robust?

Robust to irrelevant part of the observation.

LLNL-PRES-705839

34

35 of 37

Demo video

LLNL-PRES-705839

35

36 of 37

Summary

  1. Learn skills with no extrinsic reward (only curiosity)
  2. Skills generalize to novel scenarios
  3. Effective exploration when extrinsic reward is sparse
  4. Robust to irrelevant part of the observation.

LLNL-PRES-705839

36

37 of 37

LLNL-PRES-705839

37