1 of 69

From Human Language

to Agent Action

Jesse Thomason

University of Washington

Mohit Shridhar

Dieter�Fox

Luke Zettlemoyer

Yonatan Bisk

Daniel Gordon

Roozbeh Mottaghi

Winson�Han

Michael Murray

Maya Cakmak

2 of 69

We Use Language to Instruct Machine Agents

No visual connections.

Require visual grounding.

2

3 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

3

4 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

4

5 of 69

Vision-and-Language Navigation

  • MatterPort Room-to-Room.
  • Navigation-only instruction following
  • Left/right/up/down/forward.
  • Discrete navigation graph.
  • Objective: get close to the goal node.

5

[Anderson et al., CVPR’18]

6 of 69

“Turn around and exit the library, head down the…”

6

[Anderson et al., CVPR’18]

7 of 69

“Turn around and exit the library, head down the…”

7

[Anderson et al., CVPR’18]

8 of 69

Sequence-to-Sequence Model

“Turn around and exit the library, head down the stairs and across the room.”

After each action, get a new visual observation from the environment.

8

...

...

â0

â1

ân

RN

RN

RN

t0

t1

t2

LE

LE

LE

Learned, token-level Language Embedding

LSTM Encoder

LSTM Decoder

Fixed, ResNet-152 Embedding Network

[Anderson et al., CVPR’18]

We will build on this model.

9 of 69

Unimodal Model Ablations

9

Embodied QA�[Das et al., CVPR’18]

Action, vision-, and language-only models.

Beats baseline

Beats initial model!

Action, vision-, and language-only models.

Language-only model.

Vision-only model.

Room-2-Room�[Anderson et al., CVPR’18]

<EOS>

RN

LE

...

LE

<EOS>

RN

LE

<EOS>

RN

LE

Vision-only

Lang-only

Action-only

[Thomason et al., NAACL’19]

Models and data may fail to address underlying vision+language challenges.

10 of 69

“Turn around and exit the library, head down the…”

  • Small action space and short expert demonstrations.
  • 24% → 80% of human performance since ‘18.
  • Instructions are also not what we’d like to tell a robot.

10

11 of 69

“Go clean the room with a plant.”

11

iRobot

12 of 69

“Go clean the room with a plant.”

  • This instruction is underspecified.

12

iRobot

13 of 69

“Go clean the room with a plant.”

  • This instruction is underspecified.
  • This instruction is ambiguous.

13

iRobot

14 of 69

In This Talk

14

Use natural language instruction following to complete high-level goals.

15 of 69

Ask questions to get additional supervision!

  • Ask: “Should I continue into the living room or go right towards the kitchen?”

15

iRobot

16 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
    • CVDN dataset
    • Navigation from dialog history
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

16

17 of 69

17

Guidance Abstraction

Language Source

Templates

Humans

Semantic

Visual

Talk the Walk�[de Vries et al., arXiv’18]

VLNA [Nguyen et al., CVPR’19]

Vision-and-Dialog Navigation

  • Human-human dialogs
  • Both participants get an egocentric scene view.

18 of 69

Outline

  • Background - Language grounding in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
    • CVDN dataset
    • Navigation from dialog history
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

18

19 of 69

19

...

Hint: The goal room contains a mat.

Into the hall or the office?

Left into the hall.

Follow it to a living room.

...

...

Should I go upstairs?

Visible to both Navigator and Oracle

Visible Only to the Oracle

-- this target object is present in at least two rooms, but only one is correct.

-- A shortest-path planner’s next steps; up to 5 navigation nodes in the direction of the goal.

20 of 69

Navigator View

Oracle View

20

21 of 69

Dialog Enables Longer Paths

  • Path Length Average:
    • Human (25.0); Planner (17.4)
    • Room-to-Room (6.0)

CVDN

R2R

21

22 of 69

Shared Visual Context Yields Egocentric Language

22

The goal room contains a rug.

Navigator: Should I go to the left or right?

Oracle: Go left and turn right after the bathroom.

Navigator: Do I need to go in the room with the run or keep on going right?

Oracle: Turn right and take the tiny hallway on the right. You will ascend the stairs you find on the right.

Navigator: Should I go into the kitchen or to the right?

Oracle: Turn toward the front door and go up the stairs you see on the right.

Navigator: Do I go left or right?

Oracle: Go along the railing to the right. Stop at the room with a brown chair.

3 steps

4 steps

7 steps

6 steps

4s

23 of 69

Dialog Leads to Rich Language

  • Average total words:
    • CVDN (82)
    • Room-to-Room (29)

CVDN

R2R

23

Should I go to the left or right?

Go left and turn right after the bathroom.

Do I need to go in the room with the run or keep on going right?

Turn right and take the tiny hallway on the right. You will ascend the stairs you find on the right.

Should I go into the kitchen or to the right?

Turn toward the front door and go up the stairs you see on the right.

Do I go left or right?

Go along the railing to the right. Stop at the room with a brown chair.

Walk between the columns and make a sharp turn right. Walk down the steps and stop on the landing.

24 of 69

Outline

  • Background - Language grounding in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
    • CVDN dataset
    • Navigation from dialog history
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

24

25 of 69

Navigation from Dialog History

25

The goal room contains a rug.

Navigator: Should I go to the left or right?

Oracle: Go left and turn right after the bathroom.

Navigator: Do I need to go in the room with the run or keep on going right?

Oracle: Turn right and take the tiny hallway on the right. You will ascend the stairs you find on the right.

Navigator: Should I go into the kitchen or to the right?

Oracle: Turn toward the front door and go up the stairs you see on the right.

Navigator: Do I go left or right?

Oracle: Go along the railing to the right. Stop at the room with a brown chair.

3 steps

4 steps

6 steps

4s

...

Navigator

7 steps

26 of 69

Navigation from Dialog History

  • Input: History so far + visual frame per timestep.
  • Output: Navigation action per timestep.
  • Goal: Get closer to the target object room.
  • 2k dialogs → 7k histories.

26

The goal room contains a rug.

Navigator: Should I go to the left or right?

Oracle: Go left and turn right after the bathroom.

Navigator: Do I need to go in the room with the run or keep on going right?

Oracle: Turn right and take the tiny hallway on the right. You will ascend the stairs you find on the right.

Navigator: Should I go into the kitchen or to the right?

Oracle: Turn toward the front door and go up the stairs you see on the right.

Navigator: Do I go left or right?

Oracle: Go along the railing to the right. Stop at the room with a brown chair.

3 steps

4 steps

7 steps

6 steps

4 steps

27 of 69

Initial, Sequence-to-Sequence Model

27

...

...

â0

â1

ân

RN

RN

RN

mat

<EOS>

Oracle: Through the lobby. So go through the door next to the green towel. Go to the left door next to the two yellow lights. Walk straight to the end of the hallway and stop

Navigator: Are these the yellow lights you were talking about?

Target

LE

LE

LE

<TAR>

<ORA>

Yeah

,

head

up

the

stairs

.

Last Answer

LE

LE

LE

LE

LE

LE

LE

LE

<NAV>

Should

I

go

upstairs

?

Last Question

LE

LE

LE

LE

LE

LE

Navigator: Should I turn left down the hallway ahead?

Oracle: ya

<NAV>

Into

the

hall

or

the

office

?

<ORA>

Left

into

the

hall

.

Follow

it

to

a

living

room

.

All Previous Questions and Answers

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

LE

28 of 69

Evaluation - Test (epoch of best Val Unseen)

  • Unseen Envs:
    • Novel dialogs.
    • Novel houses not seen during training.

1.90

2.05

2.27

2.35

28

Adding Dialog History Helps in Unseen Environments.

29 of 69

Evaluation - Unimodal Baselines

9.52

9.76

0.91

0.52

5.72

1.74

1.58

1.40

5.92

2.35

Action-only

Vis-�only

Lang-�only

29

<EOS>

RN

LE

<EOS>

RN

LE

...

LE

<EOS>

RN

LE

Initial Model Uses Multimodal Input in Unseen Environments.

Lots of Headroom for Future Models

30 of 69

“Go clean the room with a plant.”

  • Visual semantic navigation is just a single step towards the long-term goal.

30

31 of 69

“Brown a potato slice.”

  • What we ultimately want is robots that cause change in the world.

31

32 of 69

“Brown a potato slice.”

  • Interactive worlds complicate the state space with object states.

32

33 of 69

“Brown a potato slice.”

  • Interactive worlds complicate the state space with object positions.

33

34 of 69

“Brown a potato slice.”

  • High-level instructions are underspecified.

34

35 of 69

Use both high- and low-level instructions

  • “Pick up the knife on the counter beside the utensils, then turn right to face the island…”

35

36 of 69

Outline

  • Background - Language grounding in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
    • ALFRED benchmark
    • Visual-and-Language Planning
  • Next steps

36

37 of 69

Long-term Aspiration in Robotics

“Turn left and head to the stove counter.”

rotate(-90)�forward(3m)

“Pick up the knife on the counter beside the toaster.”

Human Instructor

Robot Actions

Robot Actuation

grasp_at(coords)

Base motor accelerations

Arm and gripper control

37

38 of 69

Translating Instructions to Actions

“Turn left and head to the stove counter.”

rotate(-90)�forward(3m)

“Pick up the knife on the counter beside the toaster.”

Human Instructor

Robot Actions

grasp_at(coords)

38

Understanding language requires context

Too see lots of diverse language, can utilize a big dataset.

39 of 69

Brown a potato slice. Turn left and head to the stove counter. Pick up the knife on the counter beside the toaster. Turn right then face to the island. Slice the potato. ...

39

Room-to-Room�[Anderson et al., CVPR’18]

VirtualHome�[Puig et al., CVPR’18]

Low-level actions.

Only about navigation.

About interactive tasks.

Semantics-level actions.

ALFRED

40 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
    • ALFRED benchmark
    • Visual-and-Language Planning
  • Next steps

40

41 of 69

Action

Learning

From

Realistic

Environments (and)

Directives

41

42 of 69

Action

Learning

From

Realistic

Environments (and)

Directives

42

iRobot

Boston Dynamics

Learn new actions by chaining together reliable, low-level skills.

Reliable, static behaviors.

43 of 69

Action

Learning

From

Realistic

Environments (and)

Directives

43

44 of 69

Action

Learning

From

Realistic

Environments (and)

Directives

44

45 of 69

(Stack, Fork, Cup, CounterTop, Kitchen3)

Task Tuple

Trajectory

Trajectory

Trajectory

45

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Language Instructions

Planner with Full Observability

(x,y,z) | is_fork(x) & is_cup(y) & on(x, y) & is_counter(z) & on(y, z)

46 of 69

ALFRED Tasks

  • We define several high-level household tasks.

Pick & Place

Double Place

Stack

Examine

46

47 of 69

ALFRED Tasks

  • We define several high-level household tasks.

Heat

Cool

Rinse

47

48 of 69

ALFRED Tasks

48

Free-form, open-vocabulary annotations.

49 of 69

ALFRED has Long Demonstrations

49

Room-to-Room

VirtualHome

Touchdown

ALFRED has as many demonstrations as others (+8k), but much longer�(50 steps versus R2R’s 6).

# Demonstrations

Actions per Demonstration

[Chen et al., CVPR’19]

50 of 69

ALFRED has Long Instructions + More Annotations

50

ALFRED has more instructions (+25k) and they are as long or longer than most (80 tokens versus R2R’s 29).

# Annotations

Words per Instruction

51 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
    • ALFRED benchmark
    • Visual-and-Language Planning
  • Next steps

51

52 of 69

Action and State Space

  • Chain together low-level actions to accomplish goals.
  • Navigation + Manipulation actions:
    • Predict an interaction mask (wrapper for AI2THOR).

Put In

Toggle On

52

53 of 69

Seq2Seq Model with Progress Monitoring

  • Auxiliary Output: Estimate of progress towards goal.
  • Effective for Room-to-Room Navigation.

53

[Ma et al., ICLR’19]

Estimate 0.5 progress.

54 of 69

Seq2Seq Model

Put�In

Visual Encoding (t)

Encoding

vt

Resnet Conv5

LSTM

at - 1

ht

Action Decoder

at ; mt

Language Instructions Encoding (t=0)

Put a chilled cup of water in the cupboard. Take a step to your right then walk to the counter in front of you then turn right and walk to the fridge. Open the freezer and grab the glass from behind the egg then close the door. Turn around and walk to the counter then turn right. Put the cup in the sink then full it with water and pick it back up remembering to shut off the tap. ...

x1

x2

x3

...

xL

BiLSTM Encoder

x

Attention over x

xt

55 of 69

Seq2Seq Model with Progress Monitoring

Visual Encoding (t)

Encoding

vt

Resnet Conv5

LSTM

at - 1

ht

Action Decoder

at ; mt

Language Instructions Encoding (t=0)

Put a chilled cup of water in the cupboard. Take a step to your right then walk to the counter in front of you then turn right and walk to the fridge. Open the freezer and grab the glass from behind the egg then close the door. Turn around and walk to the counter then turn right. Put the cup in the sink then full it with water and pick it back up remembering to shut off the tap. ...

x1

x2

x3

...

xL

BiLSTM Encoder

x

Attention over x

xt

Progress Monitor

st ; pt

Predict normalized step num st/T.

Predict step progress in [0, 1], pt.

Reasonable adaptation of a model that works well for visual navigation tasks.

56 of 69

Unimodal Ablations - Seen Rooms

56

57 of 69

Unimodal Ablations - Seen Rooms

57

Vision-only can only learn the next step of an expert demonstration.

ALFRED does not exhibit powerful unimodal bias.

58 of 69

Performance in Unseen Rooms

58

59 of 69

Performance in Unseen Rooms

59

ALFRED demonstrations are non-trivially more complex than navigation.

Explicit Memory? Hierarchy? Modularity? Symbolic planning?

60 of 69

Outline

  • Background - Instruction following in visual environments
  • Vision-and-Dialog Navigation [CoRL’19]
  • Interpreting Grounded Instructions for Everyday Tasks� [in submission]
  • Next steps

60

61 of 69

We Use Language to Instruct Machine Agents

No visual connections.

Require visual grounding.

61

62 of 69

In This Talk

62

...

Hint: The goal room contains a mat.

Into the hall or the office?

Left into the hall.

Follow it to a living room.

Vision-and-Dialog Navigation [CoRL’19]

Interpreting Grounded Instructions�for Everyday Tasks [in submission]

Longer dialog context improves navigation in unseen environments.

Out of the box, successful navigation models are not sufficient for ALFRED.

Integrating language and vision is necessary for more complex tasks.

63 of 69

Long-term Aspiration in Robotics

“Turn left and head to the stove counter.”

rotate(-90)�forward(3m)

“Pick up the knife on the counter beside the toaster.”

Human Instructor

Robot Actions

Robot Actuation

grasp_at(coords)

Base motor accelerations

Arm and gripper control

63

Train agents for the ALFRED benchmark.

ALFRED also reveals some complexities for what comes next.

64 of 69

ALFRED Reveals API Ambiguity

“Turn on the sink.”

grasp_at(coords)

Arm and gripper control

Human Instructor

Robot Actions

Robot Actuation

64

65 of 69

ALFRED Reveals API Ambiguity

“Open the _____.”

grasp_at(coords)

Arm and gripper control

Human Instructor

Robot Actions

Robot Actuation

65

We will always need on the fly, arbitrarily low level language instructions!

Any API for robot control from language may miss low-level details!

Dialog-enabled agents can request this supervision as needed.

66 of 69

66

...

Hint: The goal room contains a mat.

Into the hall or the office?

Left into the hall.

Follow it to a living room.

...

...

Should I go upstairs?

Visible to both Navigator and Oracle

Visible Only to the Oracle

Navigation

Question Generation

Question Answering

67 of 69

Mohit Shridhar

Dieter�Fox

Luke Zettlemoyer

Yonatan Bisk

Daniel Gordon

Roozbeh Mottaghi

Winson�Han

Michael Murray

Maya Cakmak

From Human Language to Agent Action

  • Language can provide structure for difficult instruction following tasks.
  • Vision-and-Dialog Navigation. Jesse Thomason, Michael Murray, Maya Cakmak, Luke Zettlemoyer. CoRL’19.
  • Interpreting Grounded Instructions for Everyday Tasks.�Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, Dieter Fox. In submission.

68 of 69

“Turn around and exit the library, head down the…”

  • Train to predict the action a shortest-path�planner would take from the current state.

68

â0=F

*a0=R

RN

â0

...

LE

<EOS>

LE

[Anderson et al., CVPR’18]

69 of 69

Navigation from Dialog History

69

The goal room contains a rug.

Navigator: Should I go to the left or right?

Oracle: Go left and turn right after the bathroom.

Navigator: Do I need to go in the room with the run or keep on going right?

Oracle: Turn right and take the tiny hallway on the right. You will ascend the stairs you find on the right.

Navigator: Should I go into the kitchen or to the right?

Oracle: Turn toward the front door and go up the stairs you see on the right.

Navigator: Do I go left or right?

Oracle: Go along the railing to the right. Stop at the room with a brown chair.

3 steps

4 steps

x5

6 steps

4s

x7

...

Oracle

...

Navigator

7 steps