1 of 28

RT-1: Robotics Transformer for Real-World Control at Scale

Anthony Brohan∗, Noah Brown∗, Justice Carbajal∗, Yevgen Chebotar∗, Joseph Dabis∗,�Chelsea Finn∗, Keerthana Gopalakrishnan∗, Karol Hausman∗, Alex Herzog†, Jasmine Hsu∗,�Julian Ibarz∗, Brian Ichter∗, Alex Irpan∗, Tomas Jackson∗, Sally Jesmonth∗, Nikhil J Joshi∗,�Ryan Julian∗, Dmitry Kalashnikov∗, Yuheng Kuang∗, Isabel Leal∗, Kuang-Huei Lee‡, Sergey Levine∗, Yao Lu∗, Utsav Malla∗, Deeksha Manjunath∗, Igor Mordatch‡, Ofir Nachum‡, Carolina Parada∗, Jodilyn Peralta∗, Emily Perez∗, Karl Pertsch∗, Jornell Quiambao∗, Kanishka Rao∗, Michael Ryoo∗, Grecia Salazar∗, Pannag Sanketi∗, Kevin Sayed∗, Jaspiar Singh∗, Sumedh Sontakke‡, Austin Stone∗, Clayton Tan∗, Huong Tran∗, Vincent Vanhoucke∗, Steve Vega∗, Quan Vuong∗, Fei Xia∗, Ted Xiao∗, Peng Xu∗, Sichun Xu∗, Tianhe Yu∗, Brianna Zitkovich∗

1

2 of 28

Success of Large Data + Large Models

Bommasani, Rishi, et al. "On the opportunities and risks of foundation models." arXiv preprint arXiv:2108.07258 (2021).

2

3 of 28

RT-1’s Recipe for Robot Learning

  • Open-ended task-agnostic training

  • High-capacity architectures

  • Large, diverse robot data

3

4 of 28

Multi-Task learning

4

Multi-task learning instead of single-policy per task (RT1: 700 tasks)�^ the figure suggests lifelong-learning but it is multi-task learning

5 of 28

“Good” Model Architecture

  • Ingest large amount of data (RT1: 35M Parameters)

  • Fast at inference for robot control (RT1: 3Hz Frequency)

5

6 of 28

“Good” Robot Data

  • Large scale (RT1: 130k episodes)

  • Diverse (?) (RT1: 700 tasks)

  • Inter-connected tasks (unlike GATO)

6

7 of 28

Problem Setting

7

Universal Sentence Encoder

Behavior Cloning

8 of 28

Model Architecture

8

History of 6 Images

Image Encoder

  • ImageNet Pretrained EfficientNet-v3

  • Language condition using FiLM Layers

Universal Sentence Encoder

 

Token Learner

 

Transformer

  • 8 Self Attention Layers

  • With Positional Embeddings*

11D Discrete Action

  • 7D Arm control
  • 3D Base control
  • 1D Arm-Base Switch
  • Each into 256 bins

Predicted Actions

* Base control is only near the tabletop movement

9 of 28

[Background] FiLM Layer

Perez, Ethan, et al. "Film: Visual reasoning with a general conditioning layer." Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 32. No. 1. 2018.

9

 

10 of 28

Key Takeaways from Model Architecture

  • Language conditioning using FiLM layer

  • EfficientV3 and TokenLearner for efficient inference (3Hz)

  • Discrete Actions: ??

10

11 of 28

Definition of Task and Skill

  • Task = Instruction

    • Instruction is a combination of a verb and nouns surrounding it

    • Example: Place water bottle upright, open the top drawer

  • Skill = Group of Instructions with common verbs

    • Skill is defined by the verb in the instruction (* verb + prepositions)

    • Example: Place Object Upright, Pick Object

11

12 of 28

Tasks and Skills

12

13 of 28

Training Objects and Skills

13

Category A of objects

Category B of objects

  • Category A: 17 objects

  • Category B: ??

  • All skills are with Category A

  • Picking skill is with Category B

14 of 28

Training Data Collection

  • Robot classrooms: Segments of office kitchen

  • Total 130k Demonstrations

  • 700 distinct task instructions

  • 17 Months & 13 Robots

14

15 of 28

Evaluation

  • On the training tabletop environments

  • Two real office kitchens
    • Similar to training environment
    • Change in background, lightning.
    • Example: Cabinet instead of drawer

15

16 of 28

Baselines

  • Gato: Full transformer based, no pre-trained language encoder

  • BC-Z: No history of images and continuous actions

  • BC-Z XL: Larger version of the BC-Z mentioned (~ RT1 size)

16

17 of 28

Evaluation

  • Seen task (Training Environment):
    • 200 tasks across all 8 skill families.
    • Remove place on counter?
  • Unseen tasks (Training Environment):
    • 21 novel unseen instructions
    • Some instances of each object and skills are seen
  • Robustness (Kitchens):
    • 30 tasks for distractor robustness
    • 22 tasks for background robustness
  • Long-horizon
    • 15 tasks using SayCan + RT1
    • ~ 10 distinct instructions from training

17

18 of 28

Evaluation

18

Environment Setup

Comment

Seen task

Training

- 200 tasks across all 8 skill families. �- Remove place on counter?

Unseen tasks

Training

- 21 novel unseen instructions

- Some instances of each object and skills are seen

Robustness

Eval Kitchens

- 30 tasks for distractor robustness

- 22 tasks for background robustness

Long-horizon

Eval Kitchens

- 15 tasks using SayCan + RT-1

- Combination of ~ 10 distinct training instructions

19 of 28

Evaluations

19

20 of 28

Results

20

Note: Gato doesn’t use pretrained LLM

21 of 28

Leveraging Simulation Data

21

“Move X to Y”��- Move X to Y is seen in simulator�- Move Z to Y is seen in real-world

- X is not seen in real world

“Move X to Y”��- Pick X is seen in simulator�- Move Z to Y is seen in real-world

- X is not seen in real world

22 of 28

Leveraging Simulation Data

22

23 of 28

Leveraging other Robot Data

23

24 of 28

SayCan + RT1

24

25 of 28

Amount of Data Vs Skill Diversity

25

26 of 28

Model Architecture Ablation

26

27 of 28

27

  • Key component (+15%)
  • Pretraining on ImageNet (+13%)
  • For efficient inference (3x)
  • Reduces tokens from 81 to 8
  • Transformer over Convnet backbone (+13%)
  • To resolve multimodal behavior.
  • +29% over Multivariate normal output
  • If we condition on action values it leads to worse performance (-12%)

28 of 28

Thank you

28