RT-1: Robotics Transformer for Real-World Control at Scale
Anthony Brohan∗, Noah Brown∗, Justice Carbajal∗, Yevgen Chebotar∗, Joseph Dabis∗,�Chelsea Finn∗, Keerthana Gopalakrishnan∗, Karol Hausman∗, Alex Herzog†, Jasmine Hsu∗,�Julian Ibarz∗, Brian Ichter∗, Alex Irpan∗, Tomas Jackson∗, Sally Jesmonth∗, Nikhil J Joshi∗,�Ryan Julian∗, Dmitry Kalashnikov∗, Yuheng Kuang∗, Isabel Leal∗, Kuang-Huei Lee‡, Sergey Levine∗, Yao Lu∗, Utsav Malla∗, Deeksha Manjunath∗, Igor Mordatch‡, Ofir Nachum‡, Carolina Parada∗, Jodilyn Peralta∗, Emily Perez∗, Karl Pertsch∗, Jornell Quiambao∗, Kanishka Rao∗, Michael Ryoo∗, Grecia Salazar∗, Pannag Sanketi∗, Kevin Sayed∗, Jaspiar Singh∗, Sumedh Sontakke‡, Austin Stone∗, Clayton Tan∗, Huong Tran∗, Vincent Vanhoucke∗, Steve Vega∗, Quan Vuong∗, Fei Xia∗, Ted Xiao∗, Peng Xu∗, Sichun Xu∗, Tianhe Yu∗, Brianna Zitkovich∗
1
Success of Large Data + Large Models
Bommasani, Rishi, et al. "On the opportunities and risks of foundation models." arXiv preprint arXiv:2108.07258 (2021).
2
RT-1’s Recipe for Robot Learning
3
Multi-Task learning
4
Multi-task learning instead of single-policy per task (RT1: 700 tasks)�^ the figure suggests lifelong-learning but it is multi-task learning
“Good” Model Architecture
5
“Good” Robot Data
6
Problem Setting
7
Universal Sentence Encoder
Behavior Cloning
Model Architecture
8
History of 6 Images
Image Encoder
Universal Sentence Encoder
Token Learner
Transformer
11D Discrete Action
Predicted Actions
* Base control is only near the tabletop movement
[Background] FiLM Layer
Perez, Ethan, et al. "Film: Visual reasoning with a general conditioning layer." Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 32. No. 1. 2018.
9
Key Takeaways from Model Architecture
10
Definition of Task and Skill
11
Tasks and Skills
12
Training Objects and Skills
13
Category A of objects
Category B of objects
Training Data Collection
14
Evaluation
15
Baselines
16
Evaluation
17
Evaluation
18
| Environment Setup | Comment |
Seen task | Training | - 200 tasks across all 8 skill families. �- Remove place on counter? |
Unseen tasks | Training | - 21 novel unseen instructions - Some instances of each object and skills are seen |
Robustness | Eval Kitchens | - 30 tasks for distractor robustness - 22 tasks for background robustness |
Long-horizon | Eval Kitchens | - 15 tasks using SayCan + RT-1 - Combination of ~ 10 distinct training instructions |
Evaluations
19
Results
20
Note: Gato doesn’t use pretrained LLM
Leveraging Simulation Data
21
“Move X to Y”��- Move X to Y is seen in simulator�- Move Z to Y is seen in real-world
- X is not seen in real world
“Move X to Y”��- Pick X is seen in simulator�- Move Z to Y is seen in real-world
- X is not seen in real world
Leveraging Simulation Data
22
Leveraging other Robot Data
23
SayCan + RT1
24
Amount of Data Vs Skill Diversity
25
Model Architecture Ablation
26
27
Thank you
28