Skip to main content

> ML_LITERATURE // BROHAN-2023-RT-2-VISION-LANGUAGE-ACTION-MODELS-ROBOTIC-CONTROL_v1.0

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alex Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Jacky Liang, Xinlei Lu, Kanishka Rao, Michael Ryoo, Quan Vuong, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, Brianna Zitkovich · Conference on Robot Learning (CoRL) (2023)

seminal-architecture2023foundationalthirdPartyReproduced

Principal Contribution

Co-fine-tuned vision-language foundation models (PaLI-X, PaLM-E) directly on robotic action tokens, transferring internet-scale semantic reasoning into physical robot control.

Operational Relevance

Serves as qualified theoretical and systems foundation for task-robotics, task-multimodal.

Assumptions

  • Markovian state dynamics and stationary reward functions hold in target evaluation environments

Limitations

  • Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: