> ML_LITERATURE // BROHAN-2023-RT-2-VISION-LANGUAGE-ACTION-MODELS-ROBOTIC-CONTROL_v1.0
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alex Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Jacky Liang, Xinlei Lu, Kanishka Rao, Michael Ryoo, Quan Vuong, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, Brianna Zitkovich · Conference on Robot Learning (CoRL) (2023)
Principal Contribution
Co-fine-tuned vision-language foundation models (PaLI-X, PaLM-E) directly on robotic action tokens, transferring internet-scale semantic reasoning into physical robot control.
Operational Relevance
Serves as qualified theoretical and systems foundation for task-robotics, task-multimodal.
Assumptions
- Markovian state dynamics and stationary reward functions hold in target evaluation environments
Limitations
- Sample efficiency, exploration stability, and real-world sim-to-real transfer gaps require specialized tuning
