RoboMamba
AdvancedAn efficient VLA that replaces the Transformer language backbone with the Mamba state-space model.
RoboMamba was released by Peking University together with Zhiping AI and the Beijing Academy of Artificial Intelligence in June 2024, accepted at NeurIPS 2024. VLAs at the time mostly used Transformer-based large language models as their backbone, which made inference slow and expensive. RoboMamba instead uses Mamba, a selective state-space model whose computation scales linearly with sequence length, as the language model, attached to a vision encoder; it first goes through alignment pretraining and joint training on general and robot-instruction data to gain reasoning ability, then the whole model is frozen and only a very small policy head (about 0.1% of the model's parameters) is added to predict the end effector's SE(3) pose — its position and orientation. The paper reports roughly 3x faster inference than existing VLAs, with competitive pose prediction in both simulation and on a real robot.
ExampleRoboMamba can answer a reasoning question like 'which object on the table could be used to hold water,' and also output the position and orientation the gripper should reach after receiving a manipulation instruction.
- Also called
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- Related
- Mamba · State Space Model · Vision-Language-Action Model · Inference Latency · End-Effector Pose · Parameter-Efficient Fine-Tuning
- Sources
- arXiv 2406.04339: RoboMamba
RoboMamba 项目主页 (Chinese) - As of
- 2024-12