RDT2
AdvancedA VLA foundation model from Tsinghua, trained on tens of thousands of hours of UMI data, that deploys zero-shot on new robot arms.
RDT2 comes from Jun Zhu's lab at Tsinghua University (the RDT team), open-sourced in September 2025 with the paper following in February 2026; it is the successor to RDT-1B. The problem it addresses: switching to a different robot arm usually forces a full recollection of data and re-fine-tuning. The team improved UMI (Universal Manipulation Interface), a handheld gripper data-collection device, and used it to gather more than 10,000 hours of demonstrations across about 100 locations, mostly real homes; because the same UMI gripper is used for both data collection and deployment, the embodiment gap stays small. The model backbone is a 7-billion-parameter Qwen2.5-VL, trained in three stages: first, residual vector quantization turns actions into discrete tokens (RDT2-VQ); then this is replaced with a roughly 400-million-parameter action expert that outputs continuous actions (RDT2-FM); finally, the model is distilled for speed. The team states it is the first foundation model that can perform simple tasks like pick-and-place zero-shot on an unseen robot embodiment, and it has also been demonstrated playing table tennis.
ExampleRDT2-FM is connected directly to a robot arm it never saw during training, fitted with the same UMI-style gripper, and without any fine-tuning it can pick up and place objects following language instructions.
- Also called
- Robotics Diffusion Transformer 2, RDT2-VQ, RDT2-FM, RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
- Related
- RDT-1B · Universal Manipulation Interface · Handheld Gripper Data Collection · Cross-Embodiment · Flow Matching · Vector Quantization
- Sources
- RDT2 project page
RDT2 (arXiv 2602.03310) - As of
- 2026-02