Embodied AI Glossary中文

RDT-1B

RDTCommon

Tsinghua's 2024 diffusion foundation model for bimanual manipulation, with about 1.2 billion parameters and a unified action space.

RDT-1B (Robotics Diffusion Transformer) is a robot diffusion foundation model, about 1.2 billion parameters, released by a Tsinghua University team in October 2024. In bimanual manipulation the same situation often has several equally valid ways to act (action multimodality), and direct regression tends to average them into a wrong action; RDT instead uses a diffusion Transformer that denoises step by step to generate an action sequence, which can represent this kind of distribution. To make use of data from many different robots, it designs a “physically interpretable unified action space” that places physical quantities like joint angles and end-effector pose from different robots into fixed positions in one shared vector. The model is first pretrained on more than 1 million trajectories across 46 datasets, then fine-tuned on more than 6,000 ALOHA bimanual demonstrations, after which it can learn a new skill from just 1 to 5 demonstrations. Its successor is RDT2.

ExampleGiven only 1 to 5 demonstrations, RDT-1B can learn a new bimanual manipulation skill on an ALOHA robot, and it can even execute zero-shot on objects and scenes it has never seen.

Also called
RDT, Robotics Diffusion Transformer
Related
Diffusion Policy · Diffusion Transformer · Bimanual Manipulation · Unified Action Space · Action Multimodality · RDT2
Sources
RDT-1B 项目主页 (Chinese)
As of
2024-10

See it in the full glossary →