Embodied AI Glossary中文

Diffusion Transformer

扩散 TransformerDiTCommon

Using a Transformer instead of a U-Net as a diffusion model's denoising network architecture.

The Diffusion Transformer was proposed by William Peebles and Saining Xie in 2022 (published at ICCV 2023). Before this, diffusion models' denoising networks were mostly U-Nets, a convolutional encoder-decoder network; DiT replaces this with a standard Transformer: a VAE-compressed latent image is cut into patches and treated as tokens, and adaptive layer normalization (adaLN) injects conditioning information, such as the denoising timestep and class label, into every layer. The paper found that generation quality kept improving with more compute, following a clean scaling trend; the largest model, DiT-XL/2, reached an FID of 2.27 on 256×256 ImageNet. Video-generation models have since widely adopted this structure. Robotics uses it to generate actions too: RDT-1B (Robotics Diffusion Transformer) uses it for bimanual manipulation, and GR00T N1's action module is likewise a DiT variant.

ExampleRDT-1B is a diffusion Transformer of about 1.2 billion parameters that, conditioned on a language instruction and camera images, denoises to generate a segment of upcoming action for a dual-arm robot.

Also called
DiT
Related
Diffusion Model · Transformer · U-Net · Adaptive Layer Normalization · Latent Diffusion Model · RDT-1B
Sources
Scalable Diffusion Models with Transformers (arXiv 2212.09748)
DiT project page (William Peebles)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)

See it in the full glossary →