Embodied AI Glossary中文

Dita

Advanced

An open-source, roughly 330-million-parameter generalist VLA policy that denoises a whole action sequence directly with a single Transformer.

Dita is a generalist robot policy released in March 2025 by researchers from Shanghai AI Lab, Zhejiang University, SenseTime, CUHK, Peking University, Tsinghua University, and other institutions, published at ICCV 2025. Many diffusion-based VLA models attach a small diffusion action head — a small network dedicated to turning noise into actions — after the main model, so the action generation only ever sees compressed features. Dita instead concatenates language tokens, visual tokens from historical images, and noisy action tokens into a single causal Transformer and denoises them together, a form of in-context conditioning that lets the action align directly with raw visual detail and better capture action deltas and differences across environments. The model is pretrained on cross-embodiment data from Open X-Embodiment, with about 334 million total parameters (about 221 million trainable), and matches or approaches the best contemporary results on simulation benchmarks including SimplerEnv, LIBERO, CALVIN, and ManiSkill2. The code is open-source, and the project positions itself as a lightweight, easily reproducible diffusion VLA baseline.

ExampleWhen moved to a new physical robot and a new scene, Dita can be fine-tuned with just 10 demonstrations per task from a single third-person camera, and still complete multi-step, long-horizon tasks while handling changes in background, clutter placement, and lighting.

Also called
RoboDita, Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Related
Diffusion Transformer · Vision-Language-Action Model · Diffusion Action Head · Open X-Embodiment · Causal Attention · LIBERO Benchmark
Sources
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy (arXiv 2503.19757)
Dita 项目主页 (Chinese)
As of
2025-09

See it in the full glossary →