Embodied AI Glossary中文

Mid-training

中训练Advanced

An extra training stage between pretraining and post-training that uses more targeted data to fill in specific abilities.

Mid-training first became popular with large language models: after large-scale pretraining and before instruction tuning and reinforcement learning, the model continues training for a while on higher-quality or more target-relevant data — math, code, long documents — often paired with learning-rate annealing and context-length extension, aiming to strengthen specific abilities without losing general foundations; a dedicated 2025 survey already maps out this stage. Embodied AI borrowed the term: an off-the-shelf vision-language model has never seen robot data, so fine-tuning it directly into a VLA gives limited results, so teams insert a transitional round in between using embodiment-related data, such as spatial reasoning, robot trajectories, or human motion aligned to robot motion. Ai2's MolmoAct released a robot dataset built specifically for mid-training, and the Chinese embodied-AI company Dexmal's DM0 likewise uses a three-stage pretraining, mid-training, post-training pipeline.

ExampleEgoScale first pretrains a VLA on more than 20,000 hours of action-annotated egocentric human video, then runs one lightweight round of mid-training on a small amount of data that aligns human and robot actions; after that, it needs only minimal robot supervision to adapt to new dexterous-manipulation tasks.

Related
Pre-training · Post-training · Supervised Fine-Tuning · Vision-Language-Action Model · Data Mixture · MolmoAct
Sources
Mid-Training of Large Language Models: A Survey (arXiv 2510.06826)
MolmoAct: Action Reasoning Models that can Reason in Space (arXiv 2508.07917)
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (arXiv 2602.16710)
As of
2026-02

See it in the full glossary →