Post-training
后训练EssentialThe training stage that follows pretraining, using curated data or reinforcement learning to turn a foundation model into something actually useful.
Post-training refers to the training steps that come after a foundation model's pretraining is done — the term became common alongside large language models. A large model is first pretrained on massive amounts of web data, then goes through post-training steps like supervised fine-tuning (SFT), preference optimization, and reinforcement learning before it becomes an assistant that can actually follow instructions; Ai2's Tulu 3, for instance, published a full post-training recipe combining SFT, DPO, and reinforcement learning with verifiable rewards. Embodied AI has adopted the same division of labor: the π0 paper states that the pretraining stage is responsible for broad capability and generalization, while the post-training stage uses narrower, more carefully curated data to make the model proficient at the actual target task. Post-training is often used interchangeably with fine-tuning, but it emphasizes the idea of a “stage,” which can itself contain multiple rounds of fine-tuning and reinforcement-learning fine-tuning.
ExampleAfter pretraining, π0 goes through post-training for complex tasks like folding laundry: the simplest tasks need only about 5 hours of data, while the most complex ones use over 100 hours.
- Also called
- Posttraining
- Related
- Pre-training · Fine-tuning · Supervised Fine-Tuning · Reinforcement Fine-Tuning (RL Fine-Tuning) · Reinforcement Learning from Human Feedback · Mid-training
- Sources
- Lambert et al. 2024: Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control