Embodied AI Glossary中文

10Training & Learning Methods

How models are actually trained: imitation learning, reinforcement learning, pretraining and fine-tuning, plus tricks that make it more stable. · 203 terms

  1. 10.1Overview of learning paradigms5
  2. 10.2Training fundamentals25
  3. 10.3Loss functions and training objectives13
  4. 10.4Imitation learning7
  5. 10.5Core reinforcement-learning concepts20
  6. 10.6Classic reinforcement-learning algorithms22
  7. 10.7Where rewards come from17
  8. 10.8Offline reinforcement learning13
  9. 10.9Pretraining and representation learning24
  10. 10.10Fine-tuning and post-training25
  11. 10.11From simulation to real robots23
  12. 10.12Inference-time gains and fast adaptation9

10.1Overview of learning paradigms

With data ready, meet the basic ways to learn: with labels, without them, self-generated labels, from demonstration, or from reward.

10.2Training fundamentals

Whatever the method, training means computing loss, taking gradients, and updating parameters, while guarding against overfitting and instability.

10.3Loss functions and training objectives

Expanding on loss functions: the different objectives used for regression, classification, sequence prediction, and diffusion generation.

10.4Imitation learning

Applying supervised training to learn actions from human demonstrations, and why it tends to drift further off course over time.

10.5Core reinforcement-learning concepts

The second path besides demonstration is trial and error by reward: first, the basics of reward, return, and value.

10.6Classic reinforcement-learning algorithms

Turning those concepts into concrete algorithms: from policy gradients and PPO to DQN and SAC.

10.7Where rewards come from

With algorithms in hand, you still need the right reward: hand-written, learned from demonstrations or a model, and what to do when it’s sparse.

10.8Offline reinforcement learning

Earlier algorithms all learn through live interaction; here a policy learns from a fixed dataset instead, then fine-tunes online.

10.9Pretraining and representation learning

Shifting from reinforcement learning to the large-model approach: pretraining on massive, varied data to learn general-purpose representations.

10.10Fine-tuning and post-training

After pretraining, adapting to specific tasks: fine-tuning, supervised and RL post-training, plus preventing forgetting and distillation.

10.11From simulation to real robots

Onto real robots: large-scale training in sim, transferring to hardware, then continuing to learn there through trial and human correction.

10.12Inference-time gains and fast adaptation

After training ends, a model can still improve without changing weights much: more compute at inference, or fast adaptation to new tasks.

See it in the full glossary →