Supervised Learning
监督学习EssentialTraining a model on input–correct-answer pairs so it learns to predict the answer from the input alone.
Supervised learning is the most common machine-learning paradigm. Every training input comes paired with the correct output — a label, also called ground truth — and the model repeatedly shrinks the gap between its own prediction and that label, measured by a loss function. The ultimate goal is for the model to predict correctly on new data it has never seen, a property called generalization. When the output is a category, this is called classification; when it's a continuous number, it's regression. The contrasting paradigms are unsupervised learning, which uses no labels at all, and reinforcement learning, which learns from a reward signal instead. In embodied AI, behavior cloning is essentially supervised learning: the observations recorded during a human demonstration become the input, and the human's actions become the label, and a policy is trained to imitate them.
ExampleCollect 1,000 paired examples of “camera image → arm joint angles” through teleoperation, then train a network that takes the image as input and outputs joint angles, with the loss defined as the mean squared error between predicted and demonstrated angles. That is behavior cloning framed as supervised learning.
- Related
- Behavior Cloning · Unsupervised Learning · Reinforcement Learning · Loss Function · Ground Truth · Supervised Fine-Tuning
- Sources
- Wikipedia: Supervised learning