Embodied AI Glossary中文

Pseudo Action Labels

伪动作标签Advanced

Action labels inferred by a model from video, standing in for actions that were never actually recorded.

When video itself doesn't record what action a human or robot took, a model can “guess” the action at each step from the surrounding frames, and those guesses can be used as labels to train a policy — this is a pseudo action label. There are two common approaches. One trains an inverse-dynamics model (which infers the action between two adjacent frames) on a small amount of action-labeled data, then uses it to label a huge amount of video in bulk — OpenAI's VPT did exactly this to add keyboard-and-mouse action labels to a large collection of Minecraft videos from the internet. The other uses a latent-action model to learn an abstract action encoding directly from video. NVIDIA's DreamGen uses both approaches to add pseudo actions to video generated by a world model, producing trainable “neural trajectories.” This lets video with no action labels be used for training too, though the labels carry error and usually still need fine-tuning on real data.

ExampleVPT first had people play Minecraft while recording their keyboard and mouse actions, trained an inverse-dynamics model on that data, then used it to label pseudo actions across a large collection of internet videos, and finally ran behavioral cloning on that labeled data.

Also called
Pseudo-Action Labeling, Pseudo-Actions
Related
Inverse Dynamics Model · Latent Action Model · Action-free Video · Neural Trajectories · DreamGen · VPT
Sources
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (arXiv:2206.11795)
DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv:2505.12705)

See it in the full glossary →