Action Label
动作标签CommonThe action value a robot executed at each moment in a dataset — the “answer” imitation learning tries to fit.
An action label is the action value aligned with each frame of observation in robot data — end-effector displacement and rotation, target joint angles, gripper opening. During teleoperation, the system records it directly; handheld capture methods like UMI instead recover it afterward with algorithms such as SLAM. Behavior cloning treats the observation as input and the action label as the supervision signal, so a label's frequency, coordinate frame, and whether it's absolute or incremental all directly shape what the model learns; RT-1-X and similar models unify data from different sources into a 7-dimensional action (3 for position, 3 for rotation, 1 for the gripper). Action labels can only be captured on a robot or with special equipment, which is the main reason robot data is expensive — hence approaches that use an inverse dynamics model to add pseudo action labels to video, or that learn latent actions instead.
ExampleOpenAI's VPT first trained an inverse dynamics model on 1,962 hours of Minecraft footage recorded by contractors along with their keyboard and mouse input, then used it to add pseudo action labels to about 70,000 hours of web video, ultimately training an agent capable of crafting a diamond tool.
- Also called
- Action Annotation
- Related
- Observation-Action Pair · Behavior Cloning · Action Space · Pseudo Action Labels · Inverse Dynamics Model · Action-free Video
- Sources
- Baker et al. 2022: Video PreTraining (VPT) (arXiv 2206.11795)
Open X-Embodiment 项目页 (Chinese)
Ye et al. 2024: Latent Action Pretraining from Videos (LAPA)