Inverse Dynamics Model
逆动力学模型IDMCommonA model that looks at two frames, before and after, and infers what action happened in between.
An inverse dynamics model runs in the opposite direction from a forward dynamics model: it takes the current state s_t and the next state s_{t+1} (or a pair of consecutive video frames) and outputs the action a_t that caused the change. In classical mechanics, inverse dynamics works backward from a desired motion to the joint torques that would produce it; in robot learning, an IDM is usually a neural network trained on data instead. Its main use is labeling videos that have no action labels: OpenAI's 2022 VPT project first trained an IDM on about 2,000 hours of Minecraft footage recorded together with keyboard and mouse input, then used it to attach pseudo action labels to about 70,000 hours of online video for pretraining a game-playing agent. Because an IDM can see both the past and future frame, predicting the action is much easier for it to learn than predicting an action from the current frame alone. ‘Generate video, then infer actions’ methods such as UniPi also rely on an IDM to translate generated frames into robot actions.
ExampleGiven two frames — the first shows an open gripper hovering above a cup, the second shows the gripper closed with the cup lifted a few centimeters — an IDM outputs the action vector for ‘close the gripper and move up.’
- Also called
- IDM, Inverse Model
- Related
- Forward Dynamics Model · Latent Action Model · Pseudo Action Labels · Action-free Video · VPT · UniPi
- Sources
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (arXiv 2206.11795)
Learning Universal Policies via Text-Guided Video Generation (UniPi, arXiv 2302.00111)