Embodied AI Glossary中文

Seer

Seer(预测式逆动力学模型)PIDMAdvanced

An end-to-end manipulation policy that first predicts the upcoming frame, then infers the action from it using inverse dynamics.

Seer was released by the Shanghai Artificial Intelligence Laboratory and other institutions in December 2024, selected for an oral presentation at ICLR 2025. Common approaches either imitate actions directly from the current frame (behavior cloning), or first generate future frames with a video model and separately infer actions afterward, training the two parts apart. Seer proposes the 'Predictive Inverse Dynamics Model' (PIDM): inside one end-to-end trained Transformer, it first predicts the frame the robot is about to see, then uses an inverse dynamics model (which infers what action was taken from a current state and a target state) to compute the action from 'now' and the 'predicted future,' with visual prediction and action prediction optimized jointly. It can be pretrained on large-scale robot data such as DROID, then fine-tuned on a small amount of downstream data. The paper reports a 13% improvement over the previous best method on LIBERO-LONG, 21% on CALVIN ABC-D, and 43% on real-robot tasks.

ExampleWhile performing 'open the drawer,' Seer first predicts the frame showing the gripper approaching the handle a moment later, then computes how the arm should move from the current frame and this predicted frame.

Also called
Predictive Inverse Dynamics Model, Seer: Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Related
Inverse Dynamics Model · Video Prediction Policy · DROID (Distributed Robot Interaction Dataset) · CALVIN Benchmark · LIBERO Benchmark · Shanghai Artificial Intelligence Laboratory
Sources
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation (arXiv 2412.15109)
OpenRobotLab/Seer (GitHub)
As of
2025-01

See it in the full glossary →