Embodied AI Glossary中文

Intuitive Physics

直觉物理Common

Commonsense prediction of how objects will behave without using formulas — unsupported things fall, hidden things still exist.

Intuitive physics was originally a cognitive-science concept describing the naive understanding humans, including infants, have of the physical world: an object still exists after being hidden from view (object permanence), objects cannot pass through each other, unsupported objects fall, and shapes do not suddenly change. Researchers commonly use the “violation of expectation” paradigm to test this: show a subject a video that breaks a physical law and see whether they show “surprise.” AI researchers test models the same way: a 2025 study from Meta found that V-JEPA, which predicts video in a learned representation space rather than pixel space, shows understanding of several intuitive-physics properties, while pixel-space video prediction models and multimodal large models perform closer to chance; Google DeepMind's Physics-IQ benchmark likewise found that video generation models such as Sora produce visually realistic footage but have limited physical understanding, and that this is unrelated to how visually realistic the video looks. For robots, intuitive physics determines whether they can predict whether nudging a cup will tip it over or whether a stack of objects is stable — a foundational capability that world models and embodied reasoning both need.

ExampleShown a video of a ball rolling behind a screen and then vanishing on the other side instead of reappearing, a model that has learned object permanence shows a noticeably higher prediction error, its 'surprise,' at that moment.

Also called
Physical Commonsense
Related
World Model · Embodied Reasoning · V-JEPA 2 · Physics-IQ (Do generative video models understand physical principles?) · Joint-Embedding Predictive Architecture · Spatial Reasoning
Sources
Intuitive physics understanding emerges from self-supervised pretraining on natural videos (Garrido et al., 2025)
Do generative video models understand physical principles? (Physics-IQ)
As of
2025-02

See it in the full glossary →