01Core Concepts & Tasks
Start with the map: what embodied AI does, the tasks robots perform, and what terms like ‘generalization’ and ‘cross-embodiment’ mean. · 150 terms
- 1.1What is embodied AI12
- 1.2Basic tasks robots perform11
- 1.3Policies: from observation to action18
- 1.4Generalization: handling what you haven’t seen23
- 1.5Manipulation tasks in detail15
- 1.6Walking, navigation, and exploration18
- 1.7Factories, warehouses, and homes11
- 1.8Working alongside people and robots10
- 1.9Understanding language and the world18
- 1.10Long-term goals and origins14
1.1What is embodied AI
The starting point: what embodied AI means, how it differs from disembodied AI, its parts, and why it’s hard.
- Embodied AI具身智能Artificial intelligence with a body, able to perceive, decide, and act in the physical world.
- Disembodied AI离身智能AI with no body, processing information purely in the digital world — chatbots, image classifiers, and so on.
- Physical AI物理AIAI that can perceive, understand, and act in the physical world — robots and self-driving cars are examples.
- Letting a car perceive road conditions, make decisions, and control itself with little or no human input.
- The basic RL framework: the decision-maker is the agent, and everything outside it that it can observe and affect is the environment.
- Embodied Agent具身智能体An agent with a physical or virtual body that perceives and acts within an environment through that body.
- A robot's physical body: its specific combination of form, joints, sensors, and actuators.
- Perception-Action Loop感知-行动闭环Perception drives action, and action changes what is perceived next, in a continuous back-and-forth cycle.
- Robot Learning机器人学习Using machine learning to let robots acquire skills from data and interaction, instead of hand-written rules.
- A robot not built for one fixed task, meant to handle many tasks across many settings.
- Moravec's Paradox莫拉维克悖论For machines, high-level reasoning is easy; the perception and movement humans take for granted turn out to be the hard part.
- Sensorimotor Skills感觉运动技能The ability to turn sensory information into body movement in real time: grasping, walking, catching, twisting.
1.2Basic tasks robots perform
With the concepts in place, see what robots actually do: manipulating objects, moving themselves, navigating, and chaining it into long tasks.
- A robot using its hand or a tool to contact an object and change its position, pose, or state.
- Grasping抓取A robot picking up an object securely with a gripper, suction cup, or dexterous hand.
- Pick-and-Place抓取放置Picking an object up from one place, moving it, and setting it down at a target location — a basic manipulation task.
- A robot uses two arms together, coordinating them to complete a single task.
- Using a multi-fingered hand, coordinated finger movement and force control to grasp, flip, and handle objects.
- Locomotion运动(移动)A robot's ability to move itself from place to place, using legs, wheels, or other means.
- A robot deciding, from its own sensor readings, how to move itself to a target location or object.
- Having an agent follow a natural-language route description and use vision to reach a destination in an unfamiliar space.
- Mounting a robot arm on a mobile base so it can move and work at once, combining navigation and manipulation.
- A task that requires completing many sub-steps in sequence, over a long stretch of time, to reach the goal.
- Skill Primitive原子技能The smallest reusable unit of action, such as grasp, place, or open drawer, combined to complete complex tasks.
1.3Policies: from observation to action
Unify these tasks into one observation-to-policy-to-action loop, then look at the types of policies and how systems are layered.
- The information an agent receives from the environment at each moment, the input a policy uses to decide.
- State Space状态空间The set of all possible states of a system; for a robot, usually made up of joint angles, pose, velocity, and similar quantities.
- Action Space动作空间The set of all actions an agent can take at each step, and the numerical form those actions take.
- Policy策略The rule that decides what action to take next given the current observation — usually a neural network.
- Inference推理(前向计算)Running a trained model forward on new input to produce an output, without updating its parameters.
- Episode回合One full attempt at a task, from the environment's reset to completion, failure, or timeout.
- A sequence of positions or states and actions over time; also a full recording of one demonstration.
- Executing a pre-computed sequence of actions straight through, without checking feedback along the way.
- Acting while watching: adjusting the next action based on the latest observation at every step.
- Using what a camera sees to guide an arm and gripper's motion in real time, adjusting as it goes.
- Visuomotor Policy视觉运动策略A control policy that maps camera images directly to robot actions, usually a neural network.
- Action Multimodality动作多峰性When several different actions are all correct in the same situation, so the action distribution has more than one peak.
- Goal-conditioned Policy目标条件策略A policy that decides its action based on the goal to reach, not just the current observation.
- A robot policy that takes a language instruction as input and produces different actions depending on what it says.
- Generalist Policy通用策略(通才策略)A single control policy that can perform many tasks, and often work across many settings or robots.
- A policy trained only for a single robot, a single task, or a single setting — the counterpart to a generalist policy.
- Sense-Plan-Act感知-规划-行动范式The classic control loop where a robot senses, models, plans its next move, then acts, and repeats.
- Brain–Cerebellum Architecture大脑-小脑架构(大小脑)Splitting a robot's software into a 'brain' that plans and reasons and a 'cerebellum' that executes movement.
1.4Generalization: handling what you haven’t seen
Now judge how good the policy is: does it still work with new objects, scenes, instructions, or even a different robot?
- A model's ability to still perform correctly in new situations it never saw during training.
- A test-time situation drawn from the same distribution as the training data — something the model has effectively seen before.
- Test-time data that comes from a different distribution than the training data, such as new objects or scenes.
- Distribution Shift分布偏移(协变量偏移)When the data seen at deployment has a different distribution than the training data, hurting performance.
- The mass of individually rare situations that training data barely covers, and where models fail most often.
- Zero-shot零样本A model handling a new task or environment directly, with no training examples for it at all.
- Few-shot少样本Learning or performing a new task correctly from just a handful to a few dozen examples.
- Robustness鲁棒性A system's ability to keep performing without much degradation when inputs or the environment carry noise or small changes.
- Extra objects in a scene that are irrelevant to the current task but can throw a policy off.
- Failure Recovery失败恢复A robot noticing it made a mistake or is about to fail, and adjusting on its own to still finish the task.
- A policy's ability to still complete the same task when the object is swapped for one it never saw in training.
- Spatial Generalization位置泛化(空间泛化)A policy's ability to still succeed when an object is placed at a position or orientation absent from training.
- A policy still completing its task when moved into a room, table, lighting, or background it never saw in training.
- Still completing the task when the scene looks different: new background, lighting, colors, distractors, or camera angle.
- Still understanding and correctly acting on unfamiliar objects, concepts, or phrasing it hasn't seen before.
- A policy still succeeding when a change in the situation forces the 'correct action' itself to change.
- A policy's ability to complete a new task or new instruction that never appeared during training.
- Recombining separately learned elements to handle a combination that was never seen together during training.
- Training one model on data from many different robots so it can control more than one of them.
- Embodiment Gap本体差异The difference in shape, structure, and way of moving between different robots, or between humans and robots.
- A method or representation not tied to one particular robot's structure, usable across different robots.
- Open-vocabulary开放词汇A model that can recognize or handle objects and concepts described in arbitrary words, beyond its training categories.
- Open-world开放世界A setting where the environment is uncontrolled and things absent from training keep appearing, and the system must still work.
1.5Manipulation tasks in detail
Back to the manipulation thread: beyond grasping, bimanual, and dexterous skills, a closer look at harder, more specialized tasks.
- A robot arm fixed beside a table grasping, pushing, and arranging objects on its surface.
- Manipulating objects with movable joints, such as opening a cabinet door, pulling a drawer, or lifting a laptop lid.
- Grasping and handling objects that bend, stretch, or change shape under force, such as clothes, cables, or food.
- Getting a robot to flatten, fold, hang, and tidy clothing — the most iconic category of deformable object manipulation.
- Manipulation with tight tolerances for position and force, such as inserting a battery, threading a wire, or turning a tiny screw.
- A manipulation task that needs sustained contact with an object or surface, with careful control of contact force.
- Getting a robot to insert, screw, snap, or clip parts together into a component or finished product.
- Inserting a pin, shaft, or plug into a matching hole, a task that tests precision and force control.
- Using the fingers to adjust an object's orientation or position while still holding it, without setting it down to regrasp.
- Tool Use工具使用A robot using an external object as a tool to complete a task it cannot do bare-handed.
- Task-Oriented Grasping任务导向抓取(功能性抓取)Choosing a grasp based on what comes next — the same object gets held differently for different jobs.
- Moving an object by pushing, poking, flipping, or throwing it, instead of gripping it firmly.
- Manipulation that deliberately exploits velocity, inertia, and gravity, such as throwing, flinging, catching, or tossing.
- Extrinsic Dexterity外在灵巧性Using gravity, a tabletop, a wall, or arm swinging to let a simple gripper carry out complex manipulation.
- Fitting a flying platform such as a drone with an arm or gripper so it can physically contact objects in the air.
1.6Walking, navigation, and exploration
Now the locomotion thread: how legged robots stay balanced over rough terrain, and how they navigate and explore toward different goals.
- The control problem of getting quadruped, biped, and other legged robots to walk, run, jump, and climb slopes stably.
- A robot's ability to walk, run, and balance using only two legs — a basic skill for humanoid robots.
- Rough-terrain Locomotion复杂地形行走A legged robot's ability to move stably over uneven ground such as stairs, gravel, grass, or snow.
- A locomotion approach that uses no camera or lidar, relying only on joint and IMU signals to walk.
- A legged robot using cameras or lidar to see the terrain ahead and plan its steps accordingly.
- Parkour跑酷A legged robot's high-dynamic skill of continuously climbing, jumping, and squeezing through obstacles using vision.
- Fall Recovery跌倒恢复(摔倒起身)A humanoid or legged robot's ability to stand back up on its own from lying or sprawled positions after falling.
- Text-to-Motion文本驱动动作生成Turning a written sentence into a matching sequence of full-body human or humanoid motion.
- Loco-manipulation运动操作一体化Coordinating leg movement and arm manipulation together as one problem, so the robot works while it moves.
- Giving a robot a goal coordinate relative to its start point and having it navigate there on its own in an unfamiliar environment.
- Object-Goal Navigation物体目标导航Telling a robot only an object category, such as 'chair,' and having it find one on its own in an unfamiliar house.
- Image-Goal Navigation图像目标导航Giving a robot a photo of a destination and letting it find its own way there in an unfamiliar space.
- An agent using both 'eyes' and 'ears' together to find a sound-emitting target in a 3D environment.
- Aerial Vision-and-Language Navigation空中视觉语言导航(无人机 VLN)Having a drone understand a natural-language instruction and fly to a target location in a 3D outdoor space such as a city.
- A robot moving through crowds of people, reaching its goal while respecting human social norms.
- Embodied Visual Tracking具身视觉跟踪(目标跟随)A robot using its own camera to keep following a specified target, keeping it in view continuously as it moves.
- An agent deciding for itself where to look and what to touch, actively gathering information about the unknown.
- A task where an agent moves through a 3D environment on its own to gather information, then answers a question.
1.7Factories, warehouses, and homes
Put these tasks into real settings: from tidy, controlled factories and warehouses to cluttered, unpredictable homes.
- Structured vs. Unstructured Environment结构化 / 非结构化环境In a structured environment, objects and workflow are fixed and predictable; an unstructured one is cluttered and cannot be specified in advance.
- A human first guides the robot through a motion and records its positions, and the robot then repeats exactly what was recorded.
- The factory task of loading a workpiece into a machine and unloading it once processing is done.
- Using cameras and other sensors to have a robot check products for defects or assembly errors.
- Tote Handling料箱搬运The task of picking up, moving, and placing standard plastic totes (storage bins) in warehouses and on production lines.
- Palletizing / Depalletizing码垛 / 拆垛Stacking boxes or bags onto a pallet in a set pattern, or taking them back off one by one.
- Sorting分拣Identifying items on a conveyor or in a bin and routing each one to a different destination by type or address.
- Order Picking拣选(订单拣货)Pulling the exact items an order calls for, one by one, off warehouse shelves or out of bins.
- Bin Picking无序抓取A robot identifying and picking parts or items, one at a time, out of a cluttered bin.
- Household Tasks家务任务Everyday activities like tidying, cleaning, laundry, and cooking in a real home — the main target setting for general-purpose robots.
- Rearrangement物体重排An embodied task where a robot moves objects in an environment to match a specified goal state.
1.8Working alongside people and robots
Real settings always include people: human-robot interaction, collaboration and safety, plus how multiple robots coordinate with each other.
- The field studying how people and robots communicate, collaborate, and safely share space.
- The more human a robot looks, the more likeable it is — until it looks almost, but not quite, human, which feels unsettling.
- People and robots dividing up work and coordinating in the same space to complete a task together.
- Interaction where a person and a robot are in direct bodily contact, exchanging physical force.
- Human-Robot Handover人机物体交接A robot passing an object to a person, or taking one from a person's hand.
- Embodied Safety具身安全Ensuring a robot acting in the real world does not injure people, damage property, or get hijacked by malicious commands.
- Asimov's Three Laws of Robotics机器人三定律 / 机器人宪法Asimov's three rules of robot behavior, and their modern descendant, the rule sets large models use to keep robots safe.
- Tri-Co Robot共融机器人A robot able to safely coexist and cooperate with people, other robots, and its environment, and read their intentions.
- Multiple robots dividing up work and coordinating to finish a task that one robot can't do alone or fast enough.
- Intelligent group behavior that emerges when many simple individuals interact only through local rules.
1.9Understanding language and the world
Shift from doing to understanding: how a robot’s ‘brain’ parses instructions, reasons about space and physics, perceives, and remembers.
- A robot understanding a natural-language instruction from a person and actually carrying it out in the real world.
- Connecting the words in language and instructions to actual objects, locations, and actions in the real world.
- A model confidently producing output that does not match the facts or the scene in front of it.
- Intent Understanding意图理解(隐式指令)Figuring out what a person actually wants from a hint, without them naming the object or action directly.
- Language Corrections语言纠正(实时语言反馈)A person telling a working robot how to fix what it's doing, in words, while it keeps working.
- Steerability可引导性The ability to change exactly how a model does something using prompts or conditioning, without retraining it.
- Reasoning推理(思考)A model's ability to analyze, break down, and plan before producing an answer or action — distinct from inference.
- Making a model understand the physical world: where things are, how to grasp them, and what to do next.
- The ability to understand where objects are in 3D space, and to reason and act using that understanding.
- The ability to judge an object's position, distance, size, orientation, and relation to other objects.
- Commonsense prediction of how objects will behave without using formulas — unsupported things fall, hidden things still exist.
- Affordance可供性What an object or environment lets you do with it — a handle can be pulled, a button can be pressed.
- Studying how people grasp, push, open, and use objects — an important source for robots to learn actions from.
- Perception in service of action: observing while moving, understanding 3D space, and supporting decisions.
- Object-centric Representation以物体为中心的表示Encoding a scene as a set of separate objects, instead of squeezing the whole image into one vector.
- Partially Observable Markov Decision Process部分可观测马尔可夫决策过程The decision-making framework for when an agent can't see the true state, only noisy observations of it.
- Embodied Memory具身记忆A robot storing what it has seen and done so it can use that later for decisions and answering questions.
- The interaction an agent has with people, objects, and its surroundings in a physical or simulated space.
1.10Long-term goals and origins
Finally, the big picture: how to measure the endgame, the scale-versus-experience debate, and the ideas this field grew out of.
- Embodied AGI具身通用智能Embodied intelligence able to perform diverse open-ended real-world tasks the way a human can — the field's long-term goal.
- A grading of how much a robot can perceive, decide, and act without relying on a person.
- A Chinese industry standard that grades humanoid robot intelligence from L1 to L5 across four dimensions.
- Physical Turing Test物理图灵测试A robot passes if people cannot tell whether a piece of real-world physical work was done by a human or a machine.
- The Coffee Test咖啡测试Sending a machine into a stranger's house to find everything it needs and brew a pot of coffee.
- The Bitter Lesson苦涩的教训Rich Sutton's argument that, over time, general methods that scale with compute beat hand-engineered domain knowledge.
- Abilities absent in small models that only appear once model or data scale grows large enough.
- Silver and Sutton's view that AI's next stage will learn mainly from its own experience of interacting with the world.
- Three Schools of AI人工智能三大学派(符号主义 / 连接主义 / 行为主义)A classification of AI research, common in Chinese textbooks, into logic-symbol, neural-network, and interact-with-environment approaches.
- Symbol Grounding Problem符号落地问题How the symbols inside a machine come to refer to real things in the world, and so mean something.
- Subsumption Architecture包容架构(行为式机器人)A layered robot-control architecture wiring simple behaviors straight from sensing to action, where higher layers override lower ones.
- The view in cognitive science that thought is inseparable from the body and its interaction with the environment.
- Letting a body's own shape and material handle part of the work a controller would otherwise have to compute.
- Morphology-Control Co-design形态-控制协同设计(形态进化)Optimizing a robot's body and its control policy together, instead of fixing the body first.