Embodied AI Glossary中文

PerAct

Advanced

A multi-task manipulation policy that voxelizes the scene and uses a Perceiver Transformer to predict the next key pose.

PerAct was introduced by Mohit Shridhar, Lucas Manuelli, and Dieter Fox at the University of Washington and NVIDIA, presented at CoRL 2022. It converts RGB-D observations into a 100×100×100 voxel grid (a 3D grid of cube-shaped cells) and feeds it into a PerceiverIO Transformer together with the language instruction. Rather than outputting a continuous trajectory, PerAct predicts the 'next best voxel': which cell the end effector should move to, a discretized rotation, whether the gripper should open or close, and whether collision avoidance is needed; a motion planner then carries the arm to this key pose. A single model trains jointly on 18 RLBench tasks (249 variations) and 7 real-world tasks, each with only a handful of demonstrations. PerAct established the 'combine a 3D scene representation with keyframe action prediction' recipe, and later work such as RVT, 3D Diffuser Actor, and BridgeVLA all use it as a baseline.

ExampleFor the instruction 'open the middle drawer,' PerAct first predicts which voxel cell holds the handle and what orientation to grasp it at; once the gripper closes on the handle, it then predicts the end-effector position that pulls the drawer open.

Also called
Perceiver-Actor, Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
Related
RVT-2 · 3D Diffuser Actor · RLBench · Voxel · Keyframe Action Prediction · CLIPort
Sources
arXiv 2209.05451: Perceiver-Actor
PerAct 项目主页 (Chinese)
As of
2022-11

See it in the full glossary →