Embodied AI Glossary中文

PoliFormer

Advanced

A Transformer navigation policy trained purely with large-scale on-policy reinforcement learning in simulation, then deployed straight to real robots.

PoliFormer was released by the Allen Institute for AI (Ai2) in June 2024 and published at CoRL 2024. It uses RGB images only: a Vision Transformer encoder (DINOv2 in the code) encodes each frame, followed by a causal Transformer decoder that aggregates a fairly long history and outputs navigation actions. Training happens entirely in simulation: on-policy reinforcement learning across a huge number of procedurally generated houses from ProcTHOR, run in parallel across many machines for hundreds of millions of interactions. It reaches 85.5% success on the CHORES-S object-goal navigation benchmark, a 28.5-percentage-point absolute improvement over the previous best method, and deploys without any additional tuning to two different real robots, the LoCoBot and the Stretch RE-1; it also transfers directly to downstream tasks like object tracking and open-vocabulary navigation. PoliFormer shows that reinforcement learning combined with Transformers and large-scale simulation also benefits from scale.

ExampleTold to 'find the apple in the kitchen,' a Stretch robot running PoliFormer decides whether to move forward or turn based only on its head-camera feed, step by step, and stops next to the apple once it spots it.

Also called
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
Related
Object-Goal Navigation · On-Policy · ProcTHOR (Large-Scale Embodied AI Using Procedural Generation) · AI2-THOR · Allen Institute for AI · Sim-to-Real Transfer
Sources
arXiv 2406.20083: PoliFormer
GitHub: allenai/poliformer
As of
2024-11

See it in the full glossary →