Robot Utility Models
RUMAdvancedPolicies that perform single tasks such as opening a cabinet or a drawer in unfamiliar homes with no fine-tuning at all.
Robot Utility Models (RUM) was released by researchers at New York University, Hello Robot, and Meta in September 2024. The problem it addresses: robot policies usually need fresh data collection and fine-tuning every time they move to a new room, whereas language and vision models can be used off the shelf. The authors trained one specialist policy each for five tasks — opening a cabinet, opening a drawer, picking up a tissue, picking up a paper bag, and righting a fallen object — using Stick-v2, a handheld gripper fitted with an iPhone, to collect about 1,000 demonstrations per task across roughly 40 environments, with VQ-BeT as the policy architecture. At deployment, GPT-4o judges whether an attempt succeeded and triggers a retry on failure. The result is about 90% success on unseen environments and objects, and the policies transfer to other robots such as the xArm. The authors' main conclusion is that data diversity matters more than the training algorithm or policy architecture. Code, data, and hardware designs are all open-sourced.
ExampleRUM's cabinet-opening policy is mounted on a Hello Robot Stretch and placed in an apartment where no data was ever collected; without any fine-tuning, it goes and opens the kitchen cabinet.
- Also called
- RUM, RUMs, Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments
- Related
- Zero-shot · Scene Generalization · Handheld Gripper Data Collection · Dobb-E · Behavior Transformer · Specialist Policy
- Sources
- arXiv 2409.05865: Robot Utility Models
Robot Utility Models 项目主页 (Chinese) - As of
- 2024-09