Embodied AI Glossary中文

Scene Generalization

场景泛化Common

A policy still completing its task when moved into a room, table, lighting, or background it never saw in training.

This is one dimension of generalization: whether a trained robot policy still succeeds once it is moved to an environment absent from training — a new room, a new tabletop texture, different lighting, a shifted camera position, or a cluttered background. Robot data is mostly collected in a handful of labs, so a model can easily end up memorizing the scene's appearance itself and fail the moment the kitchen changes, making this a key metric for whether a robot can actually enter a user's home. Tianhe Yu, Chelsea Finn, and colleagues (2023) broke the contributing factors into 11 categories, including lighting and camera pose, and tested each separately; Physical Intelligence's π0.5 (2025) uses “tidying a kitchen and bedroom in an entirely new home” as its main evaluation. Common countermeasures include collecting data from a larger and more diverse set of environments, data augmentation, and co-training with web data.

Exampleπ0.5 completes long-horizon tasks such as tidying a kitchen and organizing a bedroom in real homes that never appeared in its training data.

Also called
Environment Generalization
Related
Generalization · Object Generalization · Task Generalization · Visual Generalization · Out-of-Distribution · Data Diversity
Sources
Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (Xie et al., 2023)
π0.5: a Vision-Language-Action Model with Open-World Generalization
Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments
As of
2025-04

See it in the full glossary →