Visual Generalization
视觉泛化AdvancedStill completing the task when the scene looks different: new background, lighting, colors, distractors, or camera angle.
Visual generalization is a subcategory of generalization: whether a robot policy still completes the same task under visual conditions it never saw in training, such as a new background or tablecloth, different lighting, an object with a new color or texture, extra distractor objects in the frame, or a moved camera. The task and the actions needed haven't changed, only what the scene looks like, so it is usually evaluated separately from semantic generalization (new objects, new instructions) and position generalization; OpenVLA's real-robot evaluation lists it as its own category. Imitation-learning policies that learn actions directly from pixels tend to also memorize irrelevant details like background and lighting. In 2023, Xie, Finn, and colleagues isolated these factors one by one and found that new backgrounds are the easiest to adapt to and new camera positions the hardest. Common fixes include domain randomization, data augmentation, pretrained vision encoders, and collecting data across more varied scenes.
ExampleOpenVLA's real-robot evaluation of “put the eggplant in the pot” used a pot made of papier-mâché, visually different from the pots in the BridgeData V2 training data, to test whether the policy could still recognize the pot and complete the task.
- Also called
- Visual Robustness, Appearance Generalization
- Related
- Generalization · Semantic Generalization · Spatial Generalization · Distractor Objects · Domain Randomization · Data Augmentation
- Sources
- Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (Xie et al., 2023)
OpenVLA: An Open-Source Vision-Language-Action Model
What Can RL Bring to VLA Generalization? An Empirical Study - As of
- 2025-05