Generalization
泛化EssentialA model's ability to still perform correctly in new situations it never saw during training.
Generalization is a core machine-learning concept: how well a model performs on new examples outside its training data, rather than just memorizing the training set (a model that memorizes training data and falls apart on new data is said to overfit). Generalization is especially hard in robot learning because the real world varies so much — a different tablecloth, a different cup, a shifted position, or a differently worded instruction can all make a policy fail. Researchers therefore often evaluate different kinds of variation separately: object generalization, scene generalization, position generalization, instruction generalization. STAR-Gen (2025), proposed by Jensen Gao, Dorsa Sadigh, and colleagues, splits generalization into three categories — visual (appearance and background changes), semantic (concept and instruction changes), and behavioral (requiring a different way of acting) — and found that open-source VLA models, despite being pretrained on internet-scale language data, still often struggle specifically with semantic generalization. Generalization ability is the main yardstick for judging a generalist policy.
ExampleA policy trained only to pick up a red block on a white table is tested on a wood-grain table (visual generalization), or given the instruction “pick up the thing that can hold water” instead (semantic generalization), to probe how well it generalizes.
- Also called
- Generalization Ability
- Related
- Overfitting · Out-of-Distribution · Zero-shot · Object Generalization · Visual Generalization · Semantic Generalization
- Sources
- A Taxonomy for Evaluating Generalist Robot Manipulation Policies (STAR-Gen)
- As of
- 2025-03