Embodied AI Glossary中文

Semantic Generalization

语义泛化Advanced

Still understanding and correctly acting on unfamiliar objects, concepts, or phrasing it hasn't seen before.

Semantic generalization is a generalization axis commonly used when evaluating robot policies, especially vision-language-action (VLA) models: whether a model can generalize at the level of meaning — unseen object categories, new containers, instructions phrased differently, or tasks requiring common-sense or conceptual reasoning. It is distinguished from visual generalization (changes in background, lighting, texture) and generalization at the execution level (changes in object position or starting pose). 2025's “What Can RL Bring to VLA Generalization?” builds tests along visual, semantic, and execution axes; the semantic category includes unseen objects, unseen containers, unseen instruction phrasings, and distractor containers. VLA models are seen as promising largely because they can inherit semantic knowledge from internet-scale image-text pretraining.

ExampleRT-2 can carry out instructions such as “move the coke can next to the photo of Taylor Swift” or “pick up the thing that could be used as an improvised hammer” (it chose a rock) — concepts nowhere in the robot's own training data.

Related
Generalization · Visual Generalization · Object Generalization · Compositional Generalization · Vision-Language-Action Model · RT-2
Sources
What Can RL Bring to VLA Generalization? An Empirical Study (arXiv 2505.19789)
RT-2: New model translates vision and language into action (Google DeepMind)
As of
2025-05

See it in the full glossary →