Embodied AI Glossary中文

Compositional Generalization

组合泛化Advanced

Recombining separately learned elements to handle a combination that was never seen together during training.

Compositional generalization means a model, after seeing individual elements — words, objects, skills, environmental factors — separately during training, can still handle them correctly when they are combined in a new way. This problem was first discussed in linguistics and cognitive science: after learning the new verb “dax,” a person can immediately understand “dax twice” and “sing and dax”; Brenden Lake and Marco Baroni's 2018 SCAN benchmark (ICML) showed that recurrent neural networks fail badly on tests that require genuine composition. Applied to robotics, the elements can be object categories, placements, tabletop textures, and camera viewpoints, or atomic skills such as “grasp” and “put into.” It matters because the number of possible combinations grows multiplicatively with the number of elements, making it impossible to collect data for every one. Zhou Gao and colleagues (RSS 2024) found that robot policies genuinely can compose certain environmental factors, and used this to design a data-collection scheme that needs far less data. Whether a VLA can carry out “put an unseen object into an unseen container” is also commonly tested as compositional generalization.

ExampleThe training data includes 'put the red cup in the bowl' and 'put the blue plate in the basket'; at test time the robot is asked to 'put the red cup in the basket.'

Also called
Systematic Generalization
Related
Generalization · Task Generalization · Semantic Generalization · Zero-shot · Out-of-Distribution · Skill Primitive
Sources
Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks (Lake & Baroni, ICML 2018)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization (Gao et al., RSS 2024)

See it in the full glossary →