Embodied AI Glossary中文

Object-centric Representation

以物体为中心的表示Advanced

Encoding a scene as a set of separate objects, instead of squeezing the whole image into one vector.

Object-centric representation means decomposing an image or scene into individual objects, each represented by its own set of features, such as position, shape, and category, rather than compressing the whole frame into a single feature vector. A representative method is Slot Attention (Locatello and colleagues, NeurIPS 2020), where a number of “slots” compete through attention and each ends up binding to one object, unsupervised. In robot manipulation, this kind of representation lets a policy attend only to the objects relevant to the task, making it more robust to changes in background or added distractor objects. 2022's VIOLA builds object-level representations from object proposals produced by a pretrained vision model, then uses a Transformer policy to select the relevant objects; the paper reports a 45.8% improvement in success rate over the strongest baseline.

ExampleA table holds a cup, a bowl, and a spoon; the policy first splits the scene into three object representations. When executing “put the spoon in the bowl,” it attends only to the spoon and bowl, unaffected even if the tablecloth changes color.

Also called
Object-centric Learning
Related
Representation Learning · Compositional Generalization · 3D Scene Graph · Distractor Objects · Vision Encoder · Keypoint Detection
Sources
Object-Centric Learning with Slot Attention
VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors

See it in the full glossary →