Embodied AI Glossary中文

Video Object Segmentation

视频目标分割VOSAdvanced

Continuously cutting out a pixel mask for a specified object in every frame of a video.

Given a video, video object segmentation (VOS) outputs a pixel-level mask of the target object (marking which pixels belong to it) in every frame, and keeps recognizing it as the same object even as it moves, deforms, or is occluded. It’s grouped by how much guidance is given: semi-supervised VOS is given the target’s mask on the first frame and propagates it forward; unsupervised VOS is given no hint and must find the main object automatically; interactive VOS lets a user click partway through to correct the mask. DAVIS is a classic benchmark. In 2024, Meta’s SAM 2 unified image and video segmentation using a transformer with memory, letting a single click keep tracking an object through an entire video. In robotics, VOS is commonly used to keep a mask on a manipulation target, to automatically label data, or to give a policy an object-centric input.

ExampleClicking once on a cup on the table in the first frame, SAM 2 outputs a mask for the cup in every frame of the whole grasping sequence, keeping track of it even when the gripper partially covers it.

Also called
VOS
Related
Instance Segmentation · Object Tracking · SAM 2 · Segment Anything Model · Mask · Auto-labeling
Sources
DAVIS: Densely Annotated VIdeo Segmentation
SAM 2: Segment Anything in Images and Videos (arXiv:2408.00714)
As of
2024-10

See it in the full glossary →