Embodied AI Glossary中文

Something-Something V2

Something-Something V2 数据集SSv2Advanced

About 220,000 short videos of hands manipulating everyday objects, labeled with 174 action templates.

Something-Something V2 is an action-recognition video dataset released by the company TwentyBN (20BN), now distributed for download by Qualcomm. It contains 220,847 short videos, each filmed by a crowdworker following a given action template such as “putting something into something” or “turning something upside down,” across 174 categories in total. The object in each template is abstracted as “something,” so a model can't guess the answer just by recognizing the object — it has to actually understand the motion and sequencing between the hand and the object — which is why it's commonly used to test a video model's understanding of temporal order. In embodied AI, it's treated as a source of human manipulation video: the videos carry no robot action labels, but they contain large amounts of hand-object interaction, useful for pretraining visual representations and for latent-action pretraining.

ExampleLAPA does latent-action pretraining using only the roughly 220,000 human videos in SSv2, then fine-tunes on robot data; the project page reports it outperforms OpenVLA, which was trained on Bridge data, on average.

Also called
20BN-Something-Something V2, Sth-Sth V2, SSv2
Related
Human Video Data · Action-free Video · Latent Action Pretraining · LAPA · Ego4D · EPIC-KITCHENS
Sources
The something something video database for learning and evaluating visual common sense (arXiv 1706.04261)
Hugging Face: something_something_v2 数据集卡片 (Chinese)
LAPA: Latent Action Pretraining from Videos(项目页) (Chinese)
As of
2024-10

See it in the full glossary →