Embodied AI Glossary中文

Sparsh

Advanced

A general-purpose vision-based-touch representation model from Meta, pretrained with self-supervision on a large set of tactile images.

Sparsh is a family of tactile encoders from Meta FAIR and collaborators, published at CoRL 2024, built for vision-based tactile sensors — sensors that perceive contact by having a camera image the deformation of a soft gel surface. Previously, each sensor type and each task needed its own labeled training data; Sparsh instead uses self-supervised learning, pretraining on more than 460,000 unlabeled tactile images with methods such as MAE, DINO, and JEPA, to produce a representation that generalizes across sensors including DIGIT, GelSight 2017, and GelSight Mini. The authors also released the TacBench benchmark, covering six categories of tasks from recognizing tactile properties to force estimation, slip detection, and manipulation planning. The paper reports that self-supervised pretraining outperforms task- and sensor-specific end-to-end training by 95.1% on average on TacBench. Code and weights are open-sourced.

ExampleA Sparsh encoder attached behind a DIGIT sensor lets a force-estimation or slip-detection head be trained with only a small amount of labeled data.

Also called
Self-Supervised Touch Representations for Vision-Based Tactile Sensing
Related
Tactile Representation Learning · Vision-Based Tactile Sensor · DIGIT · Self-Supervised Learning · Joint-Embedding Predictive Architecture · AnyTouch
Sources
Sparsh: Self-supervised touch representations for vision-based tactile sensing (arXiv)
facebookresearch/sparsh (GitHub)
As of
2024-10

See it in the full glossary →