Embodied AI Glossary中文

EgoDex

EgoDex 数据集Common

An Apple dataset of egocentric manipulation video captured with Vision Pro, annotated with precise 3D hand joints.

EgoDex was released by Apple's research team in May 2025, and the paper was accepted at ICLR 2026. Collectors wore an Apple Vision Pro to perform everyday tabletop manipulation; the headset's multiple calibrated cameras and on-device SLAM (simultaneous localization and mapping) computed the 3D position and orientation of the head, upper body, and 25 joints per hand, in sync, while recording. The dataset totals 829 hours, about 90 million frames, and 338,000 demonstrations, covering 194 kinds of tabletop tasks, with language descriptions generated by GPT-4. It fills a gap left by video datasets like Ego4D, which lack precise hand pose, and can be used to train hand-trajectory prediction models that then transfer to dexterous-hand robots. The data is released under a CC BY-NC-ND license, for non-commercial use only.

ExampleThe paper trains an imitation-learning policy on EgoDex that takes an egocentric frame as input and predicts the 3D motion trajectory of both hands over the following span of time, and uses this to build a hand-trajectory-prediction benchmark.

Related
Egocentric Video · Human Video Data · Apple Vision Pro · Hand Pose Estimation · Ego4D · Dexterous Manipulation
Sources
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video (arXiv 2505.11709)
apple/ml-egodex (GitHub)
As of
2026-03

See it in the full glossary →