09Data & Collection
The raw material for learning: how demonstration data is collected, what datasets exist, and how data is processed and scaled. · 177 terms
- 9.1Data’s basic unit and sources13
- 9.2Real-robot collection: teleop, teaching22
- 9.3No robot needed: handheld and wearable13
- 9.4Motion capture and humanoid motion data14
- 9.5Human video and first-person data27
- 9.6Simulated assets and synthetic data24
- 9.7Major real-robot datasets22
- 9.8Data formats and tools9
- 9.9Cleaning, labeling, and mixing21
- 9.10Scaling: from data farms to flywheels12
9.1Data’s basic unit and sources
First, see that a demonstration is made of observations and actions, then the main sources: real robots, simulation, and human video.
- The recorded observations and actions from a human performing a task, used as raw material for robot imitation learning.
- Observation-Action Pair观测-动作对One training sample: what the robot 'saw' at a given moment paired with what it 'did' next.
- Action Label动作标签The action value a robot executed at each moment in a dataset — the “answer” imitation learning tries to fit.
- Multimodal Data多模态数据Images, depth, joint state, force and touch, language, and other signals recorded in sync while a robot works.
- Tactile Data触觉数据Pressure, shear force, and contact-surface deformation recorded by a tactile sensor when it touches an object.
- Multi-sensor Time Synchronization / Timestamp Alignment多传感器时间同步(时间戳对齐)Aligning readings from cameras, encoders, IMUs, and other sensors onto one shared clock so simultaneous data lines up correctly.
- Real-Robot Data真机数据Data recorded from an actual robot physically performing tasks in the real world.
- The core bottleneck in embodied AI: usable real robot-interaction data falls far short of what's needed.
- Simulation Data仿真数据Training data generated automatically by having a virtual robot perform tasks inside a physics simulator.
- Synthetic Data合成数据Training data manufactured by computers, via simulation or generative models, instead of collected from the real world.
- Human Video Data人类视频数据Video of people doing everyday tasks, collectable without any robot, used to teach robots to understand and copy manipulation.
- The vast supply of paired images and text, plus image-based Q&A, collected from the web, teaching a model about the world.
- Data Pyramid数据金字塔A three-tier framework organizing robot training data by volume and how closely it matches a real robot.
9.2Real-robot collection: teleop, teaching
Starting with the most reliable source, real-robot data: a person pushes, guides, or remotely operates the robot while it’s recorded.
- A human controls a robot in real time while the process is recorded to produce training data.
- A human physically pushes a robot arm through a motion by hand, and the robot records it as a demonstration.
- A human moves a leader arm, and the robot's follower arm copies each joint angle in real time.
- Berkeley's open-source low-cost leader-follower teleoperation device, using a scaled-down replica arm to control the real robot.
- Teleoperation where position and force are exchanged in both directions between leader and follower, so the operator feels what the robot touches.
- The operator wears an exoskeleton matching the robot's joints, controlling it directly with their own arm motion.
- AirExoAirExo 外骨骼A low-cost bimanual exoskeleton from Shanghai Jiao Tong's Cewu Lu group that can both teleoperate a robot and collect data without one.
- SpaceMouse Teleoperation3D 鼠标遥操作Controlling a robot arm's end effector in six degrees of freedom by pushing, pulling, twisting, and tilting a 3D mouse.
- VR TeleoperationVR 遥操作Controlling a robot in real time by wearing a VR/XR headset and using hand tracking or controllers.
- Apple's head-mounted display, whose hand tracking embodied-AI researchers commonly use for teleoperation and collecting human data.
- Data Glove数据手套A wearable sensing glove that tracks each finger's bend and the hand's pose in real time.
- Haptic Glove力反馈手套A teleoperation glove that both records hand motion and pushes force back against the operator's fingers.
- Tactile Glove触觉手套A glove covered in pressure or tactile sensors on the fingers and palm, recording the contact forces of a human hand.
- Motion Retargeting动作重定向Converting a human's (or another robot's) motion into the joint motion a target robot can actually perform.
- A general vision-based teleoperation system that uses an ordinary camera to capture hand motion and drive many arms and dexterous hands.
- An open-source library that converts human hand keypoint poses into joint angles for many different robotic dexterous hands.
- An open-source system for teleoperating a humanoid robot by wearing a VR headset and seeing stereo video from the robot's own viewpoint.
- Unitree xr_teleoperate宇树 xr_teleoperate(XR 遥操作)Unitree's open-source XR headset teleoperation program, used to control its humanoid robots and record training data.
- A teleoperation system that uses Apple Vision Pro to control two dexterous hands in real time, with vibration feedback, for imitation-learning data collection.
- PICO's open-source cross-platform XR teleoperation framework, controlling many kinds of robots with a headset and controllers to collect data.
- NVIDIA Isaac TeleopIsaac Teleop 遥操作框架NVIDIA's open-source teleoperation and data-collection framework, with unified support for XR headsets, gloves, and other input devices.
- One operator simultaneously controls a robot's arms, torso, and legs or base — its whole body at once.
9.3No robot needed: handheld and wearable
Teleoperation needs a real robot and is slow and costly, so instead a person uses a gripper-like tool or wearable device directly.
- Collecting data with no real robot present, by having a human hold or wear a device shaped like the robot's end effector.
- A human holds a camera-equipped gripper and performs tasks directly, recording footage and gripper motion as robot training data.
- A handheld gripper for recording demonstrations without a real robot, producing data that trains policies deployable on actual arms.
- Training a manipulation policy on handheld-gripper human demonstrations, then deploying it on a quadruped fitted with a robot arm.
- A hardware-independent handheld-gripper data-collection system and dataset from Shanghai AI Lab, improving on UMI.
- AgileX Pika松灵 Pika 采集套件AgileX Robotics' handheld data-collection kit that records manipulation demonstrations without needing a real robot.
- Having a person wear cameras, gloves, and similar devices while doing their normal job, recording their motion as robot training data.
- Stanford's wearable hand motion-capture system that collects data by having a human use their own hands, for teaching robot dexterous hands.
- CMU's portable hand-data-collection system that lets ordinary people gather dexterous-manipulation data with their own hands in real-world settings.
- A term from China's robotics-data industry: building the capture device to match the robot's end-effector so recorded data transfers directly, with no retargeting.
- Skill Capture Glove技能采集手套Sunday Robotics's glove that lets a person do housework wearing it and directly produce robot training data.
- A wearable exoskeleton that lets a human hand collect dexterous-hand data directly, then digitally replaces the hand in footage with the robot hand.
- DEXOPDEXOP 数据采集装置MIT's passive hand exoskeleton that lets a human's own hand directly drive a robot hand to collect vision and touch data.
9.4Motion capture and humanoid motion data
From hands to the whole body: motion capture records a person’s full movement, then converts it into trajectories a humanoid can follow.
- Motion Capture动作捕捉Recording the motion of a person or object into a computer with high precision, using cameras or wearable sensors.
- Tracking markers on the body with multiple infrared cameras and triangulating them into 3D motion.
- A motion-capture method that straps a set of IMUs to the body to measure each segment's rotation and reconstruct full-body pose.
- Markerless (Video-Based) Motion Capture视频动捕(无标记动捕)Estimating a person's 3D motion directly from ordinary video, with no markers or motion-capture suit needed.
- BVH / FBX Motion Capture File FormatsBVH / FBX 动捕文件格式Two common motion-capture file formats that store a human skeleton's structure and its joint rotations frame by frame.
- AMASS (Archive of Motion Capture as Surface Shapes)AMASS 人体动捕数据集A large human-motion library that unifies 15 optical motion-capture datasets into the SMPL body-model format.
- LAFAN1LAFAN1 动捕数据集Ubisoft La Forge's public human motion-capture dataset, commonly used as a source of motion for humanoid robots to imitate.
- HumanML3DHumanML3D 数据集A text-to-motion dataset pairing 14,000 clips of 3D human motion with 45,000 natural-language descriptions.
- OMOMOOMOMO 人-物交互动作数据集About 10 hours of optical motion-capture data from Stanford, recording full-body motion while carrying everyday objects.
- Foot Skating / Floating / Penetration脚滑 / 漂浮 / 穿地(动作数据伪影)Three common physically-impossible foot errors that appear when human motion is retargeted onto a robot.
- General Motion RetargetingGMR 通用动作重定向Stanford's open-source tool that retargets human motion onto many different humanoid robots in real time.
- A data-generation method that retargets a human's interaction with objects and terrain onto a humanoid robot together, not just body pose.
- Humanoid-XHumanoid-X 数据集A humanoid-motion dataset built by extracting and captioning motion from massive amounts of internet human videos.
- PHUMAPHUMA 人形动作数据集A large-scale humanoid reference-motion dataset, filtered and retargeted under physical constraints.
9.5Human video and first-person data
Stepping back further to just filming people: video is cheap and abundant, but lacks action labels, and human hands aren’t robot hands.
- Egocentric Video第一人称视频Video filmed from a head- or glasses-mounted camera, showing the wearer's own point of view with the hands in frame.
- Exocentric Video第三视角视频Video shot from a camera outside the performer's own body, looking at a person or robot from a bystander's vantage point.
- Internet Video Data互联网视频数据The vast supply of publicly available online video — large and cheap, but without action labels a robot can use.
- Action-free Video无动作标签视频Video that has only images, no frame-by-frame action record, like human videos and web videos.
- Pseudo Action Labels伪动作标签Action labels inferred by a model from video, standing in for actions that were never actually recorded.
- Robotizing Human Videos / Human-to-Robot Video Translation人类视频机器人化(人→机视频转换)Editing footage of a human doing a task so it looks like a robot doing it, turning it into usable training data.
- Cross-Painting跨本体图像替换Digitally erasing the deployed robot from the camera feed and painting in the robot the policy was trained on, so a vision policy transfers without retraining.
- Something-Something V2Something-Something V2 数据集About 220,000 short videos of hands manipulating everyday objects, labeled with 174 action templates.
- EPIC-KITCHENSEPIC-KITCHENS 数据集A classic first-person video dataset of unscripted everyday cooking and kitchen tasks, filmed with head-mounted cameras.
- Ego4DEgo4D 数据集A roughly 3,670-hour egocentric daily-video dataset collected by Meta with 13 universities.
- Ego-Exo4DEgo-Exo4D 数据集A large-scale video dataset led by Meta, capturing skilled human activities simultaneously from first-person and third-person viewpoints.
- Project Aria GlassesProject Aria 眼镜Meta's research eyewear for collecting first-person data; it only records, it doesn't display anything to the wearer.
- NymeriaNymeria 数据集Meta's large-scale first-person dataset of everyday human motion, recorded with Aria glasses and a motion-capture suit.
- DexYCBDexYCB 数据集NVIDIA's multi-view dataset of human hands grasping objects, with 3D pose annotations for both the hand and the object.
- HOI4DHOI4D 数据集A first-person 4D hand-object interaction dataset with 2.4 million RGB-D frames, from Tsinghua and collaborators.
- OakInkOakInk 数据集A Shanghai Jiao Tong University hand-object interaction dataset, annotated with object use and two-handed manipulation.
- A motion-capture dataset of two-handed dexterous manipulation of articulated objects, with precise 3D meshes for hands and objects every frame.
- HOT3DHOT3D 数据集Meta's first-person hand-object interaction dataset, recorded with Aria glasses and Quest 3, with motion-capture ground truth.
- A framework that co-trains a single policy on Project Aria human-hand data and robot data together.
- PH2DPH2D 数据集(HAT)First-person human manipulation data collected with a VR headset, co-trained with humanoid robot data to train one policy.
- EgoDexEgoDex 数据集An Apple dataset of egocentric manipulation video captured with Vision Pro, annotated with precise 3D hand joints.
- UniHandUniHand 数据集BeingBeyond's dataset that unifies many kinds of human hand video into hand-motion annotations, for pretraining dexterous-hand VLA models.
- Project Go-BigFigure Project Go-BigFigure's plan to pretrain its humanoid robot using massive amounts of first-person human video.
- World In Your Hands (WIYH)它石 WIYH 数据集TARS Robotics's open-source dataset of over a thousand hours of real-world human manipulation, with vision, language, touch, and action.
- Egocentric-10KEgocentric-10K 数据集Build AI's open-source dataset of roughly ten thousand hours of first-person video recorded by factory workers.
- Xperience-10MXperience-10M 数据集Ropedia's open dataset of 10,000 hours of first-person, multimodal human experience, aimed at embodied AI.
- EgoVerseEgoVerse 数据集A continually growing platform of first-person human demonstration data for robot learning, built jointly by universities and companies.
9.6Simulated assets and synthetic data
No longer collecting one demo at a time: prepare object and scene assets, then generate demonstrations in bulk in sim or with generative models.
- YCB Object and Model SetYCB 物体集A standard set of purchasable everyday objects with 3D scans, letting grasping and manipulation research compare results directly.
- Google Scanned ObjectsGoogle Scanned Objects(GSO)Google's open-source library of high-precision 3D scans of 1,030 real household objects.
- ShapeNetShapeNet 3D 模型库A large-scale library of 3D object meshes organized by WordNet categories, with semantic annotations.
- ObjaverseObjaverse 3D 资产库A massive open library of 3D object models, led by the Allen Institute for AI.
- PartNet-MobilityPartNet-Mobility 数据集A library of articulated 3D object models with joint-motion annotations, ready to load directly into a simulator.
- ObjectFolderObjectFolder 多感官物体数据集A multisensory object dataset providing an object's appearance, tapping sound, and tactile readings all together.
- COCO / LVISCOCO / LVIS 数据集The most widely used object-detection and segmentation benchmark; LVIS extends COCO's images to over a thousand long-tail categories.
- ScanNetScanNet 数据集A large-scale indoor RGB-D scan dataset with 3D reconstruction and semantic annotations.
- Matterport3DMatterport3D 数据集A dataset of RGB-D indoor scans from 90 real buildings, a standard scene library for indoor navigation research.
- Habitat-Matterport 3D DatasetHM3D 数据集A dataset of 1,000 real buildings, 3D-reconstructed by Meta and Matterport, used to train navigation agents.
- GraspNet-1BillionGraspNet-1Billion 数据集A Shanghai Jiao Tong University dataset of real cluttered-scene grasping data, with over 1.1 billion annotated grasp poses.
- An NVIDIA grasp dataset labeled in bulk through physics simulation, containing about 17.7 million parallel-jaw grasps.
- DexGraspNetDexGraspNet 数据集A simulation-generated dexterous-hand grasping dataset from Peking University, with 1.32 million grasps across 5,355 objects.
- Demonstration data generated automatically by a hand-written rule-based program controlling the robot, instead of a human.
- A system that segments a few human demonstrations by object, transforms and stitches them, and auto-generates many new demonstrations.
- NVIDIA's automatic data-generation system that expands a few dozen demonstrations into tens of thousands of bimanual dexterous-hand demonstrations.
- A method that automatically turns one real-robot demonstration into many synthetic demonstrations with objects in new positions.
- A generative simulation framework where a large model proposes tasks, builds the simulated scene, and learns the skill.
- NVIDIA Physical AI DatasetNVIDIA 物理 AI 数据集NVIDIA's open collection of robotics and autonomous-driving datasets hosted on Hugging Face.
- SynGrasp-1BSynGrasp-1B 数据集A billion-frame simulated synthetic grasping dataset built by Galbot and others, used to pretrain GraspVLA.
- InternData-A1InternData-A1 数据集A large-scale simulated synthetic dataset from Shanghai AI Lab, built for pretraining general-purpose robot policies.
- Generative Data Augmentation生成式数据增强Using image or video generation models to rewrite existing robot data, producing new-scene training examples from old demonstrations.
- Synthetic robot training data generated by a video world model, with pseudo action labels added afterward.
- Real2Render2RealReal2Render2Real(R2R2R)A method that renders large amounts of robot training data from a phone scan plus one video of a human demonstration.
9.7Major real-robot datasets
Back to real robots and their public datasets: starting with OXE, which pools many robots, then key releases in order.
- Data collected from many different kinds of robots and pooled together for training.
- Open X-EmbodimentOpen X-Embodiment 数据集A Google-led open dataset pooling more than a million real-robot trajectories from 22 kinds of robots.
- Robot data from different robots, sensors, and collection methods, with mismatched formats and meanings.
- RoboNetRoboNet 数据集A 2019 dataset of multi-robot interaction video, about 15 million frames across 7 kinds of robots.
- BC-ZBC-Z 数据集A 2021 Google dataset of real-robot demonstrations across 100 tasks, collected to study zero-shot task generalization.
- Language-TableLanguage-Table 数据集Google's tabletop block-pushing dataset, with nearly 600,000 trajectories paired with natural-language instructions.
- RT-1 Robot Action DatasetRT-1 数据集Google's real-robot manipulation dataset, collected over 17 months using 13 robots.
- RH20TRH20T 数据集A multimodal real-robot manipulation dataset from Shanghai Jiao Tong University, with force, audio, and human demonstration video.
- BridgeData V2BridgeData V2 数据集A roughly 60,000-trajectory multi-task manipulation dataset collected by UC Berkeley on a low-cost WidowX arm.
- RoboSetRoboSet 数据集A real-robot kitchen-task dataset from CMU and Meta, with multi-task demonstrations and 4 camera views per frame.
- A real-robot manipulation dataset collected by 50 people across 564 scenes on three continents using Franka arms.
- All Robots In OneARIO 数据集A unified embodied-data format standard from Peng Cheng Laboratory and partners, and a roughly 3-million-item dataset built to that standard.
- A multi-embodiment real-robot manipulation dataset from the Beijing Humanoid Robot Innovation Center and Peking University, collected under a unified protocol.
- AgiBot WorldAgiBot World 数据集A large-scale real-robot manipulation dataset from AgiBot and partners, with the Beta release exceeding a million trajectories.
- Fourier ActionNet傅利叶 ActionNet 数据集Fourier's open-source teleoperated dataset of about 30,000 humanoid-plus-dexterous-hand manipulation trajectories.
- Galaxea Open-World Dataset星海图开放世界数据集A roughly 500-hour dataset Galaxea collected at real-world locations using a single model of mobile bimanual robot.
- Humanoid EverydayHumanoid Everyday 数据集A real-robot humanoid manipulation dataset from USC and the Toyota Research Institute, covering 260 tasks.
- RoboCOINRoboCOIN 数据集A cross-embodiment bimanual manipulation dataset led by BAAI, open-sourced across 15 kinds of robots.
- MolmoAct2-BimanualYAMBimanualYAM 数据集Ai2's open-source bimanual teleoperation dataset for MolmoAct2, totaling more than 720 hours.
- RoboVQARoboVQA 数据集A video question-answering dataset from Google DeepMind, built around long-horizon robot tasks.
- Baihu-VTouch Visuo-Tactile Dataset白虎-VTouch 视触觉数据集A cross-embodiment visuotactile dataset released in 2026 by China's National Humanoid Robot Innovation Center and Weitai Robotics.
- Daimon-Infinity戴盟 Daimon-Infinity 数据集A large-scale real-robot manipulation dataset with vision-and-touch signals, released by Daimon Robotics in 2026.
9.8Data formats and tools
What file formats those datasets use and what libraries read them: from HDF5 and RLDS to LeRobot.
- A general-purpose file format that stores many large arrays hierarchically in one file, commonly used for robot demonstration data.
- ZarrZarr 格式An open-source format for storing large multidimensional arrays in compressed chunks; Diffusion Policy and UMI use it for training data.
- A Google-proposed standard format and toolchain for organizing sequential-decision data as episodes made of steps.
- TensorFlow Datasets (TFDS)TFDS(TensorFlow Datasets)Google's open-source dataset library that downloads and loads a dataset into a trainable pipeline in one line of code.
- TFRecordTFRecord 格式TensorFlow's binary file format, storing a sequence of serialized records back to back.
- LeRobotDatasetLeRobot 数据集格式LeRobot's standard dataset format, storing state and actions in tables and camera footage as video.
- Apache ParquetParquet 格式An open-source column-oriented file format; LeRobot datasets use it to store numeric data like state and actions.
- WebDatasetWebDataset 格式A format that packages training samples into same-named files inside a series of tar shards, for fast large-scale sequential reads.
- MCAPMCAP 格式An open-source multimodal logging file format from Foxglove, the default recording format for ROS 2 bags.
9.9Cleaning, labeling, and mixing
Raw data still needs processing: cleaning and quality checks, adding labels, then selecting and mixing ratios before training.
- Data Cleaning数据清洗Finding and fixing or removing bad data before training, such as failed trajectories, empty actions, or misaligned frames.
- No-op (Idle) Action Filtering空操作 / 静止帧过滤Removing timesteps where the robot doesn't move from demonstrations before training, so the model doesn't learn to freeze in place.
- Checking each piece of collected robot data before training to catch and discard unusable or flawed samples.
- Re-sending a recorded action sequence to a robot or simulator to check whether the data can be reproduced.
- Blurring or removing personal information like faces, license plates, and voices from data so specific individuals can't be identified.
- Data Annotation数据标注Adding text descriptions, segment boundaries, bounding boxes, and other labels to raw data to tell a model what it is.
- Attaching a natural-language description to a robot trajectory, such as 'put the red cup in the sink,' so a model can follow instructions.
- Subtask Segmentation子任务切分Splitting one long demonstration into steps, marking each segment's start and end frame with a description.
- Auto-labeling自动标注Using pretrained models or scripts, instead of humans, to automatically add language instructions, bounding boxes, and other labels to robot data.
- Play Data玩耍数据Robot data recorded while an operator freely manipulates a scene with no specific task in mind.
- Hindsight Relabeling事后重标注After data is collected, relabeling a trajectory's goal or instruction to match whatever it actually achieved.
- Automatically writing or rewriting language instructions for robot trajectories, so the same motion gets many phrasings.
- Data Diversity数据多样性How broadly training data varies across scenes, objects, tasks, and viewpoints.
- Research into how, in robot imitation learning, generalization improves as training data grows.
- In-the-wild Data野外数据Data collected in homes, offices, outdoors, and other uncontrolled real-world settings, rather than a fixed lab bench.
- Human demonstration data that includes wasted motion, hesitation, mistakes, or inconsistent quality across operators.
- Data Curation数据筛选Picking out the subset of a large robot dataset that actually helps training, and discarding the harmful part.
- Using a handful of demonstrations for a target task to search a large dataset for similar clips, then training on the combined set.
- Data Mixture数据配比The proportion each data source is given when training on a mix of multiple sources at once.
- OXE Magic SoupMagic Soup(OXE 数据配方)The recipe Octo and OpenVLA used to select and weight subsets of Open X-Embodiment for pretraining.
- Data Leakage / Test-Set Contamination数据泄漏 / 测试集污染When information that should only appear at test time leaks into training, inflating evaluation scores beyond real-world performance.
9.10Scaling: from data farms to flywheels
Finally, how data volume is scaled up: data-collection farms, crowdsourcing, robots collecting their own data, and deployment feeding a flywheel.
- A facility built to replicate real-world settings, where large numbers of robots are teleoperated to mass-produce training data.
- A front-line worker who records robot training data through teleoperation or wearable demonstration devices.
- Data Collection SOP采集 SOPThe written procedure that specifies exactly how each demonstration should be collected and what counts as an acceptable one.
- The portion of raw collected data that passes quality control and can actually be used to train a model.
- Distributing robot data-collection tasks to large numbers of non-professional workers, who are paid for each validated, usable clip.
- Stanford's 2018 platform for crowdsourced robot demonstrations, teleoperated remotely from a smartphone.
- Letting a robot propose its own tasks, attempt them, and log the results, with little or no human teleoperation.
- Failure Data失败数据Trajectories that didn't complete the task, useful for learning to recover from mistakes and for training reward models.
- Data recorded when a human takes over from an autonomously running policy just as it's about to make a mistake.
- Demonstration data specifically recorded to show how a robot gets back on track after starting to go wrong.
- A term from China's robotics industry: feeding data generated during real-world robot deployment back into training, then redeploying the updated model.
- Data Flywheel数据飞轮A self-reinforcing loop: deployment produces data, the data improves the model, and the better model gets deployed even more widely.