Embodied AI Glossary中文

WebDataset

WebDataset 格式Advanced

A format that packages training samples into same-named files inside a series of tar shards, for fast large-scale sequential reads.

WebDataset is an open-source deep-learning data format and Python library of the same name, developed by Thomas Breuel; the official PyTorch blog has featured it, and it pairs naturally with NVIDIA's AIStore storage service. It doesn't invent a new file format — it uses standard tar archives directly: the multiple files belonging to one sample (say, 000042.jpg and 000042.json) share the same base name minus extension and sit next to each other, and the whole dataset is split into consecutively numbered shards. During training, shards are read in order with no need to decompress the whole archive, and can be streamed directly from local disk, a web server, or cloud storage, which suits huge numbers of small files and multi-GPU training while avoiding the filesystem overhead of random access. It implements PyTorch's IterableDataset interface, so it plugs straight into a DataLoader. In embodied AI, it's commonly used to package large-scale video and image data — for instance, BeingBeyond's UniHand_Preview, released on Hugging Face, uses this format.

ExampleOne million first-person videos are split into 1,000 tar shards, each containing pairs like number.mp4 plus number.json; during training, multiple machines each read a different subset of shards in parallel.

Also called
wds
Related
Apache Parquet · Hierarchical Data Format version 5 · Zarr · TFRecord · LeRobotDataset · RLDS (Reinforcement Learning Datasets)
Sources
webdataset/webdataset (GitHub)
Efficient PyTorch I/O library for Large Datasets, Many Files, Many GPUs (PyTorch Blog)
BeingBeyond/UniHand_Preview (Hugging Face)

See it in the full glossary →