D4RL
D4RL 离线强化学习基准AdvancedThe most widely used standard datasets and benchmark for offline reinforcement learning.
D4RL was released in 2020 by Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine at UC Berkeley and Google Brain. At the time, offline reinforcement learning (training a policy purely from pre-collected data, with no further interaction with the environment during training) lacked dedicated test data, making it hard to compare results across papers. D4RL designs its datasets around how the data was collected: trajectories from hand-coded controllers, human demonstrations, multi-task data, and data mixed from several different policies, covering Maze2D and AntMaze maze navigation, Gym-MuJoCo locomotion control, Adroit dexterous-hand tasks, Franka Kitchen manipulation, and even Flow traffic and CARLA driving. It provides normalized scores to make comparison across tasks easier. Offline RL algorithm papers such as CQL and IQL commonly report D4RL scores. The original repository is no longer maintained: its environments moved to Gymnasium-Robotics and its data moved to Minari.
Examplehalfcheetah-medium-v2 consists of running data for the HalfCheetah collected from a policy trained to medium skill level; an offline algorithm can only learn from this fixed batch of data, and results are reported as a normalized score via get_normalized_score.
- Also called
- Datasets for Deep Data-Driven Reinforcement Learning
- Related
- Offline Reinforcement Learning · Benchmark · Gym/Gymnasium MuJoCo Tasks · Adroit · Franka Kitchen · Conservative Q-Learning
- Sources
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning (arXiv 2004.07219)
Farama-Foundation/D4RL GitHub(已弃用,迁往 Minari) (Chinese) - As of
- 2026-09