CleanRL
AdvancedAn open-source deep reinforcement-learning library that implements each algorithm as a single, readable file.
CleanRL is an open-source deep reinforcement-learning library, primarily authored by Shengyi Huang and others, with an accompanying paper published in JMLR in 2022. Its distinguishing feature is “single-file implementation”: algorithms such as PPO, DQN, SAC, TD3, and DDPG are each written in one standalone Python file, with everything from environment creation to the network and training loop on the same page, with none of the layered abstraction found in more modular libraries. The benefit is that newcomers can read every detail of an algorithm start to finish, and researchers can copy and modify a file directly to run an experiment; it also has built-in logging to Weights & Biases and TensorBoard. Compared with a modular library like Stable-Baselines3, it's better suited to learning and research prototyping than to being called as a general-purpose toolkit.
ExampleRunning python cleanrl/ppo_continuous_action.py --env-id HalfCheetah-v4 trains PPO on a MuJoCo environment, with the curves viewable in TensorBoard.
- Related
- Reinforcement Learning · Proximal Policy Optimization · Stable-Baselines3 · Gymnasium · rsl_rl · Weights & Biases (W&B)
- Sources
- vwxyzjn/cleanrl (GitHub)
CleanRL (JMLR 2022)