Model-Free Reinforcement Learning
无模型强化学习CommonReinforcement learning that skips building an environment model and learns a policy or value function directly from trial-and-error data.
The counterpart to model-based reinforcement learning: the algorithm never tries to estimate the environment's state-transition or reward dynamics, and learns a policy or value function directly from (state, action, reward) samples gathered through interaction — essentially pure trial-and-error learning. It splits into two main families: methods that directly optimize the policy, such as policy gradient and PPO; and methods that learn a Q-function (estimating how much return ultimately follows from taking a given action in a given state), such as Q-learning and DQN, with DDPG and SAC sitting somewhere between the two. The advantage is simplicity and immunity to model error; the disadvantage is needing a huge amount of interaction, since sample efficiency is low. Robotics commonly compensates for this with GPU-parallel simulation: running thousands of environments at once in Isaac Gym or Isaac Lab to collect data, then transferring the trained policy to a real robot.
ExampleThe mainstream approach to legged-robot locomotion control: train a walking policy with PPO from the rsl_rl library inside Isaac Lab, never building an environment model at all, then deploy the trained policy to the real robot.
- Also called
- Model-Free RL
- Related
- Model-Based Reinforcement Learning · Policy Gradient · Proximal Policy Optimization · Deep Q-Network · Sample Efficiency · Massively Parallel Reinforcement Learning
- Sources
- Wikipedia: Model-free (reinforcement learning)
OpenAI Spinning Up: Kinds of RL Algorithms
RSL-RL (GitHub, leggedrobotics)