Early Termination
提前终止AdvancedEnding an episode immediately and resetting the environment as soon as a failure state, like falling, occurs during training.
Early termination is a common trick in reinforcement-learning training: an episode doesn't have to run for its full fixed length — as soon as a preset condition triggers, such as the torso touching the ground, the body tilting past some angle, or tracking error growing too large, the episode ends immediately and resets, with no more reward for the time that would have remained. Peng and colleagues' 2018 DeepMimic evaluated this specifically and found it, together with reference state initialization, to be key to letting a simulated character learn highly dynamic skills like backflips. It acts partly as an implicit penalty, since the policy learns to actively avoid failure, and partly as a way to avoid wasting samples on useless post-failure states. Implementations need to distinguish failure termination from timeout truncation — the latter isn't a failure, and value estimation typically still bootstraps through it.
ExampleIn legged_gym's legged-robot environments, an environment is judged to have fallen and reset as soon as the contact force on a designated body part exceeds a threshold; exceeding the maximum episode length is instead recorded separately as a time_out, with no termination penalty.
- Also called
- ET, Termination Condition
- Related
- Termination vs. Truncation · Reference State Initialization · Reward Shaping · Episode · DeepMimic · Massively Parallel Reinforcement Learning
- Sources
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills (arXiv:1804.02717)
legged_gym: legged_robot.py (check_termination)