Eureka
CommonHas GPT-4 write reward-function code and repeatedly refine it, automating the reward design that reinforcement learning depends on.
Eureka was released in October 2023 by NVIDIA together with the University of Pennsylvania, Caltech, and UT Austin, and published at ICLR 2024. Reinforcement learning's performance depends heavily on its reward function, and writing one by hand is slow and relies on expert intuition. Eureka gives GPT-4 the environment's source code and a task description and has it write several candidate reward functions in one pass; each is used to train a policy in parallel GPU simulation, and statistics from training on each reward term are fed back to the model (“reward reflection”) so it can rewrite an improved version for the next round, iterating like an evolutionary search. Across 10 robot embodiments and 29 open-source reinforcement-learning environments, its rewards beat human expert-designed ones on 83% of tasks, for an average normalized improvement of 52%. A follow-up, DrEureka, applies the same idea to sim-to-real transfer.
ExampleUsing a reward Eureka designed, a simulated Shadow dexterous hand learned to spin a pen rapidly and continuously between its fingers.
- Also called
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Related
- Reward Function · Reward Engineering · Large Language Model · DrEureka · Code as Policies · Isaac Gym
- Sources
- Eureka (arXiv 2310.12931)
- As of
- 2024-04