Embodied AI Glossary中文

Language to Rewards

Language to Rewards(L2R)L2RAdvanced

Has a large language model translate plain instructions into a reward function, handed to a real-time optimizer to generate robot motion.

Language to Rewards was released in June 2023 by Google DeepMind, an oral presentation at CoRL 2023. Having a large model output actions directly, or call preset skills, makes it hard to produce low-level motions like “stand up” or “moonwalk.” L2R instead has the large model write a reward function: a reward translator first expands the instruction into a description of the motion, then writes it as reward code (a set of objectives and weights); a motion controller, MuJoCo MPC (a real-time optimizer based on model predictive control), then solves for the action that satisfies that reward, replanning in real time as the user adds corrections. Across 17 tasks on a simulated quadruped and a dexterous robot hand, it completed 90%, versus 50% for a baseline built on preset skill primitives; it also demonstrated pushing an object on a real robot arm. It belongs to the same “have a large model write the reward” line of work as Eureka.

ExampleThe user tells a simulated quadruped to “walk backward like doing the moonwalk,” watches the result, then adds a few corrections; after a few rounds of this back-and-forth, the robot has learned to moonwalk.

Also called
L2R, Language to Rewards for Robotic Skill Synthesis
Related
Reward Function · Model Predictive Control · Large Language Model · Code as Policies · Eureka · MuJoCo (Multi-Joint dynamics with Contact)
Sources
Language to Rewards for Robotic Skill Synthesis (arXiv 2306.08647)
Language to Rewards 项目页 (Chinese)
As of
2023-11

See it in the full glossary →