Embodied AI Glossary中文

End-to-End Training of Deep Visuomotor Policies

端到端视觉运动策略(引导策略搜索)GPSAdvanced

A landmark 2015 Berkeley paper that used a convolutional network to output robot-arm joint torques directly from camera images.

This paper was released in April 2015 by Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel at UC Berkeley, and published in JMLR in 2016. At the time, most robots were designed with perception and control as separate components; the paper set out to test whether jointly training both inside a single network would work better. The policy is a convolutional network with about 92,000 parameters, taking a monocular image and joint state as input and outputting joint torques for a 7-DOF arm at 20Hz. Training uses guided policy search: with object positions known during training, trajectory optimization first finds good actions, and supervised learning then has an image-only policy imitate them — turning reinforcement learning into supervised learning. The spatial softmax layer introduced in this paper was later adopted by many subsequent visuomotor policies.

ExampleThe PR2 robot learned to hang a coat hanger on a rod, fit blocks into a shape-sorting cube, use the claw of a toy hammer to pull out a nail, and screw on a bottle cap, judging target positions entirely from the camera; each policy took 3 to 4 hours to train in total, of which only about 15 minutes was actual real-robot execution.

Also called
Guided Policy Search, GPS, Levine et al. 2016
Related
Visuomotor Policy · End-to-End · Spatial Softmax · Trajectory Optimization · Reinforcement Learning · Imitation Learning
Sources
End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)
JMLR 17(39):1-40, 2016

See it in the full glossary →