Embodied AI Glossary中文

QT-Opt

Advanced

Google's large-scale reinforcement learning grasping system, trained on more than 580,000 real-robot grasp attempts to learn a visual Q-function.

QT-Opt was released by Dmitry Kalashnikov, Sergey Levine, and colleagues at Google Brain and X in June 2018 and published at CoRL 2018. At the time, most grasping systems first chose a grasp point and then executed it open-loop. QT-Opt instead uses deep reinforcement learning to learn a Q-function directly from an overhead RGB camera image, re-deciding how the gripper should move at every step to achieve closed-loop grasping; because it is hard to directly maximize over a continuous action space, it uses the cross-entropy method (an iterative sampling-based optimization) to search for the best action. Training used more than 580,000 real grasp attempts, reaching a 96% success rate on objects never seen before, and the policy spontaneously learned behaviors like re-grasping and nudging an object before picking it up. QT-Opt is a landmark result for large-scale real-robot reinforcement learning, and both MT-Opt and Q-Transformer build on it.

ExampleIf an object is pushed away by a person mid-grasp, the QT-Opt policy adjusts the gripper's position based on the new camera image and re-attempts the grasp, rather than closing on empty space as originally planned.

Also called
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Related
Q-Transformer · MT-Opt · Real-World Reinforcement Learning · Q-Function · Cross-Entropy Method · Google Arm Farm
Sources
QT-Opt (arXiv 1806.10293)
As of
2018-06

See it in the full glossary →