Keyframe Action Prediction
关键帧动作预测AdvancedPredicting only a few key end-effector poses for a task, and leaving the path between them to a motion planner.
A standard visuomotor policy outputs dozens of continuous actions per second; a keyframe method instead compresses a demonstration down to a handful of key end-effector poses (such as before grasping, at the moment of gripping, before releasing), and the policy only learns 'where is the next key pose,' with a motion planner generating the path between two poses. Imperial College's Stephen James and colleagues introduced keyframe discovery in 2021's Q-attention and C2F-ARM, and 2022's PerAct follows the same approach: simple rules, such as joint velocity being near zero, automatically pick out keyframes, typically leaving just 2 to 17 per RLBench demonstration; position is then discretized into voxels and rotation into angle bins, turning prediction into a classification problem. PerAct's ablation shows that choosing frames randomly or at even intervals drops performance to zero. This setup learns fast from few examples, and is also used by 3D manipulation policies like RVT and 3D Diffuser Actor; the cost is dependence on a planner, which struggles with dynamic tasks that need continuous adjustment of force and speed.
ExampleTo have an arm open a drawer, PerAct only has to predict a few key poses in sequence: a pre-grasp pose in front of the handle, gripping the handle, and pulling out to the target position, with a motion planner filling in the path between each pair.
- Also called
- Next-Best-Pose Prediction, Next Best Action, Keypose Prediction
- Related
- PerAct · RVT-2 · 3D Diffuser Actor · Motion Planning · RLBench · Action Representation
- Sources
- Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation (arXiv:2209.05451)
Q-attention: Enabling Efficient Learning for Vision-based Robotic Manipulation (arXiv:2105.14829)
Coarse-to-Fine Q-attention (C2F-ARM, arXiv:2106.12534) - As of
- 2022-11