RT-Trajectory
AdvancedA policy that replaces language instructions with a trajectory sketch drawn on the image, showing the robot how the motion should go.
RT-Trajectory was released by Google DeepMind together with UC San Diego, Stanford, and Intrinsic in November 2023, selected as an ICLR 2024 Spotlight. A language instruction only says 'what to do,' which is often not specific enough for tasks the model never saw during training. RT-Trajectory instead conditions on a 'trajectory sketch': the path the end effector should follow is drawn as a curve overlaid on the camera image, with color encoding time and height, and marks showing where the gripper should open or close. No manual annotation is needed for training — the end-effector positions recorded in each demonstration are simply projected onto the image after the fact to generate the sketch (hindsight). The policy backbone reuses RT-1. At test time, sketches can be hand-drawn by a person, extracted from human videos, or generated by a large model. Across 7 tasks unseen during training, it clearly outperforms language-conditioned baselines like RT-1 and RT-2, as well as goal-image-conditioned baselines.
ExampleTraining data is mostly pick-and-place; at test time, a person draws a curve on the image that first grabs one corner of a cloth and then pulls it toward the opposite side, and the robot follows it to perform 'fold the cloth,' a task absent from training.
- Also called
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
- Related
- RT-1 · RT-2 · Intermediate Representation · Task Generalization · Hindsight Relabeling · Goal-conditioned Policy
- Sources
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches (arXiv 2311.01977)
RT-Trajectory project page - As of
- 2024-01