Humanoid Locomotion as Next Token Prediction
人形行走即下一个 token 预测AdvancedTreats a humanoid robot's walking control as a language-model-style “predict the next token” problem, learned with a causal Transformer.
This was released in February 2024 by Ilija Radosavovic, Koushil Sreenath, Jitendra Malik, and colleagues at UC Berkeley. It frames real humanoid robot control as a problem similar to language modeling: a causal Transformer autoregressively predicts a sensorimotor trajectory made of observations and actions, where each token predicts the next token of the same modality. This lets data lacking action labels — such as action trajectories extracted from human video — also participate in training. The data comes from simulated trajectories generated by existing neural-network policies and model-based controllers, human motion-capture data, and YouTube videos of humans. The model was deployed on Agility Robotics' full-size humanoid Digit, walking zero-shot on the streets of San Francisco; using just 27 hours of walking data is also enough to transfer to the real robot, and it generalizes to a backward-walking command that was never in the training data.
ExampleThere is no backward-walking command in the training data, yet once deployed, the model can still make Digit walk backward on command.
- Related
- Next-Token Prediction · Autoregressive Decoding · Action-free Video · Bipedal Locomotion · Agility Robotics Digit · Real-World Humanoid Locomotion with Reinforcement Learning
- Sources
- Humanoid Locomotion as Next Token Prediction (arXiv 2402.19469)
Project page - As of
- 2024-02