Behavior Transformer
BeT / VQ-BeTBeTAdvancedA Transformer-based imitation-learning method that learns several different valid behaviors at once from multimodal demonstration data.
BeT (Behavior Transformer) is a behavior-cloning method proposed in 2022 by Lerrel Pinto's group at NYU. Human demonstrations often show “action multimodality”: several equally valid ways to act in the same situation, and direct regression averages them into one wrong action. BeT first clusters continuous actions into a number of bins with k-means, has a Transformer predict which bin to pick, and then predicts a continuous offset to correct it into a precise action — preserving multiple behavior modes this way. VQ-BeT, from 2024 (NYU and Seoul National University, ICML 2024), replaces k-means with residual vector quantization, which suits high-dimensional actions and long action sequences better, running inference at roughly 5x the speed of a diffusion policy. BeT and diffusion policies are two representative lines of work from the same period addressing the same multimodality problem.
ExampleIn the Franka Kitchen simulated kitchen, demonstrators complete a set of subtasks in different orders each time, and BeT is used to test whether a model can learn and reproduce these different behavior patterns.
- Also called
- BeT, VQ-BeT, Vector-Quantized Behavior Transformer
- Related
- Action Multimodality · Behavior Cloning · Vector Quantization · Diffusion Policy · Action Tokenizer · Franka Kitchen
- Sources
- Behavior Transformers: Cloning k modes with one stone (arXiv 2206.11251)
Behavior Generation with Latent Actions (VQ-BeT, arXiv 2403.03181)
VQ-BeT 项目主页 (Chinese) - As of
- 2024-07