Embodied AI Glossary中文

Action Binning

分箱离散化Common

Dividing each continuous action dimension into equal bins and using the bin number as a discrete token.

Action binning is the simplest way to turn a continuous robot action into a discrete symbol: for each action dimension, such as end-effector x-displacement or gripper open/close, define a value range, divide it into equal bins, and use the bin a value falls into to represent it. RT-2 splits each dimension into 256 bins and maps each bin number onto a token in the language model's vocabulary, so actions can be predicted the same way text is. OpenVLA keeps 256 bins too, but uses the 1st and 99th percentile of the training data instead of the true minimum and maximum to set the range, so a few outliers don't stretch the bins and hurt precision. The downside is that every timestep and every dimension needs its own token, and at high control frequency, neighboring tokens are highly correlated, giving the model a lot of redundant, hard-to-learn tokens, which is exactly what compressive action tokenizers such as FAST were built to fix.

ExampleIf x-axis displacement ranges from −2 to 2 centimeters and is split into 256 bins, each bin is about 0.016 centimeters wide, so 0.5 centimeters would fall into bin number 160.

Also called
Action Discretization, Bucketing
Related
Action Tokenizer · Action Representation · Discrete Cosine Transform · RT-2 · OpenVLA · π0-FAST
Sources
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv:2307.15818)
FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)

See it in the full glossary →