Action / State Normalization
动作与状态归一化AdvancedScaling each action and state dimension to a common range before training, then converting model output back to real units at inference.
A robot's actions and proprioceptive state (joint angles, gripper opening, end-effector position, and so on) span very different units and scales across dimensions, and feeding them into a network unscaled lets the largest-magnitude dimensions dominate the loss and destabilize training. A common fix is to compute, from the training data, each dimension's mean and standard deviation, or its 1st and 99th percentiles (quantile normalization), and rescale the data to zero mean and unit variance, or to the range [-1, 1]. These statistics are called norm stats and must be saved alongside the model weights, since inference needs to de-normalize the output back into real commands; percentiles are more robust to outliers than a raw min/max. Physical Intelligence's FAST paper, for instance, normalizes actions using the 1st/99th percentiles; openpi requires running compute_norm_stats.py before fine-tuning, though a new task on a robot already present in the pretraining data can reuse the pretrained statistics. Mismatched normalization statistics are a common cause of erratic real-robot behavior.
Exampleπ0-FAST maps each action dimension's 1st and 99th training-set percentiles onto [-1, 1] before tokenizing actions, so data from different robots can share the same tokenizer.
- Also called
- Norm Stats, Quantile Normalization
- Related
- Action Tokenizer · π0-FAST · Proprioception · Cross-Embodiment Data · openpi (Physical Intelligence) · Fine-tuning
- Sources
- FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)
openpi(Physical Intelligence GitHub)