State / Proprioception Encoder
状态编码器CommonA small network that turns a robot's own state — joint angles, gripper opening — into a vector the model can use.
A state encoder processes a robot's proprioception — information about its own state that doesn't need a camera, such as joint angles, end-effector pose, and gripper opening. These are numeric vectors of a few to a few dozen dimensions, and they first need to be mapped into the model's embedding dimension before they can be fed into a Transformer alongside image and text tokens. The implementation is usually simple: π0 and ACT each use a single linear layer, while GR00T N1 gives each embodiment its own MLP to handle robots with different state dimensions. State input lets the policy know where the arm currently is, but it has a side effect too: the Octo team found that adding proprioceptive state often hurt performance, likely because the model over-relies on the strong correlation between state and action — a form of causal confusion.
ExampleThe bimanual ALOHA robot's state is 14 joint positions across both arms; ACT projects this 14-dimensional vector to 512 dimensions with a single linear layer and feeds it as one token into the Transformer encoder alongside the image features.
- Also called
- Proprioception Encoder, State Encoder, State Projector
- Related
- Proprioception · Projector / Connector · Embodiment-specific Head · Action / State Normalization · Causal Confusion (Causal Misidentification) · History Encoder
- Sources
- π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)