Embodiment-specific Head
本体专属头AdvancedIn a cross-embodiment model, a separate output layer per robot type that turns shared features into that robot's actions.
Different robots have different state and action dimensions: a single arm with a gripper might be 7 dimensions, while a bimanual robot with dexterous hands might have several dozen. Cross-embodiment models usually share most of their parameters, but attach a small network per embodiment at the input and output ends; the output-side piece is the embodiment-specific head, typically an MLP or a small decoder that maps the shared backbone's representation into an action with that embodiment's specific dimensions and meaning. MIT and Meta's Lirui Wang, Kaiming He, and colleagues proposed the Heterogeneous Pre-trained Transformer (HPT) in 2024, using an 'embodiment-specific stem + shared trunk + embodiment-specific head' design; NVIDIA's GR00T N1 similarly gives each embodiment an MLP to encode state and action, with a dedicated action decoder for output. A different approach is a unified action space, which pads every embodiment's action into the same large vector instead.
ExampleGR00T N1 attaches an embodiment-specific MLP action decoder after its last DiT block, letting the same backbone control everything from tabletop arms to humanoids with dexterous hands.
- Also called
- Embodiment-specific Decoder
- Related
- Cross-Embodiment · Action Head · Unified Action Space · Heterogeneous Pre-trained Transformers · NVIDIA Isaac GR00T N1 · Embodiment
- Sources
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (HPT, arXiv:2409.20537)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv HTML)