Text Encoder
文本编码器CommonA network that turns a text instruction into a sequence of vectors for the rest of the model to use.
A text encoder splits a piece of text (such as 'put the cup in the sink') into tokens and turns them into vectors, used as a conditioning input for a policy or generative model. It's usually a pretrained language model, and it's often kept frozen when training a robot policy: Octo uses a roughly 110-million-parameter T5-base, RDT-1B uses a frozen T5-XXL, and CLIP and SigLIP each come with their own text encoder aligned to their image encoder; text-to-image and text-to-video models rely on one too, to understand the prompt. It brings pretrained language knowledge into the policy, helping the model tell different instructions apart. VLA models built on a large language model backbone, such as OpenVLA and π0, don't have a separate text encoder at all — text goes directly into the language model itself as tokens.
ExampleOcto encodes instructions into 16 language tokens using a frozen t5-base, fed into the Transformer alongside image tokens; the paper tried swapping in a larger T5 or fine-tuning it, and neither improved performance.
- Also called
- Language Encoder, Instruction Encoder
- Related
- Language-conditioned Policy · CLIP · SigLIP · Large Language Model · Embedding · Feature-wise Linear Modulation
- Sources
- Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)