Embodied AI Glossary中文

PaLI-X

Advanced

A roughly 55-billion-parameter multilingual vision-language model from Google, one of the two backbones behind RT-2.

PaLI-X is a multilingual vision-language model (VLM) that Google Research released in May 2023, a scaled-up version of the earlier PaLI. Its vision encoder is ViT-22B, a 22-billion-parameter Vision Transformer; its language component is a 32-billion-parameter UL2 encoder-decoder; together they total roughly 55 billion parameters. The paper's main finding is that scaling up both the vision and language sides together keeps paying off, and training mixed prefix-completion and masked-token-completion objectives. After fine-tuning, PaLI-X set new state-of-the-art results on more than 15 benchmarks and showed emergent abilities it was never specifically trained for, such as complex object counting and object detection using non-English category names. In embodied AI, PaLI-X is best known as one of the two backbones behind RT-2: RT-2 fine-tuned PaLI-X (55B) and PaLM-E (12B) separately, training each on robot actions represented as text tokens, producing some of the earliest vision-language-action (VLA) models.

ExampleRT-2-PaLI-X-55B: PaLI-X was co-fine-tuned on web-scale image-text data together with robot trajectory data, so it could look at an image, read an instruction, and directly output discretized action tokens.

Also called
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Related
RT-2 · Vision-Language Model · PaLM-E · Vision Transformer · Vision-Language-Action Model · Encoder-Decoder
Sources
arXiv 2305.18565: PaLI-X
RT-2 项目主页 (Chinese)
As of
2023-07

See it in the full glossary →