SigLIP
CommonGoogle's image-text model trained with a sigmoid loss instead of softmax; its image encoder is used by many VLAs.
SigLIP is an image-text pretraining method proposed in 2023 by Xiaohua Zhai and colleagues at Google (ICCV 2023). Like CLIP, it trains an image encoder and a text encoder so that matched image-text pairs end up close together as vectors; the difference is the loss function. CLIP uses a softmax normalized across the whole batch, while SigLIP treats each image-text pair independently as a binary 'do they match or not' classification problem with a sigmoid loss, which doesn't depend on batch-wide normalization, so it trains well even with small batches and uses less GPU memory. Its vision encoder (such as the roughly 400-million-parameter So400m) is the image input stage for many VLMs and VLAs. SigLIP 2, released in February 2025, added objectives like captioning and self-distillation during training, improving multilingual ability, localization, and dense features.
Exampleπ0's backbone, PaliGemma, is made of a SigLIP-So400m vision encoder and a Gemma-2B language model; OpenVLA instead concatenates SigLIP features together with DINOv2 features.
- Also called
- Sigmoid Loss for Language-Image Pre-training, SigLIP 2, SigLIP-So400m
- Related
- CLIP · Contrastive Learning · Vision Encoder · PaliGemma · DINOv2 · Softmax
- Sources
- Sigmoid Loss for Language Image Pre-Training (arXiv 2303.15343)
SigLIP 2: Multilingual Vision-Language Encoders (arXiv 2502.14786)
PaliGemma: A versatile 3B VLM for transfer (arXiv 2407.07726) - As of
- 2025-02