Model reference · open weights
siglip2-so400m-patch16-naflex is an open-weight embedding model from Google. siglip2-so400m-patch16-naflex (FP32) weighs 2.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.
siglip2-so400m-patch16-naflex is a 1.1B parameter vision-language model developed by Google. It is designed for zero-shot image classification, image-text retrieval, and use as a vision encoder for vision-language models. The model is pre-trained on the WebLI dataset and is available under the apache-2.0 licence.
Summary of the google/siglip2-so400m-patch16-naflex model card, 2026-10-01
What it is
| Released by | |
|---|---|
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 1.1B |
| Runs with | transformers |
| Released | 2025-02-18 |
| Popularity | 395k downloads / month |
| Weights | 2.3 GB (siglip2-so400m-patch16-naflex (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 2.3 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
SigLIP 2 extends the pretraining objective of SigLIP with prior, independently developed techniques into a unified recipe, for improved semantic understanding, localization, and dense features.
You can use the raw model for tasks like zero-shot image classification and image-text retrieval, or as a vision encoder for VLMs (and other vision tasks).
Here is how to use this model to perform zero-shot image classification:
from transformers import pipeline
# load pipeline
ckpt = "google/siglip2-so400m-patch16-naflex"
image_classifier = pipeline(model=ckpt, task="zero-shot-image-classification")
# load image and candidate labels
url = "http://images.cocodataset.org/val2017/000000039769.jpg"
candidate_labels = ["2 cats", "a plane", "a remote"]
# run inference
outputs = image_classifier(image, candidate_labels)
print(outputs)
You can encode an image using the Vision Tower like so:
import torch
from transformers import AutoModel, AutoProcessor
from transformers.image_utils import load_image
# load the model and processor
ckpt = "google/siglip2-so400m-patch16-naflex"
model = AutoModel.from_pretrained(ckpt, device_map="auto").eval()
processor = AutoProcessor.from_pretrained(ckpt)
# load the image
image = load_image("https://huggingface.co/datasets/merve/coco/resolve/main/val2017/000000000285.jpg")
inputs = processor(images=[image], return_tensors="pt").to(model.device)
# run infernece
with torch.no_grad():
image_embeddings = model.get_image_features(**inputs)
print(image_embeddings.shape)
For more code examples, we refer to the siglip2 documentation.
SigLIP 2 adds some clever training objectives on top of SigLIP:
SigLIP 2 is pre-trained on the WebLI dataset (Chen et al., 2023).
The model was trained on up to 2048 TPU-v5e chips.
Evaluation of SigLIP 2 is shown below (taken from the paper).
@misc{tschannen2025siglip2multilingualvisionlanguage,
title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
author={Michael Tschannen and Alexey Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier Hénaff and Jeremiah Harmsen and Andreas Steiner and Xiaohua Zhai},
year={2025},
eprint={2502.14786},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2502.14786},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.