Model reference · open weights
SigLIP2-giant is an open-weight embedding model from AvitoTech. SigLIP2-giant (FP32) weighs 3.7 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the AvitoTech/SigLIP2-giant model card, 2026-10-03
What it is
| Released by | AvitoTech |
|---|---|
| Released | 2025-12-02 |
| Parameters | 1.9B |
| VRAM | 3.7 GB for the weights |
What it runs on
| Card | Runs | Memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
From the model card
Fine-tuned SigLIP2-Giant model for individual animal identification, specializing in distinguishing between unique cats and dogs. This model produces robust image embeddings optimized for pet recognition, re-identification, and verification tasks.
The model was trained on a comprehensive dataset combining multiple sources:
Total Dataset Statistics:
Training Configuration:
Loss Function: The model is trained using a combined loss function consisting of:
Combined as: L_total = 1.0 × L_triplet + 0.5 × L_var
This approach creates compact feature clusters for each individual animal while maintaining large separation between different identities.
The model has been benchmarked against various vision encoders on multiple pet recognition datasets:
| Model | ROC AUC | EER | Top-1 | Top-5 | Top-10 |
|---|---|---|---|---|---|
| CLIP-ViT-Base | 0.9821 | 0.0604 | 0.8359 | 0.9579 | 0.9711 |
| DINOv2-Small | 0.9904 | 0.0422 | 0.8547 | 0.9660 | 0.9764 |
| SigLIP-Base | 0.9899 | 0.0390 | 0.8649 | 0.9757 | 0.9842 |
| SigLIP2-Base | 0.9894 | 0.0388 | 0.8660 | 0.9772 | 0.9863 |
| Zer0int CLIP-L | 0.9881 | 0.0509 | 0.8768 | 0.9767 | 0.9845 |
| SigLIP2-Giant | 0.9940 | 0.0344 | 0.8899 | 0.9868 | 0.9921 |
| SigLIP2-Giant + E5-Small-v2 + gating | 0.9929 | 0.0344 | 0.8952 | 0.9872 | 0.9932 |
| Model | ROC AUC | EER | Top-1 | Top-5 | Top-10 |
|---|---|---|---|---|---|
| CLIP-ViT-Base | 0.9739 | 0.0772 | 0.4350 | 0.6417 | 0.7204 |
| DINOv2-Small | 0.9829 | 0.0571 | 0.5581 | 0.7540 | 0.8139 |
| SigLIP-Base | 0.9792 | 0.0606 | 0.5848 | 0.7746 | 0.8319 |
| SigLIP2-Base | 0.9776 | 0.0672 | 0.5925 | 0.7856 | 0.8422 |
| Zer0int CLIP-L | 0.9814 | 0.0625 | 0.6289 | 0.8092 | 0.8597 |
| SigLIP2-Giant | 0.9926 | 0.0326 | 0.7475 | 0.9009 | 0.9316 |
| SigLIP2-Giant + E5-Small-v2 + gating | 0.9920 | 0.0314 | 0.7818 | 0.9233 | 0.9482 |
| Model | ROC AUC | EER | Top-1 | Top-5 | Top-10 |
|---|---|---|---|---|---|
| CLIP-ViT-Base | 0.9752 | 0.0729 | 0.6511 | 0.8122 | 0.8555 |
| DINOv2-Small | 0.9848 | 0.0546 | 0.7180 | 0.8678 | 0.9009 |
| SigLIP-Base | 0.9811 | 0.0572 | 0.7359 | 0.8831 | 0.9140 |
| SigLIP2-Base | 0.9793 | 0.0631 | 0.7400 | 0.8889 | 0.9197 |
| Zer0int CLIP-L | 0.9842 | 0.0565 | 0.7626 | 0.8994 | 0.9267 |
| SigLIP2-Giant | 0.9912 | 0.0378 | 0.8243 | 0.9471 | 0.9641 |
| SigLIP2-Giant + E5-Small-v2 + gating | 0.9882 | 0.0422 | 0.8428 | 0.9576 | 0.9722 |
Metrics Explanation:
pip install transformers torch pillow
import torch
import torch.nn.functional as F
from PIL import Image
from transformers import AutoImageProcessor, AutoModel
repo = "AvitoTech/SigLIP2-giant"
processor = AutoImageProcessor.from_pretrained(repo)
model = AutoModel.from_pretrained(repo).eval()
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
image = Image.open("your_image.jpg").convert("RGB")
with torch.no_grad():
inputs = processor(images=[image], return_tensors="pt").to(device)
embedding = model.get_image_features(**inputs)
embedding = F.normalize(embedding, dim=1)
print(f"Embedding shape: {embedding.shape}") # torch.Size([1, 1536])
If you use this model in your research or applications, please cite our work:
@Article{jimaging12010030,
AUTHOR = {Kudryavtsev, Vasiliy and Borodin, Kirill and Berezin, German and Bubenchikov, Kirill and Mkrtchian, Grach and Ryzhkov, Alexander},
TITLE = {From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification},
JOURNAL = {Journal of Imaging},
VOLUME = {12},
YEAR = {2026},
NUMBER = {1},
ARTICLE-NUMBER = {30},
URL = {https://www.mdpi.com/2313-433X/12/1/30},
ISSN = {2313-433X},
ABSTRACT = {Automated animal identification is a practical task for reuniting lost pets with their owners, yet current systems often struggle due to limited dataset scale and reliance on unimodal visual cues. This study introduces a multimodal verification framework that enhances visual features with semantic identity priors derived from synthetic textual descriptions. We constructed a massive training corpus of 1.9 million photographs covering 695,091 unique animals to support this investigation. Through systematic ablation stQuoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.