Model reference · open weights
webssl-dino3b-light2b-224 is an open-weight embedding model from Meta. webssl-dino3b-light2b-224 (FP32) weighs 5.9 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | Meta |
|---|---|
| Published under | |
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 2.9B |
| Runs with | transformers |
| Released | 2025-04-15 |
| Popularity | 737 downloads / month |
| Weights | 5.9 GB (webssl-dino3b-light2b-224 (FP32), file size) |
| Licence | Non-commercial |
What it runs on
Weights 5.9 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
A 3 billion parameter Vision Transformer (ViT) trained with DINOv2 self-supervised learning on lightly filtered web-scale image data without language supervision. Introduced in "Scaling Language-Free Visual Representation Learning" (Fan et al., 2025).
Web-SSL DINO 3B is a 3 billion parameter Vision Transformer model trained using self-supervised learning on lightly filtered web images without language supervision. The "light2b" designation indicates training on a subset of images containing any textual content, retaining approximately 50.3% of the original MetaCLIP dataset. This filtering improves OCR & Chart understanding capabilities while maintaining strong performance across all vision tasks. This model demonstrates that pure visual learning, when scaled appropriately, can match or exceed the performance of language-supervised models like CLIP across various vision tasks.
from transformers import AutoImageProcessor, Dinov2Model
import torch
from PIL import Image
processor = AutoImageProcessor.from_pretrained('facebook/webssl-dino3b-light2b-224')
model = Dinov2Model.from_pretrained('facebook/webssl-dino3b-light2b-224')
# Process an image
image = Image.open('path/to/image.jpg')
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
cls_features = outputs.last_hidden_state[:, 0] # CLS token features
patch_features = outputs.last_hidden_state[:, 1:] # patch-wise token features
@article{fan2025scaling,
title={Scaling Language-Free Visual Representation Learning},
author={David Fan and Shengbang Tong and Jiachen Zhu and Koustuv Sinha and Zhuang Liu and Xinlei Chen and Michael Rabbat and Nicolas Ballas and Yann LeCun and Amir Bar and Saining Xie},
year={2025},
eprint={2504.01017},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.