Model reference · open weights

webssl-dino300m-full2b-224

Embeddings facebook Image embed 1 build Non-commercial 9k dl/mo

webssl-dino300m-full2b-224 is an open-weight embedding model from Meta. webssl-dino300m-full2b-224 (FP32) weighs 607 MB; the smallest configuration that runs it is RTX 3060 12 GB.

webssl-dino300m-full2b-224 is a 304M parameter Vision Transformer developed by Meta for image feature extraction. It was trained using self-supervised learning on 2 billion web images without language supervision. The model is released under the cc-by-nc-4.0 license.

Summary of the facebook/webssl-dino300m-full2b-224 model card, 2026-10-01

What it is

Released byMeta
Published underfacebook
TypeEmbedding models
TaskImage embed
Parameters (lead)304M
Runs withtransformers
Released2025-04-22
Popularity9k downloads / month
Weights607 MB (webssl-dino300m-full2b-224 (FP32), file size)
LicenceNon-commercial

What it runs on

Memory and cards for webssl-dino300m-full2b-224 (FP32)

Weights 607 MB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Meta says about webssl-dino300m-full2b-224

Read the model card

A 300 million parameter Vision Transformer (ViT) trained with DINOv2 self-supervised learning on web-scale image data without language supervision. Introduced in "Scaling Language-Free Visual Representation Learning" (Fan et al., 2025).

Model Details

  • Architecture: ViT (1536 width, 40 depth, 24 heads)
  • Parameters: 300M
  • Resolution: 224×224 pixels
  • Training: Self-supervised Web-DINO on 2B image samples from MetaCLIP web data

Model Descriptions

Web-SSL DINO 300M is a 300 million parameter Vision Transformer model trained using self-supervised learning on 2 billion web images without language supervision. This model demonstrates that pure visual learning, when scaled appropriately, can match or exceed the performance of language-supervised models like CLIP across various vision tasks.

Usage

from transformers import AutoImageProcessor, Dinov2Model
import torch
from PIL import Image

processor = AutoImageProcessor.from_pretrained('facebook/webssl-dino300m-full2b-224')
model = Dinov2Model.from_pretrained('facebook/webssl-dino300m-full2b-224')

# Process an image
image = Image.open('path/to/image.jpg')
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

cls_features = outputs.last_hidden_state[:, 0]  # CLS token features
patch_features = outputs.last_hidden_state[:, 1:] # patch-wise token features

Citation

@article{fan2025scaling,
  title={Scaling Language-Free Visual Representation Learning},
  author={David Fan and Shengbang Tong and Jiachen Zhu and Koustuv Sinha and Zhuang Liu and Xinlei Chen and Michael Rabbat and Nicolas Ballas and Yann LeCun and Amir Bar and Saining Xie},
  year={2025},
  eprint={2504.01017},
  archivePrefix={arXiv},
  primaryClass={cs.CV}
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms