Model reference · open weights
clip-ViT-B-16 is an open-weight embedding model from sentence-transformers. clip-ViT-B-16 (BF16) weighs 599 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | sentence-transformers |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Runs with | sentence-transformers |
| Released | 2022-04-12 |
| Popularity | 0 downloads / month |
| Weights | 599 MB (clip-ViT-B-16 (BF16), file size) |
| Licence | Licence not stated |
What it runs on
Weights 599 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This is the Image & Text model CLIP, which maps text and images to a shared vector space. For applications of the models, have a look in our documentation SBERT.net - Image Search
After installing sentence-transformers (pip install sentence-transformers), the usage of this model is easy:
from sentence_transformers import SentenceTransformer, util
from PIL import Image
#Load CLIP model
model = SentenceTransformer('clip-ViT-B-16')
#Encode an image:
img_emb = model.encode(Image.open('two_dogs_in_snow.jpg'))
#Encode text descriptions
text_emb = model.encode(['Two dogs in the snow', 'A cat on a table', 'A picture of London at night'])
#Compute cosine similarities
cos_scores = util.cos_sim(img_emb, text_emb)
print(cos_scores)
See our SBERT.net - Image Search documentation for more examples how the model can be used for image search, zero-shot image classification, image clustering and image deduplication.
In the following table we find the zero-shot ImageNet validation set accuracy:
| Model | Top 1 Performance |
|---|---|
| clip-ViT-B-32 | 63.3 |
| clip-ViT-B-16 | 68.1 |
| clip-ViT-L-14 | 75.4 |
For a multilingual version of the CLIP model for 50+ languages have a look at: clip-ViT-B-32-multilingual-v1
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works