Model reference · open weights
tipsv2-so400m14 is an open-weight embedding model from Google. tipsv2-so400m14 (FP32) weighs 1.7 GB; the smallest configuration that runs it is RTX 3060 12 GB.
tipsv2-so400m14 is a zero-shot image classification model developed by Google that uses contrastive vision-language pretraining to align spatially rich image features with text embeddings. The model contains 862M parameters, consisting of 412M vision parameters and 448M text parameters, and produces 1152-dimensional embeddings. It is distributed under the Apache 2.0 license.
Summary of the google/tipsv2-so400m14 model card, 2026-10-01
What it is
| Released by | |
|---|---|
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 862M |
| Runs with | transformers |
| Released | 2026-04-09 |
| Popularity | 267k downloads / month |
| Weights | 1.7 GB (tipsv2-so400m14 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 1.7 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
TIPSv2 (Text-Image Pre-training with Spatial awareness) is a family of contrastive vision-language models that produce spatially rich image features aligned with text embeddings. This is the SO400m variant with 412M vision params and 448M text params. Try the code snippets below or check out the GitHub repo for more use cases and visualizations, including zero-shot segmentation.
| Variant | Vision params | Text params | Embed dim | DPT Heads |
|---|---|---|---|---|
| B/14 | 86M | 110M | 768 | B/14-dpt |
| L/14 | 303M | 184M | 1024 | L/14-dpt |
| SO400m/14 | 412M | 448M | 1152 | SO400m/14-dpt |
| g/14 | 1.1B | 389M | 1536 | g/14-dpt |
pip install transformers torch torchvision sentencepiece scikit-learn
from transformers import AutoModel
model = AutoModel.from_pretrained("google/tipsv2-so400m14", trust_remote_code=True)
model.eval()
Images should be tensors in [0, 1] range (just ToTensor(), no ImageNet normalization).
from torchvision import transforms
from PIL import Image
import requests
transform = transforms.Compose([
transforms.Resize((448, 448)),
transforms.ToTensor(),
])
url = "https://huggingface.co/spaces/google/TIPSv2/resolve/main/examples/zeroseg/pascal_context_00049_image.png"
image = Image.open(requests.get(url, stream=True).raw)
pixel_values = transform(image).unsqueeze(0)
out = model.encode_image(pixel_values)
print(out.cls_token.shape) # (1, 1, 1152) — global image embedding
print(out.patch_tokens.shape) # (1, 1024, 1152) — per-patch spatial features
text_emb = model.encode_text(["a photo of a bus", "a photo of a dog"])
print(text_emb.shape) # (2, 1152) — one embedding per query
import torch.nn.functional as F
classes = ["bus", "car", "dog", "cat"]
cls = F.normalize(out.cls_token[:, 0, :], dim=-1)
text_emb = F.normalize(model.encode_text(classes), dim=-1)
similarity = cls @ text_emb.T
print(classes[similarity.argmax()]) # bus — predicted class
import numpy as np
from sklearn.decomposition import PCA
spatial = out.patch_tokens.reshape(1, 32, 32, 1152)
feat = spatial[0].detach().cpu().numpy().reshape(-1, 1152)
rgb = PCA(n_components=3, whiten=True).fit_transform(feat).reshape(32, 32, 3)
rgb = 1 / (1 + np.exp(-2.0 * rgb)) # sigmoid for [0, 1] range with good contrast
print(rgb.shape) # (32, 32, 3) — PCA of patch features as RGB
model = model.cuda()
out = model.encode_image(pixel_values.cuda())
text_emb = model.encode_text(["a city"])
[0, 1] (no ImageNet normalization)Apache 2.0
@inproceedings{cao2026tipsv2,
title = {{TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment}},
author = {Cao, Bingyi and Chen, Koert and Maninis, Kevis-Kokitsi and Chen, Kaifeng and Karpur, Arjun and Xia, Ye and Dua, Sahil and Dabral, Tanmaya and Han, Guangxing and Han, Bohyung and Ainslie, Joshua and Bewley, Alex and Jacob, Mithun and Wagner, Rene and Ramos, Washington and Choromanski, Krzysztof and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andre},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026},
url = {https://arxiv.org/abs/2604.12012}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.