Model reference · open weights
tipsv1-g14-lowres is an open-weight embedding model from Google. tipsv1-g14-lowres (FP32) weighs 3.0 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | |
|---|---|
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 1.5B |
| Runs with | transformers |
| Released | 2026-08-19 |
| Popularity | 335 downloads / month |
| Weights | 3.0 GB (tipsv1-g14-lowres (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 3.0 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
TIPS (Text-Image Pre-training with Spatial awareness, ICLR 2025) is a family of contrastive vision-language models that produce spatially rich image features aligned with text embeddings. This is the original (v1) g/14 low-res release with 1.1B vision params and 389M text params, converted from the official checkpoints.
| Variant | Vision params | Text params | Embed dim | Resolution |
|---|---|---|---|---|
| S/14 | 22M | 34M | 384 | 448 |
| B/14 | 86M | 110M | 768 | 448 |
| L/14 | 304M | 184M | 1024 | 448 |
| So400m/14 | 413M | 448M | 1152 | 448 |
| g/14 | 1.1B | 389M | 1536 | 448 |
| g/14 low-res | 1.1B | 389M | 1536 | 224 |
pip install transformers torch torchvision sentencepiece scikit-learn requests
from transformers import AutoModel
model = AutoModel.from_pretrained("google/tipsv1-g14-lowres", trust_remote_code=True)
model.eval()
Images should be tensors in [0, 1] range (just ToTensor(), no ImageNet normalization).
import requests
from PIL import Image
from torchvision import transforms
url = "https://huggingface.co/spaces/google/TIPSv2/resolve/main/examples/zeroseg/pascal_context_00049_image.png"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
transform = transforms.Compose([transforms.Resize((224, 224)), transforms.ToTensor()])
pixel_values = transform(image).unsqueeze(0)
out = model.encode_image(pixel_values)
print(out.cls_token.shape) # (1, 1, 1536) — global image embedding
print(out.patch_tokens.shape) # (1, 256, 1536) — per-patch spatial features
The second CLS token (out.register_tokens) was trained on synthetic captions; the first (out.cls_token) on web alt-text, and is the one aligned with the text tower.
text_emb = model.encode_text(["a photo of a bus", "a photo of a dog"])
print(text_emb.shape) # (2, 1536) — one embedding per query
import torch.nn.functional as F
classes = ["bus", "car", "dog", "cat"]
cls = F.normalize(out.cls_token[:, 0, :], dim=-1)
text_emb = F.normalize(model.encode_text(classes), dim=-1)
similarity = cls @ text_emb.T
print(classes[similarity.argmax()]) # predicted class
import numpy as np
from sklearn.decomposition import PCA
feat = out.patch_tokens[0].detach().cpu().numpy()
rgb = PCA(n_components=3, whiten=True).fit_transform(feat).reshape(16, 16, 3)
rgb = 1 / (1 + np.exp(-2.0 * rgb)) # sigmoid for [0, 1] range with good contrast
[0, 1], no normalization; SentencePiece tokenizer, lowercased, max 64 tokensApache 2.0
@inproceedings{maninis2025tips,
title = {{TIPS: Text-Image Pretraining with Spatial Awareness}},
author = {Maninis, Kevis-Kokitsi and Chen, Kaifeng and Ghosh, Soham and Karpur, Arjun and Chen, Koert and Xia, Ye and Cao, Bingyi and Salz, Daniel and Han, Guangxing and Dlabal, Jan and Gnanapragasam, Dan and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andre},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2025},
url = {https://arxiv.org/abs/2410.16512}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.