Model reference · open weights

tipsv1-s14

Embeddings google Image embed 1 build Open weights 279 dl/mo

tipsv1-s14 is an open-weight embedding model from Google. tipsv1-s14 (FP32) weighs 111 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byGoogle
TypeEmbedding models
TaskImage embed
Parameters (lead)56M
Runs withtransformers
Released2026-08-19
Popularity279 downloads / month
Weights111 MB (tipsv1-s14 (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for tipsv1-s14 (FP32)

Weights 111 MB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Google says about tipsv1-s14

TIPS (Text-Image Pre-training with Spatial awareness, ICLR 2025) is a family of contrastive vision-language models that produce spatially rich image features aligned with text embeddings. This is the original (v1) S/14 release with 22M vision params and 34M text params, converted from the official checkpoints.

Read the full model card
VariantVision paramsText paramsEmbed dimResolution
S/1422M34M384448
B/1486M110M768448
L/14304M184M1024448
So400m/14413M448M1152448
g/141.1B389M1536448
g/14 low-res1.1B389M1536224

Usage

pip install transformers torch torchvision sentencepiece scikit-learn requests

Load the model

from transformers import AutoModel

model = AutoModel.from_pretrained("google/tipsv1-s14", trust_remote_code=True)
model.eval()

Encode images

Images should be tensors in [0, 1] range (just ToTensor(), no ImageNet normalization).

import requests
from PIL import Image
from torchvision import transforms

url = "https://huggingface.co/spaces/google/TIPSv2/resolve/main/examples/zeroseg/pascal_context_00049_image.png"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")
transform = transforms.Compose([transforms.Resize((448, 448)), transforms.ToTensor()])
pixel_values = transform(image).unsqueeze(0)

out = model.encode_image(pixel_values)
print(out.cls_token.shape)     # (1, 1, 384) — global image embedding
print(out.patch_tokens.shape)  # (1, 1024, 384) — per-patch spatial features

The second CLS token (out.register_tokens) was trained on synthetic captions; the first (out.cls_token) on web alt-text, and is the one aligned with the text tower.

Encode text

text_emb = model.encode_text(["a photo of a bus", "a photo of a dog"])
print(text_emb.shape)  # (2, 384) — one embedding per query

Zero-shot classification

import torch.nn.functional as F

classes = ["bus", "car", "dog", "cat"]
cls = F.normalize(out.cls_token[:, 0, :], dim=-1)
text_emb = F.normalize(model.encode_text(classes), dim=-1)
similarity = cls @ text_emb.T
print(classes[similarity.argmax()])  # predicted class

Visualize spatial features

import numpy as np
from sklearn.decomposition import PCA

feat = out.patch_tokens[0].detach().cpu().numpy()
rgb = PCA(n_components=3, whiten=True).fit_transform(feat).reshape(32, 32, 3)
rgb = 1 / (1 + np.exp(-2.0 * rgb))  # sigmoid for [0, 1] range with good contrast

Model details

  • ViT-S/14 vision encoder (12 layers, patch size 14, two CLS tokens) + 12-layer transformer text encoder
  • Native resolution 448; other patch-multiple resolutions work via positional-embedding interpolation
  • Preprocessing: images to [0, 1], no normalization; SentencePiece tokenizer, lowercased, max 64 tokens

License

Apache 2.0

Citation

@inproceedings{maninis2025tips,
  title     = {{TIPS: Text-Image Pretraining with Spatial Awareness}},
  author    = {Maninis, Kevis-Kokitsi and Chen, Kaifeng and Ghosh, Soham and Karpur, Arjun and Chen, Koert and Xia, Ye and Cao, Bingyi and Salz, Daniel and Han, Guangxing and Dlabal, Jan and Gnanapragasam, Dan and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andre},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2025},
  url       = {https://arxiv.org/abs/2410.16512}
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms