Model reference · open weights

tipsv2-b14

Embeddings google Image embed 1 build Open weights 190k dl/mo

tipsv2-b14 is an open-weight embedding model from Google. tipsv2-b14 (FP32) weighs 392 MB; the smallest configuration that runs it is RTX 3060 12 GB.

tipsv2-b14 is a contrastive vision-language model developed by Google for zero-shot image classification. It features a ViT vision encoder and a Transformer text encoder, totaling 196M parameters with a 768-dimensional embedding space. The model uses a 14x14 pixel patch size and supports a context length of 64 tokens, and it is released under the Apache 2.0 license.

Summary of the google/tipsv2-b14 model card, 2026-10-01

What it is

Released byGoogle
TypeEmbedding models
TaskImage embed
Parameters (lead)196M
Context64 tokens
Runs withtransformers
Released2026-04-09
Popularity190k downloads / month
Weights392 MB (tipsv2-b14 (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for tipsv2-b14 (FP32)

Weights 392 MB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Google says about tipsv2-b14

Read the model card

TIPSv2 (Text-Image Pre-training with Spatial awareness) is a family of contrastive vision-language models that produce spatially rich image features aligned with text embeddings. This is the Base variant with 86M vision params and 110M text params. Try the code snippets below or check out the GitHub repo for more use cases and visualizations, including zero-shot segmentation.

VariantVision paramsText paramsEmbed dimDPT Heads
B/1486M110M768B/14-dpt
L/14303M184M1024L/14-dpt
SO400m/14412M448M1152SO400m/14-dpt
g/141.1B389M1536g/14-dpt

Usage

pip install transformers torch torchvision sentencepiece scikit-learn

Load the model

from transformers import AutoModel

model = AutoModel.from_pretrained("google/tipsv2-b14", trust_remote_code=True)
model.eval()

Encode images

Images should be tensors in [0, 1] range (just ToTensor(), no ImageNet normalization).

from torchvision import transforms
from PIL import Image
import requests

transform = transforms.Compose([
    transforms.Resize((448, 448)),
    transforms.ToTensor(),
])

url = "https://huggingface.co/spaces/google/TIPSv2/resolve/main/examples/zeroseg/pascal_context_00049_image.png"
image = Image.open(requests.get(url, stream=True).raw)
pixel_values = transform(image).unsqueeze(0)
out = model.encode_image(pixel_values)

print(out.cls_token.shape)     # (1, 1, 768) — global image embedding
print(out.patch_tokens.shape)  # (1, 1024, 768) — per-patch spatial features

Encode text

text_emb = model.encode_text(["a photo of a bus", "a photo of a dog"])
print(text_emb.shape)  # (2, 768) — one embedding per query

Zero-shot classification

import torch.nn.functional as F

classes = ["bus", "car", "dog", "cat"]
cls = F.normalize(out.cls_token[:, 0, :], dim=-1)
text_emb = F.normalize(model.encode_text(classes), dim=-1)
similarity = cls @ text_emb.T
print(classes[similarity.argmax()])  # bus — predicted class

Visualize spatial features

import numpy as np
from sklearn.decomposition import PCA

spatial = out.patch_tokens.reshape(1, 32, 32, 768)
feat = spatial[0].detach().cpu().numpy().reshape(-1, 768)
rgb = PCA(n_components=3, whiten=True).fit_transform(feat).reshape(32, 32, 3)
rgb = 1 / (1 + np.exp(-2.0 * rgb))  # sigmoid for [0, 1] range with good contrast
print(rgb.shape)  # (32, 32, 3) — PCA of patch features as RGB

GPU inference

model = model.cuda()
out = model.encode_image(pixel_values.cuda())
text_emb = model.encode_text(["a city"])

Model details

  • Architecture: ViT vision encoder (12 layers) + Transformer text encoder (12 layers)
  • Image preprocessing: resize to any resolution, convert to [0, 1] (no ImageNet normalization)
  • Text preprocessing: SentencePiece tokenizer, lowercased, max 64 tokens
  • Patch size: 14x14 pixels

License

Apache 2.0

Citation

@inproceedings{cao2026tipsv2,
  title     = {{TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment}},
  author    = {Cao, Bingyi and Chen, Koert and Maninis, Kevis-Kokitsi and Chen, Kaifeng and Karpur, Arjun and Xia, Ye and Dua, Sahil and Dabral, Tanmaya and Han, Guangxing and Han, Bohyung and Ainslie, Joshua and Bewley, Alex and Jacob, Mithun and Wagner, Rene and Ramos, Washington and Choromanski, Krzysztof and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andre},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2604.12012}
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms