Model reference · open weights

tips-s14

tips-s14 is an open-weight embedding model from google, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Embeddings google 1 variants 279 downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What tips-s14 is

TIPS — S/14 (v1) TIPS (Text-Image Pre-training with Spatial awareness, ICLR 2025) is a family of contrastive vision-language models that produce spatially rich image features aligned with text embeddings. This is the original (v1) S/14 release with 22M vision params and 34M text params, converted from the official checkpoints. Usage Load the model Encode images Images should be tensors in [0, 1] range (just ToTensor(), no ImageNet normalization). The second CLS token (out.registertokens) was trained on synthetic captions; the first (out.clstoken) on web alt-text, and is the one aligned with the text tower. Encode text Zero-shot classification Visualize spatial features Model details - ViT-S/14 vision encoder (12 layers, patch size 14, two CLS tokens) + 12-layer transformer text encoder - Native resolution 448; other patch-multiple resolutions work via positional-embedding interpolation - Preprocessing: images to [0, 1], no normalization; SentencePiece tokenizer, lowercased, max 64 tokens License Apache 2.0 Citation

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makergoogle
TypeEmbedding models
Parameters (lead)56M
Variants1
Runs withtransformers
Released2026-08-19
Popularity279 downloads / month
Likes1
LicenceOpen weights

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
tipsv1-s1456MBF16~0.1 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys tips-s14 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (tips-s14 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"tips-s14","input":"text to embed"}'

Details

Languages, data & research

Tags

transformers safetensors tipsv2 feature-extraction vision image-text contrastive-learning zero-shot zero-shot-image-classification custom_code

Papers

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want tips-s14 on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More embedding models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms