Model reference · open weights

granite-vision-3.3-embedding

Available as managed deployment Embeddings ibm-granite Embeddings 1 variants 642 dl/mo

granite-vision-3.3-embedding is an open-weight embedding model from ibm-granite. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byIBM
Published underibm-granite
TypeEmbedding models
TaskEmbeddings
Parameters (lead)3.0B
Context128k tokens
Runs withtransformers
Based onibm-granite/granite-vision-3.3-2b
Released2025-06-03
Popularity642 downloads / month
LicenceOpen weights

About

What granite-vision-3.3-embedding is

Model Summary: Granite-vision-3.3-2b-embedding is an efficient embedding model based on granite-vision-3.3-2b. This model is specifically designed for multimodal document retrieval, enabling queries on documents with tables, charts, infographics, and complex layouts. The model generates ColBERT-style multi-vector representations of pages. By removing the need for OCR-based text extractions, granite-vision-3.3-2b-embedding can help simplify and accelerate RAG pipelines.

Evaluations: We evaluated granite-vision-3.3-2b-embedding alongside other top colBERT style multi-modal embedding models in the 1B-4B parameter range using two benchmark: Vidore2 and Real-MM-RAG-Bench which aim to specifically address complex multimodal document retrieval tasks.

Read the full model card

NDCG@5 - ViDoRe V2

Collection \ ModelColPali-v1.3ColQwen2.5-v0.2ColNomic-3bColSmolvlm-v0.1granite-vision-3.3-2b-embedding
ESG Restaurant Human51.168.465.862.465.3
Economics Macro Multilingual49.956.555.447.451.2
MIT Biomedical59.763.663.558.161.5
ESG Restaurant Synthetic57.057.456.651.156.6
ESG Restaurant Synthetic Multilingual55.757.457.247.655.7
MIT Biomedical Multilingual56.561.162.550.555.5
Economics Macro51.659.860.260.958.3
Avg (ViDoRe2)54.560.660.254.057.7

NDCG@5 - REAL-MM-RAG

Collection \ ModelColPali-v1.3ColQwen2.5-v0.2ColNomic-3bColSmolvlm-v0.1granite-vision-3.3-2b-embedding
FinReport5566786573
FinSlides6879815579
TechReport7886888387
TechSlides9093929193
Avg (REAL-MM-RAG)7381857483
  • Release Date: June 11th 2025
  • License: Apache 2.0
  • Supported Input Format: Currently the model supports English instructions and images (png, jpeg) as input format.

Intended Use: The model is intended to be used in enterprise applications that involve retrieval of visual and text data. In particular, the model is well-suited for multi-modal RAG systems where the knowledge base is composed of complex enterprise documents, such as reports, slides, images, canned doscuments, manuals and more. The model can be used as a standalone retriever, or alongside a text-based retriever.

Usage

pip install -q torch torchvision torchaudio
pip install transformers==4.50

Then run the code:

from io import BytesIO

import requests
import torch
from PIL import Image
from transformers import AutoProcessor, AutoModel
from transformers.utils.import_utils import is_flash_attn_2_available

device = "cuda" if torch.cuda.is_available() else "cpu"
model_name = "ibm-granite/granite-vision-3.3-2b-embedding"
model = AutoModel.from_pretrained(
                      model_name,
                      trust_remote_code=True,
                      torch_dtype=torch.float16,
                      device_map=device,
                      attn_implementation="flash_attention_2" if is_flash_attn_2_available() else None
                      ).eval()
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)

# ─────────────────────────────────────────────
# Inputs: Image + Text
# ─────────────────────────────────────────────
image_url = "https://huggingface.co/datasets/mishig/sample_images/resolve/main/tiger.jpg"
print("\nFetching image...")
image = Image.open(BytesIO(requests.get(image_url).content)).convert("RGB")

text = "A photo of a tiger"
print(f"Image and text inputs ready.")

# Process both inputs
print("Processing inputs...")
image_inputs = processor.process_images([image])
text_inputs = processor.process_queries([text])

# Move to correct device
image_inputs = {k: v.to(device) for k, v in image_inputs.items()}
text_inputs = {k: v.to(device) for k, v in text_inputs.items()}

# ─────────────────────────────────────────────
# Run Inference
# ─────────────────────────────────────────────
with torch.no_grad():
    print("🔍 Getting image embedding...")
    img_emb = model(**image_inputs)

    print("✍️ Getting text embedding...")
    txt_emb = model(**text_inputs)

# ─────────────────────────────────────────────
# Score the similarity
# ─────────────────────────────────────────────
print("Scoring similarity...")
similarity = processor.score(txt_emb, img_emb, batch_size=1, device=device)

print("\n" + "=" * 50)
print(f"📊 Similarity between 

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys granite-vision-3-3-embedding for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (granite-vision-3-3-embedding below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"granite-vision-3-3-embedding","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms