Model reference · open weights
granite-4.0-vision is an open-weight language model from IBM. granite-4.0-3b-vision (BF16) weighs 9.5 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.
Summary of the ibm-granite/granite-4.0-3b-vision model card, 2026-10-02
What it is
| Released by | IBM |
|---|---|
| Released | 2026-03-03 |
| Parameters | 4.0B |
| VRAM | 9.5 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | — | — | 2K | 11.6 GB |
| RTX 4060 Ti 16 GB | 6 | 1 | 48K | 15.4 GB |
| RTX 3090 24 GB | 17 | 4 | all 128K | 23.4 GB |
| RTX 4090 24 GB | 17 | 4 | all 128K | 23.4 GB |
| RTX 5090 32 GB | 29 | 7 | all 128K | 31.0 GB |
| L40S 48 GB | 48 | 12 | all 128K | 44.0 GB |
| A100 80 GB | 99 | 24 | all 128K | 78.2 GB |
| H100 80 GB | 94 | 23 | all 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 117 | 29 | all 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 137 | 34 | all 128K | 107 GB |
| H200 141 GB | 183 | 45 | all 128K | 138 GB |
| B200 180 GB | 239 | 59 | all 128K | 176 GB |
| 2× RTX 3060 12 GB tensor parallel | 14 | 3 | 118K | 11.6 GB a card |
| 2× RTX 4060 Ti 16 GB tensor parallel | 26 | 6 | all 128K | 15.4 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 12.1 GB | 14.1 GB |
| 5 | 14.7 GB | 24.8 GB |
| 8 | 16.8 GB | 32.9 GB |
| 16 | 22.1 GB | 54.3 GB |
| 32 | 32.9 GB | 97.3 GB |
| 64 | 54.3 GB | 183 GB |
One card, with vLLM's small-card settings.
From the model card
Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:
The model is delivered as a LoRA adapter on top of Granite 4.0 Micro, with a 3.5B base LLM and 0.5B LoRA adapters. This enables a single deployment to support both multimodal document understanding and text-only workloads — the base model handles text-only requests without loading the adapter. See Model Architecture for details.
The methodology and data (ChartNet) used for this model are described in the paper ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding.
While our focus is on specialized document extraction tasks, the current model preserves and extends the capabilities of Granite-Vision-3.3 2B, ensuring that existing users can adopt it seamlessly with no changes to their workflow. It continues to support vision‑language tasks such as producing detailed natural‑language descriptions from images (image‑to‑text). The model can be used standalone and integrates seamlessly with Docling to enhance document processing pipelines with deep visual understanding capabilities.
The model supports specialized extraction tasks, each activated by a simple task tag in the user message. The chat template automatically expands tags into the full prompt — no need to write verbose instructions.
| Tag | Task | Output |
|---|---|---|
| `` | Chart to CSV | CSV table with headers and numeric values |
| `` | Chart to Python code | Python code that recreates the chart |
| `` | Chart to summary | Natural-language description of the chart |
| `` | Table extraction (JSON) | Structured JSON with dimensions and cells |
| `` | Table extraction (HTML) | HTML `` markup |
| `` | Table extraction (OTSL) | OTSL markup with cell/merge tags |
| KVP (see prompt instructions below) | Schema based Key-Value pairs extraction | JSON with nested dictionaries and arrays |
We compare Granite-4.0-3B-Vision against leading small VLMs across multiple extraction benchmarks.
We evalute chart extraction using the human-verified test-set from ChartNet. The models are evaluated using LLM-as-a-judge (GPT4o) comparing the model prediction against the ground-truth. We report the average scores 0-100 on chart2csv and chart2summary extraction tasks.
To benchmark table extraction, we construct a unified evaluation suite spanning multiple datasets and settings to assess end-to-end table extraction capabilities of vision-language models:
To unify evaluation, we replace each dataset’s original annotations (e.g., Q&A pairs) with a single instruction: extract the table(s) from the image in HTML format, using the corresponding HTML as ground truth. For full-page inputs, only tabular elements are considered; when multiple tables appear, they are aggregated into a Python list.
We report results using TEDS (Tree-Edit Distance-based Similarity), which measures structural and content similarity between predicted and ground-truth HTML tables.
Results are presented separately for cropped-table and full-page settings to highlight performance across controlled and realistic document scenarios.
We evaluate on VAREX, a benchmark for multimodal structured extraction from documents. Granite-4.0-3B-Vision achieves 85.5% exact-match accuracy (zero-shot), ranking 3rd among 2–4B parameter models as of March 2026 (view results here).
Tested with python=3.11
pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
pip install transformers==4.57.6 peft==0.18.1 tokenizers==0.22.2 pillow==12.1.1
import re
from io import StringIO
import pandas as pd
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from PIL import Image
from huggingface_hub import hf_hub_download
model_id = "ibm-granite/granite-4.0-3b-vision"
device = "cuda" if torch.cuda.is_available() else "cpu"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map=device
).eval()
# Optional: merge LoRA adaptQuoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.