Model reference · open weights

granite-4.0-vision

LLMs ibm-granite Vision + text 1 build Open weights 12k dl/mo

granite-4.0-vision is an open-weight language model from IBM. granite-4.0-3b-vision (BF16) weighs 9.5 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

  • Granite-4.0-3B-Vision is a vision-language model developed by IBM for enterprise-grade document data extraction, including chart, table, and key-value pair tasks.
  • The model features a 4.0B parameter architecture with a 131,072 token context length and supports English.
  • It is distributed under the Apache 2.0 license.

Summary of the ibm-granite/granite-4.0-3b-vision model card, 2026-10-02

What it is

Released byIBM
Released2026-03-03
Parameters4.0B
VRAM9.5 GB for the weights

What it runs on

Memory and cards for granite-4.0-3b-vision (BF16)

9.5 GBweights, file size
82 MBcache per 1K tokens
1.9 GBruntime overhead, at least
131,072 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB——2K11.6 GB
RTX 4060 Ti 16 GB6148K15.4 GB
RTX 3090 24 GB174all 128K23.4 GB
RTX 4090 24 GB174all 128K23.4 GB
RTX 5090 32 GB297all 128K31.0 GB
L40S 48 GB4812all 128K44.0 GB
A100 80 GB9924all 128K78.2 GB
H100 80 GB9423all 128K78.1 GB
RTX PRO 6000 Blackwell 96 GB11729all 128K93.8 GB
DGX Spark (GB10) 128 GB unified13734all 128K107 GB
H200 141 GB18345all 128K138 GB
B200 180 GB23959all 128K176 GB
2× RTX 3060 12 GB
tensor parallel
143118K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
266all 128K15.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
112.1 GB14.1 GB
514.7 GB24.8 GB
816.8 GB32.9 GB
1622.1 GB54.3 GB
3232.9 GB97.3 GB
6454.3 GB183 GB

One card, with vLLM's small-card settings.

From the model card

What IBM says about granite-4.0-vision

Read the model card

Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:

  • Chart extraction: Converting charts into structured, machine-readable formats (Chart2CSV, Chart2Summary, and Chart2Code)
  • Table extraction: Accurately extracting tables with complex layouts from document images to JSON, HTML, or OTSL
  • Semantic Key-Value Pair (KVP) extraction: Extracting values based on key names and descriptions across diverse document layouts

The model is delivered as a LoRA adapter on top of Granite 4.0 Micro, with a 3.5B base LLM and 0.5B LoRA adapters. This enables a single deployment to support both multimodal document understanding and text-only workloads — the base model handles text-only requests without loading the adapter. See Model Architecture for details.

The methodology and data (ChartNet) used for this model are described in the paper ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding.

While our focus is on specialized document extraction tasks, the current model preserves and extends the capabilities of Granite-Vision-3.3 2B, ensuring that existing users can adopt it seamlessly with no changes to their workflow. It continues to support vision‑language tasks such as producing detailed natural‑language descriptions from images (image‑to‑text). The model can be used standalone and integrates seamlessly with Docling to enhance document processing pipelines with deep visual understanding capabilities.

  • Developer: IBM Research
  • GitHub Repository: https://github.com/ibm-granite
  • Release Date: March 27th, 2026
  • License: Apache 2.0

Supported Tasks

The model supports specialized extraction tasks, each activated by a simple task tag in the user message. The chat template automatically expands tags into the full prompt — no need to write verbose instructions.

TagTaskOutput
``Chart to CSVCSV table with headers and numeric values
``Chart to Python codePython code that recreates the chart
``Chart to summaryNatural-language description of the chart
``Table extraction (JSON)Structured JSON with dimensions and cells
``Table extraction (HTML)HTML `` markup
``Table extraction (OTSL)OTSL markup with cell/merge tags
KVP (see prompt instructions below)Schema based Key-Value pairs extractionJSON with nested dictionaries and arrays

Model Performance

Benchmark Results

We compare Granite-4.0-3B-Vision against leading small VLMs across multiple extraction benchmarks.

Chart Extraction

We evalute chart extraction using the human-verified test-set from ChartNet. The models are evaluated using LLM-as-a-judge (GPT4o) comparing the model prediction against the ground-truth. We report the average scores 0-100 on chart2csv and chart2summary extraction tasks.

Table Extraction

To benchmark table extraction, we construct a unified evaluation suite spanning multiple datasets and settings to assess end-to-end table extraction capabilities of vision-language models:

  1. TableVQA-Extract — Converts the original visual table QA benchmark into a cropped table extraction task.
  2. OmniDocBench-tables — A document parsing benchmark over diverse PDF types with detailed annotations for layout, text, formulas, and tables. We use the subset of pages that contain one or more tables to evaluate table extraction in full-page settings.
  3. PubTablesV2 — A large-scale table extraction benchmark evaluated in both cropped-table and full-page document settings.

To unify evaluation, we replace each dataset’s original annotations (e.g., Q&A pairs) with a single instruction: extract the table(s) from the image in HTML format, using the corresponding HTML as ground truth. For full-page inputs, only tabular elements are considered; when multiple tables appear, they are aggregated into a Python list.

We report results using TEDS (Tree-Edit Distance-based Similarity), which measures structural and content similarity between predicted and ground-truth HTML tables.

Results are presented separately for cropped-table and full-page settings to highlight performance across controlled and realistic document scenarios.

Key-Value Pair (KVP) Extraction

We evaluate on VAREX, a benchmark for multimodal structured extraction from documents. Granite-4.0-3B-Vision achieves 85.5% exact-match accuracy (zero-shot), ranking 3rd among 2–4B parameter models as of March 2026 (view results here).

Setup

Tested with python=3.11

pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
pip install transformers==4.57.6 peft==0.18.1 tokenizers==0.22.2 pillow==12.1.1

Usage with Transformers

import re
from io import StringIO

import pandas as pd
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from PIL import Image
from huggingface_hub import hf_hub_download

model_id = "ibm-granite/granite-4.0-3b-vision"
device = "cuda" if torch.cuda.is_available() else "cpu"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map=device
).eval()

# Optional: merge LoRA adapt

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms