Model reference · open weights
Intern-Decision is an open-weight language model from internlm. Intern-Decision-2B (BF16) weighs 4.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | internlm |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 4.5B |
| Context | 262,144 tokens |
| Runs with | transformers |
| Based on | Qwen/Qwen3.5-4B |
| Released | 2026-09-26 |
| Popularity | 0 downloads / month |
| Weights | 4.4 GB (Intern-Decision-2B (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 4.4 GB (file size) · KV cache 12 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 39 MB of fixed state per request · runtime overhead from 2.1 GB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 36 | 11 | all 256K | 11.6 GB |
| RTX 4060 Ti 16 GB | 63 | 20 | all 256K | 15.4 GB |
| RTX 3090 24 GB | 121 | 38 | all 256K | 23.4 GB |
| RTX 4090 24 GB | 120 | 38 | all 256K | 23.4 GB |
| RTX 5090 32 GB | 175 | 55 | all 256K | 31.0 GB |
| L40S 48 GB | 268 | 84 | all 256K | 44.0 GB |
| A100 80 GB | 513 | 162 | all 256K | 78.2 GB |
| H100 80 GB | 475 | 150 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 587 | 186 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 684 | 216 | all 256K | 107 GB |
| H200 141 GB | 903 | 285 | all 256K | 138 GB |
| B200 180 GB | 1000+ | 371 | all 256K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 6.6 GB | 6.9 GB |
| 5 | 7.2 GB | 8.7 GB |
| 8 | 7.6 GB | 10.0 GB |
| 16 | 8.7 GB | 13.6 GB |
| 32 | 11.0 GB | 20.6 GB |
| 64 | 15.4 GB | 34.8 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
Builds
| Build | Parameters | Precision | Weights | Smallest configuration (1 request, 8K) |
|---|---|---|---|---|
| Intern-Decision-4B ↗ | 4.5B | BF16 | 9.1 GB | RTX 3060 12 GB |
| Intern-Decision-2B (above) ↗ | 2.2B | BF16 | 4.4 GB | RTX 3060 12 GB |
| Intern-Decision-0.8B ↗ | 853M | BF16 | 1.7 GB | RTX 3060 12 GB |
Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.
From the model card
Demo | Model Weights | GitHub
Intern-Decision-4B is a multimodal structured decision model fine-tuned from Qwen3.5-4B. It accepts a shared state, a schema of named questions, and optional images, and returns an answer distribution for every question in one model forward pass.
A, B, …, Z, a, …, z, 0, …, 9.This API performs structured candidate scoring. It does not call generate() or
sample free-form text. A request can contain multiple fields; no gold answers are
inserted into the prompt. The inference compiler uses only state, questions,
and optional images.
| Model | Jevbench-Easy | Jevbench-Original | Jevbench-Hard | Typed Decision | ToolACE | AG News | WildJailBreak | Average | Brier ↓ | ECE ↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| Jev | 100.00 | 98.61 | 72.07 | 73.35 | 91.29 | 89.57 | 96.29 | 88.74 | 0.358 | 0.095 |
| Laya | 95.83 | 72.22 | 28.83 | 35.95 | 63.87 | 92.84 | 14.84 | 57.77 | 0.804 | 0.246 |
| SemIf | 100.00 | 98.61 | 61.26 | 62.80 | 85.16 | 89.22 | 92.53 | 84.23 | 0.498 | 0.112 |
| Kev | 100.00 | 93.06 | 45.05 | 65.60 | 87.42 | 89.82 | 75.97 | 79.56 | 0.738 | 0.262 |
| JevK5 | 100.00 | 97.22 | 73.87 | 64.50 | 80.97 | 89.13 | 90.45 | 85.16 | 0.366 | 0.047 |
| Intern-Decision-0.8B | 97.92 | 80.56 | 52.25 | 77.35 | 94.52 | 88.61 | 64.48 | 79.38 | 0.530 | 0.066 |
| Intern-Decision-2B | 100.00 | 84.72 | 63.96 | 79.35 | 96.45 | 89.96 | 78.33 | 84.68 | 0.437 | 0.100 |
| Intern-Decision-4B | 100.00 | 98.61 | 73.87 | 80.55 | 96.45 | 90.82 | 89.86 | 90.02 | 0.347 | 0.065 |
Measured on a single RTX 4090 with the local HF inference path. Values are per-query end-to-end latency; they are workload and hardware dependent.
| Model | Mean | Median / P50 | P95 |
|---|---|---|---|
| Jev | 109.70 ms | 106.30 ms | 146.70 ms |
| Intern-Decision-0.8B | 33.98 ms | 33.44 ms | 37.50 ms |
| Intern-Decision-2B | 33.28 ms | 33.15 ms | 33.55 ms |
| Intern-Decision-4B | 44.16 ms | 44.03 ms | 44.60 ms |
This separate 96-case diagnostic uses exact reference distributions rather than sampled hard labels. Lower is better. The pilot was not used to fit or select the published temperature; the 4B model used its separately fitted T=1.992418.
| Category | Intern-Decision-4B before | Intern-Decision-4B after | Jev |
|---|---|---|---|
| Direct randomness and support | 0.483 / 0.181 | 0.421 / 0.129 | 0.490 / 0.216 |
| Composed events and mixtures | 0.677 / 0.254 | 0.577 / 0.150 | 0.682 / 0.274 |
| History, conditioning, and hidden state | 0.711 / 0.219 | 0.613 / 0.108 | 0.657 / 0.113 |
| Daily evidence and observation bias | 0.701 / 0.328 | 0.575 / 0.210 | 0.603 / 0.114 |
| Selective disclosure and probability puzzles | 0.540 / 0.119 | 0.510 / 0.049 | 0.483 / 0.138 |
| Sequential and combinatorial processes | 0.656 / 0.180 | 0.605 / 0.058 | 0.657 / 0.116 |
| Overall (Brier / ECE) | 0.628 / 0.213 | 0.550 / 0.089 | 0.595 / 0.130 |
Use Python 3.12+. Install requirements.txt in a suitable PyTorch/CUDA
environment, then import DecisionEngine from the downloaded model directory:
pip install -r requirements.txt
from inference import DecisionEngine
engine = DecisionEngine(device="cuda") # Load once; reuse for subsequent requests.
request = {
"state": "The customer was charged twice and asks for the extra payment back.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments and refunds",
"delivery": "Shipping and delivery",
},
},
"urgency": {
"type": "score",
"instructions": "Rate the priority.",
"criteria": ["Low", "Medium", "High"],
},
"refund_requested": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
},
}
response = engine.predict(request) # One Python dict in, one response dict out.
print(response["answers"])
predict(request) accepts one request dictionary per call and returns a
JSON-serializable Jev-compatible response. It does not read request files or mutate
the supplied dictionary. Reuse the engine for each subsequent request.
The engine defaults to the checkpoint next to inference.py. To load another
local copy of this same model, use DecisionEngine(checkpoint="./model-copy").
Use the inference module shipped with the selected size so its default calibration
matches. backend="hf" is the default and the only implemented backend. The
optional request model field does not switch checkpoints; the response model
identifies the weights actually loaded by this module.
{
"state": "The customer was charged twice and asks for the extra payment back.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.