Model reference · open weights

Intern-Decision

NEW · this week LLMs internlm Vision + text 3 builds Open weights 0 dl/mo

Intern-Decision is an open-weight language model from internlm. Intern-Decision-2B (BF16) weighs 4.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byinternlm
TypeLanguage models
TaskVision + text
Parameters (lead)4.5B
Context262,144 tokens
Runs withtransformers
Based onQwen/Qwen3.5-4B
Released2026-09-26
Popularity0 downloads / month
Weights4.4 GB (Intern-Decision-2B (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Intern-Decision-2B (BF16)

Weights 4.4 GB (file size) · KV cache 12 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 39 MB of fixed state per request · runtime overhead from 2.1 GB on a small card · context up to 262,144 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB3611all 256K11.6 GB
RTX 4060 Ti 16 GB6320all 256K15.4 GB
RTX 3090 24 GB12138all 256K23.4 GB
RTX 4090 24 GB12038all 256K23.4 GB
RTX 5090 32 GB17555all 256K31.0 GB
L40S 48 GB26884all 256K44.0 GB
A100 80 GB513162all 256K78.2 GB
H100 80 GB475150all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB587186all 256K93.8 GB
DGX Spark (GB10) 128 GB unified684216all 256K107 GB
H200 141 GB903285all 256K138 GB
B200 180 GB1000+371all 256K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
16.6 GB6.9 GB
57.2 GB8.7 GB
87.6 GB10.0 GB
168.7 GB13.6 GB
3211.0 GB20.6 GB
6415.4 GB34.8 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

Builds

Sizes, precisions & builds

BuildParametersPrecisionWeightsSmallest configuration (1 request, 8K)
Intern-Decision-4B ↗ 4.5BBF16 9.1 GBRTX 3060 12 GB
Intern-Decision-2B (above) ↗ 2.2BBF16 4.4 GBRTX 3060 12 GB
Intern-Decision-0.8B ↗ 853MBF16 1.7 GBRTX 3060 12 GB

Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.

From the model card

What internlm says about Intern-Decision

Demo | Model Weights | GitHub

Intern-Decision-4B is a multimodal structured decision model fine-tuned from Qwen3.5-4B. It accepts a shared state, a schema of named questions, and optional images, and returns an answer distribution for every question in one model forward pass.

Read the full model card

How inference works

  1. Preserve the question and option order, and map each question's options to single-token symbols A, B, …, Z, a, …, z, 0, …, 9.
  2. Render the original system prompt, state, decision schema, and a complete assistant JSON skeleton with one `` placeholder per field. Preserve the checkpoint's chat template and empty thinking block.
  3. Run one causal Hugging Face forward pass. For the masked-next-token decision objective, read logits at the position immediately before each placeholder.
  4. Take a softmax over only that field's allowed candidate-symbol logits, then apply the checkpoint's probability calibration.
  5. Map symbols back to the original option values and return typed JSON answers.

This API performs structured candidate scoring. It does not call generate() or sample free-form text. A request can contain multiple fields; no gold answers are inserted into the prompt. The inference compiler uses only state, questions, and optional images.

Benchmark results

ModelJevbench-EasyJevbench-OriginalJevbench-HardTyped DecisionToolACEAG NewsWildJailBreakAverageBrier ↓ECE ↓
Jev100.0098.6172.0773.3591.2989.5796.2988.740.3580.095
Laya95.8372.2228.8335.9563.8792.8414.8457.770.8040.246
SemIf100.0098.6161.2662.8085.1689.2292.5384.230.4980.112
Kev100.0093.0645.0565.6087.4289.8275.9779.560.7380.262
JevK5100.0097.2273.8764.5080.9789.1390.4585.160.3660.047
Intern-Decision-0.8B97.9280.5652.2577.3594.5288.6164.4879.380.5300.066
Intern-Decision-2B100.0084.7263.9679.3596.4589.9678.3384.680.4370.100
Intern-Decision-4B100.0098.6173.8780.5596.4590.8289.8690.020.3470.065

Inference latency

Measured on a single RTX 4090 with the local HF inference path. Values are per-query end-to-end latency; they are workload and hardware dependent.

ModelMeanMedian / P50P95
Jev109.70 ms106.30 ms146.70 ms
Intern-Decision-0.8B33.98 ms33.44 ms37.50 ms
Intern-Decision-2B33.28 ms33.15 ms33.55 ms
Intern-Decision-4B44.16 ms44.03 ms44.60 ms

Known-distribution calibration pilot

This separate 96-case diagnostic uses exact reference distributions rather than sampled hard labels. Lower is better. The pilot was not used to fit or select the published temperature; the 4B model used its separately fitted T=1.992418.

CategoryIntern-Decision-4B beforeIntern-Decision-4B afterJev
Direct randomness and support0.483 / 0.1810.421 / 0.1290.490 / 0.216
Composed events and mixtures0.677 / 0.2540.577 / 0.1500.682 / 0.274
History, conditioning, and hidden state0.711 / 0.2190.613 / 0.1080.657 / 0.113
Daily evidence and observation bias0.701 / 0.3280.575 / 0.2100.603 / 0.114
Selective disclosure and probability puzzles0.540 / 0.1190.510 / 0.0490.483 / 0.138
Sequential and combinatorial processes0.656 / 0.1800.605 / 0.0580.657 / 0.116
Overall (Brier / ECE)0.628 / 0.2130.550 / 0.0890.595 / 0.130

Quick start

Use Python 3.12+. Install requirements.txt in a suitable PyTorch/CUDA environment, then import DecisionEngine from the downloaded model directory:

pip install -r requirements.txt
from inference import DecisionEngine

engine = DecisionEngine(device="cuda")  # Load once; reuse for subsequent requests.
request = {
    "state": "The customer was charged twice and asks for the extra payment back.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "billing": "Payments and refunds",
                "delivery": "Shipping and delivery",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "Rate the priority.",
            "criteria": ["Low", "Medium", "High"],
        },
        "refund_requested": {
            "type": "noul",
            "instructions": "Is the customer asking for a refund?",
        },
    },
}
response = engine.predict(request)  # One Python dict in, one response dict out.
print(response["answers"])

predict(request) accepts one request dictionary per call and returns a JSON-serializable Jev-compatible response. It does not read request files or mutate the supplied dictionary. Reuse the engine for each subsequent request.

The engine defaults to the checkpoint next to inference.py. To load another local copy of this same model, use DecisionEngine(checkpoint="./model-copy"). Use the inference module shipped with the selected size so its default calibration matches. backend="hf" is the default and the only implemented backend. The optional request model field does not switch checkpoints; the response model identifies the weights actually loaded by this module.

Request format

{
  "state": "The customer was charged twice and asks for the extra payment back.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle 

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms