Model reference · open weights

Mellum2-Thinking

LLMs JetBrains Text gen 1 build Open weights 2k dl/mo

Mellum2-Thinking is an open-weight language model from JetBrains. Mellum2-12B-A2.5B-Thinking (BF16) weighs 24.3 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.

What it is

Released byJetBrains
TypeLanguage models
TaskText gen
Parameters (lead)12.1B
Context131,072 tokens
Runs withtransformers
Released2026-05-26
Popularity2k downloads / month
Weights24.3 GB (Mellum2-12B-A2.5B-Thinking (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Mellum2-12B-A2.5B-Thinking (BF16)

Weights 24.3 GB (file size) · KV cache 14 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 220 MB a request for its sliding-window layers · runtime overhead from 1.2 GB on a small card · context up to 131,072 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … RTX 4090 24 GB
4 smaller cards
———
RTX 5090 32 GB168all 128K31.0 GB
L40S 48 GB5426all 128K44.0 GB
A100 80 GB15676all 128K78.2 GB
H100 80 GB10139all 128K78.1 GB
RTX PRO 6000 Blackwell 96 GB13451all 128K93.8 GB
DGX Spark (GB10) 128 GB unified16363all 128K107 GB
H200 141 GB22888all 128K138 GB
B200 180 GB30476all 128K176 GB
2× RTX 4060 Ti 16 GB
tensor parallel
126all 128K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
5929all 128K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
5929all 128K23.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
125.8 GB26.1 GB
527.1 GB28.9 GB
828.2 GB31.0 GB
1630.9 GB36.5 GB
3236.3 GB47.5 GB
6447.1 GB69.6 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What JetBrains says about Mellum2-Thinking

[!Note] Use this model when you want explicit chain-of-thought before the final answer — complex debugging, multi-step planning, agentic workflows, and math- or reasoning-heavy tasks. For direct, low-latency answers without reasoning traces, use Instruct instead.

Read the full model card

Mellum2 Thinking Highlights

Mellum 2 Thinking is a post-trained reasoning-augmented assistant model trained by JetBrains.

The model uses a Mixture-of-Experts architecture with 64 experts and activates 8 experts per token. It uses a combination of sliding-window and full attention layers, with a context length of 131,072 tokens.

It is produced from Mellum2-12B-A2.5B-Base by supervised fine-tuning (loss computed only on the final assistant turn) followed by reinforcement learning with verifiable rewards (RLVR) on a harder data mix that includes a long-form math subset. The model emits its reasoning inside ... blocks before the final answer.

Mellum2 Model Family

This repository contains one checkpoint from the Mellum 2 family.

CheckpointDescription
Base PretrainBase checkpoint before long-context extension
BaseFinal base model
Instruct SFTSupervised instruction-tuned checkpoint
Thinking SFTSupervised thinking checkpoint
InstructRL-tuned instruction model
ThinkingRL-tuned thinking model

Model Overview

Mellum2 Thinking has the following features:

  • Number of Layers: 28
  • Hidden Size: 2304
  • Intermediate Size: 7168
  • MoE Intermediate Size: 896
  • Number of Experts: 64
  • Number of Activated Experts: 8
  • Number of Attention Heads (GQA): 32 for Q and 4 for KV
  • Context Length: 131,072
  • Sliding Window: 1,024
  • Vocabulary Size: 98,304
  • Precision: bfloat16
  • License: Apache 2.0

Serving with vLLM

# Without tool calling
vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3

# With tool calling
vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

Quickstart

Text-Only Input

from openai import OpenAI
# Configured by environment variables
client = OpenAI()

messages = [
    {"role": "user", "content": "Is 1024 a power of 2? Explain your reasoning."},
]

chat_response = client.chat.completions.create(
    model="JetBrains/Mellum2-12B-A2.5B-Thinking",
    messages=messages,
    max_tokens=81920,
    temperature=0.6,
    top_p=0.95,
    extra_body={
        "top_k": 20,
    },
)
print("Chat response:", chat_response)

Evaluation

Post-training evaluation for the thinking/reasoning variants. All values are percentages; higher is better except HarmBench, where lower is better. All values self-reported by JetBrains.

BenchmarkMellum2 Thinking SFTMellum2 ThinkingQwen3.5 (4B)Qwen3.5 (9B)OLMo-3 (7B)Ministral 3 (14B)
Coding
LiveCodeBench v675.169.959.468.359.842.7
Tool Use
BFCL v438.845.642.942.7—35.9
BFCL v360.569.473.968.5—52.2
Math
AIME20.058.468.373.461.738.3
GSM-Plus62.687.089.390.788.186.5
Knowledge
MMLU-Redux84.886.288.391.771.384.4
GPQA Diamond39.957.676.881.329.346.0
Conversational
IFEval69.176.587.189.884.759.7
JetBrains pairwise64.469.540.556.732.263.8
MixEval63.466.971.976.067.070.8
BS-Bench14.015.063.070.023.09.0
Safety
HarmBench (↓)12.220.615.96.648.770.0
XSTest90.889.696.897.693.296.8

Notes:

  • AIME

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Text GenerationLiveCodeBench v6pass@169.900
Text GenerationBFCL v3accuracy69.400
Text GenerationBFCL v4 (macro-avg of 5 subtasks)accuracy45.600
Text GenerationAIME 2025+2026 (mean, 30 questions each)exact match58.400
Text GenerationGSM-Plusexact match87
Text GenerationMMLU-Reduxaccuracy86.200
Text GenerationGPQA Diamondaccuracy57.600
Text GenerationIFEval (prompt-level strict accuracy)accuracy76.500
Text GenerationMixEvalaccuracy66.900
Text GenerationBS-Bench (detection rate)detection rate15
Text GenerationHarmBench (harmful rate, lower is better)harmful rate20.600
Text GenerationXSTest (safe compliance)safe compliance89.600
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms