Model reference · open weights
Mellum2 is an open-weight language model from JetBrains. Mellum2-12B-A2.5B-Base (BF16) weighs 24.3 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.
Mellum2 is a 12.1B parameter Mixture-of-Experts causal language model developed by JetBrains for text generation. It supports a context length of 131,072 tokens and is designed as a base checkpoint for fine-tuning, alignment, or domain adaptation. The model is trained on English text and released under the Apache 2.0 license.
Summary of the JetBrains/Mellum2-12B-A2.5B-Base model card, 2026-10-01
What it is
| Released by | JetBrains |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 12.1B |
| Context | 131,072 tokens |
| Runs with | transformers |
| Released | 2026-05-26 |
| Popularity | 16k downloads / month |
| Weights | 24.3 GB (Mellum2-12B-A2.5B-Base (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 24.3 GB (file size) · KV cache 14 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 220 MB a request for its sliding-window layers · runtime overhead from 1.2 GB on a small card · context up to 131,072 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | — | — | — | |
| RTX 5090 32 GB | 16 | 8 | all 128K | 31.0 GB |
| L40S 48 GB | 54 | 26 | all 128K | 44.0 GB |
| A100 80 GB | 156 | 76 | all 128K | 78.2 GB |
| H100 80 GB | 101 | 39 | all 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 134 | 51 | all 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 163 | 63 | all 128K | 107 GB |
| H200 141 GB | 228 | 88 | all 128K | 138 GB |
| B200 180 GB | 304 | 76 | all 128K | 176 GB |
| 2× RTX 4060 Ti 16 GB tensor parallel | 12 | 6 | all 128K | 15.4 GB a card |
| 2× RTX 4090 24 GB tensor parallel | 59 | 29 | all 128K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 59 | 29 | all 128K | 23.4 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 25.8 GB | 26.1 GB |
| 5 | 27.1 GB | 28.9 GB |
| 8 | 28.2 GB | 31.0 GB |
| 16 | 30.9 GB | 36.5 GB |
| 32 | 36.3 GB | 47.5 GB |
| 64 | 47.1 GB | 69.6 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
[!Note] Use this checkpoint as the starting point for your own fine-tuning, alignment, or domain adaptation on top of the long-context base. For instruction-following or reasoning tasks out of the box, use Instruct or Thinking instead.
Mellum2 Base is a long-context pretrained causal language model trained by JetBrains.
The model uses a Mixture-of-Experts architecture with 64 experts and activates 8 experts per token. It uses a combination of sliding-window and full attention layers, with a context length of 131,072 tokens.
This is the long-context base, produced from Mellum2-12B-A2.5B-Base-Pretrain by a layer-selective YaRN extension stage that re-maps RoPE frequencies on the global-attention layers only. It is the shared starting point for the released Instruct and Thinking variants.
This repository contains one checkpoint from the Mellum2 family.
| Checkpoint | Description |
|---|---|
| Base Pretrain | Base checkpoint before long-context extension |
| Base | Final base model |
| Instruct SFT | Supervised instruction-tuned checkpoint |
| Thinking SFT | Supervised thinking checkpoint |
| Instruct | RL-tuned instruction model |
| Thinking | RL-tuned thinking model |
Mellum2 Base has the following features:
vllm serve JetBrains/Mellum2-12B-A2.5B-Base --max-model-len 131072
Text-Only Input (base model — use the completions endpoint, not chat)
from openai import OpenAI
# Configured by environment variables
client = OpenAI()
completion = client.completions.create(
model="JetBrains/Mellum2-12B-A2.5B-Base",
prompt="def fibonacci(n):\n ",
max_tokens=81920,
temperature=0.6,
top_p=0.95,
extra_body={
"top_k": 20,
},
)
print("Completion:", completion)
Mellum2 Base pretraining results compared with similarly-sized open base models. All values are self-reported by JetBrains.
| Benchmark | Mellum2 (12B-A2.5B) | OLMo-3 (7B) | Qwen2.5 (7B) | Qwen3 (4B) | Qwen3.5 (4B) |
|---|---|---|---|---|---|
| Code Generation | |||||
| HumanEval | 41.5 | 45.1 | 55.5 | 57.3 | 50.0 |
| HumanEval+ | 37.2 | 39.6 | 47.0 | 51.2 | 43.9 |
| MBPP | 62.4 | 50.6 | 63.6 | 67.0 | 52.2 |
| MBPP+ | 61.4 | 52.9 | 64.0 | 64.5 | 55.0 |
| MultiPL-E (7 langs) | 21.0 | 10.0 | 19.2 | 26.0 | 12.1 |
| CRUXEval-I | 45.4 | 38.8 | 44.0 | 44.6 | 49.1 |
| CRUXEval-O | 43.9 | 36.6 | 42.9 | 43.5 | 43.2 |
| Knowledge & Reasoning | |||||
| MMLU | 70.9 | 62.1 | 71.8 | 71.1 | 74.2 |
| MMLU-Pro | 59.3 | 34.5 | 48.6 | 51.5 | 52.4 |
| BBH | 74.9 | 63.6 | 69.0 | 71.3 | 80.2 |
| ARC-Challenge | 53.5 | 53.6 | 51.3 | 51.2 | 54.9 |
| HellaSwag | 73.7 | 74.2 | 78.9 | 73.7 | 75.3 |
| WinoGrande | 65.5 | 69.5 | 73.3 | 71.2 | 70.8 |
| TruthfulQA MC2 | 44.5 | 47.0 | 56.4 | 53.5 | 52.1 |
| Math & Science | |||||
| GSM8K | 81.7 | 73.5 | 81.9 | 82.0 | 80.1 |
| MATH | 10.0 | 18.7 | 24.6 | 27.7 | 25.3 |
| GPQA Diamond | 31.3 | 28.8 | 32.8 | 36.9 | 41.4 |
| GPQA Main | 35.0 | 27.9 | 34.2 | 36.8 | 40.2 |
For more details, see the Mellum2 Technical Report.
Released under the Apache 2.0 license.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Text Generation | HumanEval | pass@1 | 41.460 |
| Text Generation | HumanEval+ | pass@1 | 37.200 |
| Text Generation | MBPP | pass@1 | 62.400 |
| Text Generation | MBPP+ | pass@1 | 78.310 |
| Text Generation | MultiPL-E HumanEval, 7 languages | pass@1 | 20.970 |
| Text Generation | CRUXEval-I | pass@1 | 45.380 |
| Text Generation | CRUXEval-O | pass@1 | 43.880 |
| Text Generation | MMLU | accuracy | 70.870 |
| Text Generation | MMLU-Pro | exact match | 59.310 |
| Text Generation | BBH | exact match | 74.900 |
| Text Generation | ARC-Challenge | normalized accuracy | 53.500 |
| Text Generation | HellaSwag | normalized accuracy | 73.720 |
| Text Generation | WinoGrande | accuracy | 65.510 |
| Text Generation | TruthfulQA MC2 | MC2 | 44.510 |
| Text Generation | GSM8K | exact match | 81.730 |
| Text Generation | MATH | exact match | 9.960 |
| Text Generation | GPQA Diamond | accuracy | 31.310 |
| Text Generation | GPQA Main | accuracy | 35.040 |