Model reference · open weights
EuroMoE-2512 is an open-weight language model from utter-project. EuroMoE-2.6B-A0.6B-Instruct-2512 (FP32) weighs 5.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | utter-project |
|---|---|
| Type | Language models |
| Task | Text gen · MoE |
| Parameters (lead) | 2.6B |
| Context | 32,768 tokens |
| Runs with | transformers |
| Based on | utter-project/EuroMoE-2.6B-A0.6B-2512 |
| Released | 2025-12-15 |
| Popularity | 539 downloads / month |
| Weights | 5.2 GB (EuroMoE-2.6B-A0.6B-Instruct-2512 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 5.2 GB (file size) · KV cache 25 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 582 MB on a small card · context up to 32,768 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 28 | 7 | all 32K | 11.6 GB |
| RTX 4060 Ti 16 GB | 47 | 11 | all 32K | 15.4 GB |
| RTX 3090 24 GB | 87 | 21 | all 32K | 23.4 GB |
| RTX 4090 24 GB | 87 | 21 | all 32K | 23.4 GB |
| RTX 5090 32 GB | 125 | 31 | all 32K | 31.0 GB |
| L40S 48 GB | 189 | 47 | all 32K | 44.0 GB |
| A100 80 GB | 359 | 89 | all 32K | 78.2 GB |
| H100 80 GB | 341 | 85 | all 32K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 419 | 104 | all 32K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 486 | 121 | all 32K | 107 GB |
| H200 141 GB | 638 | 159 | all 32K | 138 GB |
| B200 180 GB | 827 | 206 | all 32K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 6.0 GB | 6.6 GB |
| 5 | 6.8 GB | 9.8 GB |
| 8 | 7.4 GB | 12.2 GB |
| 16 | 9.0 GB | 18.7 GB |
| 32 | 12.2 GB | 31.6 GB |
| 64 | 18.7 GB | 57.3 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
From the model card
This is the model card for EuroMoE-2.6B-A0.6B-Instruct-2512. You can also check the pre-trained version: EuroMoE-2.6B-A0.6B-2512.
The EuroLLM project has the goal of creating a suite of LLMs capable of understanding and generating text in all European Union languages as well as some additional relevant languages. EuroMoE-2.6B-A0.6B is a 2.6B parameter model trained on 8 trillion tokens divided across the considered languages and several data sources: Web data, parallel data (en-xx and xx-en), and high-quality datasets. EuroMoE-2.6B-A0.6B-Instruct was further instruction tuned on EuroBlocks, an instruction tuning dataset with focus on general instruction-following and machine translation.
EuroMoE uses a standard MoE Transformer architecture:
For pre-training, we use 512 Nvidia A100 GPUs of the Leonardo supercomputer, training the model with a constant batch size of 4096 sequences, which corresponds to approximately 17 million tokens, using the Adam optimizer, and BF16 precision. Here is a summary of the model hyper-parameters:
| Sequence Length | 4,096 |
| Number of Layers | 24 |
| Embedding Size | 1,024 |
| Total/Active experts | 64/8 |
| Expert Hidden Size | 512 |
| Number of Heads | 8 |
| Number of KV Heads (GQA) | 2 |
| Activation Function | SwiGLU |
| Position Encodings | RoPE (\Theta=500,000) |
| Layer Norm | RMSNorm |
| Tied Embeddings | Yes |
| Embedding Parameters | 0.13B |
| LM Head Parameters | 0.13B |
| Active Non-embedding Parameters | 0.34B |
| Total Non-embedding Parameters | 2.35B |
| Active Parameters | 0.6B |
| Total Parameters | 2.61B |
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "utter-project/EuroMoE-2.6B-A0.6B-Instruct-2512"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
messages = [
{
"role": "system",
"content": "You are EuroLLM --- an AI assistant specialized in European languages that provides safe, educational and helpful answers.",
},
{
"role": "user", "content": "What is the capital of Portugal? How would you describe it?"
},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
outputs = model.generate(inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
EuroMoE-2.6B-A0.6B-Instruct-2512 has not been aligned to human preferences, so the model may generate problematic outputs (e.g., hallucinations, harmful content, or false statements).
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.