Model reference · open weights
EuroMoE is an open-weight language model from utter-project. EuroMoE-2.6B-A0.6B-Instruct-Preview (BF16) weighs 5.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | utter-project |
|---|---|
| Type | Language models |
| Task | Text gen · MoE |
| Parameters (lead) | 2.6B |
| Context | 4,096 tokens |
| Runs with | transformers |
| Based on | utter-project/EuroMoE-2.6B-A0.6B-Preview |
| Released | 2025-06-09 |
| Popularity | 1k downloads / month |
| Weights | 5.2 GB (EuroMoE-2.6B-A0.6B-Instruct-Preview (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 5.2 GB (file size) · KV cache 25 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 582 MB on a small card · context up to 4,096 tokens.
| Card | Requests at once 4K, its whole window tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 57 | — | all 4K | 11.6 GB |
| RTX 4060 Ti 16 GB | 95 | — | all 4K | 15.4 GB |
| RTX 3090 24 GB | 174 | — | all 4K | 23.4 GB |
| RTX 4090 24 GB | 174 | — | all 4K | 23.4 GB |
| RTX 5090 32 GB | 250 | — | all 4K | 31.0 GB |
| L40S 48 GB | 379 | — | all 4K | 44.0 GB |
| A100 80 GB | 719 | — | all 4K | 78.2 GB |
| H100 80 GB | 682 | — | all 4K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 838 | — | all 4K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 972 | — | all 4K | 107 GB |
| H200 141 GB | 1000+ | — | all 4K | 138 GB |
| B200 180 GB | 1000+ | — | all 4K | 176 GB |
| Requests at once | 4K, its whole window tokens each | 32K tokens each |
|---|---|---|
| 1 | 5.9 GB | — |
| 5 | 6.3 GB | — |
| 8 | 6.6 GB | — |
| 16 | 7.4 GB | — |
| 32 | 9.0 GB | — |
| 64 | 12.2 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
From the model card
⚠️ PREVIEW RELEASE: This is a preview version of EuroMoE-2.6B-A0.6B-Instruct-Preview. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroMoE-2.6B-A0.6B-Instruct-Preview. You can also check the pre-trained version: EuroMoE-2.6B-A0.6B-Preview.
The EuroLLM project has the goal of creating a suite of LLMs capable of understanding and generating text in all European Union languages as well as some additional relevant languages. EuroMoE-2.6B-A0.6B is a 22B parameter model trained on 8 trillion tokens divided across the considered languages and several data sources: Web data, parallel data (en-xx and xx-en), and high-quality datasets. EuroMoE-2.6B-A0.6B-Instruct was further instruction tuned on EuroBlocks, an instruction tuning dataset with focus on general instruction-following and machine translation.
EuroMoE uses a standard MoE Transformer architecture:
For pre-training, we use 512 Nvidia A100 GPUs of the Leonardo supercomputer, training the model with a constant batch size of 4096 sequences, which corresponds to approximately 17 million tokens, using the Adam optimizer, and BF16 precision. Here is a summary of the model hyper-parameters:
| Sequence Length | 4,096 |
| Number of Layers | 24 |
| Embedding Size | 1,024 |
| Total/Active experts | 64/8 |
| Expert Hidden Size | 512 |
| Number of Heads | 8 |
| Number of KV Heads (GQA) | 2 |
| Activation Function | SwiGLU |
| Position Encodings | RoPE (\Theta=500,000) |
| Layer Norm | RMSNorm |
| Tied Embeddings | Yes |
| Embedding Parameters | 0.13B |
| LM Head Parameters | 0.13B |
| Active Non-embedding Parameters | 0.34B |
| Total Non-embedding Parameters | 2.35B |
| Active Parameters | 0.6B |
| Total Parameters | 2.61B |
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "utter-project/EuroMoE-2.6B-A0.6B-Instruct-Preview"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
messages = [
{
"role": "system",
"content": "You are EuroLLM --- an AI assistant specialized in European languages that provides safe, educational and helpful answers.",
},
{
"role": "user", "content": "What is the capital of Portugal? How would you describe it?"
},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
outputs = model.generate(inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
EuroMoE-2.6B-A0.6B-Instruct-Preview has not been aligned to human preferences, so the model may generate problematic outputs (e.g., hallucinations, harmful content, or false statements).
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.