Model reference · open weights
Kolibri-1 is an open-weight language model from Aleph-Alpha. Kolibri-1 (FP8) weighs 78.8 GB; the smallest configuration that runs it is 2× L40S 48 GB.
Summary of the Aleph-Alpha/Kolibri-1 model card, 2026-10-05
What it is
| Released by | Aleph-Alpha |
|---|---|
| Released | 2026-10-02 |
| Parameters | 78.1B |
| VRAM | 78.8 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB … H100 80 GB 8 smaller cards | — | — | — | |
| RTX PRO 6000 Blackwell 96 GB | 10 | 4 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 26 | 10 | all 256K | 107 GB |
| H200 141 GB | 63 | 25 | all 256K | 138 GB |
| B200 180 GB | 106 | 26 | all 256K | 176 GB |
| 2× L40S 48 GB tensor parallel | 12 | 6 | all 256K | 44.0 GB a card |
| 4× RTX 4090 24 GB tensor parallel | 19 | 10 | all 256K | 23.4 GB a card |
| 4× RTX 3090 24 GB tensor parallel · FP8 without its speed-up here | 19 | 10 | all 256K | 23.4 GB a card |
| 4× RTX 5090 32 GB tensor parallel | 75 | 39 | all 256K | 31.0 GB a card |
| 2× H100 80 GB tensor parallel | 78 | 31 | all 256K | 78.1 GB a card |
| 2× A100 80 GB tensor parallel · FP8 without its speed-up here | 138 | 72 | all 256K | 78.2 GB a card |
| 2× RTX PRO 6000 Blackwell 96 GB tensor parallel | 115 | 47 | all 256K | 93.8 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 80.4 GB | 80.9 GB |
| 5 | 82.6 GB | 85.1 GB |
| 8 | 84.2 GB | 88.2 GB |
| 16 | 88.6 GB | 96.6 GB |
| 32 | 97.3 GB | 113 GB |
| 64 | 115 GB | 147 GB |
One card, with vLLM's small-card settings.
From the model card
Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.
| Model | Kolibri 1 |
|---|---|
| Model Provider | Aleph Alpha GmbH |
| Model Developer | Aleph Alpha Research GmbH |
| Architecture | Mixture-of-Experts |
| Total parameters | 78B (78,103,074,560) |
| Active parameters / token | 3.46B (3,457,573,120) |
| Languages | German, English |
| Context length | 1,048,576 tokens; we recommend ≤262,144 tokens for serving efficiency and complex tasks |
| Precision | float8_e4m3fn weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16 |
| Reasoning mode | Yes |
| Tool calling | Yes |
| License | Apache 2.0 |
| Knowledge cutoff | EN: June 18, 2026, DE: June 18, 2026This only affects implicit knowledge, the model may use more recent information through tool use. |
| Hardware requirements | Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300. |
| Best for | Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant |
| Release Date | 3rd of October 2026 |
| Code of Practice | Aleph Alpha is a signatory of the EU GPAI Code of Practice, see . |
| Training Data | Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases. |
| Training method | We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed. |
| Computing Resources | Pre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh)Mid-training: 5 days, 90k GPUh (same setup as above)Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6)FLOPS: 6.4e23 |
| Tech Report |
Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.
Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.
For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.
Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.
Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.
9.5×10² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.
For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.