Model reference · open weights

Kolibri-1

NEW · this week LLMs Aleph-Alpha Text gen 1 build Open weights 1k dl/mo

Kolibri-1 is an open-weight language model from Aleph-Alpha. Kolibri-1 (FP8) weighs 78.8 GB; the smallest configuration that runs it is 2× L40S 48 GB.

  • Kolibri-1 is a 78.1B parameter mixture-of-experts model by Aleph-Alpha designed for text generation in German and English.
  • It supports a context length of 262,144 tokens and features explicit reasoning and tool calling.
  • The model is released under the Apache 2.0 license.

Summary of the Aleph-Alpha/Kolibri-1 model card, 2026-10-05

What it is

Released byAleph-Alpha
Released2026-10-02
Parameters78.1B
VRAM78.8 GB for the weights

What it runs on

Memory and cards for Kolibri-1 (FP8)

78.8 GBweights, file size
20 MBcache per 1K tokens
378 MBwindow cache per request
1.0 GBruntime overhead, at least
262,144 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … H100 80 GB
8 smaller cards
———
RTX PRO 6000 Blackwell 96 GB104all 256K93.8 GB
DGX Spark (GB10) 128 GB unified2610all 256K107 GB
H200 141 GB6325all 256K138 GB
B200 180 GB10626all 256K176 GB
2× L40S 48 GB
tensor parallel
126all 256K44.0 GB a card
4× RTX 4090 24 GB
tensor parallel
1910all 256K23.4 GB a card
4× RTX 3090 24 GB
tensor parallel · FP8 without its speed-up here
1910all 256K23.4 GB a card
4× RTX 5090 32 GB
tensor parallel
7539all 256K31.0 GB a card
2× H100 80 GB
tensor parallel
7831all 256K78.1 GB a card
2× A100 80 GB
tensor parallel · FP8 without its speed-up here
13872all 256K78.2 GB a card
2× RTX PRO 6000 Blackwell 96 GB
tensor parallel
11547all 256K93.8 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
180.4 GB80.9 GB
582.6 GB85.1 GB
884.2 GB88.2 GB
1688.6 GB96.6 GB
3297.3 GB113 GB
64115 GB147 GB

One card, with vLLM's small-card settings.

From the model card

What Aleph-Alpha says about Kolibri-1

Read the model card

Tech report | Tech blog

Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.

Model overview

ModelKolibri 1
Model ProviderAleph Alpha GmbH
Model DeveloperAleph Alpha Research GmbH
ArchitectureMixture-of-Experts
Total parameters78B (78,103,074,560)
Active parameters / token3.46B (3,457,573,120)
LanguagesGerman, English
Context length1,048,576 tokens; we recommend ≤262,144 tokens for serving efficiency and complex tasks
Precisionfloat8_e4m3fn weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16
Reasoning modeYes
Tool callingYes
LicenseApache 2.0
Knowledge cutoffEN: June 18, 2026, DE: June 18, 2026This only affects implicit knowledge, the model may use more recent information through tool use.
Hardware requirementsModel memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300.
Best forMulti-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant
Release Date3rd of October 2026
Code of PracticeAleph Alpha is a signatory of the EU GPAI Code of Practice, see .
Training DataPre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases.
Training methodWe trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed.
Computing ResourcesPre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh)Mid-training: 5 days, 90k GPUh (same setup as above)Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6)FLOPS: 6.4e23
Tech Report

Intended use

Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.

Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.

For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.

Design goals

Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.

AI system types

Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.

Sustainability

Energy consumption

9.5×10² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.

Energy measurement

For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms