German AI company Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts language model built from the ground up for German and English. The model packs 78.1 billion total parameters but activates only 3.46 billion of them per token — about 4.4 percent — and ships under the permissive Apache 2.0 license on Hugging Face. It accepts up to 1,048,576 tokens of context and lets users set reasoning effort per request. For readers tracking the open-weights ecosystem, this lands alongside our latest AI developments.
The release is a deliberate statement about where Aleph Alpha sees its market: sovereign deployment in regulated sectors such as public administration, industry, and aerospace, where knowing exactly where a model was trained, on what data, and under which legal framework is not a nice-to-have but a procurement requirement.
What Kolibri actually is
Kolibri, referred to as Kolibri-1 in the technical materials, is a bilingual English-German MoE transformer that Aleph Alpha says was developed end to end by teams in Germany. According to the company's technical research report, the team controlled the full pipeline — data, architecture, training infrastructure, post-training, and evaluation — rather than starting from another lab's checkpoint.
Training ran on infrastructure located in Germany and Finland. The design process explicitly targeted compliance with the EU General-Purpose AI Code of Practice, the EU AI Act, and GDPR, and Aleph Alpha is a signatory of that Code. The company says its data pipeline redacts personal data before training, a detail aimed directly at public-sector buyers who cannot legally feed citizen data into models of unclear provenance.
It runs on a single GPU
Deployability was clearly a design constraint, not an afterthought. The FP8 checkpoint weighs in at roughly 78GB and runs on a single NVIDIA B200, B300, or H200 — or on two H100 SXM5 GPUs — served through vLLM with dedicated Kolibri reasoning and tool-call parsers. For a 78-billion-parameter model, that is a modest hardware footprint, and it is the direct payoff of the sparse MoE design: with only 3.46 billion parameters active per token, inference cost tracks the active parameter count rather than the total.
Sparse experts and hybrid attention
Under the hood, Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert alongside them. Expert load is balanced using two techniques with alphabetical-soup names — Exact Quantile Balancing and Load-Error Injection — that keep the router from collapsing all traffic onto a handful of specialists.
Attention is where the architecture gets more interesting. Kolibri uses grouped-query attention with 48 query heads and 4 KV heads, but the layers are not uniform: every fifth block uses full attention without positional encoding, while the other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. The practical consequence is memory efficiency — sliding-window layers hold a fixed-size KV cache, so only 10 of the 50 layers grow with context length. At matched compute, Aleph Alpha reports that the hybrid design supports sequences four times longer than an equivalent full-attention model. That is how a model of this size claims a one-million-token context window without exotic infrastructure.
A tokenizer built for German
Most frontier tokenizers treat German as an afterthought, splitting its long compound words into fragments that inflate token counts and effective costs. Aleph Alpha attacked this directly with UniBPE, a 128,000-token vocabulary that builds merges the way BPE does but scores each merge by Unigram loss.
The numbers make the case. On German text, Kolibri's tokenizer reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer — meaning 11.2 percent fewer tokens on German web text. On English, Kolibri actually comes out ahead too, at 4.58 bytes per token against GPT-5's 4.67. For enterprise customers paying by the token, that is a direct discount on every German-language workload.
Trained at frontier scale
The training run behind Kolibri is substantial. Pre-training covered 20 trillion tokens on 768 NVIDIA B200 GPUs, followed by 3.44 trillion mid-training tokens at a 65,536-token sequence length. The pipeline then moved through supervised fine-tuning and reinforcement learning stages to produce the final instruction-following, reasoning-capable model. The company has published a technical research report covering the process, an unusual level of disclosure for a European lab and one clearly aimed at building trust with regulated customers.
What it means for Europe's AI push
Kolibri enters a market where open-weight releases from American and Chinese labs dominate the conversation, and where European governments have grown increasingly vocal about digital sovereignty. Aleph Alpha's bet is that a slice of that market — ministries, defense-adjacent industry, critical infrastructure — will pay for a model that is not merely open-weight but verifiably European in its data handling, training location, and legal posture.
Whether that bet pays off depends on questions the release does not yet answer: how Kolibri benchmarks against frontier open models on standard evaluations, and how quickly an ecosystem of tooling forms around it. What the release does establish is that a full-stack, frontier-adjacent training capability now exists inside the EU, with the paperwork to match. For organizations bound by GDPR and the AI Act, that combination may matter more than leaderboard positions.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →