Aleph Alpha just did something genuinely unusual for a European AI company: it gave away the crown jewels. On October 3, German Reunification Day, because symbolism matters, the Heidelberg-based company released Kolibri-1’s full weights on Hugging Face under Apache 2.0. No gated access, no “contact sales for the real version”, no regulatory hand-wringing. Just a 78.1-billion-parameter model you can download, run on your own hardware, and modify without asking anyone’s permission.
The headline numbers are impressive. But the story here isn’t just about specs, it’s about what happens when a company built for sovereign government contracts decides to go fully open, and whether the compromises they made to satisfy European regulators cost them the performance crown.
The Architecture: Sparsity as a Serving Strategy
Let’s start with what makes Kolibri-1 technically interesting. This is a Mixture-of-Experts model with 78.1 billion total parameters, but only 3.46 billion active per token. That’s a sparsity ratio most models don’t come close to. Mixtral 8x7B, for comparison, activates roughly 13 billion of its 47 billion parameters per token. DeepSeek’s V3-era models lean into single-digit-billion active counts, but Kolibri-1 pairs that sparsity with something those models don’t have: a 1,048,576-token context window.
The architecture uses 50 layers, each carrying 384 experts plus one shared expert that every token passes through. The router sends each token to just 6 of those 384 experts. Here’s the clever part: only 10 of the 50 layers process the full context window. The other 40 use a tight 512-token sliding window. This keeps decode computation and memory bounded regardless of context length, which is how you get a 1M-token context window without needing a supercomputer to serve it.
The practical impact of this design shows up in real-world usage. One developer who ran the model on a single Nvidia RTX Pro 6000 in FP8 measured about 170 tokens per second, a speed they credited directly to the small active-parameter count. That’s not theoretical, that’s someone running a 78B-parameter model on consumer hardware and getting usable inference speeds.
Hardware requirements are equally accessible. The model footprint is roughly 78 GB in FP8 weights. You can run it on 2× A100 80 GB, 2× H100 SXM5, a single H200, or even a single B200 or B300.
The $4 Million Question: What Training Actually Costs Now
Here’s where the conversation gets spicy. Kolibri-1 was trained on 20 trillion tokens using a 768 GPU B200 cluster over 21 days. Community cost estimates put that around $4 million for compute alone, assuming roughly $8/hour/GPU.
Four million dollars. For a frontier-adjacent model with 1M-token context and Apache 2.0 licensing.
That number would have been unthinkable two years ago. The original ChatGPT reportedly cost billions between data filtering, labeling, and training. Now, as one commenter pointed out, even adding another $4-5 million for data acquisition brings the total to under $10 million, within reach of large enterprises, universities, or research institutions in most Western countries.
The comparison to the broader open-weights ecosystem matters here. There’s been a wave of open-weights AI models with permissive licensing and technical scale hitting the market recently, and the cost trajectory is making the “we can’t afford to train our own models” argument increasingly weak. When a single 78B model costs $4M to train, and that cost is projected to halve again within a year as next-gen hardware ships in volume, the barrier to entry for enterprise-scale AI isn’t compute anymore, it’s expertise and data.
The Tokenizer Nobody’s Talking About
Most coverage of Kolibri-1 focuses on parameters and context length. But the tokenizer might be the most under-appreciated piece of engineering here.
Aleph Alpha developed a new algorithm called UniBPE that combines the bottom-up approach of BPE with the Unigram training objective for selecting which merges to add to the vocabulary. The result is a 128,000-vocabulary tokenizer that respects German morphology better than anything else on the market.
The numbers tell the story. On German web text, Kolibri achieves 4.90 bytes per token. That beats GPT-5’s tokenizer (4.35), Qwen3.8’s (4.17), and Gemini’s (4.13). But the real magic is in how it splits words. Look at “Bundessozialgericht” (Federal Social Court, genitive case):
- Kolibri: Bundes / sozial / gericht / es
- GPT-5: Bund / ess / oz / ial / gericht / es
- Qwen3.8: Bund / ess / oz / ial / gericht / es
The Kolibri tokenizer follows the morphology of the language. The competitors slice across morpheme boundaries, creating fragments that carry no semantic meaning. This matters for inference cost, fewer tokens per word means fewer FLOPs per document, and it matters for German-language performance, which we’ll get to shortly.
The Sovereign AI Twist
Here’s the part that makes this release genuinely unusual: this is a model built for government contracts being released as fully open weights. The original Kolibri launch was framed as a sovereign tool for German government and industry clients running their own infrastructure. The pitch was control, compliance, and GDPR-friendly data provenance.
Now the same model is on Hugging Face under Apache 2.0. Anyone can use it, modify it, or roll it into a commercial product without any obligation to open their own code.
This connects to a broader trend. European AI ambitions and open-weight models like Mistral represent a deliberate bet that open ecosystems are the way Europe competes with US closed models. Kolibri-1 is that bet in its purest form: a sovereign model, trained on European infrastructure, with data provenance documentation that would satisfy GDPR auditors, released under the most permissive license available.
But there’s a harder truth underneath the sovereignty narrative. The community analysis of Kolibri-1’s architecture suggests some design decisions were made for legal, not technical, reasons. Specifically, the model uses a 4:1 sliding-window attention ratio with a 512-token window, a choice that multiple research papers have shown to be suboptimal compared to 3:1. The model also uses pure NoPE (no positional encoding) in the full-attention layers, which is generally worse for the context lengths they’re claiming.
Why make these choices? The leading theory is data provenance. By training everything from scratch with their own German-first data pipeline, Aleph Alpha can claim full GDPR compliance and avoid the legal uncertainty that comes with borrowing components from models like Gemma, whose training data copyright status is murky. It’s a “Germany first” decision that costs them 15-20% of training compute efficiency, but gives them a defensible legal position that closed models can’t match.
Whether that trade-off is worth it depends entirely on your regulatory exposure. For a German public sector agency, that legal certainty might be worth more than a few benchmark points. For someone trying to build the best possible coding assistant, it’s a questionable allocation of compute.
The Benchmark Reality Check
Let’s talk numbers. Aleph Alpha’s published benchmarks show Kolibri-1 performing well, but not dominating:
| Benchmark | Kolibri-1 | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B |
|---|---|---|---|
| AIME 2025 (EN) | 96.9 | 84.6 | 91.7 |
| GPQA Diamond (EN) | 84.3 | 83.4 | 78.0 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.0 |
| SWE-Bench Verified | 66.4 | 73.8 | 60.2 |
| LongBench Pro | 64.5 | 70.8 | 62.9 |
| Overall (EN) | 75.5 | 71.4 | 73.0 |
| Overall (DE) | 70.8 | 67.3 | 67.9 |
The English overall score of 75.5 leads the MoE pack, and German performance at 70.8 is strong. But here’s the complication: Qwen3.6-35B-A3B beats it on agentic benchmarks including tau²-bench Telecom and the BFCL tool-calling test, and also wins on long-context tasks, the exact area Kolibri’s architecture was built to serve.
The German performance story is more nuanced. Kolibri leads on GPQA Diamond (DE) with 81.3, but Qwen3.6 takes German knowledge average at 61.3 vs Kolibri’s 57.6. For German-language RAG and document processing, Kolibri’s specialized tokenizer and training data make it genuinely competitive. For general-purpose English tasks, it’s competitive but not best-in-class.
One critical caveat: no independent lab has verified these scores. Aleph Alpha used its own evaluation harnesses, and several early community comparisons suggest the gap between vendor-reported and independently-measured performance could be meaningful.
Where Kolibri-1 Actually Shines
The most interesting results come from benchmarks that measure what enterprises actually care about. On Honeypot, an agentic-RAG benchmark, Kolibri-1 posts 80.8, beating Nemotron 3 Super (68.8) and Mistral Small 4 (68.1). On the automotive supplier customer-proxy benchmark, it hits 99.0 vs the next best at 94.6. Industrial drive technology: 60.0, tied for first.
These are the benchmarks that matter for government contracts. Aleph Alpha built internal evaluation suites for the German public sector, aviation, manufacturing, and automotive industries, and then trained against those suites without using customer data. The result is a model that’s specifically good at the tasks German enterprises need, regardless of what generic benchmark leaderboards say.
For running large open-weights models locally for cost and control benefits, Kolibri-1 is a legitimate option. The 3.46B active parameters keep inference costs low, and the Apache 2.0 license means no usage restrictions. If you’re processing German-language documents and need long-context capabilities, this might be the best open option available.
The Long Context Elephant
Here’s the tension in the release. The model claims up to 1M tokens of context, but Aleph Alpha’s own documentation recommends staying at or under 262,144 tokens for serving efficiency. The model was trained natively at 262,144 tokens, with the 1M capability achieved through extrapolation rather than direct training.
The RULER benchmark results tell an interesting story. At 1M tokens, Kolibri Base scores 63.2, beating Nemotron 3 Nano (58.5) and Qwen3.5-35B-A3B (57.5). But the scores drop significantly from the 256K peak (69.8) to 512K (65.5) and 1M (63.2). The model can technically process a million tokens, but quality degrades measurably as you push past its native training context.
The practical implication: if you need to process a genuinely enormous document, think full contract archives or multi-year regulatory filings, you can do it without chunking. But you’ll get better results if you’re willing to stay under 256K tokens.
Serving Kolibri-1: What You Need to Know
Getting Kolibri-1 running requires the aleph-alpha-inference package, which provides the Kolibri vLLM plugin. The setup is straightforward:
pip install 'aleph-alpha-inference>=1'
Then serve with reasoning and tool-calling enabled:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
For contexts beyond 262,144 tokens:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--max-model-len 1048576 \
--hf-overrides '{"max_position_embeddings": 1048576}'
The recommended sampling parameters are temperature=1.0, top_p=0.97, and top_k=128. Kolibri supports four reasoning effort levels, none, low, medium, and high, which gives you control over the compute quality trade-off depending on the task.
One note: the model uses FP8 weights with dynamically quantized activations, evaluated with an FP8 KV cache. If you need full precision, the BF16 version is available as a separate base model.
The Verdict
Kolibri-1 is honestly a bit of a mixed bag from a pure ML research perspective. The architecture has some questionable choices (4:1 SWA ratio, NoPE in full-attention layers) that likely cost 15-20% training compute efficiency. On raw English benchmarks, it’s competitive but not leading. And the benchmarks are vendor-run, which should always invite skepticism.
But here’s the thing: this isn’t a research model. It’s a sovereign AI play designed for German government agencies and regulated industries. From that perspective, the priorities make sense. Full data provenance for GDPR compliance. A tokenizer that handles German compounds efficiently. Training on European infrastructure without foreign control. Abstention training so the model says “I don’t know” instead of hallucinating, something that’s literally a deployment requirement in regulated environments, not a nice-to-have.
The Apache 2.0 release changes the calculus. Anyone can now take this sovereign model and build on it. The open-weight decision models advancing transparent AI ecosystem just got a serious new player, and for German-language applications specifically, Kolibri-1 may be the best open option available.
The real story here isn’t whether Kolibri-1 beats Qwen3.6 on a benchmark. It’s that a German company with government contracts, GDPR obligations, and a $20 billion merger with Cohere just released a genuinely useful 78B model with 1M context under Apache 2.0, and it cost them about $4M to train. Two years ago, that was a billion-dollar research project. Now it’s a line item on a mid-cap startup’s budget.
The question is whether the rest of the industry is ready for what that means.




