
Kolibri's First Day: Aleph Alpha Opened 78B of German-English Weights, and the Reaction Is Running Ahead of the Evidence
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 100 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Twenty-four hours after Aleph Alpha published Kolibri, the reaction has hardened into a single line, and it is worth taking that line apart. The Heidelberg lab put a 78.1-billion-parameter mixture-of-experts model on Hugging Face under Apache 2.0 on 3 October 2026, announced it the same morning in its own newsroom and from its own account, and by the next day the summary circulating in AI feeds read: 78B total parameters, 3.46B active per token, open weights under Apache 2.0, built for German and English. Every clause of that sentence is checkable from the repository. What has gone largely unsaid is which parts of the release are measured facts and which are claims the authors made about their own model that nobody outside Heidelberg has run. That gap is the day-one story, and it matters more than the headline number.
Start with the timeline, because a surprising amount of the reaction treated this as a mysterious silent drop, and it was not one. The Hugging Face repository Aleph-Alpha/Kolibri-1 was created on 2 October 2026 at 05:54:02 UTC, the model card carries a stated release date of 3 October, and the vendor's newsroom post went out the same morning at 08:43:53 GMT — on the Day of German Reunification, a date the lab clearly chose. Aleph Alpha's own account posted about it within minutes. There was no press embargo and no launch event, which is probably where the "quiet release" framing came from, but a vendor newsroom post, a technical report and an official announcement are not a quiet drop. They are a normal release that the general AI press simply did not cover on the day.
The number the reaction is repeating
The reason 3.46B active is the figure everyone quotes is that it is the only number in the release that changes what a buyer has to own. Kolibri is a 50-layer transformer in which every layer is mixture-of-experts, with 384 experts per layer, one shared and six routed, over a total of 78,103,074,560 parameters. Per token it activates 3,457,573,120 of them — a sparsity ratio of roughly 22.6 to 1. That is what makes a 78-billion-parameter model servable on a pair of H100s instead of a rack, and it is the single most consequential claim in the release.
Around it sit the other headline figures, all of them from the model card rather than from an independent run:
• Context — trained at 16,384 tokens, mid-trained at 65,536, native at 262,144, and validated up to 1,048,576. The card recommends staying at or below 262,144 for latency-sensitive work. Note carefully that the flat "1M context" you will see in summaries is the validated ceiling, not the native window.
• Precision — FP8 weights in 128x128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, the LM head, the norms and the MoE router stay in bfloat16.
• Reasoning — four effort levels — none, low, medium, high — set through the chat template, with Hermes-style tool calling and a vLLM parser shipped in the same repository as the weights.
• Languages — German and English, and nothing else.
• Training — 20 trillion pre-training tokens over a filtered bilingual corpus that is roughly 62.5 per cent English, 23.9 per cent German and 13.6 per cent code, plus 3.44 trillion mid-training tokens and 201 billion for long-context extension.
• Compute and energy — 768 NVIDIA B200s across 96 HGX nodes, 21 days of pre-training for 511 hours and 392,000 GPU-hours at a reported 6.4x10^23 FLOPs, and an estimated 9.5x10^2 MWh including data-centre overhead, excluding supervised fine-tuning and reinforcement learning.
• Licence — Apache 2.0. No acceptable-use rider, no monthly-active-user threshold, no separate commercial terms.
The two claims nobody has checked
This is the part of day one that the reaction has skipped, and it is the part that decides whether Kolibri is important or merely interesting.
The first is the serving-economics claim. Aleph Alpha's framing is not that Kolibri is the strongest model available — on the lab's own comparison table it is not, and the lab says so. Its framing is that no compared model delivers more quality at the same serving cost, or the same quality at lower cost, along a Pareto frontier of quality against decoded tokens per second per GPU. That is a claim about throughput on specific hardware, and it is exactly the kind of claim that only a third party with a stopwatch and two H100s can settle. Nobody has. There is no Artificial Analysis entry for Kolibri as of 4 October 2026 and no third-party replication of the evaluation harness.
The second is the tokenizer claim. Aleph Alpha trained a bilingual German-English tokenizer with a 128,000-token vocabulary using a method it calls UniBPE, and reports 4.90 average bytes per token on German web text against GPT-5 at 4.35, Gemini at 4.13, Qwen 3.5-3.8 at 4.17, GLM 5.3 at 3.93, DeepSeek V4 at 3.72 and Kimi K3 at 3.28. Every one of those figures is the vendor's own, measured on the vendor's own corpus, against competitors the vendor chose. More text per token is a real inference-cost effect rather than a quality claim, and if it holds it compounds across any document-processing workload — but it is one lab's measurement, not a benchmark result.
German is the point, and it is the most interesting engineering in the release
Most "multilingual" model cards describe a training mix that happened to include German. This one is built the other way round, and the card is unusually specific about the mechanics.
Small proxy-model experiments led Aleph Alpha to a target of roughly 20 per cent German data in the mix, which for a 20-trillion-token run meant finding 4 trillion German tokens. Deduplicated and filtered open German datasets supplied 390 billion — an order of magnitude short. The lab closed the gap three ways: retuning a Common Crawl filter for German specifically, which produced 1.3 trillion unique organic tokens; rephrasing existing German documents into encyclopaedia entries, question-and-answer dialogues and text passages, which produced about another trillion and became the single largest German source; and translation, which was used for the predecessor model and deliberately dropped for Kolibri. German ended as a 2.4-trillion-token unique pool, 80 per cent of it curated or generated in-house, seen at 21.3 per cent of pre-training tokens after upsampling.
The filter detail is the kind of thing that only surfaces in a lab that actually trained on German. A standard language-data pipeline drops documents with too many long words — but German administrative prose routinely exceeds the English bound on mean word length, so default settings quietly delete the register that public administration writes in. A model trained on a stock English filter has no German legal-administrative vocabulary to speak of, and no benchmark will tell you so.
What the comparison table actually shows
Aleph Alpha published a post-training comparison covering fourteen models, and it is more useful than most launch tables because it does not flatter the subject. On the same harness with Kolibri at reasoning effort high, Qwen3.8 27B scores 80.2 on the English average against Kolibri's 75.5, and 79.9 on the German average against 70.8, and it leads on GPQA Diamond, LiveCodeBench v6, SWE-Bench Verified and both long-context benchmarks. Nemotron 3 Super 120B-A12B scores 73.0 English against Kolibri's 75.5. Gemma 4 26B-A4B IT scores 71.9 and 66.3 on the two averages.
So the honest summary of the release is not "state of the art". It is that Kolibri sits in the middle of a field of models activating between three and twelve billion parameters per token, that it beats the much smaller ones clearly, and that it loses to a dense 27-billion-parameter model from Alibaba on most rows of its own table while activating roughly an eighth of the parameters. Whether the economics rescue that gap is the Pareto claim, and the Pareto claim is untested.

How you would actually call it, today
There is no hosted endpoint from the vendor — the release is weights plus a technical report, with no Aleph Alpha API SKU for Kolibri in the published material. There is also no route for it on OrcaRouter: we probed the catalogue under every spelling of the vendor prefix and the model name and Kolibri is not there, so calling it today means downloading roughly 78 GB of FP8 weights and running a vLLM endpoint yourself from the aleph-alpha-inference package or the published container image.
That is a real procurement decision, not a footnote, and it is where a routing layer earns its place even for a model it does not carry. The honest comparison for anyone weighing a self-hosted 78B sovereign deployment is not benchmark rows — it is 78 GB of VRAM and a procurement conversation against a per-token line item on hardware someone else owns. If what you actually need this month is a small mixture-of-experts model you can call without buying two GPUs, the Gemma 4 26B-A4B tier is the nearest thing already on a routed endpoint, at $0.06 per million input tokens and $0.33 per million output with a 262,144-token window, and OrcaRouter routes that tier on one OpenAI-compatible key at the provider's list price with nothing added. That is the cheap way to find out whether the workload justifies owning the hardware before the hardware arrives.

What to watch, and what would change the verdict
Three things, in order of how much they would move the picture.
An independent throughput measurement at the card's recommended two-H100 configuration would confirm or sink the Pareto claim, which is the only claim in the release that is genuinely novel rather than a restatement of a spec sheet. An independent reproduction of the German and English task-suite averages would tell you whether the gap to Qwen3.8 27B is as narrow as the vendor's own table suggests or narrower. And a German-language evaluation on real administrative text — not a translated general benchmark — would test the thing the release was actually built for, which no leaderboard currently measures.
Until one of those lands, the correct framing is the one the repository itself supports. Kolibri is a genuine open-weights release from a European lab, at a scale European labs have not previously shipped, with Apache 2.0 terms no incumbent matched, a data pipeline published alongside the weights, and a two-model iteration story that is more interesting than the headline number. It is also, for now, a model whose most quoted figures come from its own authors. Both of those sentences are true on day one, and the second one is the one the reaction left out.

