
Kolibri vs Qwen 3.8: the German Model Loses on Its Own Table, and That Is the Point
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 218 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 115 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 982 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 47 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 105 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 215 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Start with the sentence most launch coverage of Kolibri left out. On the benchmark table Aleph Alpha published with its own model on 3 October 2026, the model it is measured against most often is not a European rival and not an American one. It is Qwen3.8 27B, the vendor's open-weight tier of that generation, released on 13 August 2026 — and on Aleph Alpha's own numbers it wins. Qwen3.8 27B scores 80.2 on the English average and 79.9 on the German average of the same evaluation suite that gives Kolibri 75.5 and 70.8. It leads on GPQA Diamond, 89.2 against 84.3. It leads on LiveCodeBench v6, 93.8 against 85.9. It leads on SWE-Bench Verified, 72.6 against 66.4, and on both long-context benchmarks, 76.9 against 64.5 on LongBench Pro and 81.3 against 68.3 on AA-LCR. A 27-billion-parameter dense Chinese model, evaluated by a German lab, beats that lab's 78-billion-parameter flagbearer across most of its own table. That is not a scandal. It is the most useful fact in the release, because it forces the question of what Kolibri is actually for.
One clarification before the comparison goes any further, because "Qwen 3.8" names a family rather than a model and the distinctions matter here. Qwen3.8-Max is Alibaba's hosted flagship, live since 3 August 2026 with a 1M-token context window at $2.00 per million input tokens and $6.00 output. Qwen3.8-27B is the open-weights tier, released 13 August 2026 under a 262,144-token window. Qwen3.8-Flash is the volume tier, released 26 August 2026 with the 1M-token window and much lower prices. And plain Qwen3.8 — announced on 19 July 2026 as a 2.4-trillion-parameter open-weight flagship — is not released; it exists as a catalogue tracking page and an announcement. Everything below is about the one with weights you can download and run, which is the 27B.
Where Kolibri actually wins, read off the same table
The rows where Kolibri leads Qwen3.8 27B are not scattered at random. They cluster, and the cluster is the product.
• Abstention — on RGB Negative, Kolibri scores 85.6 against 70.6. When the context does not contain the answer, Kolibri says so far more often.
• Tool-calling under constraint — on τ²-bench telecom, 94.7 against 82.5; on the automotive-supplier vertical suite, 99.0 against 97.3.
• Very long context — on SealQA with no distractors above 24k tokens, 100.0 against 94.4.
• German serving economics — a bilingual tokenizer at 4.90 average bytes per token on German web text, against 4.17 for the Qwen3.5–3.8 family on Aleph Alpha's own measurement. Fewer tokens per German document is a direct inference-cost reduction that shows up on an invoice and in no benchmark row.
Three of those four are about behaviour rather than capability. Abstention, constrained tool-calling and grounded long-context retrieval are what you need from a model that reads your organisation's own material and is expected to admit when the material does not answer the question. Aleph Alpha built the whole release around that bet, and the table shows the bet landing.

The rows where Qwen3.8 27B wins are about breadth: knowledge in both languages, graduate-level science questions, competitive programming, repository-level software engineering, long-document comprehension. If your requirement is a strong general model at a low per-token price with open weights, Qwen3.8 27B is the better model and the table says so in the vendor's own ink.
That reading is corroborated where Kolibri's is not. Artificial Analysis has measured Qwen3.8 27B independently, in its Xhigh reasoning configuration, and reports an Intelligence Index of 34 at #1 of the 142 models on that comparison, 45.4 output tokens per second, a 256,000-token context and an Apache 2.0 licence. Their note is unflattering in a familiar way — very verbose at 200 million tokens generated during the index evaluation against a median of 82 million, and slow — but the score itself is a third-party measurement of the opponent. Kolibri has no Artificial Analysis page at all. So the strongest model in this comparison has been independently verified and the weaker one is the vendor's word, which is the reverse of how a first-generation European flagship is usually framed.
Two kinds of supply-chain claim, pointed in opposite directions
Both of these releases are arguments about where a model comes from, and the arguments are mirror images.
Aleph Alpha's case is provenance as a feature. Kolibri was trained by German teams on infrastructure in Germany and Finland, under EU and German law. The card discloses the pre-training mix — 20 trillion tokens at roughly 62.5 per cent English, 23.9 per cent German and 13.6 per cent code — the mid-training and long-context budgets, the exact compute 768 NVIDIA B200s across 96 HGX nodes for 21 days, and an estimated 9.5×10² MWh of energy including data-centre overhead. It states the training method down to the layer: a 50-layer mixture-of-experts transformer with a 4:1 sliding-window to grouped-query attention split, Muon and Exact Quantile Balancing across 384 experts with 1 shared and 6 routed. Twelve thousand words of disclosures attached to a 78 GB download, and a stated signatory position on the EU's General-Purpose AI Code of Practice.
Alibaba's case is different in kind: not "we can account for every decision" but "here are the weights, take them". Apache-licensed open weights at a scale and quality that made Qwen the effective default around the world, a 248,077-token vocabulary covering well over a hundred languages, and a serving ecosystem that has been tuned around those checkpoints for long enough that almost every inference stack supports them out of the box. The sovereignty argument there is about availability: no vendor can withdraw the model, because you already have the file.
Both claims are honest and they answer different fears. A public-sector buyer in Baden-Württemberg is worried about jurisdiction and auditability, and Kolibri is addressed to that worry. A research group in Nairobi or São Paulo is worried about a model being gated behind a credit card in a currency they cannot pay in, and open Qwen weights are addressed to that one. What is not honest is pretending either is a capability comparison, because the capability table already gave its answer.

The licence gap is real but narrower than it sounds
Kolibri is Apache 2.0. Qwen3.8 27B is Apache-licensed open weights. Both come with no acceptable-use rider, no monthly-active-user threshold and no separate commercial terms — which puts them in a small group and means neither requires a legal review before commercial use.
What separates them is everything that follows the licence. Kolibri arrives with a vLLM plugin and parser package from the vendor, a container image, four reasoning effort levels configurable through the chat template, an explicit abstention behaviour, and a recommended sampling configuration. Qwen3.8 27B arrives as weights that Transformers, vLLM and SGLang already know how to serve, with a tool-calling format the wider ecosystem has had months to accommodate.
Deployment hardware differs along the same lines. Kolibri's FP8 weights are a ~78 GB footprint with a floor of two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300. Qwen3.8 27B in bfloat16 is roughly 54 GB of weights, which is a single 80 GB accelerator at full precision or a quantised build on considerably less. The German model costs more hardware and returns less benchmark for it. That is only a bad trade if you were buying benchmark.
What the price tags do and do not tell you
Kolibri has no price, because Aleph Alpha is not selling inference for it. The release is weights, a technical report and a card; there is no hosted SKU. Any cost figure for Kolibri is a hardware amortisation calculation you do yourself.
Qwen3.8 is priced in public at three tiers, and the middle one is the one that matters for a team comparing self-hosting against renting.
• Qwen3.8-Max — $2.00 per million input tokens and $6.00 output, 1M-token context, cache reads at $0.25, released 3 August 2026.
• Qwen3.8-27B — $0.33 input and $2.40 output, 262,144-token context, released 13 August 2026. These are the open weights, so this price is what you pay if you would rather not run them.
• Qwen3.8-Flash — $0.15 input and $0.47 output with the 1M-token window, released 26 August 2026.
Those are provider list prices on OrcaRouter, passed through with nothing added per token, which is the useful property when the question is whether a self-hosted deployment pays for itself: you can measure your real token volume at the hosted tiers before committing to hardware. We do not add a margin, so a vendor price cut is live on our side the same day. All three Qwen3.8 tiers are routable today on one OpenAI-compatible key with automatic failover across providers, and the upcoming 2.4-trillion-parameter Qwen3.8 has a catalogue page waiting for the weights rather than a placeholder pretending to serve them.
Kolibri is not routable. We probed the catalogue under every vendor and model spelling and it is absent — so the only way to call it is the 78 GB download. If you want to know what the German model would cost your workload, the answer is a rack and a person who can run vLLM, and no third party can currently quote you anything else.

Who should buy which
The decision does not turn on quality, because that question has been settled by the vendor's own table in Alibaba's favour. It turns on which risk you are buying insurance against.
• Choose Kolibri if the binding constraint is jurisdiction, provenance documentation or auditability — if you need to be able to answer where the training data came from, where the compute ran and under which law, and if the workload is German and English documents processed inside your own perimeter. You are paying a capability premium for that, and the premium is visible in the table above.
• Choose Qwen3.8 27B if the binding constraint is capability per dollar, breadth of language, or ecosystem support. It is the stronger model on Aleph Alpha's own evaluation, it is Apache-licensed, it runs on one accelerator, and the hosted tiers start at $0.33 per million input tokens if you would rather not run it at all.
• Consider both in the same architecture if the workload has a regulated core and a general periphery. The documents that cannot leave the building go to a self-hosted Kolibri endpoint; the summarisation, translation and code work goes to a routed Qwen3.8 tier at a third of a cent per thousand tokens. One API key, one routing rule, provider list prices passed through, and the failover handled for you — which is a considerably more useful arrangement than picking one of these two and pretending the other does not exist.
