Generated hero card for the Kolibri explainer, headed 'Kolibri: a 78B open-weight model from Germany' with the subline '78B total / 3.46B active / 1M-token context / Apache 2.0' and three flat line icons: a hummingbird, a stack of layers, and a padlock over a document. The OrcaRouter logo sits in the strip below the artwork.
Guides & Insights

Kolibri, Explained: Aleph Alpha Ships a 78B German-English Model Under Apache 2.0

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Aleph Alpha released Kolibri on 3 October 2026, a 78.1-billion-parameter mixture-of-experts model that activates 3.46 billion parameters per token, carries a validated context window of 1,048,576 tokens, and ships with full weights on Hugging Face under the Apache 2.0 licence. The Heidelberg lab published it on the Day of German Reunification, alongside a technical report, a German-and-English model card, and an unusually detailed account of how the thing was trained. The model it is measured against throughout Aleph Alpha's own documentation is Kolibri Origin — the 30.6-billion-parameter, 3.27-billion-active predecessor that finished pre-training on 11 June 2026 and was never publicly released. Two models, three months apart, and the distance between them is the most interesting thing in the release.

What makes Kolibri worth a careful read is not that it is the strongest model available. On Aleph Alpha's own comparison tables it is not, and we will get to that. What is notable is that a European lab shipped an open-weight model at this scale at all, with the training pipeline, the data provenance and the energy figure published alongside it, and set the licence to Apache 2.0 rather than a bespoke research licence. For teams whose procurement rules start with "where does the data go", that combination is the product.

The spec sheet, and the two numbers that matter

Everything below is from the model card Aleph Alpha published with the weights; none of it has been independently reproduced yet, and there is no Artificial Analysis row for Kolibri at the time of writing.

• Total parameters — 78,103,074,560, with 3,457,573,120 active per token, a ratio of roughly 22.6 to 1.

• Architecture — a 50-layer transformer, all layers mixture-of-experts, 384 experts per layer with 1 shared and 6 routed, and a 4:1 mix of sliding-window to grouped-query attention across the stack.

• Context — trained at 16,384 tokens, mid-trained at 65,536, and extended to a native 262,144. Because positional encoding is applied only in the sliding-window layers, Aleph Alpha states the window can be extended without position scaling, and has validated quality and serving efficiency up to 1,048,576 tokens. The card recommends staying at or below 262,144 for latency-sensitive work and for complex tasks.

• Precision — FP8 weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, the LM head, the norms and the MoE router stay in bfloat16.

• Reasoning — four effort levels, none, low, medium and high, set through the chat template.

• Tool calling — yes, Hermes-style, with a vLLM parser shipped in the same repository as the weights.

• Languages — German and English, and nothing else. That is stated as a design choice rather than an omission.

• Training — 20 trillion pre-training tokens on a filtered bilingual corpus, roughly 62.5 per cent English, 23.9 per cent German and 13.6 per cent code, plus 3.44 trillion mid-training tokens and 201 billion for long-context extension.

• Compute — 768 NVIDIA B200s across 96 HGX nodes, 21 days of pre-training for 511 hours and 392,000 GPU-hours, at a reported 6.4×10²³ FLOPs. Mid-training added five days and 90,000 GPU-hours; long-context extension added 13 hours and 10,000 GPU-hours.

• Energy — 9.5×10² MWh estimated including data-centre overhead, covering pre-training, mid-training and long-context training, but excluding supervised fine-tuning and reinforcement learning.

• Licence — Apache 2.0. No acceptable-use rider, no monthly-active-user threshold, no separate commercial terms.

Two of those lines are the ones a buyer should read twice. The active-parameter count is what you pay for at inference, and 3.46 billion active out of 78 billion total is an aggressive sparsity ratio — it is the whole reason a model this size can be served at all on two GPUs. And the context figure is stated as 262,144 native with 1,048,576 validated, which is a more careful claim than the flat "1M context" you will see in most coverage of this release.

Generated single-column scoreboard for Kolibri with six rows: total parameters 78.1B, active per token 3.46B, context 262,144 native and 1M validated, licence Apache 2.0, independent score 'none published', and deployment floor 2x H100 80GB. The footer reads 'All figures vendor-reported from the model card; no independent evaluation yet.'

Two models, three months apart

Kolibri is the second model out of what Aleph Alpha calls its Model Factory, and the timeline it published is the part of the release that most coverage skips.

Work on the training pipeline began in January 2026. Kolibri Origin finished pre-training at target scale on 11 June 2026 — 30.6 billion total parameters, 3.27 billion active, a 65,536-token context window, 7.51 trillion training tokens, 50 layers of which two were dense and 48 were MoE, and a single reasoning mode. It was validated internally and never released publicly. Kolibri finished pre-training on 11 September 2026.

• Parameters — 30.6B total and 3.27B active, against 78.1B total and 3.46B active.

• Context — 8,192 tokens of pre-training context on Origin, against 16,384 on Kolibri, ending at a native 262,144.

• Data — 7.51 trillion pre-training tokens on Origin, against 20 trillion, drawn from a raw pool of more than 200 trillion that the pipeline filtered, deduplicated and curated.

• Architecture — two dense layers plus 48 MoE layers on Origin, against 50 MoE layers, with a changed attention design, triple the expert count, higher sparsity and a replaced routing algorithm.

• Reasoning — one mode on Origin, against none, low, medium and high on Kolibri.

• Release — no public release for Origin, 3 October 2026 for Kolibri.

The numbers that move most are the task-suite ones, where the same eval harness scored both models. On the German average of the post-training suite Origin scored 46.4 and Kolibri scored 70.8; on the English average, 54.1 against 75.5. On AIME 2025 in German, 73.5 against 87.5. On the internal customer-proxy benchmarks the vendor publishes as a vertical set, the automotive-supplier suite went from 0.72 to 0.99, semiconductors from 0.35 to 0.80, the German public sector from 0.54 to 0.75, industrial drive technology from 0.31 to 0.60 and aerospace from 0.14 to 0.59.

Those internal vertical numbers are vendor-run on vendor-built evaluation suites designed around vendor customers, so treat them as a description of intent rather than a score. The point of publishing them is that they explain what the model was optimised for: not a leaderboard, but a set of regulated-industry documents.

Screenshot of Aleph Alpha's newsroom post 'Kolibri Has Landed: A Sovereign Open-Weight Model', showing the opening lines dating the release to the Day of German Reunification, describing Kolibri as an English-German mixture-of-experts transformer with 78B total and 3B active parameters, context lengths up to 1M tokens, full weights on Hugging Face under Apache 2.0, and the paragraph on Kolibri Origin as the 30B total, 3B active predecessor with a 65k context window.

German is a design decision here, not a language tag

Most "multilingual" model cards mean a training mix that happened to include some German. This one is built the other way round, and the card is specific about the mechanics.

Aleph Alpha found from small proxy-model experiments that around 20 per cent German data in the mix produced the best results, which for a 20-trillion-token run meant finding 4 trillion German tokens. Deduplicated and filtered open German datasets gave it 390 billion — short of the target by an order of magnitude. It closed the gap three ways: retuning a Common Crawl filter for German specifically, which produced 1.3 trillion unique organic tokens; rephrasing existing German documents into encyclopaedia-style entries, question-and-answer dialogues and text passages, which produced about 1 trillion more and became the single largest German source; and translation, used only in Kolibri Origin and dropped for Kolibri. German ended as a 2.4-trillion-token unique pool, 80 per cent of it curated or generated in-house, seen at 21.3 per cent of pre-training tokens after upsampling.

The filter detail is worth quoting because it is the kind of thing that only shows up in a lab that actually trained on German. A standard language-data pipeline drops documents with too many long words — but German administrative prose routinely exceeds the English bound on mean word length, so the default settings quietly delete the register that public administration writes in. The obvious commercial consequence is that a model trained on a stock English filter has no German legal-administrative vocabulary to speak of, and no benchmark will tell you that.

The tokenizer gets the same treatment. Aleph Alpha trained a bilingual German-English tokenizer with a 128,000-token vocabulary using a new method it calls UniBPE, which keeps the bottom-up merging of byte-pair encoding but selects merges with a unigram objective. On German web text it reports 4.90 average bytes per token, ahead of GPT-5 at 4.35, Gemini at 4.13, Qwen3.5–3.8 at 4.17, GLM 5.3 at 3.93, DeepSeek V4 at 3.72 and Kimi K3 at 3.28 — all vendor-reported figures on the vendor's own comparison. More text per token means fewer tokens per task, which is a direct inference-cost effect rather than a quality claim, and it is the sort of efficiency win that compounds across a document-processing workload.

What the vendor's table does not settle

Aleph Alpha published a post-training comparison covering fourteen models, and it is more useful than most launch tables because it does not flatter the subject.

On the same harness, with Kolibri at reasoning effort high, Qwen3.8 27B scores 80.2 on the English average against Kolibri's 75.5, and 79.9 on the German average against 70.8. It also leads on GPQA Diamond, LiveCodeBench v6, SWE-Bench Verified and the two long-context benchmarks. Nemotron 3 Super 120B-A12B scores 73.0 English against Kolibri's 75.5 — closer, and ahead on AIME 2025. Gemma 4 26B-A4B IT scores 71.9 and 66.3 on the two averages.

So the honest summary of the release is not "state of the art". It is that Kolibri sits in the middle of a field of models that activate between three and twelve billion parameters per token, that it beats the much smaller ones clearly, and that it loses to a dense 27-billion-parameter model from Alibaba on most rows of its own table while activating roughly an eighth of the parameters. Aleph Alpha's framing for that is the Pareto frontier of quality against decoded tokens per second per GPU — none of the compared models delivers more quality at the same serving cost, or the same quality at lower cost. That is the claim to interrogate, and it is a serving-economics claim, not a capability claim.

None of these figures are independently verified. There is no Artificial Analysis entry for Kolibri, no third-party replication of the eval framework, and the vertical benchmarks are vendor-built. The comparison is also structurally generous to Kolibri in one direction and ungenerous in another: all fourteen models were run through Aleph Alpha's own harness with identical prompts and few-shot settings, which is the right way to do it, but it is still one lab's harness. Treat every number in this section as vendor-reported until someone else runs it.

Screenshot of the Hugging Face model card for Aleph-Alpha/Kolibri-1, showing the licence badge apache-2.0 and the Model overview block: architecture Mixture-of-Experts, 78B total parameters given as 78,103,074,560, active parameters per token 3,457,573,120, languages German and English, context length 1,048,576 tokens with a recommendation to stay at or below 262,144 for serving efficiency and complex tasks, float8_e4m3 weights in 128x128 blocks with an FP8 KV cache and embeddings, LM head, norms and MoE router in bfloat16, reasoning mode yes, tool calling yes, and the June 18 2026 knowledge cutoff.

What you need to run it

Kolibri is not a model you try on a laptop. The FP8 weights are a model memory footprint of about 78 GB, and the card's minimum is two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300; the recommended configuration is two H100 SXM5s, two H200s, one B200 or one B300.

Serving it means installing the aleph-alpha-inference package, which provides the Kolibri vLLM plugin and pins the vLLM version it supports, or pulling the container image ghcr.io/aleph-alpha/aleph-alpha-inference. From there it is a standard vLLM invocation with the FP8 KV cache, the Kolibri reasoning parser and the Kolibri tool-call parser. Contexts beyond 262,144 tokens need an explicit maximum-model-length flag and a position-embedding override. The recommended sampling parameters are temperature 1.0, top-p 0.97 and top-k 128.

The API surface is OpenAI-compatible, which matters more than it sounds: reasoning effort is passed through the chat template as a reasoning_effort value, and tool calls are emitted in the standard tools field. A team that already speaks the chat-completions format can point existing code at a Kolibri endpoint on a machine in their own building without a wrapper layer.

What Kolibri does not have yet

Four absences are worth stating plainly, because launch coverage tends to skip them.

• No independent evaluation. Artificial Analysis has no model page for Kolibri, and no third party has published a reproduction of the vendor's benchmark table.

• No hosted endpoint from the vendor. The release is weights plus a technical report; there is no Aleph Alpha API SKU for Kolibri in the material published with the model.

• No route on OrcaRouter, and none on any aggregator we would name. We probed the catalogue under every spelling of the vendor prefix and the model name and Kolibri is not there — so if you want to call it, you are downloading 78 GB and running a vLLM endpoint yourself. We would rather say that than imply we serve it.

• No multimodal input. Text in, text out, two languages. That is the trade the lab made for depth over breadth, and for a document-processing deployment it is probably the right one, but it is a real boundary.

If what you actually need today is a small mixture-of-experts model you can call without buying two GPUs, the Gemma 4 MoE tier is the nearest thing already on a routed endpoint, at $0.06 per million input tokens and $0.33 per million output on the 26B-A4B variant, with a 262,144-token window. For anyone weighing a self-hosted 78B sovereign deployment against that, the useful comparison is not benchmark rows — it is 78 GB of VRAM and a procurement conversation against a per-token line item. OrcaRouter routes that tier on one OpenAI-compatible key with the provider's list price passed through and nothing added, which is the cheap way to find out whether the workload justifies owning the hardware.

Kolibri itself is the more interesting artifact. A 78-billion-parameter German-English model under Apache 2.0, with the data pipeline, the tokenizer method, the energy figure and a three-month iteration story published next to the weights, is a different kind of release from the usual weight drop — the disclosure is part of the product, not a blog post about it. Whether the Pareto claim holds up is now a question for people with two H100s and a stopwatch, and the answer will not come from anybody who wrote the model card.