Generated hero card for Kolibri vs Gemma 4 12B, headed 'Kolibri vs Gemma 4 12B' with the subline 'a 78B MoE against an 11.95B dense model', a hummingbird icon on the left and a diamond on the right either side of a thin divider. The OrcaRouter logo sits in the strip below the artwork.
Guides & Insights

Kolibri vs Gemma 4 12B: 78 Billion Parameters Against 12, and Only One of Them Has a Score

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Kolibri and Gemma 4 12B are the two ends of a spectrum that open-weight releases do not usually let you compare on equal terms. Aleph Alpha's Kolibri shipped on 3 October 2026: 78.1 billion total parameters, 3.46 billion active per token, a validated 1,048,576-token context window, German and English only, Apache 2.0. Gemma 4 12B shipped on 3 June 2026: 11.95 billion dense parameters, a 262,144-token context window, text, image, audio and video input, Apache 2.0. One of them has an independent, publicly reproducible score. The other one has a vendor benchmark table. That asymmetry is not a footnote to this comparison; it is the comparison, because everything else a buyer cares about — cost per task, quality per watt, whether it will work on your documents — runs through it.

We should be equally direct about the shape of the matchup, because a headline pairing a 78B model against a 12B model invites the wrong conclusion. These are not competitors. One is a 78 GB deployment aimed at German public-sector and industrial document work on racks the customer owns; the other is a 12-billion-parameter unified multimodal model that runs on a laptop and carries an independent index of 14. What follows is an account of what each is actually for, where the published numbers do and do not let you rank them, and what it costs to not have an independent score.

The one number you can trust here, and the one you cannot

Artificial Analysis has measured Gemma 4 12B, in its reasoning configuration. The findings, as published on their model page, are the ones to work from:

• Intelligence — an Artificial Analysis Intelligence Index of 14, which places it in the top class band with a #21 position across the 142 models on that comparison, against a median of 8 for comparable models.

• Speed — 112.5 output tokens per second, #21 of 142 on the speed view, and above the 91-token median for comparable models.

• Price — $0.10 per million input tokens and $0.30 per million output, which their page flags as somewhat expensive against a median of $0.03 and $0.15 for models of similar size and class.

• Specification — 262,144-token context, 12B total parameters, Apache 2.0, text, image, speech and video input, text output.

Artificial Analysis's own summary verdict is unusually blunt for a leaderboard page, and worth repeating: Gemma 4 12B is among the leading models in intelligence but somewhat expensive compared to other open-weight models of similar size. That is a real, independent, reproducible assessment from a third party with no stake in either vendor.

Kolibri has nothing equivalent. There is no Artificial Analysis model page for it, and at the time of writing no third party has reproduced the eval suite. What exists is Aleph Alpha's own comparison table, on which Kolibri scores 75.5 on the English average and 70.8 on the German average across its task suite. Those are unweighted percentage averages over a benchmark mix the vendor chose, on a harness the vendor wrote, at reasoning effort high. They are not an Intelligence Index, and the two numbers cannot be converted into each other or set side by side. Anyone presenting "Kolibri 75.5 versus Gemma 14" as a ranking is comparing a percentage on one scale to a score on another. The honest position is that Gemma 4 12B's quality is measured and Kolibri's is claimed.

What the vendor table does let you do is read across rows. Gemma 4 26B-A4B IT, Google's own mixture-of-experts sibling to the 12B, appears in Aleph Alpha's table at 71.9 English and 66.3 German, below Kolibri on both. That is the closest thing to a cross-check available, and it comes with the same caveat: vendor harness, vendor numbers. It is a weak signal, not evidence.

Screenshot of the Artificial Analysis model page for Gemma 4 12B (Reasoning), marked 'Open weights model, Released June 2026', showing an Intelligence Index of 14 against a median of 8, a #21/142 intelligence rank and a #21/142 speed rank, 112.5 output tokens per second against a 91-token median, input pricing $0.10 per million tokens against a $0.03 median and output pricing $0.30 against a $0.15 median, a 262k-token context window, the summary verdict that the model is amongst the leading models in intelligence but somewhat expensive compared with other open-weight models of similar size, and the 142 models in this class count.

Dense 12B against sparse 78B is not the mismatch it looks like

The architecture gap is larger than the parameter gap once you account for what actually gets computed per token.

• Compute per token — Gemma 4 12B activates all 11.95 billion parameters on every token. Kolibri activates 3,457,573,120 of its 78,103,074,560, roughly one twenty-second of the model, with one shared expert and six routed experts per layer out of 384.

• Memory — the trade runs the other way and it is brutal. Gemma 4 12B in bfloat16 is on the order of 24 GB of weights. Kolibri's card puts the FP8 weights at roughly 78 GB, with a minimum of two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300.

• Attention — Gemma 4 uses a hybrid of local sliding-window attention and full global attention, with the final layer always global. Kolibri uses a 4:1 sliding-window to grouped-query split across 50 all-MoE layers.

• Context — 262,144 tokens for Gemma 4 12B; 262,144 native for Kolibri, validated to 1,048,576 without position scaling.

So the honest framing of "78B against 12B" is that Kolibri wins on activation cost and loses on weight footprint, and neither advantage is free. A sparse model that activates 3.46 billion parameters per token still has to hold all 78 billion in memory, which is why the deployment floor is a pair of 80 GB accelerators rather than one consumer card. Gemma 4 12B makes the opposite trade: a larger share of the model fires every token, but it fits on hardware a person can buy.

There is a German-specific wrinkle on the Kolibri side that changes the arithmetic rather than the ranking. Its tokenizer averages 4.90 bytes per token on German web text, against 4.13 for Gemini, 4.17 for the Qwen3.5–3.8 family and 4.35 for GPT-5, on Aleph Alpha's own measurements. Fewer tokens per German document is a direct reduction in inference cost, and for a workload that is German administrative text by the million, that effect can be worth more than a few index points. It is also entirely vendor-reported.

Generated two-column scoreboard for Kolibri and Gemma 4 12B. Left column Kolibri: independent score 'none published', context '262,144 native, 1M validated', parameters '78.1B MoE, 3.46B active', modalities 'text only', licence Apache 2.0, price 'self-host, 78 GB'. Right column Gemma 4 12B: independent score 'AA Index 14', context 262,144, parameters '11.95B dense', modalities 'text, image, audio, video', licence Apache 2.0, price '$0.10 in / $0.30 out'. The footer reads 'Kolibri figures vendor-reported; Gemma 4 12B figures per Artificial Analysis.'

What each one can actually see

This is where the matchup stops being close, and it is a specification fact rather than a quality claim. Gemma 4 12B is a unified multimodal model: its card describes an encoder-free architecture that projects raw image patches and audio waveforms directly into the language model's embedding space, with text, image, speech and video in and text out. It is built to run on consumer devices with no separate encoders to provision.

Kolibri reads text, in two languages, and nothing else. Aleph Alpha frames that as depth over breadth deliberately: two languages, heavily curated, rather than a wide multilingual mix. There is no vision tower, no audio path, no video. If your workload includes scanned forms, recorded calls or photographs of equipment, Gemma 4 12B is a candidate and Kolibri is not, whatever the benchmark rows say. If your workload is a corpus of German and English documents that must never leave the building, the reverse is closer to true — but note that the "never leaves the building" requirement is satisfied by Gemma 4 12B too, since the weights are Apache 2.0 and downloadable.

Which one is cheaper depends on what you are buying

The two models are priced in currencies that do not convert. Gemma 4 12B has a market rate: $0.10 and $0.30 per million tokens as measured across providers by Artificial Analysis, and cheaper at the MoE tier on routed endpoints. Kolibri has no rate at all, because Aleph Alpha is not selling inference for it — the release is weights, a technical report and a model card.

• Gemma 4 12B on a hosted route — $0.10 per million input and $0.30 per million output on the Artificial Analysis figures. On our own catalogue, the MoE sibling Gemma 4 26B-A4B IT is listed at $0.06 in and $0.33 out with a 262,144-token window, and the larger dense Gemma 4 31B IT at $0.13 and $0.38; those are provider list prices passed through, which is the only number we add nothing to.

• Kolibri on your own hardware — 78 GB of weights, a vLLM plugin from the vendor's aleph-alpha-inference package, and a floor of one H200 or B200, or a pair of H100 SXM5s. The cost is a capital purchase, a rack slot and a person who can run a vLLM deployment, amortised over as many tokens as you generate.

Those are not the same purchase and the comparison is not "which is cheaper". It is whether the workload has a shape that justifies owning the hardware. A team processing a fixed, sensitive document corpus at high volume can come out far ahead on the self-hosted side. A team building a product feature that calls a model a few million tokens a day almost never does.

One practical middle path: the Gemma tier is on OrcaRouter today, on the same OpenAI-compatible endpoint as everything else, with failover across providers and the provider's list price passed through with nothing added per token. That makes it a cheap way to measure your actual token volume on a hosted route before deciding whether a 78 GB sovereign deployment is justified. It is worth saying plainly what it is not: we do not route Kolibri. We probed the catalogue under every spelling of the vendor and model name, and it is not there. Self-hosting is currently the only way to call it.

Screenshot of the OrcaRouter model page for google/gemma-4-26b-a4b-it, showing the 262K-token context badge, text, image and video input with text output, the Vision, Tools, JSON and Reasoning capability tags, the attribution 'Public benchmarks by Google - 2026-04-05', the description of Gemma 4 26B A4B IT as an instruction-tuned mixture-of-experts model from Google DeepMind with roughly 26B total parameters of which about 4B activate per token, the pricing tiles $0.06 and $0.33, and the OpenAI-compatible base URL https://api.orcarouter.ai/v1.

Who should pick which

The decision rule is short enough to state in three lines.

Pick Gemma 4 12B if you need anything other than German and English text, if you need a measured quality number before you commit, if the model has to run on hardware you already have, or if you would rather rent tokens than own a rack. It is the only one of the two with an independent score, a published speed, and a market price.

Pick Kolibri if your requirement is a 78-billion-parameter reasoning model whose weights, training data provenance, energy consumption and licence are all documented, deployed inside your own perimeter, with a tokenizer built for German administrative prose and an abstention behaviour that a regulated workflow can rely on. That is a narrow target and Aleph Alpha aimed at it deliberately.

Do not pick either on the strength of a benchmark row you saw in a launch post. Gemma 4 12B's 14 is a measurement someone else made and you can check. Kolibri's 75.5 is a claim its author made, and the only person who can currently falsify it is a customer with two H100s and a document corpus.