A hero title card reading 'Kolibri vs Intern-Decision 4B' in large bold type, with a single smaller line below reading 'generated text against a calibrated probability distribution', and two small flat line icons side by side — a hummingbird on the left and a small bell-curve chart on the right — separated by a thin vertical divider, set on a white background with soft blue-and-cyan gradient accents and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Kolibri vs Intern-Decision 4B: One Writes Documents, the Other Refuses to Write Anything at All

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most useful thing to notice about Kolibri and Intern-Decision-4B is that they are not the same kind of object, and treating this as a model-versus-model comparison would be a category error. Kolibri is Aleph Alpha's 78.1-billion-parameter mixture-of-experts language model, released on 3 October 2026, which activates 3.46 billion parameters per token and writes German and English text. Intern-Decision-4B is a 4.54-billion-parameter multimodal structured decision model from the InternLM team, uploaded to Hugging Face on 26 September 2026, whose card states plainly that it "does not call generate() or sample free-form text" at all. One produces prose. The other produces a probability distribution over options you supply, in a single forward pass, and has no ability to write a sentence. They only become competitors in one narrow situation, and knowing exactly which situation that is happens to be the most useful thing in either release.

Two releases, two entirely different terse claims

Kolibri's card is a specification of a large sparse language model: 50 layers, every one mixture-of-experts, 384 experts per layer with one shared and six routed, a native context of 262,144 tokens validated to 1,048,576, four reasoning effort levels, Hermes-style tool calling with a shipped vLLM parser, FP8 weights in 128×128 blocks, and Apache 2.0 terms. Its training ran 20 trillion tokens on 768 NVIDIA B200s across 21 days — 392,000 GPU-hours at a reported 6.4×10²³ FLOPs and an estimated 9.5×10² MWh including data-centre overhead, excluding supervised fine-tuning and reinforcement learning. The card's minimum serving configuration is two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300, and the FP8 footprint is about 78 GB. It is a serious piece of infrastructure, and it answers questions by generating text.

Intern-Decision-4B is the opposite kind of object and the card is refreshingly blunt about it. It is a fine-tune of Qwen3.5-4B in which the vision tower and projector are frozen and only the language backbone is trained. You hand it a state, a schema of named questions, and optionally up to eight images, and it returns a distribution over the candidate answers for every question at once. The mechanism is unusual and worth stating precisely: options are mapped to single-token symbols spanning A through Z, a through z and 0 through 9 — 62 candidates, so a maximum of 62 options per question — the prompt is rendered with one placeholder per field, and the model reads logits at the position immediately before each placeholder. Softmax runs over only that field's allowed symbol logits, a checkpoint-specific temperature calibrates the result, and the symbols map back to your original option values as typed JSON. There is no decoding loop, no sampling and no prose.

• Output — Kolibri: freely generated German or English text, tool calls, reasoning traces. Intern-Decision-4B: typed JSON answers with a probability per option.

• Parameters — Kolibri: 78,103,074,560 total, 3,457,573,120 active. Intern-Decision-4B: 4.54 billion across four safetensors shards in bfloat16, with a frozen vision tower.

• Input — Kolibri: text only. Intern-Decision-4B: text plus up to eight images.

• Question capacity — Intern-Decision-4B: up to 62 options per field, in one forward pass, with 'choice', 'score' and 'noul' (yes/no) question types. Kolibri: unbounded in principle, one token at a time in practice.

• Language — Kolibri: German and English by construction. Intern-Decision-4B: whatever Qwen3.5-4B carries, with all published evaluation suites in English.

• Latency — Intern-Decision-4B: 44.16 ms mean, 44.03 ms median and 44.60 ms P95 per query on a single RTX 4090, per its own measurement. Kolibri: no published per-query figure at all.

• Licence — both Apache 2.0; Intern-Decision-4B ships with the Qwen licence file preserved alongside it.

A two-column comparison scoreboard titled 'Kolibri vs Intern-Decision 4B'. Left column Kolibri: Output freely generated text, Parameters 78.1B MoE / 3.46B active, Input text only, Context 262,144 native / 1M validated, Latency none published, Licence Apache 2.0. Right column Intern-Decision 4B: Output typed JSON with a probability per option, Parameters 4.54B bfloat16 with a frozen vision tower, Input text plus up to eight images, Options up to 62 per field in one forward pass, Latency 44.16 ms mean on one RTX 4090, Licence Apache 2.0. A footer line reads 'Kolibri figures vendor-reported; Intern-Decision 4B figures per its model card.' with the OrcaRouter logo in the bottom-right corner.

Why a 4B scorer exists at all

The case for Intern-Decision-4B is not that it is small. It is that asking a large generative model for a probability is not the same as measuring one, and the difference is measurable.

The card's own benchmark table makes the argument with an unusual amount of self-disclosure. Across seven accuracy suites, Intern-Decision-4B averages 90.02 — against 88.74 for the strongest baseline it compares itself to, a model called Jev. But accuracy is the uninteresting half. The table also reports Brier score and expected calibration error, and there the picture is more exacting. The 4B posts a Brier of 0.347 and an ECE of 0.065, against Jev's 0.358 and 0.095. The smaller siblings are where the calibration story becomes genuinely informative. Intern-Decision-2B scores 84.68 average, well ahead of the 0.8B at 79.38 — and yet its ECE of 0.100 is worse than the 0.8B's 0.066 and worse than the 4B's 0.065, with a Brier of 0.437 against the 0.8B's 0.530. The most accurate of the two small models is the least honest about its own confidence. If you had assumed calibration improves with accuracy, or with parameter count, that table is the counterexample on both counts.

A second, separate diagnostic covers 96 cases built with exact reference distributions rather than sampled hard labels — random draws, composed events, conditional history, probability puzzles. The 4B model's error on that pilot came down from 0.628 to 0.550 on the metric reported, while the comparison model sat at 0.595. The card is explicit that this pilot was not used to fit or select the published temperature; the 4B's temperature of 1.992418 was fitted separately on a calibration set. The InternLM team also shipped the benchmark itself into the GitHub repository — the deterministic generator, the references and the offline scorer — which means the calibration claim is the rare kind a third party can actually re-run rather than take on faith.

Nothing in Kolibri's release offers an equivalent. Its comparison table reports task accuracy across fourteen models on one vendor harness, and accuracy is what a generative model can be measured on. There is no calibration figure, no Brier or ECE anywhere in the material, and no way to ask it for "the probability that this contract clause is enforceable" that does not involve sampling a sentence and reading it.

The one seam where they actually meet

Put the two side by side in a real pipeline and the shapes become obvious. A German document workflow needs both halves, and neither tool covers the other's half.

Kolibri is the half that reads and writes. A team with German contracts or technical documentation gets a 262,144-token native window, an FP8 KV cache, a bilingual tokenizer that Aleph Alpha reports at 4.90 average bytes per token on German web text, and Hermes tool calling — which is to say, a model that can summarise a filing, draft a reply, and drive a tool loop. What it cannot do is tell you its own confidence in a way you can audit, because everything it emits is a sampled string.

Intern-Decision-4B is the half that decides. Given a page of German text and a fixed set of options, it returns a calibrated distribution in 44 milliseconds on one consumer GPU, with no generation step and therefore no sampling variance at all. What it cannot do is produce the German text in the first place, and its published evaluation suites are English — its German behaviour is inherited from the Qwen3.5-4B base rather than trained for, which is a real limitation for a German-language deployment and one nobody has evaluated.

So the honest engineering answer for a regulated German-language workflow is that these are two stages of one pipeline, not two candidates for one slot. The practical question is where each one runs. Kolibri needs two H100s or a B200 and about 78 GB. Intern-Decision-4B needs one RTX 4090 and runs a query in under 45 milliseconds. The cost asymmetry is roughly two orders of magnitude on hardware, and it points at a design where a small scorer runs continuously on every case and the large generator is invoked only when a case actually needs prose. That is a cheaper and more audit-friendly shape than routing every case through the 78B model to get an answer you then have to interpret.

A screenshot of the Hugging Face model card for internlm/Intern-Decision-4B, showing the card header, the description of it as a multimodal structured decision model fine-tuned from Qwen3.5-4B that returns an answer distribution in one forward pass, and the numbered 'How inference works' steps mapping each question's options to single-token symbols.

The hosting reality for both, and what the layer is for

Neither model is available as a hosted API, and we checked rather than assumed. Intern-Decision-4B has no vendor endpoint; the release is weights, a GitHub repository and a Hugging Face Space. Its download and like counts have moved since mid-week, which tells you people are picking it up, but there is still no paid callable route anywhere we would name.

Kolibri likewise has no Aleph Alpha API SKU — the release is weights, a technical report and a container image at ghcr.io/aleph-alpha/aleph-alpha-inference. And it is not on OrcaRouter: we probed the catalogue under every spelling of the vendor prefix and the model name and it returns not found, which we would rather say than imply otherwise.

What a routing layer changes here is not access to these two models, it is the economics of deciding whether to buy the hardware for either. The honest pre-purchase test for the generative half is to run the workload against a small mixture-of-experts tier that is already routed — the Gemma 4 26B-A4B variant at $0.06 per million input tokens and $0.33 per million output with a 262,144-token window and text, image and video input — on one OpenAI-compatible key at the provider's list price with nothing added, and see whether the German document workload actually needs 78 billion parameters before the GPUs are ordered. The decision half has no such shortcut, because the calibrated-distribution behaviour is the whole point of the model and no routed tier does it. But that is itself the finding: it tells you the 4B scorer is the part you will genuinely self-host, and the generator is the part worth re-testing against something you can rent this afternoon.

What would settle it

Two measurements, and neither exists yet.

The first is Intern-Decision-4B on German input. Every published suite is English and the model is a fine-tune of a multilingual base, so its German calibration is unknown — and calibration is exactly the property that does not transfer for free across languages. A German version of the 96-case distribution pilot would be the single most informative artifact anyone could publish about this model.

The second is Kolibri's cost per decision. Its serving-economics claim is a Pareto frontier of quality against decoded tokens per second per GPU, and there is no Artificial Analysis page for Kolibri and no third-party replication of the harness. Until someone runs it, "78 billion parameters" is a specification rather than a cost, and the case for keeping a 44-millisecond scorer in front of it stays a design argument rather than a measured one.

Until then, the right way to read this pairing is not as a contest. Intern-Decision-4B is the more technically interesting release of the two, because calibration-aware structured output is a capability almost no generative model offers and its card lets you check the claim. Kolibri is the more consequential release, because a European lab shipping an Apache 2.0 model at 78 billion parameters with the data pipeline published alongside it is a supply-chain event rather than a model event. Neither replaces the other. Teams that need both should plan for both and size the hardware accordingly, and the harness around them is where the engineering effort actually goes.

A screenshot of the InternLM 'Intern Large Models' organisation page on Hugging Face, showing the organisation header, its stated affiliation with Shanghai AI Laboratory, and a recent-activity feed listing models published within the last hour.