
Kolibri vs MiniCPM5-2B: 78 Billion Parameters Against 2.5, and Why the Small One Is Not the Cheap Option
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 100 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Set Kolibri beside MiniCPM5-2B and the instinctive reading is that this is a size comparison, and that the 2.5-billion-parameter model is the economical choice. That reading is wrong in an instructive way. Kolibri is Aleph Alpha's 78.1-billion-parameter mixture-of-experts model, released on 3 October 2026, activating 3.46 billion parameters per token, with an FP8 footprint of about 78 GB and a minimum serving configuration of two A100 80 GB cards. MiniCPM5-2B is ModelBest and OpenBMB's 2.52-billion-parameter dense model, unveiled at the World Artificial Intelligence Conference on 19 July 2026 with weights landing on Hugging Face on 6 September. MiniCPM5-2B is thirty-one times smaller by parameter count. It is also the one that demands an evaluation budget before you trust it, because its most quoted numbers are vendor-run and unreproduced — which is precisely the opposite of what "small model" usually implies about risk.
Both shipped quietly; only one of them shipped recently
They have similar origin stories and very different ages, and the ages matter. Aleph Alpha's repository for Kolibri was created on 2 October 2026 with a stated release date of 3 October, and the vendor's own newsroom announcement went out that morning, on the Day of German Reunification. The general AI press largely missed it on the day, which is where the "quiet release" framing came from, but a model card, a technical report and a vendor post are not a silent drop. MiniCPM5-2B genuinely was quieter: the WAIC announcement in July was a conference unveiling of a pitch and specs, and the actual weights did not appear until 6 September, published without a launch event, alongside GGUF, MLX, GPTQ, base, SFT and draft checkpoints and the training corpora behind them.
That five-week gap between announcement and weights is worth remembering, because it is why two different sets of numbers circulate for this model. The July material touted a 512K-token context window. The shipped config sets 131,072. The released artifact is the one that counts.
What each model is actually for
Kolibri is built for German and English document work, and Aleph Alpha is unusually specific about why that took real engineering. Small proxy-model experiments set a target of roughly 20 per cent German in the training mix, which for a 20-trillion-token run meant finding 4 trillion German tokens. Deduplicated and filtered open German datasets yielded 390 billion — an order of magnitude short — so the lab retuned a Common Crawl filter for German specifically, producing 1.3 trillion unique organic tokens, and rephrased existing German documents into encyclopaedia entries, dialogue and passages for roughly another trillion, which became the largest single German source. Translation, used for the predecessor Kolibri Origin, was dropped for Kolibri. The filter detail is the one to hold on to: a standard language-data pipeline discards documents with too many long words, and German administrative prose routinely exceeds the English bound on mean word length, so default settings quietly delete the register that public administration writes in.
MiniCPM5-2B was built for the opposite end — an on-device agent. The card describes a dense transformer of 2.52 billion total parameters, 1.98 billion non-embedding, 42 layers at hidden size 2048, grouped-query attention with 16 query heads and 2 KV heads, and a standard LlamaForCausalLM architecture with no custom kernels. The training pipeline leans agentic: 200-billion-scale agent mid-training, a large agent-SFT set, agent RL alignment, and a final On-Policy Distillation stage that folds sixteen RL expert models back into one small dense network. It ships XML-style tool calls with a built-in SGLang parser, and its released config sets a 131,072-token window.
• Parameters — Kolibri: 78,103,074,560 total, 3,457,573,120 active, 22.6:1 sparsity. MiniCPM5-2B: 2,516,756,480 total, 1.98B non-embedding, dense.
• Released — Kolibri: 3 October 2026. MiniCPM5-2B: announced 19 July 2026, weights 6 September 2026.
• Context — Kolibri: 16,384 trained, 65,536 mid-trained, 262,144 native, 1,048,576 validated. MiniCPM5-2B: 131,072 in the shipped config; the 512K claim from July is not in it.
• Footprint — Kolibri: about 78 GB in FP8; two A100 80 GB, two H100 SXM5, one H200, one B200 or one B300 minimum. MiniCPM5-2B: a single BF16 shard, laptop-class.
• Languages — Kolibri: German and English, by design and nothing else. MiniCPM5-2B: English and Chinese, with all published evaluation suites in English.
• Tool calling — Kolibri: Hermes-style with a shipped vLLM parser. MiniCPM5-2B: XML-style with an SGLang parser.
• Licence — both Apache 2.0, both without an acceptable-use rider.

The scoreboards, and the sourcing behind them
There is no head-to-head number for this pairing anywhere, in either vendor's material, and we are not going to assemble one from two different harnesses and call it a comparison. What each side publishes is worth separating carefully, because the two sets of figures carry very different amounts of evidence.
Kolibri's numbers come from Aleph Alpha's own fourteen-model comparison table, run on one harness with Kolibri at reasoning effort high: 75.5 on the English average, 70.8 on the German average. That table does not flatter its subject — Qwen3.8 27B scores 80.2 and 79.9 on the same two averages and leads on GPQA Diamond, LiveCodeBench v6, SWE-Bench Verified and both long-context benchmarks, while activating roughly an eighth of Kolibri's parameters. The vendor's framing for that gap is a Pareto frontier of quality against decoded tokens per second per GPU, which none of the compared models beats. It is a serving-economics argument, not a capability argument, and no third party has measured it: there is no Artificial Analysis entry for Kolibri and no replication of the harness.
MiniCPM5-2B's numbers come from its own card: an internal average of 53.9 across a vendor-selected comparison set, AIME 2025/2026 at 86.5, MATH-500 at 94.6, and a vendor-run 46.4 on SWE-bench Verified that would be remarkable for a 2.5-billion-parameter dense model if it survives independent testing. On that card, IBM's Granite 4.2 3B sits at 42.7 against MiniCPM5-2B's 53.9 — one vendor running both models on one suite it chose. Different suites, different harnesses, different vendors. None of it is a verdict on this pairing.
What is genuinely useful on MiniCPM5-2B's side is the disclosure. OpenBMB released the UltraData training corpora alongside the weights — Ultra-FineWeb, UltraX, UltraData-Code, UltraData-Math, the agent SFT and RL sets — all Apache 2.0, which turns a black box into a recipe you can inspect and continue on your own domain. The model also ships with day-zero adaptation across nine chip families through ModelBest's FlagOS community, from Huawei Ascend to NVIDIA and ARM, which is a deployment statement rather than a benchmark one. Against that, IBM published Granite 4.2 3B's weights and a detailed technical account of the build but not its training data, and Aleph Alpha published Kolibri's data pipeline and energy accounting — 20 trillion pre-training tokens, 9.5×10² MWh including data-centre overhead and excluding supervised fine-tuning and reinforcement learning — but not the corpus itself.
Where the size gap actually shows up
The most useful way to put these two side by side is on the two axes that decide an actual deployment: what has to be in the room, and what happens when the model is wrong.
On hardware, MiniCPM5-2B wins and it is not close. A 2.52-billion-parameter dense model in one BF16 shard runs on a laptop, an edge box or a single consumer GPU, and FlagOS support across nine chip families means it runs on hardware that is not an NVIDIA accelerator at all. Kolibri's floor is two A100 80 GB cards or one B200, roughly 78 GB of FP8 weights, served through the aleph-alpha-inference vLLM plugin or the published container. For a smart-cockpit or phone-adjacent workload, the 2B is the only one of the two that qualifies, and Kolibri was never designed to compete there.
On verifiability, Kolibri wins for a subtler reason. Its most quoted claim — the serving economics — is untested, but its quality figures are at least published against thirteen named competitors on one harness with the settings disclosed, and the vendor's own table hands the win on most rows to an opponent. MiniCPM5-2B's headline numbers are the vendor's on the vendor's suite with no comparison harness disclosed, and the model carries the extra uncertainty of a checkpoint that is weeks old rather than days. Neither model has been independently benchmarked. The difference is that one vendor wrote down a comparison that makes its own model lose, and the other wrote down one that makes its own model win by a wide margin.
On purpose, there is no contest at all. Kolibri exists for German-language document processing on premises, with a bilingual tokenizer Aleph Alpha reports at 4.90 average bytes per token on German web text against GPT-5's 4.35 and Kimi K3's 3.28 — all vendor-measured, and a per-token cost effect rather than a quality claim. MiniCPM5-2B exists for tool-calling agents in English and Chinese on constrained hardware. A team with German contracts and a compliance requirement is not choosing between them.

The routing reality, and the experiment that is worth running
Neither model sits in a hosted catalogue today, and we checked both rather than assuming. Kolibri has no vendor API SKU — the release is weights, a technical report and a container image — and it is not on OrcaRouter: we probed the catalogue under every spelling of the vendor prefix and the model name and it returns not found. MiniCPM5-2B has no hosted endpoint from ModelBest either, and it is not routed with us. Both are self-hosted propositions and we would rather say so than imply we serve either one.
Where this gets interesting is the experiment. A 2.5-billion-parameter agentic model and a 78-billion-parameter document model solve different problems, and the only way to know which your workload needs is to run both against it — which is exactly the situation a routing layer is built for, even when it carries neither model. Point a test path at a self-hosted MiniCPM5-2B build through one OpenAI-compatible endpoint with automatic failover, with production left on a proven routed model, and the new stack can stall or produce nonsense without taking anything down. Provider list prices on the models you compare against pass through at zero markup, so the A/B stays cheap and reversible. And if what the experiment actually reveals is that your German workload needs neither self-hosted model, the same key reaches a small mixture-of-experts tier that is already routed — the Gemma 4 26B-A4B variant at $0.06 per million input tokens and $0.33 per million output, with a 262,144-token window and text, image and video input — which is the cheapest way to find out before you buy two GPUs.
Which one, and what would change the answer
Run MiniCPM5-2B if the constraint is the device. At 2.52 billion parameters it is the only one of the two that fits a laptop, a cockpit or an edge box, and if your workload is an interactive tool-calling agent in English or Chinese, that is the model aimed at it. Budget time to reproduce its numbers first — the 53.9 average and the 46.4 SWE-bench claim are ModelBest's alone, and the training corpora being open makes that reproduction genuinely possible rather than theoretical.
Run Kolibri if the corpus is German, the hardware exists, and the weights may not leave the building. A 262,144-token native window, an FP8 KV cache, four reasoning effort levels and Hermes tool calling are the right shape for regulated document work, and the Apache 2.0 terms plus the published data pipeline are the part that survives a procurement review.
Two things would settle this faster than any argument. An independent SWE-bench Verified run on MiniCPM5-2B would confirm or sink the single most surprising claim either card makes. And an independent throughput measurement of Kolibri at its recommended two-H100 configuration — or an Artificial Analysis entry, which does not exist today — would turn "78 billion parameters" from a specification into a cost, which is the number anyone weighing this comparison against hardware is actually looking for. Until then the size gap is real, the purpose gap is larger, and the cheap-looking option is the one carrying the unverified claim.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
