
Kolibri vs Granite 4.2 3B: A 78B Sparsity Bet, and the 3B Model That Refuses to Play the Same Game
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 100 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Put Kolibri and Granite 4.2 3B next to each other and the first honest observation is that this is not a fair fight, and not in the direction you would assume. Kolibri is Aleph Alpha's 78.1-billion-parameter mixture-of-experts model, released on 3 October 2026, activating 3.46 billion parameters per token. Granite 4.2 3B is IBM's roughly three-billion-parameter dense reasoning model, whose weights reached Hugging Face on 7 August 2026 with the model card and technical blog following on 25 August. One is twenty-six times the other by total parameters. The reason they belong in the same decision is that both are Apache 2.0, both are text-only, both are self-hosted by design, and both were aimed at the same buyer: a team that wants document intelligence on hardware it controls, with the provenance to survive a procurement review. The interesting question is not which one is better. It is what the twenty-six-times parameter budget actually buys, and what it costs to carry.
The mismatch, stated plainly
Granite 4.2 3B is a standalone dense model post-trained from Granite-4.1-3B-Base, part of the Granite 4.2 family IBM shipped through August. It has 40 layers with grouped-query attention, a 128K native context that IBM extends to 512K in a fifth pre-training phase, and three per-query thinking modes: full thinking by default, a low-effort path, and a non-thinking path. IBM's card is unusually candid about what the 3B is not. Unlike its 8B and 30B siblings, it deliberately skipped the specialized agentic reinforcement-learning block trained in SWE-agent, terminal and search environments, which is why IBM lists no SWE-bench figure for it at all and frames the model as a reasoning specialist rather than an agent.
Kolibri goes the other way on every axis. Fifty layers, every one of them mixture-of-experts, 384 experts per layer with one shared and six routed, and a sparsity ratio of about 22.6 to 1. Its context is natively 262,144 tokens, validated to 1,048,576, with the card recommending you stay at or below 262,144 for latency-sensitive work. Four reasoning effort levels. Hermes-style tool calling with a vLLM parser shipped in the same repository as the weights. And a footprint of roughly 78 GB in FP8, with the card's minimum configuration being two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300.
• Parameters — Kolibri: 78,103,074,560 total, 3,457,573,120 active per token. Granite 4.2 3B: approximately 3B dense, all parameters active on every token.
• Context — Kolibri: 16,384 trained, 65,536 mid-trained, 262,144 native, 1,048,576 validated. Granite 4.2 3B: 128K native, extended to 512K.
• Thinking modes — Kolibri: none, low, medium, high, set through the chat template. Granite 4.2 3B: full, low-effort and non-thinking, per query.
• Tool calling — Kolibri: Hermes-style, with a shipped parser. Granite 4.2 3B: yes, but not the agentic-RL-trained path its larger siblings got.
• Languages — Kolibri: German and English, by design and nothing else. Granite 4.2 3B: English-first, with IBM's broader multilingual training behind it.
• Footprint — Kolibri: about 78 GB in FP8, two GPUs minimum. Granite 4.2 3B: roughly 6–8 GB in bfloat16, under 2 GB quantized, laptop-class.
• Licence — both Apache 2.0, both without an acceptable-use rider or a monthly-active-user threshold.
• Serving — Kolibri: the aleph-alpha-inference vLLM plugin or the published container image. Granite 4.2 3B: vLLM, SGLang, Transformers, GGUF and Ollama under granite4.2:3b.

What the extra 75 billion parameters buy
Three things, and it is worth being precise about which of them are established and which are claimed.
The first is German. This is the sharpest real difference between the two models and it is not a benchmark row. Aleph Alpha built Kolibri around a German-English bilingual corpus, targeting roughly 20 per cent German in a 20-trillion-token pre-training run, and ended with a 2.4-trillion-token German pool 80 per cent of which the lab curated or generated itself. The card explains why that took work: deduplicated open German datasets supplied only 390 billion tokens, an order of magnitude short, so the lab retuned a Common Crawl filter for German specifically and rephrased existing German documents into encyclopaedia entries, dialogue and passages. The filter detail is the one to remember. A standard language-data pipeline drops documents with too many long words, and German administrative prose routinely exceeds the English bound on mean word length — so default settings quietly delete the register that public administration writes in. IBM did not build Granite for that corpus. Granite 4.2 3B will handle German; it was not designed around German legal and administrative register, and no leaderboard will tell you the difference.
The second is long context that survives real documents. Granite's 512K ceiling is genuinely large, but the two models got there differently and Kolibri's positional design is the more conventional long-context argument. Treat IBM's RULER figures — 67.52 at 64K and 55.30 at 128K in the 4.2 family's published material — as the honest disclosure of how much retrieval quality decays by 128K, and remember that Kolibri has no equivalent published degradation curve at all.
The third is raw reasoning headroom, and this is where the honest answer is "not as much as the parameter ratio suggests". Kolibri's training pipeline bought it a 20-trillion-token pre-training run on 768 NVIDIA B200s over 21 days, and Aleph Alpha's own comparison table, with Kolibri at reasoning effort high, puts it at 75.5 on the English average and 70.8 on the German average of a fourteen-model comparison — where it loses to a dense 27-billion-parameter model from Alibaba on most rows. Granite 4.2 3B's headline claims are its own: AIME 2025 at 78.33, GPQA at 54.80, LiveCodeBench v6 at 69.71 and MMLU-Pro at 67.84, all IBM-reported and unreproduced. Different suites, different harnesses, different vendors. There is no same-harness number for this pairing anywhere, and we are not going to invent one.
Where each one actually wins
Run Granite 4.2 3B if your constraint is the machine. At 6–8 GB in bfloat16, or under 2 GB quantized, it fits a laptop, a single workstation GPU, or an air-gapped edge box that will never see two H100s. It serves through five different runtimes including Ollama and GGUF, which matters when the deployment target is someone else's laptop rather than your own rack. And its card is unusually trustworthy precisely because IBM wrote down what it left out. A vendor that declines to claim SWE-bench results for a 3B is telling you where the model stops.
Run Kolibri if the corpus is the point and the hardware exists. A team with German-language contracts, technical documentation or administrative filings, an existing two-GPU node, and a requirement that the weights never leave the building is exactly who this release was designed for. The 262,144-token native window, the FP8 KV cache, the four effort levels and the Hermes tool-calling path all point at document workflows rather than chat. The bilingual tokenizer is part of the same argument: Aleph Alpha reports 4.90 average bytes per token on German web text against 4.35 for GPT-5 and 3.28 for Kimi K3, all vendor-measured on the vendor's corpus, and more characters per token is a direct inference-cost effect rather than a score. If it holds, it compounds across every page you process.
The asymmetry nobody advertises is data. IBM published Granite 4.2 3B's weights and a detailed technical account of how the model was built, but not the training data. Aleph Alpha published the pipeline, the data provenance and the energy figure alongside Kolibri's weights — 20 trillion pre-training tokens, 9.5×10² MWh including data-centre overhead and excluding supervised fine-tuning and reinforcement learning. For a team that has to answer "where did this model's text come from", that difference is not cosmetic.

The routing reality for both
Neither model sits in a hosted catalogue today. Kolibri has no vendor API SKU at all — the release is weights plus a technical report — and Granite 4.2 3B ships as weights for five runtime stacks with no hosted endpoint from IBM. Both are self-hosted propositions, and the practical question for most teams is not which one to adopt but whether the workload justifies owning either.
That is where a routing layer earns its place even for models it does not carry. We probed the OrcaRouter catalogue under every spelling of both model names and neither Kolibri nor Granite 4.2 3B is there, so we will not pretend otherwise. What OrcaRouter does give you is the cheap way to find out whether the hardware decision is justified before you make it: point a test path at a small mixture-of-experts tier that is already routed — the Gemma 4 26B-A4B variant at $0.06 per million input tokens and $0.33 per million output, with a 262,144-token window — and see whether your German document workload actually needs what Kolibri provides, on one OpenAI-compatible key, at the provider's list price with nothing added. If it does, you buy the GPUs with evidence rather than a hunch. If it does not, you just saved a hardware order.
The verdict, and what would change it
This is not a matchup with a winner. Kolibri and Granite 4.2 3B answer different questions at different price points, and the only honest ranking is by constraint: if the constraint is hardware, Granite 4.2 3B is the only one of the two that qualifies. If the constraint is German regulatory-register document work on premises, Granite 4.2 3B was never in the running and Kolibri is the more interesting artifact — a legitimate 78B open-weights release, Apache 2.0, with the data pipeline and tokenizer published next to the weights.
Two things would settle the comparison. An independent run of Kolibri on a German document-QA task would test the claim the release was actually built to make, since no leaderboard currently measures it. And an independent reproduction of Granite 4.2 3B's reasoning numbers would tell you whether a laptop-class box can hold its own on the subset of your workload that never needed 78 billion parameters in the first place. Until one of those lands, buy by constraint, not by parameter count.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
