
Kolibri vs Solar Mini 4: The Same 3B Active Budget, Two Very Different Bets
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 116 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 51 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 100 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two models shipped seventeen days apart, both tuned to spend roughly three billion active parameters per token, and they could hardly be less alike to actually adopt. Aleph Alpha's Kolibri is 78.1 billion parameters of Apache-2.0 weights that you download and serve yourself: no hosted endpoint exists, and as of today there is not a single independent score of it anywhere. Upstage's Solar Mini 4 is 35 billion parameters that you cannot download at all — it exists only as a metered endpoint, and it already carries an Artificial Analysis Intelligence Index of 24.05. The active-parameter count tells you what each one costs to run. Nothing else about that number tells you what it costs you to start.
The reason the totals diverge so far — 78.1B against 35B with almost the same active slice — is that both are sparse mixture-of-experts designs, and the two labs made opposite calls about how much expert capacity to keep in reserve. Kolibri carries 384 experts and routes to 1 shared plus 6 of them per token. Solar Mini 4 discloses a 35B/3B split and nothing about its expert count. When two models are advertised on the same active figure, the parameter total, not the active figure, is what decides your hardware bill — and here it is a 2.2× difference on the same stated compute budget.
What each one actually is
• Total parameters — Kolibri 78.10B vs Solar Mini 4 35B
• Active parameters per token — Kolibri 3.46B vs Solar Mini 4 3B
• Weights — Kolibri Apache 2.0, full weights on Hugging Face vs Solar Mini 4 closed; API only
• Context — Kolibri 1,048,576 tokens validated vs Solar Mini 4 512K per Upstage's docs, 1,048,576 per Artificial Analysis's model card
• Max output — Kolibri not capped by the vendor vs Solar Mini 4 128,000 tokens
• Languages — Kolibri German and English vs Solar Mini 4 English, Korean and Japanese
• Knowledge cutoff — Kolibri 18 June 2026 vs Solar Mini 4 February 2026
• Delivery — Kolibri self-hosted (vLLM) vs Solar Mini 4 hosted on Upstage's Console, Playground and on-premises
• Independent evaluation — Kolibri none published vs Solar Mini 4 AA Intelligence Index 24.05
Kolibri: 78 billion parameters, and no one outside the lab has scored it
Aleph Alpha published Kolibri on 3 October 2026, deliberately on the Day of German Reunification, as the first of a new family and the successor to Kolibri Origin. The model card is unusually forthcoming about what went into it: 20 trillion tokens of pre-training, 3.44 trillion of mid-training and 201 billion of long-context training, run across 768 B200 GPUs in 21 days for roughly 6.4×10²³ FLOPs and about 950 MWh. The vendor also volunteers the failure count — 38 unplanned interruptions, about one per 10,000 GPU-hours — which is the kind of number labs usually leave out.

The architecture is a clean sparse-MoE story. Fifty layers with a 4:1 ratio of sliding-window attention (512-token windows) to global attention, which fires only on every fifth layer. FP8 weights with 128×128 block scaling and an FP8 KV cache. Kolibri Origin, for comparison, used 128 experts routed 8-wide, a 768-wide expert hidden dimension and full attention on every layer, on English data cut off in September 2024. Kolibri moves to 384 experts routed 6-wide with a 512-wide hidden dimension, and pushes both language cutoffs to 18 June 2026.
The German pipeline is the part that is genuinely hard to buy elsewhere. Aleph Alpha trained its own UniBPE tokenizer and reports German text at 4.90 bytes per token, against 4.35 for GPT-5, 4.17 for Qwen3.5 and 4.13 for Gemini. That is a tokenizer claiming better German compression than any of the frontier multilingual models, and it is the concrete mechanism behind the sovereign-deployment pitch: fewer tokens per German document means fewer billed tokens and a longer effective window on the same hardware.
Every performance figure that exists for Kolibri comes from Aleph Alpha's own model card, and it should be read that way. The vendor-reported table has Kolibri at an overall 75.5 on English evals and 70.8 on German, 84.3 GPQA Diamond in English against 81.3 in German, 21.5 on Humanity's Last Exam in English against 15.9 in German, 66.4 on SWE-Bench Verified, 85.9 on LiveCodeBench v6 and 68.3 on AA-LCR. Those are unaudited numbers. There is no Artificial Analysis page for Kolibri — the tracker's model URL returns a 404 — so nothing in that table has been independently reproduced, and the article you are reading cannot make it more solid than it is.
What the card does make solid is the serving story. The hardware floor is two A100 80GBs, or two H100 SXM5s, or a single H200, B200 or B300. A published vLLM recipe runs it with --kv-cache-dtype fp8 and dedicated reasoning and tool-call parsers, at temperature 1.0, top-p 0.97 and top-k 128, with a reasoning_effort dial spanning none, low, medium and high. A single B300 serving a 78B model at a 1M-token validated context is the concrete shape of the offer: you buy the box, and then you own the marginal cost of every token.
Solar Mini 4: you rent the 3B active slice instead of buying it
Solar Mini 4 arrived on 22 September 2026 — snapshot solar-mini4-260922, alias solar-mini4 — as Upstage's agent-oriented small model: 35B total, 3B active, a February 2026 cutoff, English, Korean and Japanese, with chat, reasoning, structured-output and tool-calling modes. It is available through Upstage's Console and Playground, and for on-premises deployment, but there is no weight download and no licence to inspect. Upstage's docs also record "Data not collected" for it, which is the compliance line that matters if you are putting it in front of real traffic.

Unlike Kolibri, Solar Mini 4 has been through an independent harness. Artificial Analysis puts it at an Intelligence Index of 24.05, with a measured 1.01% on Terminal-Bench 4.0, 47.6% on SciCode, 83.3% on AA-LCR and 25.8% on Humanity's Last Exam. Upstage's launch post frames the 24.1 as the highest Index among models with 3B active parameters, ahead of Qwen3.6 35B A3B at 18 and Nemotron 3.5 Lightning at 13 — a framing that is fair as far as it goes, and worth noting is exactly as narrow as it sounds, since it is a claim about the 3B-active class rather than about small models generally.
The measurement that should shape how you budget for it is not the Index. Artificial Analysis's per-task cost breakdown puts Solar Mini 4 at roughly $0.36 per Intelligence-Index task, and about $0.30 of that — some 82% — is prompt-cache writes. Cache reads are $0.03 and generated output is only $0.035. The model emits about 88,300 output tokens per task, of which some 71,600 are reasoning tokens. In other words: this is a cheap model whose bill is dominated by re-writing a long context, not by the tokens it writes back. Per-eval costs run from $1.63 for Terminal-Bench 4.0 down to $0.95 for AA-Briefcase. Median output speed is 198 tokens per second at a 1.58-second time to first chunk.
The price of each path
Solar Mini 4 is not priced at list today. Upstage's own pricing page runs a dated ladder for it: $0.03 per million input, $0.003 cached and $0.12 output — a 70% launch discount — through 10 October 2026 UTC, then $0.05 / $0.005 / $0.20 from 11 October, and finally the $0.10 / $0.01 / $0.40 list from 23 October. Upstage's launch post describes the discount more loosely than the pricing page's own schedule does; the ladder above is the version with dates on it, and it is the one worth planning against. Note where the steepest discount sits: cached input falls to a tenth of the input rate, which rewards exactly the long-stable-context workload the cache-write cost profile above charges for.
Kolibri's price is a purchase order. There is no per-million-token rate to compare because there is no hosted meter — you pay for GPUs and electricity, and the vendor's only public cost signal is the training run itself: about 950 MWh to produce the model. If you already own inference capacity, Kolibri's marginal cost per token approaches the cost of the electricity; if you do not, the entry ticket is a two-A100 box or better. That is a fundamentally different shape of commitment from a model you can call by the million tokens and stop calling tomorrow, and it is the actual decision the two options present.
Where they overlap, and where the comparison breaks
• Active compute — near-identical: 3.46B against 3B per token
• Everything that follows from it — not comparable, because only one of the two has any verified evidence
• Best measured score — Solar Mini 4 has an audited Index of 24.05; Kolibri has a vendor table and no audited score at all
• German — Kolibri is the only one of the two trained for it, at 4.90 bytes per token; Solar Mini 4 does not list it
• Korean and Japanese — Solar Mini 4 lists both; Kolibri lists neither
• Long context — Kolibri validates a full 1,048,576 tokens; Solar Mini 4's own docs say 512K while its Artificial Analysis card says 1M, so treat the ceiling as vendor-dependent
• Time to first token — Solar Mini 4 is minutes, because you sign up; Kolibri is days, because you rack a box
• Cost model — per-token, falling with the discount ladder, against capital plus power
The honest summary is that this is not a matchup you settle on a leaderboard. Kolibri's vendor table even includes its own comparison set — GLM-4.7 Flash 30B-A3B at 64.7 overall English, Nemotron 3 Nano 30B-A3B at 65.6, Qwen3.5 35B-A3B at 74.7, Qwen3.6 35B-A3B at 71.4, Gemma 4 26B-A4B IT at 71.9, all against Kolibri's own 75.5 — and every one of those rows is Aleph Alpha's own number, run by the same lab on its own harness. It is a useful map of the 3B-active neighbourhood. It is not an independent ranking.
If what you want is a 3B-active MoE on a key today
Neither model in this comparison is on OrcaRouter, and it is worth being precise about why. Kolibri has no hosted inference of any kind yet — it is weights and a vLLM recipe, so there is nothing for a router to point at. Solar Mini 4 is served by Upstage and by Upstage's own on-premises channel; we do not route it, and we do not route any Upstage model. If you want either one, you get it from the vendors themselves.
What we do route is the surrounding class — the sparse small-MoE bracket that both of these models are competing inside. Gemma 4 26B-A4B IT is on the catalogue at $0.06 in and $0.33 out per million tokens with a 262,144-token window. Qwen3.5 35B-A3B is at $0.057 / $0.459 and Qwen3.6 35B-A3B at $0.248 / $1.485 with a 262,144-token window and 65,536 tokens of max output. Those are the same trade the two models above are making — a 3B-to-4B active slice bought for the price of a small model — available on one API key with the provider's list price passed through at 0% markup, so a vendor price change lands on your bill the same day it is announced rather than at the next contract renewal. Automatic failover matters more here than usual: when you are evaluating an unproven sparse model, a second route behind the same key is what keeps a bad afternoon from becoming an outage.

What to do, and what to watch
Choose Kolibri if German or English is your workload, if the data cannot leave your own rack, and if you have — or are willing to buy — the GPUs to serve 78B at FP8. In that configuration it is the only one of these two models that is even an option, and the 1M-token validated context with a 4.90-byte-per-token German tokenizer is a real, if vendor-measured, advantage. Choose Solar Mini 4 if you want to be in production this week, if Korean or Japanese is on the roadmap, and if your workload is a long stable context that the cache discount rewards; budget for cache writes, not output tokens.
The open question is the same for both, and it is an evidence question rather than a capability one. Kolibri's 75.5 and 84.3 are Aleph Alpha's own numbers on Aleph Alpha's own harness, and until an independent tracker runs it, anyone quoting them is quoting the lab. Solar Mini 4's 24.05 is real and reproduced, but the Index is a snapshot of a model that is still mid-discount-ladder, with list pricing not arriving until 23 October. Watch for Kolibri's first independent evaluation, and watch what Solar Mini 4 actually costs once the ladder runs out — because at $0.10 / $0.40 the economics of renting the 3B slice look different from the $0.03 / $0.12 you are being quoted today.
