A generated hero title card for Solar Mini 4 headlined 'Solar Mini 4 — the independent run', with the subtitle 'Intelligence Index 24, and roughly five times the cost per task of GPT-6 Luna' and two badges reading 'Released 22 September 2026' and 'First third-party evaluation, 30 September 2026'.
Guides & Insights

Solar Mini 4 Scored 24 on the Artificial Analysis Index — and Costs Five Times More Per Task Than GPT-6 Luna

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Solar Mini 4 charges the same per input token as GPT-6 Luna, and roughly five times as much to actually finish a job. That is the whole story of the first independent evaluation of Upstage's new model, and it is not the story the launch materials told. Solar Mini 4 went live on 22 September 2026 as a 35-billion-parameter sparse mixture-of-experts that activates about 3 billion parameters per token, priced at $0.10 per million input tokens and $0.40 per million output. GPT-6 Luna, OpenAI's efficiency tier, lists at $0.10 and $0.50, with a long-context tier above 272,000 tokens that roughly doubles both. On a rate card they look like the same purchase. On Artificial Analysis's Intelligence Index, where both models were run through the same ten-evaluation suite, Solar Mini 4 costs $0.36 per completed task against $0.07 for GPT-6 Luna (max) — a factor of 5.4 that comes entirely from how the two models behave during a task rather than from what they charge for a token.

On 30 September 2026, Artificial Analysis published its write-up of the model, which scores 24 on the Intelligence Index v4.3.2. Everything in this article that is attributed to Artificial Analysis is that organisation's own measurement; everything attributed to Upstage is the vendor's claim. The distinction matters more here than usual, because the two sets of numbers do not always agree.

What the independent run actually says

A scoreboard titled 'Solar Mini 4 — the scoreboard' with the rows: Intelligence Index 24.05 against a median of 12; Cost per index task $0.36 against GPT-6 Luna (max) at $0.07; Rate card $0.10 in / $0.01 cached / $0.40 out per 1M; Output per index task 88,300 tokens, 71,600 of them reasoning; Output speed 204 tokens per second and 430 seconds per index task; Context Upstage says 512K while Artificial Analysis lists 1.05M; Weakest result Terminal-Bench 4.0 at 1%.

Solar Mini 4 lands at 24.05 on the Intelligence Index, against a median of 12 among the reasoning models in its price tier. That places it well above average for its weight class, and Upstage's own framing — "highest overall score among the 3B-active models in the Artificial Analysis Intelligence Index comparison" — survives independent testing on that specific point. Qwen3.6 35B A3B, which also activates 3B parameters, scores 18.2 on the same index. K2 Horizon MoVA 36B A4B, with 4B active, scores 25.3 — and it does that on a shorter context, 524K tokens against Solar Mini 4's 1.05M. Against Inkling, a 975-billion-parameter open-weights MoE released in July, Solar Mini 4 comes in at 24.1 against 25.0 — within a point of a model nearly thirty times its size that costs $1.00/$4.05 per million tokens.

The profile inside that score is uneven, and the unevenness is the useful part:

• Long-context reasoning is the relative strength. Solar Mini 4 scores 83% on AA-LCR v1.1, level with MiniMax-M3 and GPT-6 Luna (max), and ahead of Gemini 3.8 Flash (high) and GPT-6 Astra (max) at 81%.

• Scientific coding holds up. 48% on SciCode, one point ahead of MiniMax-M3 and Inkling (xhigh).

• Agentic coding collapses. 1% on Terminal-Bench 4.0. Not "weak" — 1%. On AutomationBench-AA, which runs multi-step workflows across real SaaS applications, 22.3%.

• Knowledge accuracy is low, but the model knows it is guessing. AA-Omniscience returns −11 with 18% accuracy; the model abstains on roughly half the questions it is asked. Its non-hallucination rate is 64%, well ahead of Inkling (xhigh) at 32% and GPT-6 Luna (max) at 23%. A model that says "I don't know" half the time and is right two-thirds of the time it does answer is a different instrument from a model that confabulates confidently.

On agentic knowledge work the figures are 1072 Elo on GDPval-AA and 872 Elo on AA-Briefcase, both close to Inkling (xhigh).

The five-times number, unpacked

Artificial Analysis's cost-per-task figure is the one that will decide most procurement conversations, and it deserves to be read as a breakdown rather than quoted as a headline. Two things drive it.

First, volume. Solar Mini 4 emits about 88,300 output tokens per Intelligence Index task; 71,600 of those are reasoning tokens. At $0.40 per million output tokens that is roughly $0.035 of the bill — cheap, because output is cheap. What it costs instead is time: at a measured 204 tokens per second, Artificial Analysis logs 430 seconds of work for an average index task, or about 7.2 minutes. GPT-6 Luna (max) emits 50,000 output tokens per task at 127 tokens per second and takes 396 seconds. Inkling (xhigh), an open-weights model of 975 billion parameters, emits 34,700 tokens and finishes in 169 seconds. Solar Mini 4 is slower than both, but by a margin that has nothing to do with the five-fold cost gap.

Second — and this is the whole story — cache-write input. Artificial Analysis's cost breakdown attributes about $0.30 of Solar Mini 4's $0.36 per task to tokens written into the prompt cache, against $0.03 for cache reads, $0.035 for output and a rounding error for genuinely uncached input. For GPT-6 Luna (max), cache write is $0.012 of $0.068 — about 17% of the bill, against 82% for Solar Mini 4. Cached input is the cheapest line on Upstage's rate card at $0.01 per million, and Artificial Analysis's model does use it; the problem is how much has to be written before it can be read. An agent that re-sends a growing conversation on every turn pays the write cost far more often than a model that answers in a short, closed turn.

That is the finding worth carrying into production: for a model like this, the per-token price predicts almost nothing about the per-task bill. Cache behaviour does.

Where Upstage's own numbers differ

A screenshot of Upstage's Console documentation page for Solar Mini 4, showing the Intelligence, Speed and Price gauges, the '$0.10 / 1M tokens' input rate, the description of a cost-efficient compact language model at 35B total and 3B active parameters built for agentic use cases, English, Japanese and Korean language support, tool calling, a 128K max output, the platforms Upstage Console and on-premises, and the version string solar-mini4-260922 with a training data cut-off of February 2026.

Upstage's launch post, published 1 October 2026 under the title "Solar Mini 4: Built for High-Volume Agent Workloads", reports 24.1 on the Artificial Analysis Intelligence Index — effectively the same figure — and makes several claims that sit alongside, rather than on top of, the independent data.

The vendor cites 22.3% on AutomationBench-AA, matching Artificial Analysis exactly, and 47.2 on τ³-Banking, a policy-and-tool-use benchmark the index suite has since dropped. It reports 83.3% on AA-LCR and 47.6% on SciCode, both within a rounding error of the independent figures. On Humanity's Last Exam it reports 19.6%, while the Artificial Analysis page carries 25.8% for the same model; both are plausible readings of a notoriously volatile evaluation, but they are not the same number and neither should be presented as the other.

Two specification lines also disagree. Upstage documents a 512K-token context window with up to 128K output tokens. Artificial Analysis lists a 1.0M-token context window and a 262K maximum output. Upstage's blog is explicit that the 512K figure is the shipping contract; the discrepancy is the kind of thing worth resolving against an API response before anyone builds a retrieval pipeline on it.

The vendor also makes the one comparison it can make on its own terms: an internal test across three Korean-language agent tasks run three times per model, in which completing 1,000 three-task sets cost an estimated $1.25 with Solar Mini 4 and $2.20 with MiMo-V2.5. That is a vendor-run test on vendor-chosen tasks with latency excluded, and it points the opposite way from the Artificial Analysis result. It is not wrong — different tasks exercise different token budgets — but it is a marketing figure, and the model's published launch discount is a separate reason to re-run your own numbers rather than adopt anyone's.

Pricing, as of the live Upstage console documentation: $0.10 per million input tokens, $0.01 per million cached input tokens, $0.40 per million output tokens, with a launch discount currently advertised at 50% off through 22 October 2026 UTC, which puts the effective rates at $0.05 input, $0.005 cached input and $0.20 output until then.

What this means if you are the one paying

The interesting thing about Solar Mini 4 is not that it is a bad model. It is that it is a genuinely good small model whose economics are dominated by behaviour a rate card cannot show you. A 24 on the index at three active parameters is a real engineering result, and Korean-language workloads — the model's stated strong suit, alongside English and Japanese — are exactly where that result is most likely to be worth its price.

The awkward part is everything downstream of the first call. Long assistant turns, low cache reuse, and a task that takes seven minutes of decode per index run add up to a number that looks nothing like $0.10/$0.40. If you are evaluating it, the honest experiment is a cost-per-completed-job test on your own workload, with prompt caching enabled in the way your production stack would actually use it — not a token-price comparison.

This is also the class of problem a router exists to solve, and the reason OrcaRouter passes provider list prices through at 0% markup so a vendor price change reaches your bill the same day. We do not route Upstage models, so neither Solar Mini 4 nor Solar Pro 4 is available through us — they come from Upstage's own API and third-party platforms. What we do route is the other half of this comparison: GPT-6 Luna, Qwen3.8-Max, Qwen3.8-27B, Qwen3.8-Flash and more than 200 other models behind a single key, which is what makes it practical to keep a cheap model and an expensive one on the same task and let automatic failover and per-request routing decide which one answers. When the difference between two models is a five-fold cost-per-task gap that no price list reveals, the ability to move a single request between them matters more than any benchmark table.

A screenshot of the Artificial Analysis article dated September 30, 2026, headlined 'Korean AI Lab Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token prices'. The visible key-results text covers the 3B-active Pareto claim, the 6-point lead over Qwen3.6 35B A3B at the same active size, 83% on AA-LCR v1.1 matching MiniMax-M3 and GPT-6 Luna (max), 48% on SciCode, and fast output at 208 tokens per second with slow tasks.

The short version: Solar Mini 4 is the new Pareto point Artificial Analysis draws for models under three billion active parameters, and it is roughly five times more expensive per completed task than a model that lists at the same input price. Both of those statements are true, and the second one is the one that shows up on an invoice.