
Solar Pro 4 Scores 42: Upstage's Agent-First Flagship, Now Independently Measured
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
Upstage Solar Pro 4 went live on August 10, 2026, and the first number worth staring at was the price: $0.30 per million input tokens and $1.20 per million output tokens on Upstage's own API, with a launch promo that cuts that to $0.03 and $0.12 — 90% off — on current listings. Four days later the second number worth staring at arrived. On August 12, 2026, Artificial Analysis published the first independent evaluation of Solar Pro 4 and scored it 42 on its Intelligence Index, against 14 for Solar Pro 3 — a threefold gain on the predecessor's measured score and a 27-point jump by the lab's own arithmetic — confirming Upstage's positioning of Solar Pro 4 as its proprietary flagship, the model that replaces Solar Pro 3 at the top of the line. Either way, cheap; now, independently measured too.
What actually shipped
Solar Pro 4 is the commercial flagship of Upstage's agent-era pivot. The company has spent 2026 rebuilding the Solar line around agentic workloads — its open-weight sibling Solar Open 2 shipped in late July — and Solar Pro 4 is the hosted product that pivot was pointing at. Upstage frames it in one sentence: a model for tasks that need sustained context, multi-step execution, and end-to-end completion rather than single-shot answers.
Concretely, the model page describes it as Upstage's flagship model specialized for agentic use, suited to agentic workflows, office productivity, document-intensive work, and coding. On paper it is the largest context the line has shipped — Solar Pro 3 offered 128K — and the only Solar to date with a published reasoning story and tool-calling surface aimed at autonomous agents.
Specs as listed on launch:
• Context window — 524,288 tokens (524K) on the vendor's listing, roughly four times the previous generation's 128K; Artificial Analysis's independent measurement records 384K, so treat the bigger figure as Upstage's claim
• Price — official Upstage API rate $0.30 per 1M input tokens / $1.20 per 1M output tokens, with cached input at $0.06; a launch promo prices it 90% off at $0.03 / $0.12 on current listings
• Capabilities — reasoning, function calling, JSON mode, streaming, and web search, with vision input supported
• Weights — proprietary, not open; no Hugging Face release (contrast with the open Solar Open 2)
• Access — Upstage Console API plus several third-party platforms; a chat app at solar-chat.upstage.ai for hands-on testing
Those figures are the vendor's listing, not an audited spec sheet — and the first independent measurement now disagrees with one of them. Artificial Analysis records a 384K-token context window and a 256K max output for Solar Pro 4, against the 524K and 128K–131K the vendor lists. The capabilities list and Upstage's own context figure come from the model page; the AA numbers are the independent read.


What changed versus Solar Pro 3 — the first independent numbers
Upstage's own claims for Solar Pro 4 center on three areas of improvement over Solar Pro 3, all of them agent-relevant: long-document reasoning, multi-turn tool use, and terminal task performance. Those are exactly the dimensions where an agent model earns its keep — keeping a thread straight across a long document, sustaining a sequence of tool calls without losing the plot, and finishing jobs that run inside a shell.
The first independent run backs the direction of every one of those claims. Artificial Analysis scored Solar Pro 4 at 42 on its Intelligence Index against Solar Pro 3's 14, and the sharpest gains sit precisely where Upstage pointed:
• Terminal-Bench v2.1 (terminal tasks): 12% → 57%
• AA-LCR (long-context reasoning): 31% → 71%
• τ³-Banking (multi-turn tool use): 9% → 23%
• GDPval-AA v2 (real-world agentic work): Elo 498 → 1,277, above the 1,000 human baseline and slightly ahead of Qwen3.7 Max (1,272) and MiMo-V2.5-Pro (1,266)
The index placement is the headline: 42 sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43). A model priced like a budget tier is now measuring like a serious agent.
Two caveats came with the score. The AA-Omniscience improvement (-53 → -1) is mostly abstention: Solar Pro 4 attempts only 41% of questions against Solar Pro 3's 92%, its hallucination rate dropped from 88% to 24%, but accuracy was flat at 19%. And the intelligence gain cost latency — about 8.6 minutes per Intelligence Index task against Solar Pro 3's 6.0, while token use improved from roughly 52k to 43k output tokens per task. It is a smarter, more careful, slower model.
Why the price matters more than the specs
The case for Solar Pro 4 is its price-to-measured-intelligence ratio. Even at the official list rate, a 384K–524K-context model at $0.30 per million input tokens undercuts essentially the entire field — roughly one-sixteenth the input price of a frontier flagship — and the 90%-off promo takes that to one-sixteenth of one-sixteenth. That lands in a band where the market currently sells small fast models, not agent-capable ones with a verified mid-tier score.

The skeptical case is no longer that Solar Pro 4 is unproven — the AA run settles that — but it is real. The omniscience gain is partly a refusal pattern, not just knowledge. The measured latency is a step back from Solar Pro 3. Artificial Analysis's 384K context sits below the vendor's 524K headline. And the "90% off" badge means the go-forward price may move once the promotional period ends — the durable rate is the $0.30 / $1.20 list, not the $0.03 / $0.12 sticker.
Who should try it, who should wait
Try it now if: you are prototyping an agent and the dominant cost is tokens, you need a long context window and your workload tolerates slower thinking, or you are evaluating Korean-language and East-Asian document workloads where Upstage has consistently trained well.
Wait if: your workload is production-critical and cannot absorb a regression while the field calibrates expectations, you need open weights for self-hosting or fine-tuning — Solar Pro 4 is not that model; Solar Open 2 is — or you are buying on the 524K context number, which the first independent measurement does not reproduce.
A middle path exists and is worth naming: put Solar Pro 4 behind a failover layer rather than betting a whole pipeline on it. A routing gateway that can fall back to a second model on timeout or error lets you capture the upside of a cheap new model without the single-point-of-failure risk — the pattern OrcaRouter's automatic failover exists for, and because OrcaRouter passes provider list prices through at zero markup, whatever Upstage charges is what you pay, with no reseller margin on top.
What to watch next
The "no independent numbers yet" chapter of this model is closed — it took four days. What matters now is what the first measurement implies. The post-promo price: the 90% badge is an opening bid, not a commitment, and the list rate behind it is $0.30 / $1.20. Whether Upstage responds to the measured latency and the 384K-versus-524K context discrepancy, both the most actionable findings in the AA run. And real agent workloads — the best test of a tool-calling model is long-running jobs in the wild, not a leaderboard. As of August 13, 2026, the accurate summary of Solar Pro 4 is narrower and stronger than a week ago: a very cheap, agent-shaped flagship from a serious lab, confirmed as Upstage's replacement for Solar Pro 3, with an independent score of 42 to argue over.
