
GPT-6.1 Sol vs GPT-6 Luna: Thirteen Index Points, Ten Times the Money, and Two Evals Where It Doesn't Hold
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 300 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
OpenAI released GPT-6.1 Sol on 29 September 2026, one week after GPT-6 Luna arrived on 22 September in the same keynote cycle. They are the two ends of the same generation: Luna is the cheap tier, Sol is the mid-tier above it, and GPT-6 Astra sits above both. The price difference between the two is an order of magnitude — $0.10 and $0.50 per million tokens against $2.00 and $10.00, cached input $0.01 against $0.10 — and the capability difference on the independent board is 13.7 Intelligence Index points, 51.8 against 38.1 at maximum reasoning effort. Ten times the money for thirteen points is roughly the worst exchange rate in the current OpenAI lineup. What makes the comparison worth writing about is the forty-dollar question that follows: which thirteen points?
The two models, and the week between them
GPT-6 Luna is the small model of the GPT-6 generation — the one OpenAI describes as built for high-volume, latency-sensitive work: classification, extraction, routing, bulk summarisation, the inner loop of an agent. It carries a 1,050,000-token context window and a 128,000-token output ceiling, takes text, images and files as input, and keeps the full six-rung effort ladder including none, which the 6.1 tier dropped. The rate card is the reason it exists: $0.10 in and $0.50 out per million tokens, cached input at $0.01 and cache writes at $0.125, with a long-context step above 272,000 input tokens to $0.20, $0.75 and $0.02 cached.
GPT-6.1 Sol is the mid-tier refresh released a week later. Same window, same output cap, same input modalities, an April 30, 2026 knowledge cutoff against Luna's, and an effort ladder that starts at low because none and minimal are rejected outright. It is priced at $2.00 and $10.00 with cached input at $0.10 — twenty times Luna's input rate and twenty times its output rate, with a cache line that is ten times more expensive per hit.
Both are on sale, both are in ChatGPT Work and Codex as well as the API, and both are callable through the same key on our side. Nothing about this pairing requires choosing.
What thirteen index points actually buys
The composite is the least informative number in the comparison, because it averages over tasks that behave very differently. Scored evaluation by evaluation at maximum reasoning effort on Artificial Analysis v4.3.2, the shape is a split rather than a slope:
• AA-LCR v1.1, long-context reasoning — GPT-6 Luna 0.833 vs GPT-6.1 Sol 0.830. Luna is marginally ahead.
• SciCode — GPT-6 Luna 0.546 vs GPT-6.1 Sol 0.542. Again effectively level, and again the cheaper model is nominally in front.
• Terminal-Bench 4.0 — GPT-6 Luna 0.126 vs GPT-6.1 Sol 0.561. A 43-point gap; the two models are in different leagues on terminal work.
• Humanity's Last Exam — GPT-6 Luna 0.385 vs GPT-6.1 Sol 0.529.
• AutomationBench-AA — GPT-6 Luna 0.532 vs GPT-6.1 Sol 0.649.
• GDPval-AA v2.1 — GPT-6 Luna Elo 1437.1 vs GPT-6.1 Sol Elo 1575.1, a 138-point Elo gap on professional knowledge work.
• GDP.pdf — GPT-6 Luna 0.228 vs GPT-6.1 Sol 0.310.
• CritPt — GPT-6 Luna 0.194 vs GPT-6.1 Sol 0.317.
• AA-Omniscience — GPT-6 Luna 0.65 vs GPT-6.1 Sol 41.5. On a scale where zero means as many correct answers as incorrect ones, Luna is at the waterline and Sol is meaningfully above it.
• MMMU-Pro — GPT-6 Luna 0.797 vs GPT-6.1 Sol 0.860.
• Cost per index task, independent — GPT-6 Luna $0.068 vs GPT-6.1 Sol $0.72; about a tenfold gap in the cheap model's favour
• Output tokens on the index run — GPT-6 Luna 144 million vs GPT-6.1 Sol 67 million, against a class median of 100 million for Luna and about 81 million for Sol
Two of ten evaluations are a tie in the cheap model's favour. Eight go to the expensive one, and in six of those eight the margin is not marginal. That is the honest read of the thirteen points: it is not a uniform quality gradient, it is two capabilities — sustained agentic execution and factual reliability — where Luna is simply not the same class of model, and a broad middle where the difference is real but small enough to be invisible in your own evaluation set.
The AA-LCR result deserves one caveat, because it is the number that most flatters Luna. Long-context reasoning at 0.833 is a strong score and the two models are inside half a point of each other, but the benchmark measures retrieval and reasoning over long inputs, not the ability to act over a long session. The gap on Terminal-Bench 4.0 is what "act over a long session" looks like when it is scored, and it is 43 points wide.

The cost-per-task arithmetic, and where it stops being true
Artificial Analysis prices each index run from measured token consumption, so the cost figure is not a rate-card inference. GPT-6 Luna's run cost $0.068 per task; GPT-6.1 Sol's cost $0.72. The ratio is 10.7×, which is slightly worse than the nominal twenty-fold price ratio only because Luna generated 144 million output tokens on the run against Sol's 67 million — the cheap model is more than twice as chatty, and it eats into its own advantage. Even so, the ordering is not close, and on a short-answer workload it would widen: less output volume means Luna's verbosity penalty shrinks while its price advantage stays fixed.
Where the arithmetic stops being true is any workload the index does not resemble. If your task is a repository-scale edit that needs twenty tool calls held coherent, a $0.07 model that fails the task costs you the whole retry loop plus the wall clock, and the $0.72 model that finishes once is cheaper. GPT-6.1 Sol's own launch framing is entirely built on this idea — cost per completed task — and it is the correct frame: the relevant comparison is not the token rate, it is the expected cost per outcome, and on tasks Luna cannot complete that expectation is unbounded.
Two structural details sharpen the boundary. First, the long-context step: both models reprice above 272,000 input tokens, Luna to $0.20 and $0.75 and Sol to $4.00 and $15.00, so the absolute gap widens at the boundary rather than closing, but Luna's repriced rate is still an order of magnitude below Sol's standard one. Second, and easy to miss: the effort ladder is not the same shape. Luna accepts none; Sol does not. A cascade that routes to Sol first and falls back to Luna on error or timeout must not carry the request's reasoning.effort across unchanged, because a configuration that is legal on one model returns an error on the other. That is a routing-layer concern, not a model-quality one, and it is exactly the kind of thing a single endpoint with a routing rule is supposed to absorb.
Running both is the point, not a hedge
Both models are on OrcaRouter at OpenAI's own list prices with nothing added: GPT-6 Luna at $0.10 and $0.50 with cached input at $0.01, GPT-6.1 Sol at $2.00 and $10.00 with cached input at $0.10, and the 272,000-token step on each passed through exactly as the vendor lists it. Because the platform passes provider list price through rather than marking it up, the twenty-fold ratio in this article is the ratio you actually pay, on both sides.

The reason that matters here specifically is that this pair is a cascade waiting to be written. A classifier, a router, a bulk extraction job and a repo-scale coding agent do not belong on the same model, and the price gap is wide enough that picking wrong in either direction is expensive — twenty times too much per call in one direction, a failed task in the other. One key, two models, and a rule that sends easy traffic to Luna and escalates to Sol when the task class changes or when Luna's output fails a check, is a configuration change on our side rather than a second vendor contract. Automatic failover covers the case where one of the two is degraded, which for a model that shipped eight days ago is still a live possibility.

The decision, in one pass
• Classification, extraction, routing, bulk summarisation, agent inner loops — GPT-6 Luna, and stop reading. At $0.10/$0.50 with a one-cent cached read, thirteen index points of general capability are not worth ten times the invoice on work where the answer is short and checkable.
• Repo-scale coding, terminal agents, multi-step execution — GPT-6.1 Sol. The Terminal-Bench 4.0 gap of 43 points is the whole argument; this is the capability the price is buying.
• Knowledge-work questions where a wrong answer is expensive — GPT-6.1 Sol, on the AA-Omniscience and GDPval-AA gaps. Luna's factual-reliability score is at the waterline.
• Long inputs with a short generated answer — GPT-6 Luna. It reasons over long context at parity with a model costing twenty times as much, and its 144-million-token verbosity penalty never gets a chance to apply.
• Anything already setting none — stay on Luna, or budget for the change. That rung does not exist on GPT-6.1 Sol.
• Latency-sensitive surfaces — GPT-6 Luna, on the board's own measurements: 129.4 output tokens per second and 100 seconds to first chunk against Sol's 59.7 and 332.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
