A generated hero title card reading 'GPT-6 Sol and GPT-6 Luna' under the overline 'Model launch - September 2026', subtitled 'Half the price per token is not half the cost per task.', with four chips reading 'GPT-6 Sol: $2.00 / $10.00 per 1M', 'GPT-6 Luna: $0.10 / $0.50 per 1M', 'Sol: index 48 at $1.06 per task' and 'Luna: index 37 at $0.07 per task', and a footnote reading 'Prices per OpenAI; index and cost per task per Artificial Analysis, Sept 22, 2026.' The real OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

GPT-6 Sol and GPT-6 Luna: The Cheaper Token Is Not Always the Cheaper Task

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The vendor shipped GPT-6 Sol and GPT-6 Luna on September 22, 2026, and the headline everywhere is the price: Sol at $2.00 per million input tokens and $10.00 output, Luna at $0.10 and $0.50, both roughly half what GPT-5.6 Sol and GPT-5.6 Luna charge under current promotional rates, and both sitting well below the flagship GPT-6 Astra at $10.00/$50.00. The vendor's own announcement leads with "we've also made caching and inference more efficient." The trouble with reading that as a discount story is that the unit that actually lands on your invoice is not a token. It is a task. And the first independent measurements, published the same day, show the two new models drifting in opposite directions on exactly that unit — with GPT-6 Luna, the cheaper of the two by a factor of twenty on input, doing the worse of the two against its own predecessor. One of these models is a straight upgrade and one is a genuine trade, and the vendor's announcement does not distinguish between them.

Before the numbers, the sourcing rules for this piece, because the two piles of evidence here deserve very different weight. Prices, context window, cache rates and knowledge cutoffs come from OpenAI's own API documentation and pricing tables, read on September 23, 2026, with the service-tier and cache-write lines corroborated against third-party cost trackers rather than taken from the announcement alone. OpenAI's benchmark tables — AutomationBench, DeepSWE 1.1, OSWorld 2.0, Agents' Last Exam, and an internal factuality evaluation — are vendor-reported and unreproduced, and every figure drawn from them below is labelled as such. The independent numbers come from Artificial Analysis, which ran both models on its Intelligence Index and Coding Agent Index on launch day; those are the ones to trust for a decision. Where the two disagree, the article says so rather than averaging them.

The launch, in one paragraph

Sol and Luna are mid- and low-tier members of the GPT-6 family, trained with the methods behind GPT-6 Astra and positioned as faster, cheaper ways to reach most of its capability. OpenAI's framing is cost per task rather than raw capability: it says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol on its internal factuality evaluation, and that GPT-6 Luna at higher reasoning effort matches GPT-5.6 Sol's factuality at roughly one-hundredth of the task cost. Both accept text and image input and produce text, both carry a 1,050,000-token context window with a 128K output ceiling — Artificial Analysis lists 872k for the same model, so treat the upper end of that window as vendor-stated — and both expose reasoning effort from none through max, a level GPT-6 Astra does not offer, because Astra has no none. Model IDs are gpt-6-sol and gpt-6-luna. Availability runs through the API, through ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu accounts, and through Amazon Bedrock, which announced general availability of both the same day. Free and Go users get Luna in the desktop app; neither model is in plain Chat mode yet.

What the four independent numbers say, and why they disagree

Start with GPT-6 Sol, because its story is simple. Artificial Analysis scores it 48 on the Intelligence Index, 18th of 212 models measured, at $1.06 per index task against $1.99 for GPT-5.6 Sol — roughly half the cost for a level index score, and a cost-per-task improvement that holds up rather than evaporating. Its Coding Agent Index rises to 57 from 55, with Terminal-Bench 4.0 moving 37% to 43% and SWE-Atlas-QnA 54% to 58%, at $2.99 per task on the Pareto frontier. Its hallucination rate falls from 92% to 60%, and its abstention behaviour changes shape entirely: it now answers 83% of questions where the previous generation answered 99%, and in exchange gets a quarter fewer wrong answers, with accuracy slipping from 59% to 54%. That is a model that has been taught to shut up when it does not know, and the AA-Omniscience Index moving from 22 to 27 is that same fact stated as a score.

Now GPT-6 Luna, and the story changes. Artificial Analysis scores it 37 on the Intelligence Index — the identical figure GPT-5.6 Luna scores. Its cost per index task falls from $0.18 to $0.07, about 60% less, and that saving is real. But its Coding Agent Index goes down, 43 to 41, with SWE-Atlas-QnA falling 49% to 44% and DeepSWE 1.1 falling 66% to 64%. Its hallucination rate improves from 93% to 77%, and its abstention behaviour barely moves — accuracy 44% against 43%, AA-Omniscience from −10 to 1. Read together: the same measured intelligence, at roughly 40% of the task cost, with a worse coding agent and essentially unchanged answer quality.

That is not a contradiction, and the reason matters. A cheaper token buys you less only if the model spends more tokens to finish the same job — and the independent data says these models spend more, not less. Output per index task rose from 29k to 31k tokens for Sol and from 41k to 51k for Luna. GPT-6 Luna's 60% saving comes entirely from the price cut, not from running leaner. The tokens-per-task drift is small enough that the arithmetic still works in your favour here; it is not small enough to ignore as a category, because it is the mechanism by which a future price cut can be neutralised by a chattier model. OpenAI's own framing has moved to cost per task for the same reason.

Two regressions sit outside the two headline indexes and deserve naming, because neither appears in OpenAI's announcement. In GDPval-AA v2.1, which measures knowledge-work deliverables, GPT-6 Sol loses roughly 100 Elo against its predecessor and GPT-6 Luna roughly 75. GPT-6 Luna also sheds around 45 Elo in AA-Briefcase v1.1, where Sol holds steady — Artificial Analysis attributes the gap to weaker presentation and missing rubric items. If your use of the cheap tier is bulk summarisation, extraction and classification, none of that touches you. If it is drafting something a person signs off on, it does.

A generated two-column scoreboard titled 'GPT-6 Sol vs GPT-6 Luna - the scoreboard'. The left column, GPT-6 Sol, lists Intelligence Index 48, cost per index task $1.06, Coding Agent Index 57, hallucination rate 60%, GDPval-AA about 100 Elo lower, and price per million tokens $2.00 / $10.00. The right column, GPT-6 Luna, lists Intelligence Index 37, cost per index task $0.07, Coding Agent Index 41, hallucination rate 77%, GDPval-AA about 75 Elo lower, and price per million tokens $0.10 / $0.50. A footer line credits Artificial Analysis for the index, cost and hallucination figures as of September 22, 2026 and OpenAI for the prices. The real OrcaRouter logo is composited in the bottom-right corner.

The caching change is the real API news

Buried under the rate card is a change that matters more to a production bill than the headline cut, and it is the part of the announcement OpenAI's own post gestures at with one clause. Three things changed, and the third is the one to read twice.

The first is the discount itself: cached input reads bill at 10% of the input rate, a 90% discount, which is the same term the family already carried rather than a new one. Sol's cached input is $0.20 per million and Luna's is $0.01; cache writes carry a 1.25x premium, $2.50 and $0.125 respectively, with a single 30-minute time-to-live.

The second is control. Explicit caching through prompt_cache_options and prompt_cache_breakpoint — the parameters that let you declare where a reusable prompt prefix ends rather than hoping the provider guesses — is available on both models. This is not new API surface; what is new is that it is worth using, because the default hit rates went up.

The third is the one that changes architecture. Changing reasoning effort or switching tools on and off no longer invalidates the cached prefix. Before this, an agent loop that dialled effort up for a hard step and back down for an easy one paid to reprocess its whole context at every dial turn, and so did one that added or removed a tool mid-conversation. That tax is now gone on these two models. OpenAI cites GitHub, which reported that the changes cut the share of prompt tokens needing fresh processing by more than half across billions of requests. Treat that as a vendor-supplied figure — it is credible and it is not independently audited.

Do the arithmetic on a long agent loop and the ranking of the two announcements flips. A loop that carries a 30,000-token system prompt and tool schema across 200 turns moves 6 million tokens of context. Without caching, on GPT-6 Sol, that is $12.00 of input. With the prefix cached at $0.20 per million, it is $1.20 plus the one-off write — an order of magnitude, from a parameter change, before the rate cut is applied at all. The rate cut is what the announcement sold. The cache behaviour is what actually moves the number, and it is the reason a $2.00 headline rate behaves like something much cheaper on the workloads OpenAI is pitching these models for.

The rate card, including the parts the announcement skips

The headline rates are not the whole rate card, and three tiers change the decision for specific workloads.

Base — GPT-6 Sol $2.00 in / $10.00 out; GPT-6 Luna $0.10 in / $0.50 out, per million tokens, for prompts up to 272K input tokens.

Long context — above 272K input tokens the entire request bills at 2x input and 1.5x output: $4.00/$15.00 for Sol and $0.20/$0.75 for Luna. This is a cliff, not a marginal rate. If your prompts run to 300K tokens you are paying double on the whole request, not on the 28K over the line, and splitting a document into two calls below the threshold is often cheaper than sending it once above.

Service tiers — Flex halves the bill ($1.00/$5.00 for Sol, $0.05/$0.25 for Luna) in exchange for slower turnaround; Priority doubles it. Both are selected through the service_tier parameter. Flex is where batch-ish jobs should live and almost never do.

Cached input — $0.20 per million on Sol and $0.01 on Luna, a 90% discount, with a 30-minute TTL and a 1.25x cache-write premium.

Context and output — 1,050,000-token window, 128K maximum output, text and image in, text out, reasoning effort none to max with medium as the default.

Knowledge cutoffs — April 20, 2026 for GPT-6 Sol and May 18, 2026 for GPT-6 Luna, which is a month apart inside the same family and worth knowing before you assume the two models share a world view.

What OpenAI's benchmark table says, and what it does not

OpenAI's published comparisons are all cost-per-task, all against Anthropic, and all vendor-run. On AutomationBench 1.0.6, a 47-tool agentic evaluation, the company reports GPT-6 Sol at xhigh effort scoring 33.2% at $0.27 per task, against Claude Opus 5 at 26.9% and roughly 11 times the cost per task. On DeepSWE 1.1 it reports Sol at max effort scoring 68.8% against Claude Fable 5 at 69.9% — a loss on score at roughly 80% lower cost per task — and Luna at 66.6%. On OSWorld 2.0, Sol scores 60.5% at xhigh against Opus 5's 60.3%, again at around 80% less per task. On Agents' Last Exam, Sol reaches 56.4%, which OpenAI says beats Opus 5's best result at 60% lower cost per task.

Every one of those figures is vendor-reported and unreproduced, and two caveats are worth carrying forward rather than filing away. First, OpenAI benchmarked against Claude Opus 5, not the Claude Opus 5.5 released hours before the announcement, so no same-harness comparison yet establishes which model actually delivers the lower cost per successful task. Second, cost per task is not cost per correct task: a model that answers 83% of the time and abstains on the rest will look cheap per task on a benchmark that counts abstentions as tasks. The independent abstention data above is what lets you price that honestly — GPT-6 Sol answering 83% of questions where the old model answered 99% is a deliberate trade, and whether it is a good one depends entirely on what a wrong answer costs you.

OpenAI reports that GPT-6 Luna at higher effort matches GPT-5.6 Sol's factuality at roughly one-hundredth of the task cost. That is a factuality comparison, not a capability comparison, and the independent Coding Agent Index puts GPT-6 Luna two points below GPT-5.6 Luna, let alone GPT-5.6 Sol.

Screenshot of the Artificial Analysis model page for GPT-6 Sol, captured 23 September 2026, showing the model's Intelligence Index of 48 alongside its cost per task and its output-token and speed figures for the measured reasoning-effort settings.

What to actually do with this

GPT-6 Sol is the straightforward one. Half the task cost at a level index score, a better coding agent, a large drop in hallucination and a genuine change in abstention behaviour add up to an upgrade you can take on a schedule rather than a gamble. If you are on GPT-5.6 Sol, this is a migration, and the only real question is which of your prompts depend on the old model's willingness to answer everything.

GPT-6 Luna is the one to test before you move. It is not a strictly better GPT-5.6 Luna; it is a differently-shaped one — cheaper per task, safer about what it will claim, and measurably worse at agentic coding. For classification, extraction, routing and summarisation at volume, the trade is close to free money. For anything that feeds a coding agent, or produces a document a person signs off on, the two-point Coding Agent regression and the GDPval slide are the reason to keep GPT-5.6 Luna in the path until you have run your own tasks through both.

That is also the argument for not hard-wiring either one. Both models are OpenAI-only right now, and the workloads these tiers serve — routing, extraction, cheap classification — are precisely the ones where a second provider behind the same endpoint turns a model regression into a non-event. OrcaRouter passes provider list prices through at 0% markup, so a vendor rate change lands on our side the same day it lands on the vendor's, and the routing DSL lets you put GPT-6 Luna behind a cheaper model for the easy share of traffic while the harder share goes to GPT-6 Sol — with automatic failover if either path degrades. You do not have to pick a winner on launch day to benefit from the launch.

The honest caveat: GPT-6 Sol and GPT-6 Luna are not routable through OrcaRouter as of September 23, 2026. We do not host them yet, and nothing above should be read as a claim that we do. What is on our side today is the comparison set — GPT-5.6 Sol at $4.00/$20.00, GPT-5.6 Luna at $0.20/$1.20, and GPT-6 Astra at $10.00/$50.00 — which is exactly what you need to price the migration before you commit to it.

Screenshot of the OrcaRouter model page for GPT-5.6 Luna, captured 23 September 2026, showing the model id openai/gpt-5.6-luna with its context window, input and output pricing per million tokens, measured time to first token and recent traffic.

What to watch next

Three things would settle what this launch actually was. The first is whether GPT-5.6 Luna's $0.20/$1.20 rate survives the quarter, because a permanent price from one vendor has historically been the opening move rather than the end of it — and the same-day Claude Opus 5.5 release suggests the pressure that produced this cut has not gone away. The second is whether the abstention shift holds up in the field: GPT-6 Sol declining to answer one question in six is the kind of change that reads as an improvement in a lab and as a broken product in a pipeline that assumed an answer. The third is whether the cache behaviour is real on your traffic, which is measurable this week and answerable only by measuring — the discount is large enough that a single hour spent instrumenting cache hit rates on your longest prompts is worth more than another benchmark table.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily