Generated hero title card headlined 'GPT-6 Luna vs GPT-6 Sol' with the subtitle 'Two rate cards, one shape, a twentyfold number', a footer reading 'Prices per OpenAI; index figures per Artificial Analysis.', and the OrcaRouter logo composited bottom-right.
Guides & Insights

GPT-6 Luna vs GPT-6 Sol: Two Rate Cards With the Same Shape and a Twentyfold Number

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-6 Luna lists at $0.10 per million input tokens and $0.50 per million output. GPT-6 Sol lists at $2.00 and $10.00. The vendor shipped both on September 22, 2026, as the two lower tiers of the GPT-6 generation beneath the flagship GPT-6 Astra, and the twentyfold gap between them is the first thing anyone reads off the page. What is harder to see, and more useful, is how little else separates them. Same 1,050,000-token context window. Same 922,000-token maximum input. Same 128,000-token output ceiling. Same reasoning ladder from none through low, medium, high, xhigh and max, with medium as the default. Same knowledge-cutoff family, a month apart. Same clause that reprices an entire request once it crosses 272,000 input tokens.

Two models that share every structural line of a rate card are not differentiated by features. They are differentiated by a score, and by what that score costs you to obtain. On the independent board that measures both, GPT-6 Sol sits eleven points above GPT-6 Luna at matched maximum effort and costs about fifteen times as much to complete the same evaluation. Eleven points and fifteen times is a different sentence from twenty times, and the difference between those two sentences is where this comparison actually lives.

What each of these two is for

GPT-6 Sol is the daily driver of the family — the tier OpenAI positions for complex coding, agentic workflows and professional work that a person will read. It accepts text and image, produces text, and is available to anyone with an API key. It is the model most teams will reach for by default when they want GPT-6-class capability without the flagship's rate card or its access gate.

GPT-6 Luna is the volume tier. OpenAI's own description is narrow and deliberately unglamorous: the most efficient model for focused, high-volume tasks. Summarisation, extraction, classification, routing, the inner loop of an agent that makes a thousand small calls on the way to one task. Same window, same output cap, same effort dial — a cheaper engine in an identical chassis.

Both were trained with methods derived from GPT-6 Astra rather than built as a separate line, which is why the spec sheets line up so exactly. The price is the only place the two diverge by design.

The rate card, line by line

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs GPT-6 Sol $2.00

• Output per 1M — GPT-6 Luna $0.50 vs GPT-6 Sol $10.00

• Cached input per 1M — GPT-6 Luna $0.01 vs GPT-6 Sol $0.20; both are a tenth of their own uncached input rate, so the twentyfold ratio survives the discount intact

• Cache writes per 1M — GPT-6 Luna $0.125 vs GPT-6 Sol $2.50, a 25% premium on the input rate in both cases

• Long-context tier — identical clause on both cards: above 272,000 input tokens the whole request bills at 2× input and cache and 1.5× output, taking GPT-6 Luna to $0.20 and $0.75 and GPT-6 Sol to $4.00 and $15.00

• Context window — 1,050,000 tokens for both, with a 922,000-token maximum input and a 128,000-token output ceiling

• Reasoning effort — none, low, medium (default), high, xhigh, max for both; the dial is the same part in both models

• Service tiers — Flex halves the bill and Priority doubles it on both cards, selected through the same service_tier parameter

• Knowledge cutoff — GPT-6 Luna May 18, 2026 vs GPT-6 Sol April 20, 2026, which is a month apart inside one family and worth knowing before assuming the two share a world view

• Intelligence Index, independent — GPT-6 Luna 37 at max effort vs GPT-6 Sol 48 at max effort, same index revision

• Default-effort score — GPT-6 Luna 29 at its default medium vs GPT-6 Sol 48 at max, the setting Artificial Analysis measured it at

• Cost per Index task, independent — GPT-6 Luna about $0.07 vs GPT-6 Sol $1.06

• Output tokens on the index run — GPT-6 Luna 150M vs GPT-6 Sol 77M, against a board median of 88M

• Coding Agent Index, independent — GPT-6 Luna 41, down two points on GPT-5.6 Luna vs GPT-6 Sol 57, up two on GPT-5.6 Sol

• Hallucination rate on AA-Omniscience — GPT-6 Luna 77%, down from 93% vs GPT-6 Sol 60%, down from 92%; both models got more willing to decline rather than more knowledgeable, and Sol got further

• On OrcaRouter — neither GPT-6 Luna nor GPT-6 Sol is in our catalogue as of this writing; both are reached through OpenAI's own API

Generated two-column scoreboard titled 'GPT-6 Luna vs GPT-6 Sol - the scoreboard'. Left column GPT-6 Luna rows: Price in/out $0.10 / $0.50, Cached input $0.01, AA Index at max effort 37, Cost per index task $0.07, Context 1,050,000, Max output 128,000. Right column GPT-6 Sol rows: Price in/out $2.00 / $10.00, Cached input $0.20, AA Index at max effort 48, Cost per index task $1.06, Context 1,050,000, Max output 128,000. Footer: 'Prices per OpenAI; index figures per Artificial Analysis.'

The effort dial is the same part in both models, and that changes the arithmetic

Because the reasoning ladder is identical, the headline index figures for both models are maximum-effort figures. GPT-6 Luna's 37 is what it scores when it is allowed to think as long as it likes. Its default is medium, where it scores 29. GPT-6 Sol's 48 is a max-effort number too, and it is the setting the independent board ran it at.

This matters more here than in most comparisons, because it means the two models are being quoted at the same rung. A team that deploys GPT-6 Luna and changes nothing is not running a 37 against a 48; it is running a 29 against a 48. That is a nineteen-point gap rather than an eleven-point one, and the size of it is a property of the default, not of the model.

The same dial is also the cheapest optimisation available on either card, and it runs in both directions. Dropping GPT-6 Sol from max to medium on work that a human reviews is a larger saving than switching a fraction of your traffic from Sol to Luna, and it costs one parameter change rather than a migration. Raising GPT-6 Luna to xhigh on the one call per task that actually decides the outcome is a rounding error against the total. The dial exists on both models specifically so that the tier choice and the effort choice can be made separately, and most teams only make one of them.

The cliff is identical, which is the point

Both models carry the same long-context clause, and because the clause is proportional it does not narrow the gap between them — it doubles both sides of it. A 300,000-token request on GPT-6 Luna bills at $0.20 per million input for all 300,000 tokens, not for the 28,000 over the line. The same request on GPT-6 Sol bills at $4.00. Split either document into two calls below the threshold and the cost halves with no change to the prompt.

What the identical clause does change is the size of the mistake. On the volume tier, a mis-sized request costs a fifth of a cent per thousand tokens. On the daily tier it costs four dollars per million. The teams most likely to have a 300,000-token working context are the ones running the agentic pipelines both models are sold into, which means the tier that can least afford the mistake is not the tier that makes it most often.

Where Sol's eleven points are, and where they are not

The gap is real and it is not distributed evenly across the tasks you might run.

GPT-6 Sol's advantages cluster in coding and in agentic execution. Its Coding Agent Index of 57 is sixteen points above GPT-6 Luna's 41, on a board where the previous generation's equivalents were 55 and 43 — so both models moved in opposite directions on that measure, which is worth knowing before assuming the new generation is an upgrade on the old one across the board. On the vendor's own published figures, Sol reports 68.8% on DeepSWE v1.1 at max effort against Luna's 66.6%, and roughly 43% on Terminal-Bench 4.0 against 37%. Those are OpenAI-reported and unreproduced, and the gap they describe is narrower than the index gap suggests.

Where Sol pulls away is on the evaluations that reward a long chain of correct decisions: GPT-6 Sol reports about 62% on AutomationBench-AA against GPT-6 Luna's 53%, and a hallucination rate of 60% against 77% on the same omniscience board. A seventeen-point difference in how often a model invents an answer is not a capability margin; for anything with a downstream consumer, it is the whole decision.

Where the eleven points are not is general text work. On extraction, classification, routing and summarisation — the tasks GPT-6 Luna is priced for — both models will produce the same usable answer the overwhelming majority of the time, and the difference will not survive contact with your own eval set. Artificial Analysis flags one regression on both models at max effort: GPT-6 Sol loses roughly 100 Elo on GDPval-AA v2.1 against its predecessor, and GPT-6 Luna roughly 75. If your work looks like GDPval, neither tier is the upgrade you were promised, and the eleven points are a comparison between two models that both went backwards.

Choosing between two tiers you already have

The unusual thing about this pairing is that it is not a purchase decision. Both models run on the same account, the same endpoint and the same SDK. There is no second contract, no second rate negotiation, and no integration work to move traffic from one to the other. That makes it a routing decision, and routing decisions are the ones worth making per request rather than once.

Neither GPT-6 Luna nor GPT-6 Sol is in our catalogue as of this writing — both are reached through OpenAI's own API, so the split has to be expressed in your application rather than in a router config. That is a reason to be deliberate about it, because the cost of getting it wrong is a code change rather than a settings change. The shape of the split is not complicated: GPT-6 Luna for the thousand small calls an agent makes on the way to a task, and GPT-6 Sol for the call that produces the artefact a person will read. Set the effort parameter explicitly on both sides, because the default is medium and every published figure for both models is a maximum-effort number.

Teams that do route across vendors — and most of the ones we see do — can hold GPT-6 Astra, Grok 4.6, Kimi K3 and Qwen3.8-Max on one key through OrcaRouter at provider list price, with 0% markup and failover between upstreams. Adding GPT-6 Luna and GPT-6 Sol to the same picture means a second key rather than a second architecture, which is the smaller of the two problems this pairing creates.

Screenshot of OpenAI's developer documentation page for GPT-6 Luna, showing the model selector, the description 'Our most efficient model for focused, high-volume tasks', reasoning set to High, speed Fast, price $0.1 · $0.5, text and image input with text output, a 1,050,000-token context window, 128,000 maximum output tokens, a May 18, 2026 knowledge cutoff, and pricing cards of $0.10 input, $0.01 cached input, $0.125 cache writes and $0.50 output per 1M tokens.

Which one, and when

Take GPT-6 Luna unless you can name the capability you are buying with the twenty times. It is cheaper on every line of the card including cached input, it is the same window and the same output ceiling, it can be told to stop reasoning entirely, and its cost per completed task on the independent board is roughly a fifteenth of GPT-6 Sol's. Set the effort explicitly, keep long calls under 272,000 input tokens, and check the sixteen-point Coding Agent Index gap against your own repository before you assume it applies to you.

Take GPT-6 Sol when the task is the product and the failure mode is a quietly wrong answer rather than a slow one. Coding agents operating a real repository, long-horizon professional analysis, anything where a hallucinated fact costs more than the tokens saved. Accept that you are paying fifteen times per completed task for eleven index points, and that a meaningful share of those eleven points sits in the hallucination column rather than the capability column — which, for this particular pair, is the better half to be paying for.

The mistake to avoid is treating the twentyfold sticker as the answer. It is the answer to what a token costs, and it says nothing about what a finished task costs or how often the cheap model's answer is wrong in a way you will not notice.

Screenshot of the Artificial Analysis comparison view for GPT-6 Luna and GPT-6 Sol, showing Luna's Intelligence Index of 37 at maximum effort against Sol's 48, pricing of $0.10/$0.50 per 1M against $2.00/$10.00, a 1,050,000-token context window and 128K maximum output for both, Luna's measured output speed of 153.9 tokens per second, and cost per index task of about $0.07 against $1.06.