Generated title card reading "GPT-6 vs GPT-6 Luna", with a chip reading "GPT-6 is a family name, not a model id" and a chip reading "GPT-6 Luna: $0.10 in / $0.50 out, powers ChatGPT Free and Go", and a footer reading "Models shipped 22 Sep 2026; tier rule per OpenAI's own announcement." The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

GPT-6 vs GPT-6 Luna: A Dime a Million Tokens, and the Tier That Answers the Free Plan

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is a version of this comparison where the answer is one line: you cannot call GPT-6, and you can call GPT-6 Luna for ten cents per million input tokens. But the more useful version is the one that explains why Luna is now the model most people have actually used — OpenAI put it behind the Free and Go tiers of ChatGPT on 8 October 2026, which makes it the cheapest tier in the generation and, by user count, the busiest. What follows is the tier ladder explained once, and then what a dime per million really buys.

The family, and where Luna sits in it

OpenAI ships the GPT-6 generation as three model ids released on 22 September 2026, with no openai/gpt-6 endpoint behind the family name:

• GPT-6 Astra — the flagship, $10.00 in / $50.00 out per million tokens, listed on OpenAI's rate card as the top of the generation
• GPT-6 Sol — the middle tier, $2.00 / $10.00, and the model OpenAI names for chat on the paid ChatGPT tiers
• GPT-6 Luna — the cheap tier, $0.10 / $0.50, and the model OpenAI names for the Free and Go tiers

That last line is the part worth pinning down, because it is the reason Luna's traffic dwarfs everything else in the family. OpenAI's own announcement for the 7–8 October rollout states that "GPT-6 in ChatGPT is powered by GPT-6 Sol for Plus, Pro, Business, and Enterprise tiers, and GPT-6 Luna for Free and Go tiers." A model at $0.10 / $0.50 answering the free plan is a model serving more questions per day than any premium tier, and it is the same checkpoint you can call over the API. The consumer surface is a distribution channel, not a separate SKU.

A caveat that has to be repeated here because this blog has published on it: the same announcement says Free and Go get GPT-6 Luna, while OpenAI's help-centre page for GPT-6 in ChatGPT described Free and Go as running GPT-5.6 Luna on the day it was read. Those are not compatible statements about the same tier, OpenAI has not said which is current, and if your work depends on which cheap model sits behind the free plan, treat the help-centre entry as the more specific page and re-check before you build against either.

What ten cents per million actually buys

The headline is the least interesting number on Luna's card. Here is the whole card, and the tier step underneath it:

• Short context, up to 272,000 input tokens — $0.10 per million in, $0.50 per million out
• Above 272,000 input tokens — $0.20 per million in, $0.75 per million out, applied to the whole request
• Cached input — $0.01 per million at the short tier, a 90% discount; $0.02 above the line
• Cache writes — $0.125 per million at the short tier, $0.25 above it
• Context and output — 1,050,000 tokens in, 128,000 tokens out
• Input modalities — text, image and file, the same multimodal input as Sol and Astra
• Reasoning — a configurable effort parameter, alongside tools, JSON output and seeded sampling

Two things follow. The first is that Luna is not a stripped model. It takes images and files, it exposes the same tool-calling and structured-output surface as its two more expensive siblings, and the only thing that separates it from Sol on the spec sheet is depth rather than capability. The second is the step. A prompt above 272,000 input tokens doubles Luna's input rate and raises its output rate by half, across the entire request rather than only the overflow. That is a gentle step at these prices — a 300,000-token request costs about six cents of input instead of three — so the long-context clause that dominates budget planning for Astra and matters for Sol is nearly irrelevant here.

The number that does move money at scale is the cached read. At one cent per million, a workload that re-sends the same large context on every turn is paying almost nothing for the repeated portion. On a high-volume extraction or classification pipeline with a stable system prompt and a stable document, that is the difference between a cheap model and a nearly free one, and it is the specific reason Luna's real cost per finished answer sits far below what the rate card suggests.

What it costs you in capability

Luna is the cheapest tier because it is the least capable one, and the independent board is unambiguous about how much less. Read off Artificial Analysis via the OrcaRouter catalogue entry for each model on 8 October 2026:

• AA Intelligence Index — GPT-6 Sol 47.6, rank 14; GPT-6 Luna 38.1, rank 37
• Humanity's Last Exam — GPT-6 Sol 47.9 vs GPT-6 Luna 38.5
• SciCode — GPT-6 Sol 57.6 vs GPT-6 Luna 54.6
• Long-context recall — GPT-6 Sol 83.7 vs GPT-6 Luna 83.3
• Terminal-Bench 4.0 — GPT-6 Sol 43.9 vs GPT-6 Luna 12.6
• Traffic over seven days — GPT-6 Luna roughly 574 million tokens; GPT-6 Sol roughly 2.1 million
• Median time to first token — GPT-6 Luna 1,494 ms vs GPT-6 Sol 6,785 ms
• Median output speed — GPT-6 Luna 132.9 tokens/s vs GPT-6 Sol 179.6 tokens/s

The shape of that is more informative than the headline gap. Luna is 9.5 index points behind Sol, and the loss is not spread evenly: long-context recall is a 0.4-point tie, and SciCode — code comprehension rather than code production — is only 3 points apart. What collapses is agentic, multi-step execution: Terminal-Bench 4.0 at 12.6 against Sol's 43.9. A model that can retrieve from a million-token document about as well as its bigger sibling, and write a competent single-shot answer, is a different instrument from one that can drive a tool loop to completion.

Two caveats. The index figures come from Artificial Analysis, not from OpenAI, and that evaluator does not guarantee that two model pages are rendered against the same index revision on the same day — treat sub-point differences as direction, not precision. And the serving figures are our own routing measurements over a rolling seven-day window; they drift between reads and are useful for the shape of the difference rather than as a service-level promise.

There is one serving number worth stopping on. Luna's median time to first token is 1,494 ms against Sol's 6,785 ms — roughly 4.5 times faster to the first token — and its throughput is materially lower once streaming starts. For a chat surface where the user is waiting on the opening line, that inversion is the entire point of the tier, and it is why the free plan can feel fast while running the cheapest model in the family.

The fresh competition inside the family

Luna's price advantage inside the generation has narrowed since it shipped. GPT-6.1 Sol arrived on 29 September at Sol's $2.00 / $10.00 with an index of 51.8, and it cut the cached-input rate to $0.10 per million. That does not touch Luna's short-context rate — a dime in and fifty cents out is still twenty times cheaper on input than any Sol — but it does mean the tier above Luna got cheaper to run on cached traffic, which is precisely the workload where Luna's advantage was widest. The gap is still large. It is no longer as large as the rate cards alone imply.

A generated two-column comparison scoreboard titled "GPT-6 Sol vs GPT-6 Luna — the scoreboard". The left column, GPT-6 Sol, reads Input $2.00, Output $10.00, AA Intelligence 47.6, Terminal-Bench 4.0 43.9, First token 6,785 ms, Cached input $0.20. The right column, GPT-6 Luna, reads Input $0.10, Output $0.50, AA Intelligence 38.1, Terminal-Bench 4.0 12.6, First token 1,494 ms, Cached input $0.01. A footer line reads "Index figures per Artificial Analysis; serving figures are a rolling seven-day window from OrcaRouter." The OrcaRouter logo is composited in the bottom-right corner.

Where the tier line falls in practice

Split the work by what the task needs, not by which model scores higher.

Luna is the right call for high-volume chat, classification, extraction, moderation and summarisation over a long document — the row where it ties its bigger sibling on retrieval is exactly the row that governs whether a large context is usable. Its first-token latency makes it the better choice for anything interactive that is not reasoning-heavy. And at a cent per million cached, a stable-context pipeline is close to free on the repeated portion.

Sol, or now GPT-6.1 Sol, is the right call the moment the task requires a tool loop, a long-horizon plan, or a chain of steps that has to complete. The Terminal-Bench spread — 43.9 against 12.6 — is the single clearest signal in this comparison, and it maps to the real failure mode: a cheap model that answers the first step well and then loses the thread. The effort parameter means you can ask Luna for more reasoning, but effort raises the token count rather than the ceiling; it does not turn a 12.6 into a 43.9.

And Astra stays the answer for the work that only the flagship can do — which, at $10.00 in and $50.00 out against Luna's ten cents, is a decision you should be able to justify per call rather than per team.

The one-API version of this choice

The awkward part of a three-tier family is that the tier boundary moves per feature, and it moves per request — the same application usually wants Luna for one endpoint and Sol for the next. That is a routing problem rather than a procurement one. All three GPT-6 tiers sit in the same OrcaRouter catalogue, called through one OpenAI-compatible API in front of 200+ models, with the provider's list price passed through at 0% markup, so a vendor price move on any tier is live here the same day rather than at your next reconciliation. The routing DSL composes the split explicitly — by request size, by endpoint, or by share — so the classification path can run Luna while the agentic path runs Sol, and automatic failover across providers covers a route that degrades rather than leaving you to discover it from a failed run.

One thing this does not do is give you a Luna that behaves like Sol. Tiers exist because the models differ, and the numbers above are the size of the difference. What a single key and a routing rule buy you is the ability to put each request on the tier that actually fits it, instead of putting everything on the most expensive one for safety or everything on the cheapest one for budget.

Screenshot of the OrcaRouter model page for openai/gpt-6-luna, showing the OpenAI vendor label, a 2026-09-22 catalogue release date, a 1.05M-token context window, a 128K maximum output, text, image and file input with Vision, Tools, JSON and Reasoning badges, and a rate strip reading $0.10 per million input, $0.50 per million output, a 1.49 s median time to first token, a 5.06 s p95, and 574.0M tokens routed in seven days.

The answer, by the only question that matters

If the task is retrieve-then-answer over a large document, chat, or bulk text processing, GPT-6 Luna is the tier and the ten-cent card is not a compromise — the long-context recall row says it retrieves as well as the model costing twenty times more. If the task needs a multi-step agent to finish, Luna is the wrong tier and the honest move is GPT-6 Sol or GPT-6.1 Sol, where the price is still low in absolute terms and the agentic evaluations stop collapsing.

What "GPT-6" is on its own, again, is nothing you can call. It is a family name covering three ids at three price points and three different capability ceilings, and Luna is the one at the bottom — the cheapest, the fastest to first token, the one answering the free plan, and the one whose ceiling shows up the moment a task needs more than a single confident answer.

Screenshot of the OrcaRouter model page for openai/gpt-6-sol, showing the OpenAI vendor label, a 2026-09-22 catalogue release date, a 1,050,000-token context window (shown as 1.05M-token context in the model blurb), a 128K maximum output, text, image and file input, and a rate strip reading $2.00 per million input, $10.00 per million output, a 6.79 s median time to first token, a 10.00 s p95, and 2.1M tokens routed in seven days.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily