Generated hero title card headlined 'GPT-6 Luna vs Grok 4.6' with the subtitle 'Twenty times the price, and a cliff at 200,000 tokens', a footer reading 'Prices per OpenAI and xAI; index figures per Artificial Analysis.', and the OrcaRouter logo composited bottom-right.
Guides & Insights

GPT-6 Luna vs Grok 4.6: The Price Gap Is Twenty Times and the Window Is the Real Difference

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put GPT-6 Luna and Grok 4.6 on the same page and the first comparison writes itself. Luna lists at $0.10 per million input tokens and $0.50 per million output. Grok 4.6 lists at $2.00 and $6.00 with a $0.50 cached read. Twenty times on input, twelve times on output, and a cached-input rate that is fifty times higher on the vendor's model. That is the sticker, and on the sticker the answer is not close.

The reason it is worth writing the rest of this piece is that the two models disagree about something more structural than price. Grok 4.6 carries a 500,000-token context window and reprices the entire request once the prompt passes 200,000 tokens. GPT-6 Luna carries 1,050,000 tokens and reprices at 272,000. Both thresholds are cliffs rather than marginal rates — cross one and the whole request bills at the higher number, not the part above the line. Where those two cliffs sit, and what happens to a long-context workload on the wrong side of each, decides more of this comparison than the twentyfold headline does.

Two models, two very different postures

Grok 4.6 is xAI's model from August 12, 2026, and it is a general-purpose frontier release: text in, text out, a 500,000-token window, a published per-task cost of $1.86 on the independent board at high effort, and a measured speed of 57.7 output tokens per second. It is the kind of model a team adopts because it is good at a lot of things and the vendor is shipping fast.

GPT-6 Luna is the opposite posture. OpenAI released it on September 22, 2026 as the volume tier of the GPT-6 family, beneath GPT-6 Sol at $2.00/$10.00 and the flagship GPT-6 Astra at $10.00/$50.00. Text and image in, text out, 1,050,000-token window with a 922,000-token maximum input, a 128,000-token output cap, and a reasoning ladder running from none through low, medium, high, xhigh and max with medium as the default. It is not a model anyone adopts for its breadth. It is a model you put underneath a workload.

So the honest framing is not "which of these is better." It is "which of these is priced for the shape of work you have," and the two shapes are genuinely different.

The cards, line by line

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs Grok 4.6 $2.00

• Output per 1M — GPT-6 Luna $0.50 vs Grok 4.6 $6.00

• Cached input per 1M — GPT-6 Luna $0.01 vs Grok 4.6 $0.50, a tenth of its own input rate against a quarter of xAI's

• Long-context threshold — GPT-6 Luna 272,000 input tokens vs Grok 4.6 200,000 prompt tokens

• Long-context rates — GPT-6 Luna reprices the whole request to $0.20 in, $0.75 out, $0.02 cached vs Grok 4.6 reprices the whole request to $4.00 in, $12.00 out, $1.00 cached

• Context window — GPT-6 Luna 1,050,000 tokens vs Grok 4.6 500,000 tokens

• Maximum output — GPT-6 Luna 128,000 tokens vs Grok 4.6 not published as a separate ceiling on the card

• Intelligence Index, independent — GPT-6 Luna 37 at max effort vs Grok 4.6 44 at high effort on index revision v4.3.2

• Default-effort score — GPT-6 Luna 29 at its default medium vs Grok 4.6 43 at medium, where the board also measures it

• Cost per Index task, independent — GPT-6 Luna about $0.07 vs Grok 4.6 $1.86 at high effort

• Output tokens on the index run — GPT-6 Luna 150M vs Grok 4.6 94M, against a board median of 88M

• Output speed, independent — GPT-6 Luna 153.9 tokens/sec vs Grok 4.6 57.7 tokens/sec at high

• Time to first token, independent — Grok 4.6 38.63s, end-to-end 47.31s; GPT-6 Luna returns a first token in a time that permits a user to be waiting

• Cyber capability, independent — Grok 4.6 scores 57 on the board's cyber evaluation; no comparable GPT-6 Luna figure is published

• On OrcaRouter — GPT-6 Luna not in our catalogue vs Grok 4.6 routable at xAI's list price

Generated two-column scoreboard titled 'GPT-6 Luna vs Grok 4.6 - the scoreboard'. Left column GPT-6 Luna rows: Price in/out $0.10 / $0.50, Cached input $0.01, Long-context threshold 272,000, Context 1,050,000, AA Index 37 at max effort, Cost per index task $0.07. Right column Grok 4.6 rows: Price in/out $2.00 / $6.00, Cached input $0.50, Long-context threshold 200,000, Context 500,000, AA Index 44 at high effort, Cost per index task $1.86. Footer: 'Prices per OpenAI and xAI; index figures per Artificial Analysis.'

200,000 against 272,000 is the whole argument

Read the two thresholds next to the two windows and the comparison reorganises itself.

Grok 4.6's cliff sits at 200,000 tokens inside a 500,000-token window — 40% of the way in. GPT-6 Luna's sits at 272,000 inside 1,050,000 — 26% of the way in. So not only does the cheaper model start repricing later, it has more than twice as much room after the threshold before it runs out of window at all.

Concretely: a 450,000-token request fits inside both models. On GPT-6 Luna it bills at the long-context tier of $0.20 and $0.75. On Grok 4.6 it bills at $4.00 and $12.00. That is twenty times on input and sixteen times on output for the same prompt — and it is a request that Grok 4.6 can serve, so the choice is real rather than theoretical.

The reverse case is the one that catches people out. A 190,000-token request sits below both thresholds and bills at the standard rates on both: $0.10/$0.50 against $2.00/$6.00, a clean twentyfold and twelvefold. A 210,000-token request sits above Grok 4.6's line and below GPT-6 Luna's. It bills at $4.00/$12.00 on Grok 4.6 and at $0.10/$0.50 on Luna — a fortyfold and twenty-fourfold gap produced entirely by a 20,000-token difference in prompt length, with no change to either model.

That is not a reason to avoid Grok 4.6. It is a reason to know where your prompts actually land, because a workload that averages 180,000 tokens and spikes to 220,000 is paying two different rate cards and will look like a billing error the first month it happens.

What the price gap buys on the independent board

Grok 4.6 scores 44 on the Artificial Analysis Intelligence Index at high effort. GPT-6 Luna scores 37 at maximum effort. Seven points, on a composite where a single point is inside the noise of a re-run — and the two models are separated by roughly twenty-six times on cost per completed task, $1.86 against about $0.07.

Two things about that pair of numbers are worth separating.

The first is the effort default. GPT-6 Luna's 37 is a maximum-effort figure and its default is medium, where it scores 29. Grok 4.6's 44 is a high-effort figure, and the board also measures it at medium, where it scores 43. So the like-for-like comparison at default settings is 29 against 43 — fourteen points, not seven, and the fourteen-point version is the one a team that deploys both and changes nothing will actually experience. Grok 4.6 loses one point going from high to medium. GPT-6 Luna loses eight. The effort dial costs the cheap model far more than it costs the expensive one, and that asymmetry is invisible in the headline index figures.

The second is verbosity, and here the direction is the opposite of what the price ratio suggests. GPT-6 Luna generated 150 million output tokens completing the index; Grok 4.6 generated 94 million. The model that costs a twelfth as much per output token spent 60% more tokens getting to a lower score. On an output-priced card that is the difference between a twenty-sixfold cost-per-task gap and a smaller one, and it is the reason any comparison built on sticker prices alone will overstate the case for the cheap tier.

Grok 4.6 also publishes a cyber capability score of 57 on the same board. There is no comparable GPT-6 Luna figure, so that dimension has one side — but it is the kind of line that matters if your workload touches security tooling, and its absence on the cheaper model is information rather than a gap to be filled by assumption.

Latency, where the two are not close and not in the same direction

Grok 4.6's independent latency figures are 38.63 seconds to first token and 47.31 seconds end to end at high effort. That is a figure for a model doing serious reasoning before it speaks, and it is not compatible with a person watching a cursor. Our own listing for the model records a p50 time to first token of 6.71 seconds and a p95 of 10.00 seconds over a seven-day window, which measures a different point in the response and is not comparable to the independent figure — the two numbers are not in disagreement, they are answering different questions.

GPT-6 Luna runs at 153.9 output tokens per second and returns its first token fast enough to sit behind an interactive surface. It is nearly three times Grok 4.6's throughput at high effort. For a batch job that difference is a wall-clock convenience. For anything a user is waiting on, it is the difference between a product and a prototype.

The combination is unusual and worth naming plainly: the cheap model is the fast one, the expensive model is the slow one, and the expensive model is seven points ahead on the composite at matched effort. That is the profile of a model built to think hard once, and a model built to answer quickly a great many times. If your workload is one call per task, Grok 4.6's latency is affordable. If it is a thousand calls per task, it is not.

What you are actually buying on the xAI side

Grok 4.6 costs roughly twenty-six times as much per completed task as GPT-6 Luna and is seven index points ahead at matched effort. That is a poor exchange rate for a general-purpose workload, and it is a reasonable one for a specific class of work.

Three things justify it. The cyber score of 57 is a capability GPT-6 Luna does not publish a counterpart to. The 94-million-token index run against Luna's 150 million is a real concision advantage on any task where the output is consumed by a person rather than a parser. And the model has been in production since August 12, six weeks longer than GPT-6 Luna has existed at all — which is not a benchmark, but it is a track record, and for a team that has already built against it, the migration cost of leaving is part of the price of leaving.

Against that, GPT-6 Luna's case is not primarily the twentyfold rate. It is the 272,000-token threshold inside a million-token window, the ability to be told to stop reasoning entirely, a cached input rate of one cent per million, and a throughput figure that makes it usable as a component rather than an endpoint. Those are structural properties of the cheap tier, and no amount of index advantage on the other side replaces them.

Running one of them, or both

Grok 4.6 is on OrcaRouter at xAI's list price, under pass-through pricing that puts a vendor rate change on our side the same day rather than waiting on a reseller to renegotiate. Automatic failover is the part that earns its place for this model specifically: a 500,000-token window with a 200,000-token cliff is a workload where a request occasionally lands on the wrong side of the line, and a route that moves between upstreams without a deployment is worth more than a fraction of a cent per token.

GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API. So a team that wants both runs one key for Grok 4.6 alongside whatever it already uses for OpenAI. The routing DSL is the piece that fits this pairing: a rule that sends prompts above a token threshold to the model whose cliff they clear, and everything below it to the model that is a twelfth of the price, is expressible as configuration rather than as a branch in your application — which matters when the thresholds are as close to common prompt sizes as these two are.

Screenshot of OpenAI's developer documentation page for GPT-6 Luna, showing the model selector, the description 'Our most efficient model for focused, high-volume tasks', reasoning set to High, speed Fast, price $0.1 · $0.5, text and image input with text output, a 1,050,000-token context window, 128,000 maximum output tokens, a May 18, 2026 knowledge cutoff, and pricing cards of $0.10 input, $0.01 cached input, $0.125 cache writes and $0.50 output per 1M tokens.

Which one, and when

Take GPT-6 Luna if your workload is high-volume, program-consumed and short. Classification, extraction, routing, bulk summarisation, the inner loop of an agent. At $0.10/$0.50 with a one-cent cached read, against a model that is twenty-six times the cost per completed task for seven index points at matched effort, the arithmetic is not close. Set the effort parameter explicitly — the default is medium and the headline 37 is a max-effort number — and keep your long calls under 272,000 tokens.

Take Grok 4.6 if you need the cyber capability, if your output is read by a person and its concision is worth paying for, or if you have a 300,000-to-500,000-token workload that GPT-6 Luna can serve but where you would rather not be the first team to find out how it behaves at that length. Budget for the 200,000-token cliff explicitly, and do not put it behind an interactive surface at high effort — 38 seconds to first token is not a latency figure, it is a statement about what the model is for.

The mistake this comparison invites is treating the twentyfold rate as the decision. It is the decision for one workload shape. For the other, the decision is made at 200,000 and 272,000 tokens, and the price ratio is a detail you read afterwards.

Screenshot of the OrcaRouter model page for Grok 4.6 showing the xAI vendor label, list pricing of $2.00 input and $6.00 output per 1M tokens, a 500,000-token context window, recorded latency statistics, an uptime chart, and the provider routing section.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily