A generated hero title card reading 'GPT-6 Luna vs GPT-6 Astra', subtitled 'A hundred times the price on the rate card. Forty-six times per finished task.', with chips showing Luna at $0.10/$0.50 per 1M with index 37, Astra at $10.00/$50.00 per 1M with index 53, and cost per index task of $0.07 against $3.26, with the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6 Luna vs GPT-6 Astra: A Hundred Times the Price for Sixteen Index Points

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The vendor shipped both of these models inside the same month. GPT-6 Astra went out on September 3, 2026 at $10.00 per million input tokens and $50.00 per million output, positioned as the company's most capable system for long-horizon professional work. GPT-6 Luna followed on September 22, 2026 at $0.10 and $0.50 — one hundredth of the input price, one hundredth of the output price — as the family's efficiency tier. Sixteen index points separate them on the independent board that measures both. Sixteen points for a hundredfold price is the whole comparison in one line, and it is not the interesting part.

The interesting part is that the hundredfold sticker shrinks to something closer to forty-sixfold once you account for how each model actually behaves on a task, and that the remaining gap is not general intelligence at all. It is computer use, long-horizon autonomy, and a safety classification that decides whether your organisation is allowed to call the model in the first place. That last item is the one no benchmark table captures, and for a lot of teams it settles the question before price is mentioned.

Two tiers of one family, nineteen days apart

Astra is OpenAI's flagship and the model the company describes as the world's most intelligent and aligned. It accepts text, image and file input, carries a 1,050,000-token context window with a 128,000-token output cap, and has always-on reasoning with configurable effort from low through max. It is the first OpenAI model to cross the Critical cybersecurity capability threshold under the company's Preparedness Framework, and the rollout followed from that classification rather than from capacity: limited organisations first, enterprise access switched off by default and administrable from September 9, with safeguard tiers widening afterwards.

Luna is the other end of the same family. Text and image in, text out, the same 1,050,000-token context window and the same 128,000-token output ceiling, and a reasoning ladder that runs from none through low, medium, high, xhigh and max with medium as the default. No gating, no enterprise switch, no safety tier to qualify for. It is generally available to anyone with an API key, and OpenAI's own description is narrow and honest: the most efficient model for focused, high-volume tasks.

The two models share a context window, an output cap, a knowledge-cutoff family and a long-context pricing clause. They share almost nothing else, and the reason is that they were built for different halves of the same job.

The rate card says one hundred times. The task cost says forty-six.

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs GPT-6 Astra $10.00

• Output per 1M — GPT-6 Luna $0.50 vs GPT-6 Astra $50.00

• Cached input — GPT-6 Luna $0.01 per 1M vs GPT-6 Astra $1.00 per 1M, both a 90% discount on their own uncached rate

• Cost per index task — GPT-6 Luna about $0.07 vs GPT-6 Astra about $3.26

• Cost to run the full index — GPT-6 Luna $122.39 vs GPT-6 Astra $5,324.10

• AA Intelligence Index — GPT-6 Luna 37 at max effort vs GPT-6 Astra 53 at max effort

• Output tokens on the index run — GPT-6 Luna 150M vs GPT-6 Astra 60M, against a board median of 88M

• Output speed — GPT-6 Luna 153.9 tokens/sec vs GPT-6 Astra 51.4 tokens/sec at xhigh

• Time to first token — GPT-6 Luna fast enough to sit behind an interactive surface vs GPT-6 Astra 208.73s at xhigh

• Reasoning off switch — GPT-6 Luna none through max, with none as an option vs GPT-6 Astra always on, no non-reasoning mode

• Computer use and autonomy — GPT-6 Luna tool calling and agentic loops vs GPT-6 Astra built for operating a computer end-to-end

• Access — GPT-6 Luna generally available to anyone vs GPT-6 Astra gated, enterprise off by default, admin controls required

• On OrcaRouter — GPT-6 Luna not in our catalogue vs GPT-6 Astra routable at OpenAI's list price

A generated single-panel scoreboard card headed 'A hundred times on the rate card' with six rows: GPT-6 Luna $0.10 in / $0.50 out per 1M with index 37 at max effort; GPT-6 Astra $10.00 in / $50.00 out per 1M with index 53 at max effort; cost per completed index task about $0.07 against about $3.26; output tokens on the index run Luna 150M against Astra's 60M and a board median of 88M; output speed Luna 153.9 tokens per second against Astra's 51.4 at xhigh; and access, Luna open to anyone with an API key against Astra gated with enterprise off by default. Footer: 'Index, cost per task and speed per Artificial Analysis; prices per OpenAI.'

Read the cost-per-task row against the rate-card ratio and the gap collapses by more than half. Luna spends 150 million output tokens completing the index; Astra spends 60 million. Astra is two and a half times more concise, and on an output-priced model that is worth a great deal — it is the reason a hundredfold rate-card difference becomes a forty-sixfold task-cost difference, and the reason any comparison built on sticker prices alone will overstate the case for the cheap tier by a factor of two.

The speed rows run the other way and are not close. Luna generates three times as many tokens per second as Astra and returns its first token in a time that permits a user to be waiting. Astra's 208.73-second time to first token at xhigh is not a latency figure; it is a statement about what the model is for. Nothing that a person is watching can run on Astra at that setting.

What sixteen points actually buys

The index gap is real, and the useful question is what is inside it, because the answer is not "sixteen percent more intelligence."

Astra's published capabilities cluster in a specific place. OpenAI reports 72.6% on OSWorld 2.0 and 69% on AutomationBench-AA — both measures of a model operating software the way a person does, across many steps, over a long stretch of time. On the knowledge side the figures are 96.1 on GPQA Diamond, 97.6% on FrontierMath Tier 4 and 57.2% on Humanity's Last Exam with tools, and long-context retrieval is a highlighted strength: 100% on OpenAI's own eight-needle benchmark between 256K and 512K tokens, 96.3% between 512K and 1M. These are vendor-reported and unreproduced, and the harness matters — the same model scores 99.9% on ARC-AGI-3 under OpenAI's custom provider adapter against 62.7% on the standard harness, which is a swing that says as much about the harness as the model.

Read together, that is not a general-purpose advantage. It is an advantage concentrated in tasks that take a long time, use a computer, and have a verifiable endpoint. Astra's competitive framing is not against another chatbot; it is against a working session, and the models it is measured beside on the independent board are the other flagships at the same tier — it ties Claude Fable 5.1 at 53 on the Intelligence Index at roughly 40% of the cost per task.

Luna's sixteen-point deficit is therefore not distributed evenly. On a short, well-specified, program-consumed task, the two models will frequently produce the same usable answer and the difference will be invisible. On a task that requires forty sequential actions in a browser with a checkpoint at the end, the deficit is the difference between completing and not — and that is the task Astra was built for and Luna was not.

The gate nobody prices in

Astra is the first OpenAI model to cross the Critical cybersecurity threshold under the company's Preparedness Framework, and the practical consequence is that it ships switched off for enterprise accounts. An administrator has to enable it, under an applicable rate card and agreement, before users see it at all. Access opened to a limited set of organisations on September 3 and widened from there; the admin machinery arrived on September 9.

Luna has none of that. It is available to anyone with an API key on day one, with no qualification step, no admin toggle, and no safeguard tier standing between a developer and a first call.

For a team in a regulated industry, a defence contractor, or anywhere with a security review in the procurement path, this is the whole decision. The hundredfold price difference is a line item; an access gate is a project. And for a team outside those constraints, the gate is irrelevant and Astra is simply an expensive model with an unusually long time to first token.

It is worth adding that the gating is a considered position rather than an inconvenience. A model that operates a computer autonomously and finds vulnerabilities well is a model a lab should be careful about handing out, and the deployment shape follows the capability finding. That does not make it easier to procure, but it does mean the constraint is unlikely to be relaxed by a support ticket.

Same cliff, same family, twice

Both models carry the identical long-context clause, which is the most expensive thing to miss on either rate card. Requests that exceed 272,000 input tokens are billed at twice the input and cache rates and 1.5x the output rate — for the entire request, not the portion above the line. On Astra that turns a $10.00/$50.00 call into $20.00/$75.00. On Luna it turns $0.10/$0.50 into $0.20/$0.75.

Because the clause is proportional, it does not change the ratio between the two models — it doubles both. What it changes is the absolute size of the mistake, and the mistake is more likely on Astra, because the workloads Astra exists for are the ones that carry a million-token working context. A single long-context Astra call costs $75 per million output tokens; splitting the same work into two sub-272K calls halves it with no change to the prompt. On a model whose entire value proposition is long-horizon work, that arithmetic is worth running before the first production call, not after.

The decision is a routing decision

What makes this pairing unusual among the comparisons on this blog is that both models belong to the same vendor, run on the same account, and are reached through the same endpoint. There is no second contract to sign and no second SDK to learn. The choice between them is not a purchase decision at all — it is a routing decision, and it is one most teams should make per request rather than once.

The shape of the split is clear enough. Luna for the thousand small calls an agent makes on the way to a task: classification, extraction, tool selection, summarisation of an intermediate result. Astra for the task itself, once, at the end, where the answer is the product. That is a two-tier pipeline inside one application, and the cost difference between running the whole thing on Astra and running the inner loop on Luna is the difference between a viable unit economics and an unviable one.

GPT-6 Astra is on OrcaRouter at OpenAI's list price, under pass-through pricing that means an OpenAI rate change is live on our side the same day. GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API, so a team that wants both runs one key for GPT-6 Astra alongside whatever it already uses for OpenAI. This is also the pairing where model fusion earns its keep: a panel of a cheap high-throughput model and an expensive careful one answering together, with the cheap member doing the bulk of the work and the expensive one adjudicating, is a configuration the routing layer can express without the application knowing either model exists.

A screenshot of the Artificial Analysis page for GPT-6 Luna (max), showing an Intelligence rank of 6th of 183 models, a Speed rank of 35th, a Cost rank of 20th and a Verbosity rank of 36th, input pricing of $0.10 and output of $0.50 per 1M tokens, an Intelligence Index score of 37 against a comparable-model median of 12, 150M output tokens generated during the index run, 153.9 tokens per second, and a total of $122.39 spent evaluating the model.

Which one, and when

Take GPT-6 Luna unless you have a specific reason not to. It is a hundredth of the price on the rate card and a forty-sixth of the price per completed task, it is three times faster, its first token arrives in a time a user can wait for, and it can be told to stop reasoning entirely — which is what makes it usable as a classifier. Set the effort parameter explicitly, because the default is medium and the headline 37 is a max-effort number. Keep long calls under 272,000 input tokens. Check the two-point Coding Agent Index regression against your own eval rather than a vendor chart.

Take GPT-6 Astra when the task is the long one and the answer is the product. Computer use across many steps, professional analysis with a verifiable endpoint, anything where sixteen index points is the difference between finishing and quietly stopping. Accept the time to first token — 208 seconds at xhigh is not a figure you can put behind a human — and accept the procurement gate if your organisation has one. Then run everything that leads up to that task on the cheap tier, because the inner loop is where the money goes.

The mistake this comparison invites is treating it as a choice between two models. It is a choice about where in a pipeline each one sits, and the answer is almost always both.

A screenshot of the OrcaRouter model page for GPT-6 Astra (openai/gpt-6-astra), showing the Vision, Tools, JSON and Reasoning badges, a 2026-09-04 catalogue date, 1M-token context with 128K maximum output and text, image and file input, list pricing of $10.00 input and $50.00 output per 1M tokens, a p50 time to first token of 7.81s, and the description of GPT-6 Astra as OpenAI's flagship model for long-horizon end-to-end work with always-on reasoning from low through max.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily