A title card for an analysis of GPT-5.6 Luna's 80% price cut, showing a large downward arrow labelled 80% cut, the pre-cut rate of $1.00 per million input tokens and $6.00 per million output, and the current rates of $0.20 input and $1.20 output.
Guides & Insights

GPT-5.6 Luna: What OpenAI's 80% Price Cut Actually Changed

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-5.6 Luna has been 80% cheaper than its launch price since late July, and Open​AI has started treating that as an argument rather than a discount. The model — the fast, low-cost tier of the GPT-5.6 family — went generally available on July 9, 2026 at $1.00 per million input tokens and $6.00 per million output. It is billed today at $0.20 and $1.20. On September 9, Open​AI's chief financial officer Sarah Friar told the Goldman Sachs Communacopia conference that the cut produced roughly a tenfold increase in usage, and that deploying Luna is cheaper than running Z.ai's GLM 5.3 on a cloud layer.

That last claim is Open​AI's own, made by an executive at an investor conference, and it deserves the same treatment as any other vendor number: worth reporting, worth checking, not worth repeating as settled fact. The price itself is a different matter. It is public, and it has not moved back.

What actually changed, and when

Nothing about GPT-5.6 Luna's capabilities, context window or availability changed in September. What changed is how Open​AI is selling it, and that shift has five dates behind it.

June 26, 2026 — Open​AI announced the GPT-5.6 family (GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna) in limited preview for a small group of partners, following coordination with US government agencies.

July 9, 2026 — general availability. Sol, Terra and Luna reached the API at $5/$30, $2.50/$15 and $1/$6 per million input/output tokens respectively.

Late July 2026 — the 80% cut on Luna alone, to $0.20 in and $1.20 out. Cached input reads run at roughly a 90% discount. Above the long-context threshold the rate roughly doubles, to about $0.40/$1.80.

August 6–7, 2026 — Luna replaced GPT-5.5 Instant as the default model for ChatGPT Free and Go, with unlimited text chats and a "Think" button for deeper reasoning arriving the following week. Open​AI's internal evaluation put Luna's rate of responses containing at least one factual error about 62% below GPT-5.5 Instant on financial, medical and legal prompts — a vendor-reported figure, not an independent one.

September 9, 2026 — Friar's remarks, which put public numbers on the cut's effect for the first time: roughly 10x usage, enterprise revenue up 32% from June to July against 20% overall annualized growth, and Codex at 25 million users.

Read together, those dates describe a two-month conversion of a cheap tier into a volume strategy. Luna was always the tier meant to absorb high-volume, low-complexity work — classification, extraction, routing, first-pass drafts. The cut made the arithmetic behind that positioning much harder to argue with.

The number that matters more than the token price

Screenshot of the Artificial Analysis model page for GPT-5.6 Luna (max), showing an Intelligence Index of 38 ranked 4th of 177 models, output speed of 109.4 tokens per second, input pricing of $0.20 and output pricing of $1.20 per million tokens, a cache discount of 90%, and a cost of $0.18 per Intelligence Index task.

Sticker prices are where most comparisons stop, and for Luna they flatter it. At $0.20/$1.20 it undercuts Anth​hropic's Claude Haiku 4.5 — the closest thing it has to a direct competitor — by roughly 5x on input and 4x on output.

Artificial Analysis measures something less convenient: cost per completed Intelligence Index task, which folds in how many tokens a model actually burns to finish the same work. On the current index, GPT-5.6 Luna's reasoning configuration comes in at $0.18 per task. Claude Haiku 4.5 in its reasoning configuration lands at $0.21.

That is a gap of about 14%, not 5x. The reason is verbosity. Completing that evaluation set took Luna roughly 150M output tokens against Haiku's 78M. Luna reasons at length, and output tokens are billed at six times the input rate.

The practical read splits cleanly. For short-output, high-volume work — classification, tagging, extraction, request routing, first-pass drafts — the 80% cut is real, large, and shows up directly on the invoice. For work where the model reasons before it answers, a substantial share of the discount is consumed by extra output tokens. Both statements are true at once, and the second is the part the cut's coverage left out.

It is also worth knowing what the index says about capability, because the tiers are not close. Artificial Analysis currently scores GPT-5.6 Luna at 38, fourth of 177 models in its class. Claude Haiku 4.5's reasoning configuration scores 18, 138th of 200. Two caveats on those numbers. First, the index was re-scored in September 2026, so launch-time figures — including the 51 and 52 that still circulate for Luna — are quoted from a retired scale. Second, Artificial Analysis has adopted GPT-5.6 Luna as the sole grader for several of its own evaluations, which is worth noting when reading its scores.

Why Open​AI is making this argument now

An 80% cut announced in late July did not need a CFO to relitigate it in September. The fact that she did tells you what the cut was for.

Open-weight Chinese models have spent 2026 pressing on exactly the axis Luna occupies: adequate quality at a price that makes per-token accounting the deciding factor. Undercutting that on a cloud deployment is a harder claim to make than undercutting it on weights, and it is the claim Open​AI chose — naming GLM 5.3 specifically rather than gesturing at "open models" in general.

The enterprise numbers Friar paired with it point the same direction. Enterprise revenue growing 32% month-over-month against 20% overall is what a company selling into procurement cycles looks like, and procurement cycles are where a line item gets questioned. She also confirmed Open​AI is experimenting with outcome-based pricing, tying fees to business results rather than token consumption. That is a meaningful admission for a company whose entire revenue model has been per-token: it suggests that at the cheap end of the market, the per-token number is becoming a weaker sales argument than what the model finishes.

None of which makes the GLM comparison verified. Friar's claim rests on a cloud deployment comparison Open​AI did not publish in detail, and "cheaper than running X yourself on a cloud layer" depends entirely on which cloud, which instance, and which utilisation assumption you pick. Treat it as a positioned claim from an interested party, and check it against your own workload before it changes anything you build.

What it changes if you already route across models

Screenshot of the OrcaRouter model page for GPT-5.6 Luna, showing the model identifier openai/gpt-5.6-luna, a release date of 2026-07-09, a 1M-token context window, 128K maximum output, input pricing of $0.20 per million tokens, output pricing of $1.20, a median time to first token of 1.83 seconds and a 95th percentile of 10.00 seconds.

Three things matter here more than the headline price.

Price cuts land the day they happen. OrcaRouter passes provider list price through at 0% markup, so when a vendor moves a rate, the rate on our side moves with it — no support ticket, no renegotiated contract, no second billing relationship. That is the whole reason the July cut is worth writing about rather than simply noting: for anyone calling Luna through one API for 200+ models, it was live the same day.

The alias is a trap. The unsuffixed gpt-5.6 identifier resolves to the flagship Sol tier, not to Luna, so any integration that wants Luna's economics has to pin gpt-5.6-luna explicitly. This is documented in the family's developer guides and it is still one of the most common ways a "cheap tier" migration quietly bills at flagship rates.

Cheap tiers fail over well. A model at this price point is often doing work where a brief 503 is expensive — thousands of classification calls, not one important completion. Routing the same request across multiple providers of GPT-5.6 Luna with automatic failover is the difference between a retry queue and an outage, and it is the feature that makes a low-cost tier safe to put in a production path.

You can see the current rate and the p50/p95 time-to-first-token figures on the GPT-5.6 Luna on OrcaRouter page — worth a look, because vendor pricing pages lag cuts and third-party trackers lag both.

What to watch next

Three things would turn this from a pricing story into a different one.

A further cut. Chinese-language financial coverage in early September framed Open​AI as preparing to adjust pricing again to meet open-weight competition. That reporting is forward-looking and unconfirmed, but it is consistent with a company that has already accepted an 80% reduction on its volume tier and is now arguing publicly about cost per deployment.

Outcome-based pricing becoming real. If Open​AI starts selling Luna by business result rather than by token, the per-token comparison that this entire market runs on stops being the relevant scoreboard — and every cost-per-task estimate, including the ones above, needs re-deriving.

A capable successor at the same price. GPT-6 Astra already exists above the GPT-5.6 line, and Luna's own position depends on Open​AI leaving it where it is. The moment a newer small tier arrives at $0.20/$1.20, the interesting question stops being whether Luna is cheap and becomes whether it is still the cheapest thing worth routing to.

A summary card for GPT-5.6 Luna listing its July 9, 2026 release date, its current price of $0.20 per million input tokens and $1.20 per million output, its launch price of $1.00 and $6.00, an Artificial Analysis Intelligence Index of 38 ranked 4th of 177, a cost of $0.18 per index task, and Claude Haiku 4.5's comparative figure of $0.21 per index task.

For now the situation is unusually simple. GPT-5.6 Luna is a two-month-old model whose price, not its capability, is the news — and the token price is the most flattering way to look at it, not the most accurate one.