
OpenAI Cut GPT-5.6 API Prices: What Changed, Why, and How to Get It
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 985 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
On July 30, 2026, OpenAI cut the API prices of its two lower-cost GPT-5.6 models — sharply reducing the cheapest tier and trimming the mid-tier. Three weeks later it turned to the flagship: on August 21, 2026, OpenAI announced that GPT-5.6 Sol API prices fall by more than 20% for the next three months as the company works to run the model more efficiently. This article explains exactly what changed across both rounds, the engineering behind them, how the new prices stack up against rivals, the hidden costs to watch, how to cut your bill further — and how the lower prices are already live on OrcaRouter at 0% markup.
Every figure below is labeled by source. Pricing is per million tokens (input / output). The Luna and Terra rates reflect OpenAI's July 30, 2026 rate card as reported by outlets including Unite.AI; the Sol reduction reflects OpenAI's August 21, 2026 announcement, with per-token prices as reported by outlets including Runtimewire and Reuters. OrcaRouter's pass-through model pages show the same numbers; competitor figures are attributed. Prices change — verify before you build.
TL;DR. OpenAI dropped GPT-5.6 Luna to $0.20 / $1.20 (from $1 / $6, ~80% off) and GPT-5.6 Terra to $2 / $12 (from $2.50 / $15, ~20% off) on July 30. On August 21 it extended the discount to the flagship: GPT-5.6 Sol API prices fall over 20% for the next three months — to $4 / $20 per million tokens from $5 / $30, per OpenAI — while token-based Codex and ChatGPT Work credits buy more and subscription usage stays unchanged. Two efficiency wins paid for the July cut: rewritten GPU kernels (about 20% lower serving cost) and a redesigned speculative-decoding draft model (15%+ faster token generation). If you route through a vendor-neutral, 0%-markup endpoint like OrcaRouter, the new rates are already live — same price, one OpenAI-compatible API, with every rival model beside it to compare and route between.
Key takeaways
• GPT-5.6 Luna: now $0.20 / $1.20 per 1M tokens, down from $1 / $6 — roughly an 80% cut.
• GPT-5.6 Terra: now $2 / $12, down from $2.50 / $15 — about a 20% cut.
• GPT-5.6 Sol (flagship): now $4 / $20 for the next three months, down from $5 / $30 — a >20% cut announced August 21, 2026.
• Codex & ChatGPT Work token-based credits buy more: Sol input drops to 100 credits and output to 500 credits per 1M tokens (from 125 / 750).
• Subscriptions unchanged: Plus, Pro, and Business included usage, five-hour/weekly limits, and legacy credit rates are unaffected.
• Why: OpenAI framed both rounds as efficiency dividends — rewritten production GPU kernels (~20% lower serving cost) plus a speculative-decoding draft-model redesign (15%+ token-generation efficiency), per OpenAI's July 29 engineering post.
• How to get it: the cuts are already reflected on OrcaRouter (0% markup) — Luna at $0.20 / $1.20, Sol at $4 / $20 — through one OpenAI-compatible endpoint.
What exactly changed
OpenAI's GPT-5.6 family runs as a tier ladder: Luna (fast, cheapest), Terra (mid), and Sol (flagship). The July 30 cut hit the two cheaper tiers. Per the reported rate card, GPT-5.6 Luna dropped from $1 / $6 to $0.20 / $1.20 per million input/output tokens — about an 80% reduction on both input and output — and GPT-5.6 Terra dropped from $2.50 / $15 to $2 / $12, roughly 20%. GPT-5.6 Sol, the flagship built for the hardest reasoning and agentic coding, was untouched that day — but on August 21, 2026, OpenAI announced a temporary cut for it too: API prices fall more than 20% for the next three months, to $4 per million input tokens and $20 per million output tokens, with cached input down from $0.50 to $0.40. The same announcement said the credits you buy go further in Codex on token-based plans, and the usage included in your subscription stays the same.
The base per-token rate isn't the whole story. GPT-5.6 pricing also has input-length tiers (very long contexts move to a higher rate), a prompt-caching discount (cached input is far cheaper than fresh input), and a batch option (asynchronous jobs at a discount). All three carry through the cut, so the effective price for a cache- or batch-heavy workload can be lower still. Watch the long-context tier in the other direction: feeding huge prompts can bump you up a tier.

The two optimizations behind the cut
This wasn't a loss-leader gesture — OpenAI framed it as an efficiency dividend. In a July 29, 2026 engineering post, members of OpenAI's technical staff described optimizations across inference and the agent harness behind Codex and ChatGPT Work. Two stand out. First, GPU kernel rewrites across production infrastructure — with Sol, running inside Codex, used to rewrite the company's own kernels — which OpenAI says cut end-to-end serving costs by about 20%. Second, a redesigned speculative-decoding draft model that reportedly improved token-generation efficiency by more than 15%. Lower serving cost, passed to customers as lower prices on the high-volume tiers, is the story — and it's a glimpse of a feedback loop the whole industry is chasing: using frontier models to optimize the very infrastructure that serves them.
The Sol cut announced on August 21 carries the same framing. Per OpenAI, the reduction — more than 20% for the next three months — comes as the company makes the model more efficient to run, and it is explicitly temporary rather than a permanent price reset.
How the new prices compare
The cut is aimed squarely at the competition. According to Unite.AI's reporting, Luna's new $0.20 / $1.20 undercuts Anthropic's Haiku 4.5 by roughly 5x on input and about 4x on output, and Terra's new $2 / $12 falls below Claude Sonnet 5's standard pricing (reported at $3 / $15 taking effect after August 31, 2026). Against the open-weight field, Luna now sits close to the cheap-inference floor established by DeepSeek's V4 line and Zhipu's GLM-5.2 (roughly $1.20 / $4.10), with Alibaba's Qwen and Google's Gemini Flash tiers in the same neighborhood. Sol's temporary cut to $4 / $20 keeps the flagship well above its own volume tiers but trims the premium — and the one-third reduction on output tokens matters most for output-heavy agent workloads. Coverage of the move notes it comes as OpenAI faces growing competition from Anthropic and Chinese AI models.
Inference keeps getting cheaper
Zoom out and this cut fits a multi-year trend: the cost of a given level of capability has fallen sharply, generation after generation, as labs wring more throughput from the same hardware. OpenAI's own older tiers illustrate the drift downward over time, and each efficiency breakthrough — better kernels, speculative decoding, cheaper attention — tends to show up as a price cut within months. For buyers, the practical implication is to architect for a world where inference gets cheaper on a schedule: don't over-commit to one model's pricing, and keep switching cheap.
The hidden cost to watch: reasoning tokens
GPT-5.6 models are reasoning models, and that shapes real-world spend. When reasoning is enabled, the model emits internal reasoning tokens that are billed as output — so your effective output cost can be higher than a naive "words in the answer" estimate suggests, especially on hard prompts with high reasoning effort. The fix is to tune reasoning effort to the task (low or off for simple calls), cap output where you can, and measure actual token usage rather than assuming. A cheaper per-token rate helps, but controlling how many tokens you generate helps more.
What it means for developers
If you use GPT-5.6 Luna or Terra for high-volume work — classification, extraction, routing, lightweight agents, chat — your bill just dropped, in some cases by most of its value, with no code change required. That makes the volume tiers even more attractive for cost-sensitive workloads and narrows the price gap versus cheap open models. The flagship math has changed too, at least for now: Sol at $4 / $20 for the next three months lowers the cost of reasoning-heavy and agentic work on the top tier, and token-based Codex and ChatGPT Work credits stretch further. It's still a premium choice for the hardest tasks — just a cheaper one than it was. The winning pattern is the same: route by difficulty — cheap tiers (or cheap open models) for the easy 80%, the flagship only where it earns its cost.
How to cut your bill further
The price cut is a floor to build on, not the finish line. The biggest additional levers: prompt caching (reuse a stable system prompt or long context and pay the much lower cached-input rate); the batch option (run non-urgent jobs asynchronously at a discount); tier and model routing (send each request to the cheapest model that can handle it, including open models); and reasoning control (don't pay for deep reasoning on trivial calls). Stacked together, these routinely cut real bills far more than any single rate change. As a concrete anchor: at 10 million tokens a month with a 70% input share, Luna's new rates work out to roughly $5 a month (about $4.37 with caching); the same volume on Sol is about $125 a month — the clearest argument for routing by difficulty.
How to capture the lower prices immediately
Here's the part that ties it together. OrcaRouter charges 0% markup and passes provider pricing straight through, so OpenAI's cuts are already live on OrcaRouter: GPT-5.6 Luna shows $0.20 / $1.20, Terra $2 / $12, and GPT-5.6 Sol $4 / $20 on the model pages right now — the same rates as OpenAI's own card, reachable through a single OpenAI-compatible endpoint alongside 185+ other models. You get the cuts with no migration, and you can price-compare Sol, Luna, and Terra against DeepSeek, GLM, Qwen, Gemini, and Claude in one place, then route each request to whatever is cheapest for the job. There's also a free tier and an Offers page for additional savings.

The cuts are already live on OrcaRouter: GPT-5.6 Luna at $0.20 / $1.20 and GPT-5.6 Sol at $4 / $20 per 1M tokens, passed through at 0% markup.
Route by difficulty to save even more
The cheapest bill isn't one model — it's the right model per request. Because OrcaRouter exposes the whole field through one endpoint, you can send easy, high-volume calls to Luna or a cheap open model and reserve Sol or another flagship for the hard ones, switching with a config change rather than a re-integration. A price cut on Luna makes that route-by-difficulty strategy cheaper still, and the temporary Sol discount lowers the cost of the hard calls too — without you re-plumbing anything.

FAQ
How much did OpenAI cut GPT-5.6 prices?
GPT-5.6 Luna fell from $1 / $6 to $0.20 / $1.20 per million tokens (about 80%), and GPT-5.6 Terra from $2.50 / $15 to $2 / $12 (about 20%). On August 21, 2026, OpenAI also cut GPT-5.6 Sol from $5 / $30 to $4 / $20 for the next three months.
When did the price cut take effect?
The Luna and Terra cuts took effect July 30, 2026, live on the published rate card. OpenAI announced the temporary Sol reduction on August 21, 2026, for the next three months.
Why did OpenAI lower prices?
OpenAI attributes the July cut to inference optimizations — rewritten production GPU kernels (about 20% lower serving cost) and a redesigned speculative-decoding draft model (15%+ token-generation efficiency) — per a July 29, 2026 engineering post. It framed the Sol cut the same way, as a step taken while the model gets cheaper to run, in a market with rising competition.
Did the flagship (Sol) get cheaper?
Yes — temporarily. The July 30 cut left GPT-5.6 Sol at $5 / $30, but on August 21, 2026, OpenAI announced a price reduction of more than 20% for the next three months, to $4 / $20 per million tokens. Codex and ChatGPT Work token-based credits also buy more, while subscription usage stays unchanged.
How do the new prices compare to Anthropic and open models?
Per reporting, Luna now undercuts Anthropic's Haiku 4.5 by roughly 5x input / 4x output, and Terra falls below Claude Sonnet 5's standard pricing ($3 / $15). Luna also sits near the cheap open-model floor set by DeepSeek and GLM-5.2.
What's the catch with reasoning tokens?
GPT-5.6 are reasoning models; internal reasoning tokens are billed as output, so effective output cost can exceed a naive estimate. Tune reasoning effort and cap output to control it.
How do I get the new lower prices?
They're already live on OpenAI's API and, at 0% markup, on OrcaRouter — where GPT-5.6 Luna shows $0.20 / $1.20 and GPT-5.6 Sol $4 / $20 — through one OpenAI-compatible endpoint, no code change needed.
Bottom line
OpenAI's July 30 cut made GPT-5.6 Luna and Terra dramatically cheaper for high-volume work, and its August 21 announcement extends the story to the flagship: GPT-5.6 Sol drops to $4 / $20 for the next three months while Codex token credits buy more and subscription usage stays unchanged. Both rounds are efficiency dividends — kernel rewrites and speculative-decoding gains on the one hand, cheaper-to-run serving on the other — and a clear response to a real price war. Route by difficulty, lean on caching and batching, and control reasoning tokens to save the most. And because OrcaRouter passes provider pricing through at 0% markup, the lower rates are already live there — Luna at $0.20 / $1.20, Sol at $4 / $20 — same price, one OpenAI-compatible endpoint, the whole model field in one place. Grab the cuts without changing a line of code, and let each request pick the cheapest model that can do the job.
