GPT-5.6-prijzen na de verlaging: Luna vs Terra vs Sol, en welk niveau je moet gebruiken
Guides & Insights

GPT-5.6-prijzen na de verlaging: Luna vs Terra vs Sol, en welk niveau je moet gebruiken

Auteur

Jim Song

Publicatiedatum

Nieuwste modellen · 20Bekijk alle modellen
Benchmarks: Artificial Analysis · dagelijks bijgewerkt
Terug naar alle berichten

After O​penAI's July 30, 2026 price cut, the GPT-5.​6 family — GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol — has never been cheaper to run at scale. A week later, on August 6, O​penAI pushed the same tiers into ChatGPT itself: GPT-5.6 Luna is becoming the default model for free and Go accounts with unlimited text chats, while GPT-5.6 Sol was updated for Plus and Pro. But the three tiers are priced very differently, and picking the right one (plus using caching, batch, and routing well) is the single biggest lever on your API bill. This guide lays out the full post-cut pricing, worked cost examples, how each tier compares to rivals, and how to route between them through one endpoint.

Prices are per million tokens (input / output), reflecting the post-cut rate card as reported and OrcaRouter's pass-through model pages; tiered and cache rates come from OrcaRouter's model pages, and competitor figures are attributed. The ChatGPT rollout details in this piece are as O​penAI stated them on August 6, 2026 — vendor-reported, not independently measured. Prices change — verify before building.

TL;DR

GPT-5.6 Luna ($0.20 / $1.20) is the high-volume workhorse; Terra ($2 / $12) is the balanced middle; Sol ($5 / $30) is the flagship for the hardest reasoning and agentic coding. All three share a ~1.05M-token context and 128K max output. Prompt caching and the batch option cut costs further, while very long inputs move you to a higher tier. On the consumer side, O​penAI announced August 6 that GPT-5.6 Luna is becoming the default for Free and Go ChatGPT accounts with unlimited text-only chats, and that an updated GPT-5.6 Sol is rolling out to Plus and Pro. Use Luna for the easy majority, Terra when you need more, and Sol only where it earns its cost — and route by difficulty through one O​penAI-compatible endpoint to minimize spend.

Belangrijkste conclusies

• Luna: $0.20 / $1.20 per 1M tokens — snel, goedkoop, voor grootschalig, latentiegevoelig werk.

• Terra: $2 / $12 — de gebalanceerde middenklasse voor zwaardere taken die het vlaggenschip niet nodig hebben.

• Sol: $5 / $30 — het vlaggenschip voor diep redeneren, grootschalig coderen en agenten met een lange horizon.

• Caching en batch verlagen de kosten verder; lange inputs brengen je naar een duurdere long-context-tier.

• ChatGPT (O​penAI-stated, Aug 6): Luna becomes the free/Go default with unlimited text chats; Sol gets a reliability update for Plus/Pro.

• Grootste besparingshefboom: routeer op basis van moeilijkheid — betaal geen Sol-tarieven voor werk dat Luna kan doen.

The new GPT-5.​6 base pricing

Here's the post-cut base pricing per million input/output tokens: Luna $0.20 / $1.20 (down from $1 / $6), Terra $2 / $12 (down from $2.50 / $15), and Sol $5 / $30 (unchanged). The two cheaper tiers were cut on July 30, 2026; the flagship held. Put simply, the low tiers are now priced for scale, while the flagship stays premium.

Getrapte, gecachte en batchprijzen

The base rate is only the starting point. GPT-5.​6 pricing has three modifiers that materially change your effective cost:

• Tiers voor lange context. Prijzen lopen op voor zeer lange inputs. Op de doorvoerpagina's van OrcaRouter kost Luna $0,20 / $1,20 tot aan een grote context-tier en $0,40 / $1,80 daarboven; Sol kost $5 / $30 op de basistier en $10 / $45 op de grootste. De tier wordt gekozen op basis van het aantal invoertokens van elk verzoek, dus een paar enorme prompts kunnen uw gemiddelde prijs stilletjes verhogen.

• Prompt-caching. Hergebruikte context wordt tegen een scherpe korting gefactureerd — de cache-read van Luna kost ongeveer $0,02 per miljoen tokens (tegenover $0,20 vers), met cache-writes rond de $0,25; de cache-read van Sol kost ongeveer $0,50 (tegenover $5 vers). Voor agents en chat met een stabiele system prompt is caching vaak de grootste besparing.

• Batch. Niet-dringende taken draaien asynchroon tegen een gereduceerd tarief — ideaal voor bulkclassificatie, evaluaties en offline generatie waarbij latentie er niet toe doet.

Uitgewerkte voorbeelden: wat het daadwerkelijk kost

Concrete cijfers maken de niveaus echt. Neem 10 miljoen tokens per maand met een verdeling van 70% input / 30% output (een typische chat/agent-mix):

• Luna: ongeveer $5 per maand (ruwweg $4,37 met prompt caching).

• Terra: ongeveer $50 per maand.

• Sol: ongeveer $125 per maand (ongeveer $109 met caching).

Dat is een spreiding van 25x tussen Luna en Sol voor hetzelfde volume — en dat is precies waarom het matchen van elke aanvraag aan de goedkoopste geschikte laag belangrijker is dan welk enkel tarief dan ook. (Deze komen overeen met de on-page kostencalculator van OrcaRouter, die op basis van de catalogusprijs schat; je werkelijke cijfers hangen af van caching en je invoer/uitvoer-mix.)

Waar elke tier eigenlijk voor is

GPT-5.6 Luna — het volumewerkpaard

Luna is the fast, cost-efficient tier, tuned for high-volume, latency-sensitive workloads: chat, classification, extraction, routing, and lightweight agentic tasks, with a p50 time-to-first-token around 1.65 seconds. After the cut, at $0.20 / $1.20 it's priced to compete directly with cheap open models — the default for the easy majority of calls. It is also the tier O​penAI is steering consumer ChatGPT toward: per its August 6 announcement, GPT-5.6 Luna is becoming the default model for Free and Go accounts, with text-only chats going unlimited and a new Think button for higher reasoning on harder questions arriving the following week (file, image, and voice limits remain). O​penAI says Luna makes 62% fewer factual errors than the GPT-5.5 Instant it replaces — a vendor-reported figure, but a clear signal that the volume tier is where the company is pointing most of its traffic.

GPT-5.6 Terra — het evenwichtige midden

Terra zit tussen volume en vlaggenschip in: capabeler dan Luna voor zwaardere redeneringen en codering, maar veel goedkoper dan Sol. Voor $2 / $12 is het een verstandige standaardkeuze wanneer Luna niet helemaal genoeg is, maar je het vlaggenschip niet nodig hebt — extractie van gemiddelde complexiteit, opstellen en meerstappentaken die nog steeds op schaal draaien.

GPT-5.6 Sol — het vlaggenschip

Sol is built for the hardest work: deep multi-step reasoning, large-scale software engineering, and long-horizon agentic workflows, staying coherent across a ~1.05M-token context and up to 128K output. At $5 / $30 (base) it's a premium choice — reserve it for tasks that genuinely need it, like complex multi-file coding or long agent runs. On August 6, O​penAI also updated GPT-5.6 Sol in ChatGPT for Plus and Pro users: the company says the chat version is more reliable with facts — 68% fewer factual errors, per O​penAI — and gives more focused answers, with a new slider to control how much reasoning effort it applies. The update is chat-only (the Sol behind Work and Codex is unchanged), and the API tier stays at $5 / $30.

De kanttekening over reasoning-tokens

GPT-5.​6 are reasoning models, so effective output cost can exceed a naive estimate: when reasoning is on, internal reasoning tokens are billed as output. On hard prompts with high reasoning effort, that can dominate your bill. Tune reasoning effort to the task (low or off for simple calls), cap output tokens where you can, and measure actual usage. Cheaper per-token rates help; generating fewer tokens helps more.

Hoe elk niveau zich verhoudt tot de concurrentie

The cut repositioned GPT-5.​6 against the field. Per reporting, Luna's $0.20 / $1.20 now undercuts Claude Haiku 4.5 by roughly 5x on input and 4x on output, and Terra's $2 / $12 falls below Claude Sonnet 5's standard pricing (reported at $3 / $15). Against open weights, Luna sits near the floor set by Deep​Seek's V4 line and Zhipu's GLM-5.2 (about $1.20 / $4.10), with Qw​en and Gemini Flash tiers nearby. The upshot: for cost-sensitive work, Luna is now competitive with the cheapest capable models rather than a premium alternative to them — while Sol remains a genuine premium tier for capability you can't get cheaply.

De echte besparingshefboom: routeer op moeilijkheid

The cheapest bill isn't a single tier — it's matching each request to the least expensive model that can do it. In practice: send easy, high-volume calls to Luna (or a cheap open model), step up to Terra for harder tasks, and use Sol only for the genuinely difficult minority. Combined with caching and batch, this routinely cuts costs far more than any single price change. The catch is operational: you don't want to re-integrate three O​penAI tiers plus open-model alternatives separately.

Toegang tot alle drie (en goedkopere concurrenten) via één endpoint

This is where a vendor-neutral router helps. OrcaRouter exposes GPT-5.6 Luna, Terra, and Sol — at the same post-cut prices, 0% markup — through one O​penAI-compatible endpoint, alongside cheaper open models like DeepSeek V4 Pro, GLM-5.2, and Qw​en. Switching tiers (or A/B testing Luna against an open model) is a config change, not a re-integration, and you can route each request to whatever is cheapest.

The post-cut Luna rate card is visible on OrcaRouter's own model page at list price — $0.20 / $1.20 per million tokens — because the router passes provider prices through at 0% markup, with an on-page cost calculator for a typical monthly bill. There's also a free tier and an Offers page for additional savings.

Veelgestelde vragen

What are the GPT-5.​6 prices after the cut?

Luna $0,20 / $1,20, Terra $2 / $12 en Sol $5 / $30 per miljoen invoer-/uitvoertokens. Luna en Terra zijn op 30 juli 2026 verlaagd; Sol is ongewijzigd.

Which GPT-5.​6 tier should I use?

Luna voor latentiegevoelig werk met hoge volumes; Terra voor zwaardere taken die het vlaggenschip niet nodig hebben; Sol voor de moeilijkste redeneer- en agentische codeertaken. Routeer op moeilijkheidsgraad om de kosten te minimaliseren.

Is GPT-5.6 Luna free in ChatGPT?

O​penAI announced on August 6 that GPT-5.6 Luna is becoming the default model for free and Go ChatGPT accounts, with unlimited text-only chats; limits remain on file uploads, images, and voice tools, and a Think button for higher reasoning is rolling out separately. These are O​penAI-stated plans, not independently verified.

How much does GPT-5.​6 cost per month?

Bij 10M tokens/maand (70% input): ongeveer $5 op Luna, $50 op Terra en $125 op Sol — een spreiding van 25x, vóór caching. Je werkelijke kosten hangen af van caching en je input/output-mix.

Hebben alle drie de niveaus hetzelfde contextvenster?

Ja — ongeveer een context van 1.05M tokens en tot 128K uitvoertokens, met ondersteuning voor vision, tools, JSON en redeneren.

Hoe beïnvloedt prompt-caching de kosten?

Caching verlaagt de kosten voor herhaalde context aanzienlijk — Luna's cache-uitlezing kost ongeveer $0,02 per miljoen tokens tegenover $0,20 voor verse tokens — dus hergebruik gecachte prompts waar mogelijk. Lange invoer kan je daarentegen naar een duurder long-context-niveau brengen.

Hoe verhouden de niveaus zich tot Anthropic of open modellen?

Per reporting, Luna undercuts Claude Haiku 4.5 (~5x input / 4x output) and Terra falls below Claude Sonnet 5 ($3 / $15). Luna is also near the cheap open-model floor (Deep​Seek, GLM-5.2 ~$1.20/$4.10).

Waar kan ik alle drie de lagen samen gebruiken?

Through OrcaRouter's single O​penAI-compatible endpoint at 0% markup, alongside cheaper open models for difficulty-based routing.

Kortom

After the cut, GPT-5.6 Luna ($0.20 / $1.20) and Terra ($2 / $12) are compelling for high-volume and mid-tier work, while Sol ($5 / $30) remains the premium flagship — a 25x cost spread that rewards smart routing. The consumer rollout reinforces the split: O​penAI is making Luna the free default while reserving Sol's reliability update for paid plans, which for API builders is one more reason to route easy traffic to Luna. The biggest savings come from routing by difficulty, leaning on caching and batch, and controlling reasoning tokens rather than defaulting to one tier. Doing that is easiest through a single 0%-markup endpoint like OrcaRouter, where all three tiers (at the new prices) sit next to the cheaper open models you'll want to compare them against.

Vergeleken in dit artikel2

Herkend uit dit artikel · Benchmarks: Artificial Analysis · dagelijks bijgewerkt

© 2026 OrcaRouter

Voor aanbieders

Beheer je een inferentieplatform? Zet je modellen op OrcaRouter.

Neem contact op

Word lid van de community

DiscordEmailXGitHubYouTube