
GPT-5.6 Luna vs Claude Haiku 4.5: Cheap Tier, Two Very Different Bills
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
GPT-5.6 Luna and Claude Haiku 4.5 are the cheap tiers of two different frontier families, and the gap between them depends entirely on which number you look at. On sticker price this is a rout: $0.20 per million input tokens against $1.00. On Artificial Analysis' cost per completed Intelligence Index task, it is $0.18 against $0.21 — a difference of about 14%. Same two models, two honest answers, and the reason they diverge is the single most useful thing to understand before you pick one.
The short version: Luna is the stronger model on paper and the faster one in throughput, but it thinks at length and bills for it. Haiku 4.5 is older, smaller in scope, and considerably more predictable in what a request will cost. Which of those matters more is a property of your workload, not of the models.
At a glance
• Vendor — OpenAI vs Anthhropic
• Released — July 9, 2026 vs October 15, 2025
• Context window — 1M tokens vs 200K tokens
• Max output — 128K tokens vs 64K tokens
• Input — text, image and file vs text, image and PDF
• Price per million tokens — $0.20 in / $1.20 out vs $1.00 in / $5.00 out
• Artificial Analysis Intelligence Index — 38, 4th of 177 vs 18, 138th of 200 (both in reasoning configuration)
• Cost per Intelligence Index task — $0.18 vs $0.21
• Output speed — 109.4 tokens/sec vs 84.7 tokens/sec
• Reasoning control — an effort dial from none to max vs extended thinking enabled manually, with no effort parameter
• Reliable knowledge cutoff — February 2026 vs February 2025 (training data through July 2025)
• Retirement commitment — none published vs not sooner than October 15, 2026
Where GPT-5.6 Luna actually wins

Capability, by a wide margin. An Intelligence Index of 38 against 18 is not a close call. Luna outranks Haiku on the composite that Artificial Analysis builds from reasoning, knowledge, maths and coding evaluations, and it does so while being the cheaper model per token. Anyone who last compared these two families on the pre-September index — where Luna was showing 51 and 52 — should note that the whole scale was re-scored in September 2026, and Luna's advantage survived the rescoring.
Context, by five times. A 1M-token window against 200K decides a category of work on its own: whole-repository passes, long document sets, extended agent transcripts that would need summarisation on Haiku. If your application is defined by how much context it must hold, the comparison ends here.
Output headroom, by twice. 128K maximum output tokens against 64K matters for long-form generation and for agentic loops that return structured plans rather than one-line answers.
Throughput. 109.4 output tokens per second against 84.7 is a real difference for streaming interfaces, though it is not the same thing as responsiveness — see the latency section below.
Price, on the sticker. Five times cheaper on input and roughly four times cheaper on output is not a rounding error, and for short-output work it is the entire decision.
Where Claude Haiku 4.5 still earns its place
Predictable cost. Haiku 4.5's reasoning is extended thinking that you switch on manually. Left off, it answers directly and bills predictably. GPT-5.6 Luna's effort dial goes up to max, and every step up is more output tokens at the output rate. A team that wants a cost ceiling it can state in advance will find Haiku easier to reason about, even at a higher sticker price.
The non-reasoning configuration is genuinely cheap to run. Artificial Analysis scores Claude 4.5 Haiku's non-reasoning configuration at an Intelligence Index of 15 with no published cost-per-task figure, and it is a reasonable fit for jobs where the bar is simply "parse this correctly" — intent classification, field extraction, routing decisions, moderation triage. On those tasks the index gap between 15 and 38 is mostly irrelevant, because neither model is being asked to think.
Caching and batch discounts are mature. Haiku 4.5 carries a 90% prompt-cache read discount and a 50% discount on the Batch API. Luna has comparable cache economics, so this is closer to a tie than an advantage — but if your pipeline already batches through Anthhropic's tooling, the migration cost of moving off Haiku is a real cost, and it will not show up on any benchmark.
An operational track record. Haiku 4.5 has been in production since October 2025 and carries a published retirement commitment of not sooner than October 15, 2026. Luna is two months old. For a workload where the model is a detail rather than the product, an eleven-month production history is worth something that a benchmark cannot express.
It is the more truthful model, on the one axis that measures it. Artificial Analysis' head-to-head puts Luna ahead on almost every capability row — AA-Briefcase 1,339 against 614, GDPval-AA v2 1,489 against 854, AutomationBench-AA 50% against 3%, Humanity's Last Exam 39% against 10%, GDPpdf 24% against 4%. The exception is AA-Omniscience, which penalises confident wrong answers on knowledge questions: Luna scores -10 there, against Haiku's -4. A negative score means the hallucination penalty outweighs the accuracy credit, and Luna's penalty is larger. A higher composite index is not the same thing as a more reliable answer, and on the one evaluation built to separate those two things, the older model does better.
The cost math on a real workload
Take a pipeline that consumes 100M input tokens and emits 20M output tokens a month.
• GPT-5.6 Luna — 100 × $0.20 = $20, plus 20 × $1.20 = $24. Total $44 per month.
• Claude Haiku 4.5 — 100 × $1.00 = $100, plus 20 × $5.00 = $100. Total $200 per month.
At those ratios Luna is about 4.5x cheaper, and that is the calculation most comparisons publish. Now change one assumption. If Luna's reasoning effort is turned up, its output tokens rise, and output is the expensive side — at $1.20 it costs six times what input does. Push Luna to the point where it emits 150M output tokens for the same volume of work, the way it does on Artificial Analysis' evaluation set, and the $44 becomes comfortably over $180. The 4.5x collapses toward parity.
The lesson is not that Luna is expensive. It is that Luna's price advantage is a function of how much it talks, which is a parameter you control. Turn reasoning down for extraction and classification and the discount is real. Leave it on max for everything and you have bought a more capable model at a price that no longer resembles its sticker.
Latency, and a number you should not trust
Published time-to-first-token figures for these two models are close to useless, and it is worth understanding why. On Artificial Analysis, the reasoning configurations of both models show first-token latencies in the tens or even hundreds of seconds, because the harness does not emit a token until the model has finished reasoning. That figure measures the evaluation setup, not the model's responsiveness.
Routing telemetry is a better source, because it measures the model in production. OrcaRouter's own figures put GPT-5.6 Luna at a median 1.83 seconds to first token and a 95th percentile of 10.00 seconds, against Claude Haiku 4.5 at 3.81 seconds median and 9.47 seconds at the 95th percentile. Read properly, that says the two are closer than the leaderboards suggest: Luna wins on the typical request, and the tails are within about half a second of each other. If your concern is the worst case rather than the average, there is very little between them.
Which one to pick
Pick GPT-5.6 Luna if you are holding long context, if the task benefits from real reasoning, if output is long or structured, or if you want the strongest model available at the bottom of the price range. It is also the more natural choice if the rest of your stack already runs on OpenAI-family behaviour.
Pick Claude Haiku 4.5 if your work is short-output and high-volume, if you need a cost ceiling you can state before the request runs, if your pipeline is already built around Anthhropic's batching and caching, or if you value a model with a production history and a published lifecycle over one that is two months old.
Do not pick either for a task where a small reasoning model or a fine-tuned classifier would do. A large part of the traffic these two tiers absorb is classification and extraction that would run more cheaply on something narrower — and the cheapest model is always the one you did not need.
Running both on one key

The matchup stops being exclusive once both are reachable through a single endpoint. On OrcaRouter, Claude Haiku 4.5 and GPT-5.6 Luna sit behind the same OpenAI-compatible base URL — one API for 200+ models, one key, one billing line, no second contract, and no code change beyond the model identifier. For a tier decision you are not yet sure about, that turns a migration into a config value.
Two features do real work here. Automatic failover across providers means a low-cost tier can carry production traffic without a single-vendor dependency — useful when the model is cheap enough that you are sending it thousands of calls an hour rather than one important one. And the routing DSL lets you compose a call out of several models: route the straightforward majority of requests to Luna, escalate the remainder to GPT-5.6 Terra, and keep Haiku 4.5 in the pool as a fallback when you want a second vendor's answer rather than a retry.
Check the live per-token rate on GPT-5.6 Luna and Claude Haiku 4.5 before committing either to a production path — both rates are passed through at 0% markup, so they move the day the vendor moves them, and third-party price trackers lag by days to weeks.
Questions worth actually answering
Is GPT-5.6 Luna really five times cheaper than Claude Haiku 4.5? On input tokens, yes — $0.20 against $1.00. On the cost of finishing the same evaluated task, no: $0.18 against $0.21. The gap between those two statements is Luna's verbosity. It produces roughly twice the output tokens Haiku does to complete the same work, and output tokens cost six times what input tokens do. For short answers the sticker is broadly honest. For reasoning-heavy work it is not.
Can I tune cost the same way on both? No, and this is the most under-reported difference between them. Luna exposes a reasoning effort dial with settings from none up to max, so cost and quality trade off explicitly per request. Haiku 4.5's extended thinking is either enabled, with a token budget, or it is not — there is no effort parameter. If per-request cost control matters to you, only one of these two gives you that lever.
What about a high-volume classifier doing a few million calls a day? Neither model needs reasoning for this. Run Haiku 4.5 with extended thinking off, or Luna with effort set to none, and compare on latency and output length rather than on the composite index — at that point the relevant questions are how fast the first token arrives and how many tokens the answer takes, and the intelligence benchmarks stop discriminating between them.

Which one will still be here in a year? Claude Haiku 4.5 carries a published commitment of no retirement before October 15, 2026. OpenAI publishes no equivalent date for GPT-5.6 Luna, and the GPT-5.6 line already has a newer generation above it in GPT-6 Astra. Neither is a reason to avoid Luna, but if your roadmap assumes an eleven-month planning horizon, it is a reason to keep the model identifier in configuration rather than hard-coded.
Live per-token rate for GPT-5.6 Luna on the same key, passed through at 0% markup with no second billing relationship.
Live per-token rate for Claude Haiku 4.5 on the same endpoint, so the two tiers can be compared without a second contract.
