A title card comparing GPT-5.6 Luna and Claude Haiku 4.5, showing Luna with an Intelligence Index of 38, a 1M-token context window and pricing of $0.20 input and $1.20 output per million tokens, against Haiku 4.5 with an Intelligence Index of 18, a 200K-token context and pricing of $1.00 input and $5.00 output.
Guides & Insights

GPT-5.6 Luna vs Claude Haiku 4.5: Cheap Tier, Two Very Different Bills

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-5.6 Luna and Claude Haiku 4.5 are the cheap tiers of two different frontier families, and the gap between them depends entirely on which number you look at. On sticker price this is a rout: $0.20 per million input tokens against $1.00. On Artificial Analysis' cost per completed Intelligence Index task, it is $0.18 against $0.21 — a difference of about 14%. Same two models, two honest answers, and the reason they diverge is the single most useful thing to understand before you pick one.

The short version: Luna is the stronger model on paper and the faster one in throughput, but it thinks at length and bills for it. Haiku 4.5 is older, smaller in scope, and considerably more predictable in what a request will cost. Which of those matters more is a property of your workload, not of the models.

At a glance

Vendor — Open​AI vs Anth​hropic

Released — July 9, 2026 vs October 15, 2025

Context window — 1M tokens vs 200K tokens

Max output — 128K tokens vs 64K tokens

Input — text, image and file vs text, image and PDF

Price per million tokens — $0.20 in / $1.20 out vs $1.00 in / $5.00 out

Artificial Analysis Intelligence Index — 38, 4th of 177 vs 18, 138th of 200 (both in reasoning configuration)

Cost per Intelligence Index task — $0.18 vs $0.21

Output speed — 109.4 tokens/sec vs 84.7 tokens/sec

Reasoning control — an effort dial from none to max vs extended thinking enabled manually, with no effort parameter

Reliable knowledge cutoff — February 2026 vs February 2025 (training data through July 2025)

Retirement commitment — none published vs not sooner than October 15, 2026

Where GPT-5.6 Luna actually wins

Screenshot of an Artificial Analysis head-to-head between GPT-5.6 Luna (max) and Claude 4.5 Haiku (Reasoning). The Intelligence Index row reads 38 against 18, AA-Briefcase 1339 against 614, GDPval-AA v2 1489 against 854, AutomationBench-AA 50% against 3%, Terminal-Bench v4.0 12% against 0%, SciCode 54% against 42%, Humanity's Last Exam 39% against 10%, GDPpdf 24% against 4%, CritPt 21% against 0%, AA-Omniscience -10 against -4, and AA-LCR v1.1 84% against 74%. The Cost section shows a blended price per 1M tokens of $0.174 against $0.77.

Capability, by a wide margin. An Intelligence Index of 38 against 18 is not a close call. Luna outranks Haiku on the composite that Artificial Analysis builds from reasoning, knowledge, maths and coding evaluations, and it does so while being the cheaper model per token. Anyone who last compared these two families on the pre-September index — where Luna was showing 51 and 52 — should note that the whole scale was re-scored in September 2026, and Luna's advantage survived the rescoring.

Context, by five times. A 1M-token window against 200K decides a category of work on its own: whole-repository passes, long document sets, extended agent transcripts that would need summarisation on Haiku. If your application is defined by how much context it must hold, the comparison ends here.

Output headroom, by twice. 128K maximum output tokens against 64K matters for long-form generation and for agentic loops that return structured plans rather than one-line answers.

Throughput. 109.4 output tokens per second against 84.7 is a real difference for streaming interfaces, though it is not the same thing as responsiveness — see the latency section below.

Price, on the sticker. Five times cheaper on input and roughly four times cheaper on output is not a rounding error, and for short-output work it is the entire decision.

Where Claude Haiku 4.5 still earns its place

Predictable cost. Haiku 4.5's reasoning is extended thinking that you switch on manually. Left off, it answers directly and bills predictably. GPT-5.6 Luna's effort dial goes up to max, and every step up is more output tokens at the output rate. A team that wants a cost ceiling it can state in advance will find Haiku easier to reason about, even at a higher sticker price.

The non-reasoning configuration is genuinely cheap to run. Artificial Analysis scores Claude 4.5 Haiku's non-reasoning configuration at an Intelligence Index of 15 with no published cost-per-task figure, and it is a reasonable fit for jobs where the bar is simply "parse this correctly" — intent classification, field extraction, routing decisions, moderation triage. On those tasks the index gap between 15 and 38 is mostly irrelevant, because neither model is being asked to think.

Caching and batch discounts are mature. Haiku 4.5 carries a 90% prompt-cache read discount and a 50% discount on the Batch API. Luna has comparable cache economics, so this is closer to a tie than an advantage — but if your pipeline already batches through Anth​hropic's tooling, the migration cost of moving off Haiku is a real cost, and it will not show up on any benchmark.

An operational track record. Haiku 4.5 has been in production since October 2025 and carries a published retirement commitment of not sooner than October 15, 2026. Luna is two months old. For a workload where the model is a detail rather than the product, an eleven-month production history is worth something that a benchmark cannot express.

It is the more truthful model, on the one axis that measures it. Artificial Analysis' head-to-head puts Luna ahead on almost every capability row — AA-Briefcase 1,339 against 614, GDPval-AA v2 1,489 against 854, AutomationBench-AA 50% against 3%, Humanity's Last Exam 39% against 10%, GDPpdf 24% against 4%. The exception is AA-Omniscience, which penalises confident wrong answers on knowledge questions: Luna scores -10 there, against Haiku's -4. A negative score means the hallucination penalty outweighs the accuracy credit, and Luna's penalty is larger. A higher composite index is not the same thing as a more reliable answer, and on the one evaluation built to separate those two things, the older model does better.

The cost math on a real workload

Take a pipeline that consumes 100M input tokens and emits 20M output tokens a month.

GPT-5.6 Luna — 100 × $0.20 = $20, plus 20 × $1.20 = $24. Total $44 per month.

Claude Haiku 4.5 — 100 × $1.00 = $100, plus 20 × $5.00 = $100. Total $200 per month.

At those ratios Luna is about 4.5x cheaper, and that is the calculation most comparisons publish. Now change one assumption. If Luna's reasoning effort is turned up, its output tokens rise, and output is the expensive side — at $1.20 it costs six times what input does. Push Luna to the point where it emits 150M output tokens for the same volume of work, the way it does on Artificial Analysis' evaluation set, and the $44 becomes comfortably over $180. The 4.5x collapses toward parity.

The lesson is not that Luna is expensive. It is that Luna's price advantage is a function of how much it talks, which is a parameter you control. Turn reasoning down for extraction and classification and the discount is real. Leave it on max for everything and you have bought a more capable model at a price that no longer resembles its sticker.

Latency, and a number you should not trust

Published time-to-first-token figures for these two models are close to useless, and it is worth understanding why. On Artificial Analysis, the reasoning configurations of both models show first-token latencies in the tens or even hundreds of seconds, because the harness does not emit a token until the model has finished reasoning. That figure measures the evaluation setup, not the model's responsiveness.

Routing telemetry is a better source, because it measures the model in production. OrcaRouter's own figures put GPT-5.6 Luna at a median 1.83 seconds to first token and a 95th percentile of 10.00 seconds, against Claude Haiku 4.5 at 3.81 seconds median and 9.47 seconds at the 95th percentile. Read properly, that says the two are closer than the leaderboards suggest: Luna wins on the typical request, and the tails are within about half a second of each other. If your concern is the worst case rather than the average, there is very little between them.

Which one to pick

Pick GPT-5.6 Luna if you are holding long context, if the task benefits from real reasoning, if output is long or structured, or if you want the strongest model available at the bottom of the price range. It is also the more natural choice if the rest of your stack already runs on Open​AI-family behaviour.

Pick Claude Haiku 4.5 if your work is short-output and high-volume, if you need a cost ceiling you can state before the request runs, if your pipeline is already built around Anth​hropic's batching and caching, or if you value a model with a production history and a published lifecycle over one that is two months old.

Do not pick either for a task where a small reasoning model or a fine-tuned classifier would do. A large part of the traffic these two tiers absorb is classification and extraction that would run more cheaply on something narrower — and the cheapest model is always the one you did not need.

Running both on one key

Screenshot of the OrcaRouter model page for Claude Haiku 4.5, showing the model identifier anthropic/claude-haiku-4.5, a release date of 2025-10-15, a 200K-token context window, 64K maximum output, input pricing of $1.00 per million tokens, output pricing of $5.00, a median time to first token of 3.81 seconds and a 95th percentile of 9.47 seconds.

The matchup stops being exclusive once both are reachable through a single endpoint. On OrcaRouter, Claude Haiku 4.5 and GPT-5.6 Luna sit behind the same Open​AI-compatible base URL — one API for 200+ models, one key, one billing line, no second contract, and no code change beyond the model identifier. For a tier decision you are not yet sure about, that turns a migration into a config value.

Two features do real work here. Automatic failover across providers means a low-cost tier can carry production traffic without a single-vendor dependency — useful when the model is cheap enough that you are sending it thousands of calls an hour rather than one important one. And the routing DSL lets you compose a call out of several models: route the straightforward majority of requests to Luna, escalate the remainder to GPT-5.6 Terra, and keep Haiku 4.5 in the pool as a fallback when you want a second vendor's answer rather than a retry.

Check the live per-token rate on GPT-5.6 Luna and Claude Haiku 4.5 before committing either to a production path — both rates are passed through at 0% markup, so they move the day the vendor moves them, and third-party price trackers lag by days to weeks.

Questions worth actually answering

Is GPT-5.6 Luna really five times cheaper than Claude Haiku 4.5? On input tokens, yes — $0.20 against $1.00. On the cost of finishing the same evaluated task, no: $0.18 against $0.21. The gap between those two statements is Luna's verbosity. It produces roughly twice the output tokens Haiku does to complete the same work, and output tokens cost six times what input tokens do. For short answers the sticker is broadly honest. For reasoning-heavy work it is not.

Can I tune cost the same way on both? No, and this is the most under-reported difference between them. Luna exposes a reasoning effort dial with settings from none up to max, so cost and quality trade off explicitly per request. Haiku 4.5's extended thinking is either enabled, with a token budget, or it is not — there is no effort parameter. If per-request cost control matters to you, only one of these two gives you that lever.

What about a high-volume classifier doing a few million calls a day? Neither model needs reasoning for this. Run Haiku 4.5 with extended thinking off, or Luna with effort set to none, and compare on latency and output length rather than on the composite index — at that point the relevant questions are how fast the first token arrives and how many tokens the answer takes, and the intelligence benchmarks stop discriminating between them.

A two-column comparison scoreboard for GPT-5.6 Luna and Claude Haiku 4.5. Luna's column reads Intelligence Index 38, context window 1M tokens, price $0.20/$1.20, cost per index task $0.18, output speed 109.4 tokens per second, reasoning control an effort dial from none to max. Haiku 4.5's column reads Intelligence Index 18, context window 200K tokens, price $1.00/$5.00, cost per index task $0.21, output speed 84.7 tokens per second, reasoning control extended thinking on or off.

Which one will still be here in a year? Claude Haiku 4.5 carries a published commitment of no retirement before October 15, 2026. Open​AI publishes no equivalent date for GPT-5.6 Luna, and the GPT-5.6 line already has a newer generation above it in GPT-6 Astra. Neither is a reason to avoid Luna, but if your roadmap assumes an eleven-month planning horizon, it is a reason to keep the model identifier in configuration rather than hard-coded.

Live per-token rate for GPT-5.6 Luna on the same key, passed through at 0% markup with no second billing relationship.

Live per-token rate for Claude Haiku 4.5 on the same endpoint, so the two tiers can be compared without a second contract.