
Claude Haiku 5.5 vs GPT-6 Luna: Same Sticker Price, Three Times the Bill
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Claude Haiku 5.5 and GPT-6 Luna are priced identically on paper: $0.10 per million input tokens and $0.50 per million output tokens, the same two numbers down to the cent, from two vendors who rarely agree on anything. Anthropic's launch post even prints Luna's scores in the column next to its own, which is not a thing vendors do about models they consider unthreatening.
The two cards diverge on everything the sticker hides. Luna holds that rate out to 272,000 tokens before it steps up; Claude Haiku 5.5 steps at 100,000. Claude Haiku 5.5 scores 43.40 on the Artificial Analysis Intelligence Index v4.3.2 against Luna's 38.12 — and costs $0.21 per finished Index task against Luna's $0.07. Same price, three times the bill, five points more intelligence. That is the whole trade, and it is a cleaner one than most head-to-heads in this band.
Two cards, side by side
Per million tokens in US dollars, from Anthropic's and OpenAI's published rates as reflected on our own catalogue, with independent measurements attributed to Artificial Analysis. Cached 2026-10-08.

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens, $0.50 above it. GPT-6 Luna $0.10 up to 272,000 tokens, $0.20 above it.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it. GPT-6 Luna $0.50 up to 272,000 tokens, $0.75 above it.
• Cached input and writes — both at $0.01 and $0.125 under the first threshold, with Claude Haiku 5.5 moving to $0.05 and $0.625 past 100,000 tokens and GPT-6 Luna to $0.02 and $0.25 past 272,000.
• Context and output — 1M tokens for both; 128,000 maximum output for both. These two genuinely are the same shape.
• Thinking — adaptive on both, both with an effort parameter. Claude Haiku 5.5's default effort is medium; not all of the published scores are at that setting.
• Independent score — Claude Haiku 5.5 (Max) 43.40, second of 182 models. GPT-6 Luna (Max) 38.12, eighth of 182.
• Independent cost per task — $0.21 against $0.07.
• Speed — 241.9 output tokens per second against 129.4, measured by the same evaluator.
The threshold is where the difference is
Both vendors built a two-tier rate card. They did it at very different places, and the placement matters far more than the number at the top.
Anthropic's step sits at 100,000 tokens. Under it you pay $0.10 and $0.50; over it, five times and five times — $0.50 and $2.50. Anthropic says prompts under the threshold made up around 90% of requests to its previous Haiku model, which is the vendor's own traffic mix and a reasonable description of short-call workloads generally.
OpenAI's step sits at 272,000 tokens. Under it, $0.10 and $0.50 — identical to Anthropic. Over it, $0.20 and $0.75, a doubling and a 50% increase. That is a much gentler step and it starts almost three times further out.
Take a 200,000-token prompt, which is a plausible size for a long-document pipeline and entirely reasonable for a 1M-token window. On Claude Haiku 5.5 that request is billed at the second tier: $0.50 in, $2.50 out. On GPT-6 Luna it is still first tier: $0.10 in, $0.50 out. Five times the input rate and five times the output rate, on the same request, between two models with the same headline price. That is not a nuance — for a long-prompt workload it is the entire comparison.
Why the same rate produces three times the bill
Cost per task is the number that should decide a high-volume pipeline, and here the two models part company in a way the rate card cannot explain.
Artificial Analysis puts GPT-6 Luna (Max) at $0.07 per Intelligence Index task and Claude Haiku 5.5 (Max) at $0.21 — a factor of three, on identical sticker rates, for a 5.28-point difference in score. The cause is on the same page and it is not subtle. Claude Haiku 5.5 generated 440 million output tokens while running the Index against a 100 million median, and the evaluator calls it very verbose. GPT-6 Luna generated 140 million against the same 100 million median, described as somewhat verbose. Three times the output tokens is three times the bill when output is the meter that dominates.
Anthropic's own numbers point the same direction. The vendor's cost-versus-accuracy charts plot Haiku 5.5 across the effort range, and the pull-quote from the launch is that Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding work, with Haiku 5.5 positioned for compaction, summarisation and subagents instead. The smaller the job, the less of that verbosity problem you buy — which is precisely the workload the model was built for, and precisely the workload where a threefold cost gap is worth checking against your own logs rather than a board's.
One more entry on the ledger, from Anthropic's documentation rather than the evaluator's: Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on the previous Haiku. The rate is the same as Luna's; the tokeniser is not.

What Anthropic's own table says about Luna
Anthropic's launch page runs GPT-6 Luna as a column in its benchmark table, and the honest way to report it is with the attribution attached: these are Anthropic's evaluations, run on Anthropic's harness, against a competitor's model. They are vendor-reported, not independently reproduced, and a vendor running its own suite against a rival's scores is a comparison to weigh rather than accept.
• Knowledge work — GDPval-AA v2.1 at 1620 for Claude Haiku 5.5 against 1437 for GPT-6 Luna. AA-Briefcase v1.1 at 1578 against 1336.
• Computer use — OSWorld 2.1 offline subset at 72.4% against 48.9%.
• Agentic coding — Terminal-Bench 4.0 at 39.2% against 16.4%; FrontierCode 1.1 (Main) at 46.4% against 42.4%.
• Visual reasoning — Chartography at 46.4% without tools against 29.1%.
On the vendor's own harness Claude Haiku 5.5 leads every row. On the independent board it leads by a smaller margin — 43.40 against 38.12 — which is the usual relationship between a vendor's table and a third party's, and the reason both are printed here. The two agree on direction and disagree on size.
The bill is the deciding vote
Line the four facts that actually matter up against each other and the choice stops being close for most workloads.
GPT-6 Luna is cheaper per completed task by a factor of three on the independent measurement, its first price tier reaches eighteen times further up the context window, its over-threshold rates are lower on both meters, and it is the only one of the two whose long-context penalty is a doubling rather than a fivefold step. Claude Haiku 5.5 is ahead on measured intelligence by a margin that is real and modest, faster by a factor of nearly two on output, and — on the vendor's own harness, where the gap is widest — ahead on every published benchmark row.
If your calls are short, if throughput is the constraint, or if the quality difference shows up in your own eval, pay the three times. If your calls are long, if they are machine-generated and never read by a human, or if you are running a classification tier that fires a million times a day, the same $0.10 and $0.50 will buy you a bill a third the size on the other side, and the capability you give up is five points on a board both models sit near the top of.
Calling either of them from one place
The useful thing about a comparison this close is that you should not have to take anyone's word for it, including ours. GPT-6 Luna is on the OrcaRouter catalogue today at the provider's rate with 0% markup — the same rate structure printed above, passed through rather than blended — and so are Claude Sonnet 5.5 and Claude Opus 5.5, which makes an escalation test from a cheap leg up to a mid tier a configuration rather than a project. One key, one SDK, automatic failover underneath, and the routing DSL if the escalation belongs in the route rather than in your application.
Claude Haiku 5.5 is not on the catalogue yet. Stating that plainly is more useful than implying otherwise: the Luna side of this comparison is callable from us this morning, and the Anthropic side has to be reached at Anthropic until it is not.

What would change the answer
Two things, and both are worth watching rather than estimating.
The first is whether the verbosity moves. Claude Haiku 5.5's cost per task is a function of how many tokens it spends per answer at max effort, and the effort parameter is the lever on that. A team that pins effort lower — the model's own default is medium, not max — will see a cost-per-task figure closer to Luna's and a score closer to Luna's too. The trade is available; what it is worth is a question about your eval, not about the model.
The second is whether the independent cost-per-task numbers hold as both models accumulate more measured runs. A single board placement is a snapshot, and a threefold gap that appears in one snapshot is worth confirming against your own invoice before a pipeline is rebuilt around it. That is what the shared endpoint is for.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
