
Claude Haiku 5.5 vs GPT-6 Luna Pro: A Model Against a Mode
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
One side of this comparison is a model, and the other is a setting. Claude Haiku 5.5 went into general availability on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output, on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. GPT-6 Luna Pro is not a model at all: the way to reach it is to call GPT-6 Luna — live since September 22, 2026 at the same $0.10 and $0.50 — with reasoning.mode set to pro instead of standard. There is no gpt-6-luna-pro identifier on OpenAI's model page. So the honest framing of this matchup is not "which model wins" but "which is the better way to spend ten cents per million input tokens": a small model with an effort dial, or a bigger model run in a mode that does more work and bills you for the tokens it does not show you. For short, high-volume work the answer is usually the first. For hard, occasional work it can be the second — and the number that decides it, OpenAI has not published.
What "pro" is, mechanically
OpenAI's reasoning guide states that GPT-5.6 and GPT-6 models support two modes in the Responses API, standard (the default) and pro, selected through reasoning.mode. Mode is a different control from reasoning.effort: mode decides how much work the model does before committing to an answer, effort decides how much reasoning happens inside whichever mode you chose. GPT-6 Luna's documented effort values are none, low, medium (the default), high, xhigh and max, and any of them can be combined with pro mode.
On billing the guide is explicit and slightly counter-intuitive. Pro mode aggregates all the model work behind the final answer and charges those tokens "at the selected model's standard token rates". There is no pro premium on the rate card. The cost appears as volume, because pro mode "performs more model work than standard mode, increasing token usage and cost" — and reasoning tokens are billed as output tokens while surfacing only in the usage object under output_tokens_details.reasoning_tokens. If you do not read that field, the increase is invisible until the invoice arrives.
Two comparisons follow from that and they are worth keeping apart. Against GPT-6 Luna in standard mode, pro mode is the same model with an unknown multiplier on token consumption. Against Claude Haiku 5.5, it is also a different model at a different scale — 1,050,000 tokens of context against 1M, a 922,000-token input ceiling and a May 18, 2026 knowledge cutoff, on a tier Anthropic's own launch documentation positions slightly above the Haiku on hard agentic work. The two differences pull the same way on cost and opposite ways on capability, which is why neither can be read off a price list.
The rate cards, side by side
Both models are billed in two tiers, and the thresholds are in different places — which matters more than the headline rate, because both headlines are identical.
• Input — Claude Haiku 5.5 $0.10 per million up to 100,000 tokens, $0.50 above it. GPT-6 Luna $0.10 up to 272,000 tokens, $0.20 above it.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it. GPT-6 Luna $0.50 up to 272,000 tokens, $0.75 above it.
• Caching — Claude Haiku 5.5 cache reads $0.01 and writes $0.125 under the line, $0.05 and $0.625 over it. GPT-6 Luna cache reads $0.01 and writes $0.125 under 272,000 tokens, $0.02 and $0.25 over it.
• Context and output — 1M tokens and 128,000 maximum output for Claude Haiku 5.5; 1,050,000 tokens with a 922,000-token input cap and 128,000 maximum output for GPT-6 Luna.
• Thinking — adaptive on both. Claude Haiku 5.5 defaults to medium effort; GPT-6 Luna defaults to medium effort and standard mode, with pro mode as an opt-in on top.
• Batch and caching discounts — Claude Haiku 5.5 takes 50% off on the Message Batches API; GPT-6 Luna takes 50% off on Batch and Flex, doubles on Fast mode, and adds a 10% premium for regional processing.
The one line that is not the same shape is the tokenizer. Anthropic's documentation notes that Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on Claude Haiku 4.5. That does not change the rate; it changes how quickly a prompt reaches the 100,000-token step, and it is a real cost difference that no rate card shows. A 90,000-token prompt measures roughly 117,000 and pays $0.50 per million input rather than $0.10 — a fivefold rate on a prompt that looks like it is comfortably under the line.
What pro mode costs, in the only units anyone can use
Because the multiplier is unpublished, the useful thing to state is what each additional token costs, and then what the plausible ranges do at volume.
• At GPT-6 Luna's $0.50 per million output tokens, every extra 10,000 output tokens per task costs $0.005. Run 100,000 tasks a day and that increment is a billion extra tokens — about $500 a day, or $15,000 a month, on a model whose headline price is a dime per million input.
• For scale, GPT-6 Luna at maximum effort already consumes 140 million output tokens running the Intelligence Index against a 100 million median, which the evaluator describes as somewhat verbose. Anthropic's launch table, run on Anthropic's own harness, puts GPT-6 Luna at 16.4% on Terminal-Bench 4.0 and 42.4% on FrontierCode 1.1, against 39.2% and 46.4% for Claude Haiku 5.5 — vendor-reported, one vendor's suite, a competitor's model.
• On the independent board the ordering is narrower. Artificial Analysis scores Claude Haiku 5.5 at Max effort 43 on Intelligence Index v4.3.2, second of 182, and GPT-6 Luna (Max) at 38, eighth of 182. Both numbers are standard-mode, max-effort. There is no independent evaluation of GPT-6 Luna in pro mode: its page carries no pro entry, and Artificial Analysis does not list one.
The structural problem for pro mode is the interaction between those last two points. GPT-6 Luna's entire cost advantage over Claude Haiku 5.5 comes from token efficiency — 140 million output tokens against 440 million for the same suite, at identical rates, which is why the same eval costs $0.07 per task on one side and $0.21 on the other. Pro mode is, by definition, a mode that emits more tokens. Turning it on is turning off the thing that made the model cheap in this comparison, and the multiplier that would tell you by how much is not published. That is not an argument that pro mode is a bad deal — on difficult tasks more work is usually better work — it is an argument that pro mode competes on capability, not on price, and should be evaluated against the mid-tier models it is priced near rather than against the cheap tier it sits inside.

Where each one actually wins
Strip it to the decision and the split is clean, though neither side is free of caveats.
Claude Haiku 5.5 owns the cheap end. It is faster in both senses that matter: 241.9 output tokens per second measured by Artificial Analysis against 123.0 for GPT-6 Luna, and a 295-second time to first token on the Index run, which is a product of how long it thinks before answering at Max effort rather than a network figure. Its rate card reaches the second tier at 100,000 tokens, and its tokenizer inflates prompts by roughly 30%, so long-prompt workloads pay more than the sticker suggests on both counts. What it gives back is a five-point intelligence advantage on the independent index and an effort dial — Low through Max — that is the single most useful cost control either vendor shipped with these launches. Pinning effort down is how you trade part of that five-point lead for a cost-per-task figure closer to GPT-6 Luna's.
GPT-6 Luna Pro owns the hard end, with an unquantified bill. The model underneath is cheaper per token above 272,000 input tokens than Claude Haiku 5.5 is above 100,000, by a factor of two and a half on input and three and a third on output, and its first tier reaches eighteen times further up the context window. Standard mode is measurably cheaper per completed task. Pro mode trades that for more work per answer at the same rates, and whether it beats Claude Haiku 5.5 at Max effort or a mid-tier model on a given task is a question no published number answers. The way to answer it is to run the task both ways and read the reasoning-token field.

Trying either one without betting on it
The awkward part of this comparison is that its most important number — pro mode's real cost — can only be produced by running your own traffic through it, and the smaller model's effective cost depends on where your prompts land relative to a 100,000-token line after re-tokenisation. Both are measurement problems, and both are cheaper on a shared endpoint than across two vendor clients.
GPT-6 Luna is on OrcaRouter's catalogue today at the provider's list price with 0% markup passed through, so the standard-mode baseline for this comparison is callable from one key, and a vendor price change lands on our side the same day rather than at the next billing cycle. Claude Haiku 5.5 is not on the catalogue yet; the Anthropic tier we do route today is Claude Haiku 4.5, Claude Sonnet 5.5 and Claude Opus 5.5, which makes an escalation test from a cheap leg up through a mid tier a routing rule rather than a code change. Automatic failover underneath is what lets a two-day-old model carry production traffic while you are still measuring it: a provider that starts failing gets routed around instead of taking the request down, which is the difference between a pilot you can run this week and a cutover you have to schedule.
What settles it
Three things, in order of how much they would change the answer. OpenAI publishing a pro-mode token multiplier, or a pro slug with its own price line, would convert the central unknown into arithmetic. Artificial Analysis re-running GPT-6 Luna in pro mode at matched effort would do it from the independent side. And a second independent run of Claude Haiku 5.5 at medium rather than Max would tell you what the model actually costs at its default setting, which is the configuration most people will ship — the $0.21 per task on the board is a Max-effort figure, and Max is not the default.
Until then the accurate summary is narrow. Claude Haiku 5.5 is a model you can call, price and benchmark today, and at the volumes its tier is built for it is the cheaper answer with a documented way to get cheaper still. GPT-6 Luna Pro is a configuration of a model you can call, its rate card is not the problem, its token consumption is unaudited, and the whole case for it rests on tasks hard enough that spending more tokens is the point. Buy it for those tasks, measure it before you do, and do not put it on a high-volume path until someone publishes the multiplier.

