A generated hero title card reading 'GPT-6 Luna vs DeepSeek V4 Pro', subtitled 'One index point apart. Five times the price. One of them is a file you can download.', with chips showing Luna at $0.10/$0.50 per 1M with closed weights, V4 Pro at $0.66/$1.98 off-peak under an MIT licence, and index scores of 37 against 36 at max effort, with the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6 Luna vs DeepSeek V4 Pro: One Index Point, Five Times the Price, and the Case for Owning Your Weights

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Here is the number that should decide this matchup and does not: GPT-6 Luna scores 37 on the Artificial Analysis Intelligence Index at maximum effort. DeepSeek V4 Pro scores 36 on the same index revision, also at maximum effort. One point apart, on a composite where a single point is inside the noise of a re-run — and the two models are separated by roughly five times on price. GPT-6 Luna lists at $0.10 per million input tokens and $0.50 per million output. DeepSeek V4 Pro 0813 lists at $0.66 and $1.98 off-peak, doubling to $1.32 and $3.96 during its peak window.

If that were the whole story, this article would be one paragraph long and it would end with "use GPT-6 Luna." It is not the whole story, and the reason is the one dimension where the comparison inverts completely: DeepSeek V4 Pro's weights are open and MIT-licensed. The model that costs five times more per token is the model you are allowed to download, modify, fine-tune, and run on your own hardware forever. That is not a footnote to the price argument — for a specific class of buyer it is the entire argument, and the September 22 launch of GPT-6 Luna sharpened it rather than settling it.

What shipped, and when

GPT-6 Luna is OpenAI's small tier, released September 22, 2026, alongside GPT-6 Sol at $2.00/$10.00 and beneath the flagship GPT-6 Astra at $10.00/$50.00. OpenAI describes it as the company's most efficient model for focused, high-volume tasks. It carries a 1,050,000-token context window with 922,000 maximum input and a 128,000-token output cap, accepts text and image, and exposes reasoning effort from none through low, medium, high, xhigh and max — with medium as the default.

DeepSeek V4 Pro 0813 is the GA build of DeepSeek's flagship fourth-generation model, published on August 13, 2026, replacing the April preview under the same deepseek-v4-pro identifier. It is a text-only mixture-of-experts model with 1.6 trillion total parameters and 49 billion active, a 1M-token context window, and an unusually generous 384,000-token output ceiling — three times GPT-6 Luna's. The weights went up on Hugging Face as deepseek-ai/DeepSeek-V4-Pro-0813 roughly half a day after the API went live: 66 fp8 shards, ungated, MIT. Our own catalogue page for the model still carries the April preview date rather than the August GA date, which is worth knowing if you are reading dates off product pages rather than release notes.

Both are reasoning-first and long-context. Only one of them is a file you can hold.

The rate cards, honestly

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs DeepSeek V4 Pro $0.66 off-peak / $1.32 peak

• Output per 1M — GPT-6 Luna $0.50 vs DeepSeek V4 Pro $1.98 off-peak / $3.96 peak

• Cached input — GPT-6 Luna $0.01 per 1M, a 90% discount off its own rate vs DeepSeek V4 Pro a 97% cache discount, the deeper percentage off a higher base

• Peak window — GPT-6 Luna none, flat all day vs DeepSeek V4 Pro doubles between 01:00-04:00 and 06:00-10:00 UTC

• Long-context clause — GPT-6 Luna doubles input and cache and lifts output 1.5x above 272K input, applied to the whole request vs DeepSeek V4 Pro no long-context surcharge published on its card

• Context and output — GPT-6 Luna 1,050,000 in / 128,000 out vs DeepSeek V4 Pro 1,048,576 in / 384,000 out

• AA Intelligence Index — GPT-6 Luna 37 at max effort vs DeepSeek V4 Pro 36 at max effort, same index revision

• Default-effort score — GPT-6 Luna 29 at medium vs DeepSeek V4 Pro 36 at its max setting, the effort Artificial Analysis measured

• Reasoning off switch — GPT-6 Luna none through max vs DeepSeek V4 Pro low, high and max only

• Cost to run the index — GPT-6 Luna $122.39 vs DeepSeek V4 Pro $1,122.27

• Weights — GPT-6 Luna closed, API only vs DeepSeek V4 Pro open, MIT-licensed, self-hostable

• On OrcaRouter — GPT-6 Luna not in our catalogue vs DeepSeek V4 Pro routable at DeepSeek's list price

A generated single-panel scoreboard card headed 'One index point, five times the price' with six rows: GPT-6 Luna $0.10 in / $0.50 out per 1M with index 37 at max effort; DeepSeek V4 Pro $0.66 in / $1.98 out per 1M off-peak doubling at peak with index 36 at max effort; maximum output Luna 128K tokens against V4 Pro's 384K; long context, Luna doubling its rates above 272K input against V4 Pro carrying no published surcharge clause; weights, Luna closed and API-only against V4 Pro open and MIT licensed; and reasoning off switch, Luna none through max against V4 Pro's low, high and max only. Footer: 'Index figures per Artificial Analysis; prices per OpenAI and DeepSeek.'

Read those lines together and the shape is unusual: the closed model is cheaper on every per-token line except cached input, and the open model wins on output length, on long-context behaviour, and on the one thing that cannot be priced per token at all.

Why the one-point gap is not the story

A one-point difference on a composite index is not a capability difference you can act on, and it is worth saying plainly that this is the least interesting number in the article. Two things about it do matter, though.

The first is that the index revision moves. Earlier snapshots of the same evaluation put DeepSeek V4 Pro's max-effort score materially higher than the 36 our capture shows today, and GPT-6 Astra's figures drifted across 52.8, 53 and 54.7 in material published within a single month. Any comparison that leans on the precise gap between two models on a rolling index is leaning on something that will not hold still. The honest reading is that these two models are in the same capability band at their ceilings, and the index cannot tell you which is better for your task.

The second is the effort default, and here the gap is real. GPT-6 Luna's 37 is a maximum-effort figure. Its default is medium, where it scores 29. DeepSeek V4 Pro's 36 is its maximum-effort figure too, and it is also the setting Artificial Analysis ran it at. A team that deploys GPT-6 Luna and changes nothing is running a 29 against a 36 — a seven-point deficit in the wrong direction, and the kind of thing that shows up as "the cheap model is worse" in a production post-mortem when the actual problem is a default.

What the open weights are actually worth

This is where the price comparison stops being arithmetic and becomes a question about your business. MIT-licensed weights buy you four things, and each has a price tag that never appears on a rate card.

The first is a floor on cost. Per-token pricing is a rental; weights are an asset. If your volume is large enough that a fixed GPU footprint beats a metered bill, DeepSeek V4 Pro is the only one of these two models where that option exists at all. At GPT-6 Luna's rates the crossover is far away — a one-cent cached input rate is genuinely hard to beat with rented silicon — but "far away" is a different statement from "impossible," and it is a one-way door: you can move from self-hosting to an API any day, and you cannot move the other direction.

The second is independence from a vendor's roadmap. GPT-5.6 Luna's promotional rates carried a stated expiry. GPT-6 Luna's are permanent, per OpenAI — but permanence is a promise about a model that a vendor can retire whenever it ships a successor. A model you host cannot be deprecated out from under you. For anyone who has been through a forced migration on a closed model, that is worth more than a factor of five on tokens.

The third is data. If the prompt contains material that cannot leave your infrastructure, the comparison ends here and it does not matter what GPT-6 Luna costs.

The fourth is that you can fine-tune it. GPT-6 Luna is not available for fine-tuning at all — it is Chat Completions, Responses and Batch, and that is the list. DeepSeek V4 Pro's weights can be trained on your own distribution, which for a narrow high-volume task is frequently a larger win than any general-purpose index point.

Against all of that, GPT-6 Luna's case is that it is roughly five times cheaper on a blended 3:1 input-to-output mix off-peak and closer to ten times cheaper during DeepSeek's peak hours, that its cached-input rate is a tenth of its already-low input rate, that it can be told to stop reasoning entirely, and that its list price does not double twice a day. For a workload with no data-residency constraint, no fine-tuning requirement, and no appetite for operating inference infrastructure, that is a straightforward answer.

The migration is smaller than the price gap suggests

Switching between these two is closer to a config change than a port. Both speak the OpenAI-compatible chat completions shape, both carry a 1M-class context window, and the failure modes are the ones you would expect rather than structural.

Two of them are worth flagging before you move traffic. GPT-6 Luna's long-context clause is the expensive one: cross 272,000 input tokens and the entire request reprices to double input and cache and 1.5x output, so a 300,000-token call costs twice what the same call costs split in two. DeepSeek V4 Pro has no such clause, which means that for genuinely enormous single requests the five-times price gap is smaller than it looks and can invert on a peak-hour call. And the output ceiling is a hard constraint rather than a preference: 128,000 tokens against 384,000. A pipeline that generates long structured artifacts may not fit in GPT-6 Luna at all, and no amount of price advantage fixes that.

DeepSeek V4 Pro is on OrcaRouter at DeepSeek's list price, with the pass-through pricing that puts a vendor rate change on our side the same day — relevant for a model whose rates already move between peak and off-peak. GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API, so a team that wants both runs one key for DeepSeek V4 Pro alongside whatever it already uses for OpenAI. Automatic failover is the part that earns its place in this particular pairing: a self-hosted or third-party DeepSeek endpoint degrading is exactly the scenario where you want the route to move without a deployment.

A screenshot of the Artificial Analysis page for GPT-6 Luna (max), showing an Intelligence rank of 6th of 183 models, a Speed rank of 35th, a Cost rank of 20th and a Verbosity rank of 36th, input pricing of $0.10 and output of $0.50 per 1M tokens, an Intelligence Index score of 37 against a comparable-model median of 12, 150M output tokens generated during the index run, 153.9 tokens per second, and a total of $122.39 spent evaluating the model.

Who should pick which

Take DeepSeek V4 Pro if any of the following is true. The prompt cannot leave your infrastructure. You need to fine-tune. You generate outputs longer than 128,000 tokens. Your requests routinely exceed 272,000 input tokens. Or your volume is large enough that you would rather buy GPUs than rent tokens, and you want the option to make that switch later without rewriting the application.

Take GPT-6 Luna if none of those apply and your workload is high-volume, program-consumed, and short. Classification, extraction, routing, bulk summarisation, the inner loop of an agent. At $0.10/$0.50 with a one-cent cached input rate and Batch at half price, the arithmetic is not close, and the one-point index difference at max effort will not be visible in your eval. Run it at medium, check the two-point Coding Agent Index regression against your own tests rather than a vendor chart, and keep your long calls on the right side of the 272,000-token line.

The mistake to avoid is treating the five-times price gap as the answer. It is the answer to one question — what does a token cost — and it is silent on the four questions that decide whether you can still run this workload in two years.

A screenshot of the OrcaRouter model page for DeepSeek V4 Pro (deepseek/deepseek-v4-pro), showing the Tools, JSON and Reasoning badges, list pricing of $0.66 input and $1.98 output per 1M tokens, a p50 time to first token of 3.56s, 4033.0M tokens of traffic, and the description of the model as DeepSeek's flagship text-generation model processing up to 1,048,576 tokens of context and generating up to 384,000 tokens of output.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily