
GPT-6 Luna vs GLM-5.3: The Open-Weight Model You Still Can't Download
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
GLM-5.3 is the better model. It scores 45 on the Artificial Analysis Intelligence Index at maximum effort against GPT-6 Luna's 37, it matches Mythos 5 on selected cybersecurity evaluations according to Z.ai, and it is the leading open-weights model on several of the composite boards. It is also, right now, an API you rent rather than a model you own: Z.ai delayed the weights release by roughly two weeks over cybersecurity concerns, and the download that is supposed to be the entire point of an open-weights flagship is not there yet.
That delay is the story of this matchup, and it changes the arithmetic in a way the price tables miss. GPT-6 Luna arrived on September 22, 2026 at $0.10 per million input tokens and $0.50 per million output, generally available, no strings. GLM-5.3 arrived on August 18, 2026 at $1.40 and $4.40 — fourteen times Luna's input rate and nearly nine times its output rate — with its weights held back. So the comparison is not open against closed. It is a closed model you can call today at a tenth of the cost against an open model you can only call today, priced as though you were buying the thing you cannot yet have.
Two models, one month apart, two very different deals
GPT-6 Luna is the small tier of OpenAI's GPT-6 generation, shipped alongside GPT-6 Sol at $2.00/$10.00 and beneath the flagship GPT-6 Astra at $10.00/$50.00. It carries a 1,050,000-token context window, a 128,000-token output cap, text and image input, and a reasoning ladder that runs from none through low, medium, high, xhigh and max — with medium as the default. OpenAI calls it the most efficient model in the family for focused, high-volume work, and the rate card backs that up: cached input at $0.01 per million is a tenth of an already-low uncached rate.
GLM-5.3 is Z.ai's current flagship for complex software engineering and long-horizon agentic work, released August 18, 2026. It is text-in, text-out, carries a 1M-token context window with a 128,000-token output cap, and reasons always on — there is no non-reasoning mode. Z.ai describes roughly a 50% improvement in coding experience over GLM-5.2 and positions it for repo-scale coding and autonomous multi-step engineering. The weights, when they land, will be the headline feature; today the API is what is live.
What the two rate cards actually say
One line per dimension, both sides on each:
• Input per 1M — GPT-6 Luna $0.10 vs GLM-5.3 $1.40
• Output per 1M — GPT-6 Luna $0.50 vs GLM-5.3 $4.40
• Cached input — GPT-6 Luna $0.01 per 1M vs GLM-5.3 $0.26 per 1M, an 81% discount on its own input rate
• AA Intelligence Index — GPT-6 Luna 37 at max effort vs GLM-5.3 45 at max effort
• Default-effort score — GPT-6 Luna 29 at its medium default vs GLM-5.3 always-on reasoning, measured at max
• Output tokens on the index run — GPT-6 Luna 150M vs GLM-5.3 210M, against a board median of 140M
• Cost to run the index — GPT-6 Luna $122.39 vs GLM-5.3 $2,503.48
• Cost per index task — GPT-6 Luna about $0.07 vs GLM-5.3 about $0.68
• Context and output — GPT-6 Luna 1,050,000 in / 128,000 out vs GLM-5.3 1M in / 128,000 out
• Long-context clause — GPT-6 Luna 2x input and cache, 1.5x output above 272K input, whole request vs GLM-5.3 no equivalent clause on the published card
• Reasoning off switch — GPT-6 Luna none through max vs GLM-5.3 always on, no non-reasoning mode
• Weights — GPT-6 Luna closed, API only vs GLM-5.3 open by commitment, release delayed

The verbosity tax is the number that decides this
The eight-point index gap is real and the price gap is larger, and neither is the figure that matters most. The one that matters is output tokens per finished task, because that is where an eight-point capability edge either pays for itself or does not.
Artificial Analysis's index run is the cleanest controlled comparison available: the same suite, the same harness, run against both models. GLM-5.3 generated 210 million output tokens completing it. GPT-6 Luna generated 150 million. Against a board median of 140 million, both are chatty — GLM-5.3 by half again the median, GPT-6 Luna by roughly a tenth. But the direction of that difference compounds against Z.ai's model on price: it emits 40% more tokens and charges 8.8 times as much for each one.
Put the two together and the cost per completed index task comes out at about $0.68 for GLM-5.3 against about $0.07 for GPT-6 Luna — roughly a tenfold gap, larger than the raw rate-card ratio, because the more expensive model is also the more verbose one. That is the honest shape of this pairing. GLM-5.3 buys you eight index points for ten times the money per finished job, and whether that is a good purchase depends entirely on whether your workload is one where eight points is the difference between working and not.
For a lot of production traffic it is not. For a coding agent working a repo at the edge of what the model can do, it frequently is, and no amount of price advantage closes a capability gap on a task that fails.
Why the open weights change the answer, and why the delay changes it back
The standard argument for an open-weights flagship is that the per-token price is temporary. You rent until your volume justifies buying, then you download the model and the marginal cost of a token becomes electricity. On that argument GLM-5.3's $1.40/$4.40 is not really fourteen times GPT-6 Luna's $0.10 — it is a bridge price on the way to zero, and the closed model's cheaper rate is a rental with no exit.
That argument is correct, and it is currently suspended. Z.ai held the weights back by roughly two weeks over cybersecurity findings, which means the exit ramp that justifies the premium does not exist yet. Until it does, GLM-5.3 is a metered API at ten times the cost per task of a model that is four index points behind it on the general board and one point behind on the coding board — and the buyer is paying open-weights prices for closed-weights access.
There is a second-order point worth naming. The delay is not a scandal; it is a vendor taking a capability finding seriously, and a model that finds vulnerabilities well is exactly the kind of model a lab should be careful about releasing as a downloadable file. But it does mean the release date of the weights is now the single most important date in this comparison, and it is not published. A team choosing between these two models on a twelve-month horizon is choosing between a model whose terms can change next month and a model whose terms cannot — and only one of those is priced accordingly.
The migration surface is small
Both models speak an OpenAI-compatible chat completions surface, both carry a 1M-class context window, and both cap output at 128,000 tokens, so a port between them is closer to a configuration change than a rewrite. The differences that will actually bite are behavioural rather than structural.
GPT-6 Luna's long-context clause is the expensive one: requests above 272,000 input tokens are billed at twice the input and cache rates and 1.5x the output rate for the entire request, so a 300,000-token call costs double what the same work costs split in two. GLM-5.3's published card carries no equivalent clause, which on genuinely large single requests narrows the price gap considerably and can invert it on an output-heavy job.
The reasoning configuration cuts the other way. GLM-5.3 reasons always on and gives you no way to switch it off, which is a hard constraint for anything latency-sensitive — a classifier or a router cannot use it. GPT-6 Luna offers none, low, medium, high, xhigh and max, and the default is medium, where it scores 29 rather than 37. That single parameter is the difference between GPT-6 Luna being eight index points behind Z.ai's model and being sixteen behind, and a team that migrates without setting it explicitly is running the second configuration while believing it bought the first.
GLM-5.3 is on OrcaRouter at Z.ai's list price, under pass-through pricing that puts a vendor rate change on our side the same day rather than waiting on a reseller to renegotiate — which matters more than usual for a model whose headline number is expected to move when the weights land. GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API. A team running both, which is the rational setup here given how different the two are at the edges, runs one key for GLM-5.3 alongside whatever it already uses for OpenAI, and moves traffic between a five-cent classifier and a repo-scale coding agent by changing a route rather than a deployment.

The call
Take GPT-6 Luna for anything program-consumed and high-volume. Classification, extraction, routing, bulk summarisation, the inner loop of an agent. At $0.10/$0.50, with cached input at a cent and Batch at half the standard rate, it is an order of magnitude cheaper per finished task than GLM-5.3 and the eight-point index gap will not be visible in your eval. Set the effort parameter explicitly and keep your long calls under 272,000 tokens.
Take GLM-5.3 when the task is the hard one and the answer matters more than the invoice. Repo-scale software engineering, autonomous multi-step work, anything where eight index points is the difference between a task completing and a task failing. The verbosity is a real tax and the ten-times task cost is real, but a model that finishes is cheaper than a model that has to be retried.
And watch the weights. The day Z.ai publishes the GLM-5.3 download is the day this comparison's central fact changes, and it is the one date on this page that is not yet on the calendar.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
