A generated hero title card reading 'GPT-6 Luna vs GLM-5.3', subtitled 'Eight index points for ten times the cost per task, and the weights are still not published.', with chips showing Luna at $0.10/$0.50 per 1M with index 37, GLM-5.3 at $1.40/$4.40 per 1M with index 45, and cost per index task of $0.07 against $0.68, with the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6 Luna vs GLM-5.3: The Open-Weight Model You Still Can't Download

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GLM-5.3 is the better model. It scores 45 on the Artificial Analysis Intelligence Index at maximum effort against GPT-6 Luna's 37, it matches Mythos 5 on selected cybersecurity evaluations according to Z.ai, and it is the leading open-weights model on several of the composite boards. It is also, right now, an API you rent rather than a model you own: Z.ai delayed the weights release by roughly two weeks over cybersecurity concerns, and the download that is supposed to be the entire point of an open-weights flagship is not there yet.

That delay is the story of this matchup, and it changes the arithmetic in a way the price tables miss. GPT-6 Luna arrived on September 22, 2026 at $0.10 per million input tokens and $0.50 per million output, generally available, no strings. GLM-5.3 arrived on August 18, 2026 at $1.40 and $4.40 — fourteen times Luna's input rate and nearly nine times its output rate — with its weights held back. So the comparison is not open against closed. It is a closed model you can call today at a tenth of the cost against an open model you can only call today, priced as though you were buying the thing you cannot yet have.

Two models, one month apart, two very different deals

GPT-6 Luna is the small tier of OpenAI's GPT-6 generation, shipped alongside GPT-6 Sol at $2.00/$10.00 and beneath the flagship GPT-6 Astra at $10.00/$50.00. It carries a 1,050,000-token context window, a 128,000-token output cap, text and image input, and a reasoning ladder that runs from none through low, medium, high, xhigh and max — with medium as the default. OpenAI calls it the most efficient model in the family for focused, high-volume work, and the rate card backs that up: cached input at $0.01 per million is a tenth of an already-low uncached rate.

GLM-5.3 is Z.ai's current flagship for complex software engineering and long-horizon agentic work, released August 18, 2026. It is text-in, text-out, carries a 1M-token context window with a 128,000-token output cap, and reasons always on — there is no non-reasoning mode. Z.ai describes roughly a 50% improvement in coding experience over GLM-5.2 and positions it for repo-scale coding and autonomous multi-step engineering. The weights, when they land, will be the headline feature; today the API is what is live.

What the two rate cards actually say

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs GLM-5.3 $1.40

• Output per 1M — GPT-6 Luna $0.50 vs GLM-5.3 $4.40

• Cached input — GPT-6 Luna $0.01 per 1M vs GLM-5.3 $0.26 per 1M, an 81% discount on its own input rate

• AA Intelligence Index — GPT-6 Luna 37 at max effort vs GLM-5.3 45 at max effort

• Default-effort score — GPT-6 Luna 29 at its medium default vs GLM-5.3 always-on reasoning, measured at max

• Output tokens on the index run — GPT-6 Luna 150M vs GLM-5.3 210M, against a board median of 140M

• Cost to run the index — GPT-6 Luna $122.39 vs GLM-5.3 $2,503.48

• Cost per index task — GPT-6 Luna about $0.07 vs GLM-5.3 about $0.68

• Context and output — GPT-6 Luna 1,050,000 in / 128,000 out vs GLM-5.3 1M in / 128,000 out

• Long-context clause — GPT-6 Luna 2x input and cache, 1.5x output above 272K input, whole request vs GLM-5.3 no equivalent clause on the published card

• Reasoning off switch — GPT-6 Luna none through max vs GLM-5.3 always on, no non-reasoning mode

• Weights — GPT-6 Luna closed, API only vs GLM-5.3 open by commitment, release delayed

A generated single-panel scoreboard card headed 'Same suite, ten times the task cost' with six rows: GPT-6 Luna $0.10 in / $0.50 out per 1M with index 37 at max effort; GLM-5.3 $1.40 in / $4.40 out per 1M with index 45 at max effort; cost per completed index task about $0.07 against about $0.68; output tokens on the index run Luna 150M against GLM-5.3's 210M and a board median of 140M; reasoning configuration Luna none through max against GLM-5.3 always on with no off switch; and weights, Luna closed and API-only against GLM-5.3 open by commitment with the release delayed. Footer: 'Index and cost per task per Artificial Analysis; prices per OpenAI and Z.ai.'

The verbosity tax is the number that decides this

The eight-point index gap is real and the price gap is larger, and neither is the figure that matters most. The one that matters is output tokens per finished task, because that is where an eight-point capability edge either pays for itself or does not.

Artificial Analysis's index run is the cleanest controlled comparison available: the same suite, the same harness, run against both models. GLM-5.3 generated 210 million output tokens completing it. GPT-6 Luna generated 150 million. Against a board median of 140 million, both are chatty — GLM-5.3 by half again the median, GPT-6 Luna by roughly a tenth. But the direction of that difference compounds against Z.ai's model on price: it emits 40% more tokens and charges 8.8 times as much for each one.

Put the two together and the cost per completed index task comes out at about $0.68 for GLM-5.3 against about $0.07 for GPT-6 Luna — roughly a tenfold gap, larger than the raw rate-card ratio, because the more expensive model is also the more verbose one. That is the honest shape of this pairing. GLM-5.3 buys you eight index points for ten times the money per finished job, and whether that is a good purchase depends entirely on whether your workload is one where eight points is the difference between working and not.

For a lot of production traffic it is not. For a coding agent working a repo at the edge of what the model can do, it frequently is, and no amount of price advantage closes a capability gap on a task that fails.

Why the open weights change the answer, and why the delay changes it back

The standard argument for an open-weights flagship is that the per-token price is temporary. You rent until your volume justifies buying, then you download the model and the marginal cost of a token becomes electricity. On that argument GLM-5.3's $1.40/$4.40 is not really fourteen times GPT-6 Luna's $0.10 — it is a bridge price on the way to zero, and the closed model's cheaper rate is a rental with no exit.

That argument is correct, and it is currently suspended. Z.ai held the weights back by roughly two weeks over cybersecurity findings, which means the exit ramp that justifies the premium does not exist yet. Until it does, GLM-5.3 is a metered API at ten times the cost per task of a model that is four index points behind it on the general board and one point behind on the coding board — and the buyer is paying open-weights prices for closed-weights access.

There is a second-order point worth naming. The delay is not a scandal; it is a vendor taking a capability finding seriously, and a model that finds vulnerabilities well is exactly the kind of model a lab should be careful about releasing as a downloadable file. But it does mean the release date of the weights is now the single most important date in this comparison, and it is not published. A team choosing between these two models on a twelve-month horizon is choosing between a model whose terms can change next month and a model whose terms cannot — and only one of those is priced accordingly.

The migration surface is small

Both models speak an OpenAI-compatible chat completions surface, both carry a 1M-class context window, and both cap output at 128,000 tokens, so a port between them is closer to a configuration change than a rewrite. The differences that will actually bite are behavioural rather than structural.

GPT-6 Luna's long-context clause is the expensive one: requests above 272,000 input tokens are billed at twice the input and cache rates and 1.5x the output rate for the entire request, so a 300,000-token call costs double what the same work costs split in two. GLM-5.3's published card carries no equivalent clause, which on genuinely large single requests narrows the price gap considerably and can invert it on an output-heavy job.

The reasoning configuration cuts the other way. GLM-5.3 reasons always on and gives you no way to switch it off, which is a hard constraint for anything latency-sensitive — a classifier or a router cannot use it. GPT-6 Luna offers none, low, medium, high, xhigh and max, and the default is medium, where it scores 29 rather than 37. That single parameter is the difference between GPT-6 Luna being eight index points behind Z.ai's model and being sixteen behind, and a team that migrates without setting it explicitly is running the second configuration while believing it bought the first.

GLM-5.3 is on OrcaRouter at Z.ai's list price, under pass-through pricing that puts a vendor rate change on our side the same day rather than waiting on a reseller to renegotiate — which matters more than usual for a model whose headline number is expected to move when the weights land. GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API. A team running both, which is the rational setup here given how different the two are at the edges, runs one key for GLM-5.3 alongside whatever it already uses for OpenAI, and moves traffic between a five-cent classifier and a repo-scale coding agent by changing a route rather than a deployment.

A screenshot of the Artificial Analysis page for GPT-6 Luna (max), showing an Intelligence rank of 6th of 183 models, a Speed rank of 35th, a Cost rank of 20th and a Verbosity rank of 36th, input pricing of $0.10 and output of $0.50 per 1M tokens, an Intelligence Index score of 37 against a comparable-model median of 12, 150M output tokens generated during the index run, 153.9 tokens per second, and a total of $122.39 spent evaluating the model.

The call

Take GPT-6 Luna for anything program-consumed and high-volume. Classification, extraction, routing, bulk summarisation, the inner loop of an agent. At $0.10/$0.50, with cached input at a cent and Batch at half the standard rate, it is an order of magnitude cheaper per finished task than GLM-5.3 and the eight-point index gap will not be visible in your eval. Set the effort parameter explicitly and keep your long calls under 272,000 tokens.

Take GLM-5.3 when the task is the hard one and the answer matters more than the invoice. Repo-scale software engineering, autonomous multi-step work, anything where eight index points is the difference between a task completing and a task failing. The verbosity is a real tax and the ten-times task cost is real, but a model that finishes is cheaper than a model that has to be retried.

And watch the weights. The day Z.ai publishes the GLM-5.3 download is the day this comparison's central fact changes, and it is the one date on this page that is not yet on the calendar.

A screenshot of the OrcaRouter model page for GLM-5.3 (z-ai/glm-5.3), showing the Tools, JSON and Reasoning badges, a 2026-08-18 catalogue date, 1M-token context with 128K maximum output, list pricing of $1.26 input and $3.96 output per 1M tokens, a p50 time to first token of 3.99s, 1807.1M tokens of traffic, and the description of GLM-5.3 as Z.ai's flagship model for complex software engineering and long-horizon agentic tasks.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily