Generated title card reading "GPT-6 vs GLM 5.3", with a chip reading "GLM 5.3: open weights, released 2026-08-18" and a chip reading "GPT-6: closed weights, shipped 2026-09-22", and a footer reading "AA peer groups differ: open-weights and proprietary scores are not subtractable." The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

GPT-6 vs GLM 5.3: One of These You Can Download, and the Price Gap Is Smaller Than the Index Gap

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The reason to put GPT-6 against GLM 5.3 is not the benchmark table. It is that GLM 5.3 is the one model in this price bracket whose weights you can take home, and GPT-6 is a closed generation whose three members are only reachable through an API. That difference decides more procurement questions than any index delta, and it is also the reason the two scores are not directly comparable — which is the part most published comparisons of this pair get wrong.

The two sides, precisely

GLM 5.3 is Z.ai's flagship, released 18 August 2026, with weights published about a week later. It is a text-only model — no image input — with a 1,000,000-token context window and a 128,000-token output ceiling, listed by Z.ai at $1.40 per million input tokens and $4.40 per million output, with a cached-input rate of $0.234 per million. It is served under the model id z-ai/glm-5.3.

GPT-6 is a generation rather than a model. GPT-6 Astra, GPT-6 Sol and GPT-6 Luna are three ids at three price points, all released 22 September 2026, all multimodal on the input side (text, image and file), all with a context window near 1,050,000 tokens and a 128,000-token output ceiling. There is no openai/gpt-6 endpoint. When someone says "GPT-6" in a price conversation they almost always mean GPT-6 Sol, the middle tier at $2.00 / $10.00 — so that is the member this comparison uses wherever a single id is needed.

• Weights — GLM 5.3 open, downloadable and self-hostable; GPT-6 closed, API only
• Input modalities — GLM 5.3 text only; all three GPT-6 tiers accept text, image and file
• Input price — GLM 5.3 $1.40 per million vs GPT-6 Sol $2.00 per million
• Output price — GLM 5.3 $4.40 per million vs GPT-6 Sol $10.00 per million
• Cached input — GLM 5.3 $0.234 per million vs GPT-6 Sol $0.20 per million
• Context — GLM 5.3 1,000,000 tokens vs GPT-6 Sol 1,050,000 tokens
• Long-request step — GLM 5.3 has none on its published card; GPT-6 Sol reprices the whole request above 272,000 input tokens to $4.00 / $15.00
• Licence — GLM 5.3 ships under Z.ai's own open licence; GPT-6 has no licence file because there is nothing to license

Note the caveat on the price line, because it is a real one and this blog has hit it before: an OrcaRouter model page is not always in sync with the vendor's own rate card, and on GLM 5.3 specifically our catalogue carries $1.26 / $3.96 while Z.ai's published figure is $1.40 / $4.40. Where the two disagree, quote the vendor's number in prose — which is what the bullets above do — and read the page figure as the page's own value. Both are legitimate; they are just not the same number.

Why the index comparison needs a warning label

Here is the trap. On the Artificial Analysis Intelligence Index, GLM 5.3 scores 44.8 and GPT-6 Sol scores 47.6 — 2.8 points apart. That looks like a near-tie between a closed model and an open one at roughly a fifth of the output price, and it is the sentence most comparisons of this pair stop at.

It is not a valid subtraction. Artificial Analysis ranks open-weights models only against other open-weights models of the same size class, and proprietary models across a proprietary band using a blended 3:1 input-to-output price ratio. GLM 5.3's 44.8 is an open-weights score. GPT-6 Sol's 47.6 is a proprietary-band score. They sit on different scales, and the gap between them is not a measurement of capability. The same caveat applies to the two figures on the OrcaRouter catalogue entries, which is where both numbers above come from.

What you can compare across that line is cost accounting and the rate cards, and secondarily the rows that are the same evaluation run the same way. Read on both pages:

• AA Coding Index — GLM 5.3 74.8 vs GPT-6 Sol 60.0
• GPQA Diamond — GLM 5.3 91.7
• Humanity's Last Exam — GLM 5.3 42.3 vs GPT-6 Sol 47.9
• Terminal-Bench 2.1 — GLM 5.3 83.9
• Terminal-Bench 4.0 — GLM 5.3 41.9 vs GPT-6 Sol 43.9
• τ-bench banking — GLM 5.3 50.3
• SciCode — GLM 5.3 59.0 vs GPT-6 Sol 57.6
• Long-context recall — GLM 5.3 79.7 vs GPT-6 Sol 83.7

Read as a shape rather than as a score difference: GLM 5.3 is strong on the agentic and coding evaluations — τ-bench banking and SciCode both lean its way, and its Coding Index is far above Sol's — while GPT-6 Sol holds the long-context recall and Humanity's Last Exam rows. Neither of those readings is a verdict. They are two different sets of strengths, measured by a third party, on a model you can download and a model you cannot.

What open weights are actually worth here

The reason to care about GLM 5.3's licence is not ideology. It is that three specific things stop being vendor decisions once the weights are yours.

The first is residency. A self-hosted GLM 5.3 runs where you put it, and no data-processing agreement governs what leaves your network. GPT-6 has no equivalent option — the only way to call it is through an API, and the data goes somewhere you do not control. For a regulated workload this is often the whole decision, and it is a decision the index never captures.

The second is the price curve. An API rate card is a vendor's number that the vendor can move; the published GPT-6 rates are vendor-stated as permanent rather than promotional, but that is a statement about intent, not a contract. Owned weights have a fixed cost that stops scaling with tokens, and the crossover is a function of your volume rather than of anyone's pricing decision. That is why the $1.40 / $4.40 card is not really the comparison — the comparison is $1.40 / $4.40 against your own hardware amortised across your own traffic.

The third is the failure mode. A hosted model can be deprecated, rate-limited, or degraded by a serving change. GLM 5.3 has been callable since 18 August with a public checkpoint behind it; if Z.ai changes something you disagree with, the checkpoint you validated against is still on disk.

The cost of that freedom is real too, and worth stating plainly. GLM 5.3 is text-only, so a workload that feeds screenshots or PDFs into the model has no path on the open-weights side of this comparison at all. Serving 1,000,000-token contexts yourself is not a laptop exercise. And the model you download is the model you are responsible for — evaluation, safety filtering, uptime and upgrades all become yours, which is a real engineering budget that the API price does not include.

Where GPT-6 earns the premium

Against that, the case for the closed side is narrower than its price suggests but it is not empty.

Multimodal input is the concrete one. All three GPT-6 tiers accept images and files natively, and GLM 5.3 does not accept either. A pipeline that reads a chart, a screenshot or a scanned document and reasons over it in the same call cannot be built on GLM 5.3 as published.

Long-context recall is the second. GPT-6 Sol leads GLM 5.3 on that row by four points, and it is the row that governs whether a very large context is actually usable or merely accepted. A 1,000,000-token window that loses the middle of the document is a larger window than a 100,000-token one that does not.

And the generation has more than one price point. The $2.00 / $10.00 Sol card is the one usually compared against GLM 5.3, but GPT-6 Luna sits at $0.10 / $0.50 — below GLM 5.3's input rate by an order of magnitude and below its output rate by nearly nine times — and GPT-6.1 Sol, shipped 29 September, scores 51.8 on the same index at Sol's price. If the question is purely "cheapest competent token", the GPT-6 generation has an answer that GLM 5.3 does not.

A generated two-column comparison scoreboard titled "GPT-6 Sol vs GLM 5.3 — the scoreboard". The left column, GPT-6 Sol, reads Weights closed, Input $2.00, Output $10.00, AA Intelligence 47.6, AA Coding 60.0, Long-context recall 83.7. The right column, GLM 5.3, reads Weights open, Input $1.40, Output $4.40, AA Intelligence 44.8, AA Coding 74.8, Long-context recall 79.7. A footer line reads "AA ranks open-weights and proprietary models on different peer scales; the two Intelligence figures are not subtractable." The OrcaRouter logo is composited in the bottom-right corner.

Running them together, which is the answer most teams land on

The practical resolution of this comparison is that it is not exclusive. GLM 5.3 has been pulling very large volume — the OrcaRouter catalogue recorded roughly 1.9 billion tokens routed to it in a seven-day window, against about 2.1 million for GPT-6 Sol — which is what an open-weight model with a cheap output rate looks like when teams put bulk work on it. GPT-6's traffic in the same window is concentrated on the higher tiers, where the work is the kind that needs the multimodal input or the premium model.

Both are live routes in the same catalogue on OrcaRouter: one API in front of 200+ models, called through an OpenAI-compatible endpoint, with the provider's list price passed through at 0% markup so a vendor price change is live on our side the same day rather than at your next billing cycle. That matters more than usual in this particular comparison, because the whole GLM 5.3 price question has a moving part — the catalogue figure and the vendor figure already disagree by about 10% — and a pass-through route means you are billed the vendor's number rather than a markup that hides it.

The routing DSL is what makes the split explicit: send the bulk, text-only classification and extraction traffic to GLM 5.3 and the multimodal or long-recall work to a GPT-6 tier, per request rather than per quarter, and let automatic failover cover either side. The open-weights-vs-API question stops being an all-or-nothing architectural commitment and becomes a routing rule you can change next week.

Screenshot of the OrcaRouter model page for z-ai/glm-5.3, showing the Z.ai vendor label, a 2026-08-18 catalogue release date, a context of 1M tokens, a 128K maximum output, text-only input, Tools, JSON and Reasoning badges, and a rate strip reading $1.26 per million input tokens, $3.96 per million output tokens, a 4.54 s median time to first token, a 10.00 s p95, and 1,903.6M tokens routed in seven days.

The verdict, which is a question about your constraints

If you can self-host, need text-only throughput at scale, or have a residency requirement that an API cannot satisfy, GLM 5.3 is the answer and the price difference is large enough to fund the serving work. If you need image and file input, if long-context recall is load-bearing, or if you have no appetite for operating a model yourself, GPT-6 is the answer and the premium buys something specific rather than a brand.

What neither answer is: "GPT-6 scores 2.8 points higher". That subtraction is not available, because the two figures come from different peer groups on the same board — and a comparison built on it would be quoting a number as if it measured something it does not. The rate cards are comparable. The licence position is comparable, in the sense that it is starkly different. The index rows are not.

Screenshot of the OrcaRouter model page for openai/gpt-6-sol, showing the OpenAI vendor label, a 2026-09-22 catalogue release date, a 1,050,000-token context window (shown as 1.05M-token context in the model blurb), a 128K maximum output, text, image and file input, and a rate strip reading $2.00 per million input, $10.00 per million output, a 6.79 s median time to first token, a 10.00 s p95, and 2.1M tokens routed in seven days.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily