A hero title card for "Qwen 4 Max vs GLM-5.3" with the subtitle "Open weights, custom licences, and the tier that actually ships", three pill badges reading "$1.26 and $3.96", "text-only vs multimodal" and "the 27B is the one to watch", a footer line reading "Alibaba announced September 22, 2026; GLM-5.3 listed August 18, 2026", and the OrcaRouter logo in the bottom-right corner.
Engineering & Research

Qwen 4 Max vs GLM-5.3: open weights, custom licences, and the tier that actually ships

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Qwen 4 Max was previewed at the Yunqi conference in Hangzhou on September 22, 2026 as the flagship of a four-model line that also includes Qwen 4 Flash, Qwen 4 Plus and an open-weight Qwen 4 27B tier. Z.ai's GLM-5.3 has been listed since August 18, 2026, text-only, with a 1-million-token context window and a 128,000-token output ceiling, at $1.26 per million input tokens and $3.96 per million output tokens on our catalogue, against the $1.40 and $4.40 list rates Z.ai publishes. Both labs come out of the open-weights tradition and both now sell frontier-tier models through hosted APIs, so the usual closed-versus-open framing does not apply here. The axis that does apply is the one nobody puts in a comparison table: what the licence actually permits, and which tier in each family is the one you are allowed to do anything interesting with.

Two labs, two answers to the same question

Alibaba and Z.ai are the two most consistent open-weights publishers among the frontier-adjacent labs, and they have arrived at different postures on how much of that openness survives into the flagship tier.

Alibaba's Qwen 3.8 generation published weights for its largest model — Qwen3.8-2.4T-A95B, in BF16 and FP8 — on August 12, 2026. It published them under a custom Qwen licence, not Apache 2.0. That distinction is the whole point: the weights exist, which means fine-tuning, quantisation and self-hosting are possible in a way they never will be for a closed model, but the terms are Alibaba's own and they are not the permissive default that the phrase "open weights" tends to imply.

Z.ai's GLM line comes from a lab with an open-weights lineage, and the flagship is no longer an exception to it. GLM-5.3's own weights were published on August 28, 2026 — the 743-billion-parameter flagship, in 141 safetensors shards — under a custom licence that Z.ai wrote for it rather than under MIT. The terms are roughly an MIT grant with a security-review condition attached for model-as-a-service businesses above a large revenue threshold. Note what that means for the family: GLM-5.2 and GLM-5.3-Flash carry MIT, and the flagship does not.

As listed, GLM-5.3 is a text-only model with a published context window, a published output ceiling and documented tool use, JSON mode and reasoning.

Put the two side by side and the practical contrast is this:

• Flagship weights — Qwen 3.8 flagship published under a custom Qwen licence versus GLM-5.3 weights published August 28, 2026 under a custom Z.ai licence

• Input price — Qwen3.8-Max $2.00 per 1M tokens versus GLM-5.3 $1.26 per 1M tokens on our catalogue, $1.40 on Z.ai's own list

• Output price — Qwen3.8-Max $6.00 per 1M tokens versus GLM-5.3 $3.96 per 1M tokens on our catalogue, $4.40 on Z.ai's own list

• Context — Qwen3.8-Max 1M tokens versus GLM-5.3 1M tokens

• Output ceiling — Qwen3.8-Max 128K tokens versus GLM-5.3 128K tokens

• Modality — Qwen3.8-Max text, image and video input versus GLM-5.3 text-only, documented as such

• Capability tags — Qwen3.8-Max vision, tools, JSON and reasoning versus GLM-5.3 tools, JSON and reasoning

• Qwen 4 Max — announced September 22, 2026, no price, no context, no weights, no date versus GLM-5.3 shipped and serving traffic

The multimodal row is what pays for the price gap. Qwen's flagship takes image and video input and charges $2.00 in and $6.00 out; GLM-5.3 takes text and charges $1.26 and $3.96. If your workload is text, you are paying a vision-encoder premium you never invoke, and that is the most concrete reason a text-only specialist beats a multimodal flagship on cost for text work.

A two-column scoreboard titled "Qwen 4 Max vs GLM-5.3 - the scoreboard". Left column Qwen 4 Max: Status Max announced, no date; Input price $2.00 per 1M tokens; Output price $6.00 per 1M tokens; Modality text, image, video in; Flagship weights custom Qwen licence; Output ceiling 128K tokens. Right column GLM-5.3: Status shipping since Aug 18 2026; Input price $1.26 per 1M tokens; Output price $3.96 per 1M tokens; Modality text-only; Flagship weights custom Z.ai licence, Aug 28 2026; Output ceiling 128K tokens. Footer reads "GLM-5.3 pricing from its catalogue listing; Qwen3.8-Max figures from its published model card, September 22 2026.", with the OrcaRouter logo bottom-right.

The 27B tier is the one that decides this

Qwen 4 Max is the headline and the 27B tier is the story. Every Qwen generation has shipped its small open tier under terms that let people do things, and the downstream evidence is unambiguous: within days of the Qwen 3.8 27B weights landing, the ecosystem had produced MLX conversions, NVFP4 quantisations and uncensored fine-tunes, none of which required anyone's permission. That is what a permissive small-model release buys, and it is a form of distribution no hosted API can match.

The flagship tier is a different matter. A 2.4-trillion-parameter model published under a custom licence is open in the sense that the bytes are downloadable and closed in the sense that what you may do with them is defined by the publisher. Both things are true at once, and the Qwen 4 announcement did not change either. If the Qwen 4 27B tier follows the pattern of its predecessors, it will be the release that matters to the largest number of people, and it is also the release with the most predictable delivery date, because small open tiers are the easiest thing on the roadmap to finish.

Which brings the comparison back to something usable. GLM-5.3 is a hosted model at a published price with a documented capability set, and it is available now. Qwen 4 Max is a tier name. If you want the openness argument to be concrete rather than philosophical, the model to wait for is the 27B, not the Max.

The routing layer is where the licence question stops mattering

There is a version of this decision that avoids the licence question entirely, and it is worth being honest that it is a real option rather than a compromise. If your requirement is to call a model rather than to own one, the licence on the weights is irrelevant and the terms that matter are the price, the latency and the failure behaviour.

OrcaRouter carries GLM-5.3 as z-ai/glm-5.3 at $1.26 input and $3.96 output per million tokens — below the $1.40 and $4.40 Z.ai publishes, with the model page showing those list rates beside the rate we bill at, and no markup added on top of what we pay the provider. The same OpenAI-compatible endpoint carries the Qwen 3.8 family, Qwen3.8-Max, Qwen3.8-Flash and Qwen3.8-27B, alongside close to 200 other models, with automatic failover when a provider path degrades and a routing DSL that lets you express the substitution policy explicitly rather than editing a model string in application code. When a vendor changes its rate, the change lands on your key the same day, because there is no margin layer for it to get stuck behind.

What OrcaRouter does not carry is Qwen 4 Max or any Qwen 4 tier. The catalogue has no entry for the family, which is the correct state for a model that was announced on a stage and has not shipped. Qwen 4 Max will be reachable, when it is reachable at all, through the vendor's own API and several third-party platforms. Until then, the Qwen model in this comparison that you can actually call is Qwen3.8-Max, and it is on the same catalogue as GLM-5.3 — which makes the live version of this matchup a routing-policy decision rather than a procurement one.

A screenshot of the OrcaRouter model page for GLM 5.3 (z-ai/glm-5.3, English UI), showing the 1M token context badge, the model ID z-ai/glm-5.3, a 128K maximum output, text input, the Tools, JSON and Reasoning tags, the listing date 2026-08-18 credited to Z.ai, a p50 time-to-first-token of 4.41s, the OpenAI-compatible code samples against api.orcarouter.ai/v1, and the opening of the model description calling GLM-5.3 Z.ai's latest flagship for complex software engineering and long-horizon agentic tasks.

What the licence question actually decides

GLM-5.3 is cheaper on both axes, its context window and output ceiling match the Qwen flagship's, its modality limits are documented rather than left to inference, and it is serving traffic today. Qwen 4 Max has a conference slot and a tier name, and the only part of the Qwen 4 roadmap with a predictable delivery is the 27B tier.

The decision that survives the launch is not which model is better. It is whether you need to own the thing. If you need to self-host, fine-tune or modify, the flagship weights from either lab come with terms you should read before you build a business on them, and the small open tier is where the permissive options have historically lived. If you need to call a model, GLM-5.3 is the cheaper answer today for text work, Qwen3.8-Max is the answer if you need image or video input, and both are one endpoint and one bill away. Revisit the flagship question when Alibaba publishes a model card, and read the licence before the benchmark table.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (qwen/qwen3.8-max, English UI), showing the 1M token context badge, the model ID qwen/qwen3.8-max, text plus image plus video input with text output, the Vision, Tools, JSON, Reasoning and Thinking capability tags, the listing date 2026-08-03, a p50 time-to-first-token of 3.29s, and the model description naming Qwen3.8-Max Alibaba's newest flagship and highest-capability tier to date.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily