Une carte de titre principale pour GLM-5.3 avec un badge « RAPPORT DE FUITE », le sous-titre « Le fleuron 'Epic Plus' de Z.ai — Ce que nous savons jusqu'à présent », et trois chips indiquant « Pas encore publié », « Piste de fuite : 3 août » et « Successeur de GLM-5.2 ».
Guides & Insights

GLM-5.3 lancé : la fuite était réelle — le vaisseau amiral post-entraîné de Z.ai pour le codage et la cyberdéfense

Auteur

Rowan Sterling

Date de publication

Derniers modèles · 20Voir tous les modèles
Benchmarks : Artificial Analysis · mis à jour quotidiennement
Retour à tous les articles

The leak was real, the launch is here, and GLM-5.3 now has an independent score to argue about. Artificial Analysis' Intelligence Index — measured by the lab, not by Z.ai — puts GLM-5.3 at 60, tied with Kimi K3 for the top open-weights score on the board and 7 points clear of GLM-5.2's 53. The API went live this week at the same price as its predecessor, the open weights are confirmed for Friday, August 28, and the coding and cyber-defense claims Z.ai has been making since the August 14 announcement are starting to become testable. This page first tracked GLM-5.3 from its August 3 leak traces; this is the launch report, updated in place with what the launch, the API, and the first independent benchmark actually confirmed.

De la fuite au lancement

The four traces that surfaced on August 3 — a "ZCode for GLM-5.3" harness page, an official docs page reachable for roughly an hour, a Bing index entry reading "GLM-5.3 Official Harness," and a commit adding a "glm-5.3" entry with JSON Schema support to Zhipu's official Java SDK — all pointed at a real, named release. Z.ai co-founder Tang Jie's "sooooooon" reply and the "epic-level plus" framing are now confirmed by an actual product rather than a rumor. The "roughly a week" timing signal this page tested — posted on X by @teortaxesTex after DeepSeek V4 Pro shipped on August 13 — held to the day: Z.ai formally announced GLM-5.3 on August 14 under the slogan "Built to Code. Ready for Cyber Defense." The follow-on came this week: on August 19 Z.ai said the GLM-5.3 API was live and open for calls, priced the same as GLM-5.2, with the model already wired into ZCode, AutoClaw, and the GLM Coding Plan.

Ce que le post-entraînement a réellement rapporté

L'affirmation concernant l'architecture est la partie qu'il vaut la peine de cerner, car le détail le plus frappant de la fuite — {{1}}des paramètres dépassant mille milliards{{/1}} — est faux. GLM-5.3 n'est pas un modèle plus grand. Z.ai affirme qu'il réutilise exactement la même base Mixture-of-Experts de 743B que GLM-5.2 (environ 40 milliards de paramètres actifs par jeton), conserve la même fenêtre de contexte de 1M de jetons et une sortie maximale d'environ 128K, et obtient tous ses gains d'un post-entraînement à plus grande échelle : plus d'environnements de tâches à horizon long, plus de types d'environnements et des sessions d'entraînement plus longues, le tout reposant sur le framework long-contexte IndexShare, la RL asynchrone SAO et le framework open-source slime qui ont déjà produit GLM-5.2. Cette présentation « pas de réentraînement, tout en post-entraînement » est propre à Z.ai et n'a pas été auditée de manière indépendante. La taille est l'autre gros titre : avec {{2}}743B de paramètres au total et seulement environ 40B actifs par jeton{{/2}}, GLM-5.3 est assez léger pour être auto-hébergé sur un cluster modeste et assez économique pour être servi à grande échelle — le positionnement « plus petit, moins cher, ouvert » que les commentaires du jour du lancement lui ont attribué, et un contraste marqué avec les fleurons fermés de l'IA frontière auxquels il est comparé.

Codage et agents : les chiffres que Z.ai avance

En codage, Z.ai rapporte — tous rapportés par le fournisseur et non reproduits — Terminal-Bench 3.0 passant de 4,6 à 28,3, ce que Z.ai qualifie de meilleur score open-weights sur ce banc d'essai ; DeepSWE v1.1 passant de 46,2 à 66,9 ; SWE-Marathon doublant approximativement de 19,4 à 42,5 ; et Agents' Last Exam (CLI) passant de 23,8 à 28,5. Sur le banc de test de code interne de Z.ai, GLM-5.3 a obtenu 31,4 % avec environ 50K jetons de sortie par tâche à effort élevé, contre 29,5 % pour Claude Opus 4.8 avec environ 120K jetons — le point de Z.ai étant que GLM-5.3 atteint un résultat comparable tout en dépensant beaucoup moins de jetons de sortie. Claude Fable 5 mène toujours cette évaluation interne avec 39,5 % à effort maximal, et Z.ai concède que GLM-5.3 reste derrière GPT-5.6 Sol et Claude Fable 5 sur plusieurs évaluations de codage plus difficiles. Considérez tous ces chiffres comme des données du fournisseur jusqu'à ce qu'un banc d'essai indépendant les reproduise.

Cyberdéfense : la capacité que personne n'a vue venir

Les chiffres en cybersécurité constituent la véritable actualité, et ils sont également entièrement rapportés par les fournisseurs. Sur CyberGym, un benchmark de découverte et de validation de vulnérabilités en boîte blanche, Z.ai rapporte GLM-5.3 à 84,5 %, en hausse par rapport aux 77,2 % de GLM-5.2 et devant Mythos 5 d'Anthropic (83,8 %) et GPT-5.6 Sol (83,6 %). Sur ExploitBench, qui exige à la fois une analyse des causes profondes et un exploit fonctionnel, GLM-5.3 a plus que doublé, passant de 24,4 % à 54,4 %, bien que Mythos 5 (78,0 %) reste en tête. Sur ExploitGym, Z.ai rapporte 105 tâches accomplies dans un budget de 2 heures et 130 en 6 heures, contre 29 et 39 pour GLM-5.2 — encore une fois derrière Mythos 5 (181 et 247). Z.ai présente la capacité cyber comme une propriété émergente d'un post-entraînement à l'échelle — « la capacité n'a cessé de s'accroître à mesure que l'entraînement montait en échelle », selon les termes de l'entreprise — plutôt que comme une cible délibérée.

Z.ai ajoute une affirmation concrète pour accompagner les benchmarks : lors de tests avec des équipes de sécurité, GLM-5.3 a identifié 2 436 vulnérabilités dans 269 projets open source, dont 1 097 jugées critiques ou à haute sévérité, la plus ancienne découverte remontant à 1981 et une « durée de vie » moyenne de 26,6 ans. C’est le chiffre le plus frappant de l’annonce, et aussi le moins vérifiable de manière indépendante. L’entreprise a associé cette capacité à un Security Disclosure Ledger pour la divulgation coordonnée, à un programme « trusted access » qui limite les fonctions cyber sensibles aux utilisateurs vérifiés, et à une initiative « Open Source Shield » visant à auditer en continu les principaux projets open source.

Le baseline GLM-5.3 devait être battu.

GLM-5.2 is the reference point the whole story hangs on. It shipped in June 2026 as a 743B Mixture-of-Experts model with roughly 40B active parameters per token, a 1M-token context window, a 128K max output, an MIT license, and open weights on Hugging Face. Independently, Artificial Analysis' Intelligence Index puts GLM-5.2 at 53 — the highest open-weights score on the index until this week. GLM-5.3 now clears it by 7 points: Artificial Analysis measures GLM-5.3 at 60 on the same index (v4.1.1), tying Kimi K3 for the top open-weights position and landing it in the frontier band alongside closed flagships like Claude Fable 5 and GPT-5.6 Sol. That is the first independent number attached to GLM-5.3, and it is consistent with the direction — if not every detail — of Z.ai's own claims. On long-horizon coding, the OrcaRouter harness measures 77.9 on Terminal-Bench 2.1, while Z.ai's best-reported GLM-5.2 figure is 82.7, which would be the first open-weight score above 80 but is vendor-reported and unreproduced. The list price is $1.40 per million input and $4.40 per million output tokens.

A two-column comparison scoreboard titled "GLM-5.3 vs GLM-5.2 — the scoreboard". Left column GLM-5.3 (rumored): Status "Leaked, not yet released", Size ">1T params (rumored)", Context "unconfirmed", Modality "text-first (rumored)", AA Index "~57-60 (projected)", License "unconfirmed". Right column GLM-5.2 (shipped): Status "Shipped June 2026", Size "753B MoE / 40B active", Context "1M tokens", Modality "text-only", AA Index "53", License "MIT". Footer reads "GLM-5.3 figures are unverified rumors; GLM-5.2 baseline per Artificial Analysis."

The scoreboard above is the leak-era projection this page published before launch — the ">1T params (rumored)" row, the unconfirmed context and license, the projected AA index. The launch corrected the biggest cell: GLM-5.3 reuses the same 743B base as GLM-5.2, so there is no parameter jump. The context window is confirmed at 1M, and the license stays unconfirmed because the open weights have not shipped yet. The projected index cell — this page's own guess of ~57–60 — was the rare projection that came in on the nose: the real number is 60, and the open question now is what happens when that score is reproduced against the actual weights.

A screenshot of the Artificial Analysis page for GLM-5.2 (max) showing an Intelligence Index of 53 (ranked #26), $1.40 per 1M input tokens and $4.40 per 1M output tokens, text input and text output, and a 1,000,000-token context window.

The capture above is the independent baseline GLM-5.3's claims are measured against. GLM-5.2 tops the open-weights leaderboard at an Artificial Analysis Intelligence Index of 53. The first test of whether GLM-5.3's post-training deltas move that number has now arrived: Artificial Analysis measures GLM-5.3 at 60 on the same index — tied with Kimi K3 for the open-weights lead, 7 points ahead of GLM-5.2, and reported by the lab as independently measured.

Tarifs et disponibilité

GLM-5.3 is priced identically to GLM-5.2: $1.40 per million input and $4.40 per million output tokens (¥8 / ¥28 in the domestic listing), with cached-input reads at $0.26 / ¥2 per million. Z.ai announced the API was open on August 19, and it is reachable through Z.ai's own API, ZCode, AutoClaw, the GLM Coding Plan, and several partner gateways. One behavior change matters for API callers: requests now require "thinking" enabled across three effort levels — low, high, and max — with no off switch, a breaking change for existing integrations.

Same price does not mean same bill. GLM-5.3 runs roughly 20% more tokens per task than GLM-5.2 did on the same workloads, which a cost-per-task reading puts at about $0.68 against GLM-5.2's $0.44 — still under Kimi K3 (about $0.84) and GPT-5.6 Sol (about $1.23). That per-task math is a derived estimate from observed token usage, not a vendor figure, but it is the number that decides whether the flat $1.40 / $4.40 rate card actually saves you money.

Ce que le lancement change pour vous

For API callers already on GLM-5.2, the practical step is a model-name change, not a project: GLM-5.2 is OpenAI-compatible and the integration carries over, with the thinking-effort caveat above. For self-hosters, the timeline is now a date rather than a guess: Zhipu promised the weights "two weeks after release" on August 14, which lands on Friday, August 28, and the open question is whether the license stays permissive. For anyone comparing models in the DeepSeek V4 Pro, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol, and Claude Fable 5 tier, GLM-5.3 is now a live, independently scored variable in that ranking instead of a rumor.

The launch-day argument around GLM-5.3 is that the coding frontier has converged: for most everyday tasks, the story goes, few users can reliably tell GPT-5.6 Sol, Claude Fable 5, Kimi K3, GLM-5.2, and Qwen3.8-Max apart. If that convergence is real, the deciding factors stop being raw capability and become price, openness, and switching cost — which is exactly the corner GLM-5.3 is staking out at $1.40 / $4.40 per million on a self-hostable 743B base with weights confirmed for August 28. Whether coding models are genuinely interchangeable is an opinion, not a benchmark; the prices, the parameter count, and the weight date are not.

On the routing side, GLM-5.3 went live on OrcaRouter on August 18, the same day Z.ai's API opened — at the first-party list price, $1.40 / $4.40 per million, passed through with zero markup. The screenshot below shows GLM-5.2's page, which is exactly the shape GLM-5.3 now has: same price, same 1M-token context, same 128K max output. Routing a slice of real traffic to GLM-5.3 with automatic failover to GLM-5.2 or another proven model is a configuration change, not a rewrite — same key, no second contract. If the new model regresses on your workload, the router falls back before a page turns, and you get a quality signal on your own traffic instead of a vendor's slide. For a model whose flagship claims are still mostly vendor-reported, that is the low-risk way to find out for yourself.

The OrcaRouter model page for z-ai/glm-5.2 showing the model id, Tools, JSON and Reasoning capability chips, a 1,000,000-token context window, a 128,000-token max output, text input and text output, $1.40 per 1M input tokens and $4.40 per 1M output tokens, and a p50 time-to-first-token of 5.95 seconds.

Que regarder ensuite

• The weights, on Friday, August 28, and the license line on the model card — permissive MIT like GLM-5.2, or something narrower. Zhipu's cyber-safety hardening is the stated reason for the two-week delay, and the "trusted access" program suggests some functions will be gated regardless.

• Whether the cyber claims hold up outside Z.ai's own harness. The 2,436-vulnerability real-world claim and the CyberGym lead are the numbers independent labs will probe first; the AA Intelligence Index measures general capability, not security.

• Where the index lands once the weights are out. The 60 is scored against the served API; the self-hosted version, with a license attached, is the one teams will actually redeploy.

• DeepSeek V4 Flash's announced price increase, which sets the pricing envelope GLM-5.3 is being judged against, and GPT-5.6 Sol's one-point lead at 61.

FAQ

Est-ce que GLM-5.3 est un modèle plus grand que GLM-5.2 ?

Non. Z.ai indique que GLM-5.3 utilise la même base Mixture-of-Experts de 743 milliards de paramètres que GLM-5.2, avec la même fenêtre de contexte de 1M de jetons et environ 40 milliards de paramètres actifs par jeton. Tous les gains rapportés proviennent d'un post-entraînement à plus grande échelle, et non d'une augmentation des paramètres — ce qui corrige directement la rumeur de l'époque des fuites selon laquelle la base dépassait un billion. Cette affirmation sur l'architecture est propre à Z.ai et n'a pas été vérifiée de manière indépendante.

Quand les poids ouverts de GLM-5.3 seront-ils disponibles ?

Friday, August 28. Zhipu promised the weights "two weeks after release" when it announced GLM-5.3 on August 14, and said "next Friday" when the API went live on August 19 — both readings land on the same date. The license has not been confirmed, and Z.ai has said sensitive cyber functions will be restricted to a verified-user "trusted access" program.

Comment dois-je traiter les chiffres de référence ?

Split the list. The coding jumps — Terminal-Bench 3.0 at 28.3, DeepSWE v1.1 at 66.9, SWE-Marathon at 42.5 — and the cyber results — CyberGym 84.5%, ExploitBench 54.4% — all come from Z.ai's own announcement and remain vendor-reported until an independent harness reproduces them. The Artificial Analysis Intelligence Index of 60 is the first independent measurement, and it is the number to weigh against everything Z.ai claims.

The leak was real, and the launch confirmed the name, the framing, and the timing — while correcting the one specific the rumor mill got loudest about. GLM-5.3 is the same base, post-trained hard, and now it carries an independent score to hold its vendor claims against: 60 on the Artificial Analysis Intelligence Index, tied with Kimi K3, seven ahead of GLM-5.2. The coding and cyber numbers are still Z.ai's own, the weights land on August 28, and the license line and a genuinely independent probe of the security claims are what's left to settle. Until then, the low-risk way to form your own view is a slice of real traffic and a failover to something proven.

Comparés dans cet article1

Détecté à partir de cet article · Benchmarks : Artificial Analysis · mis à jour quotidiennement

© 2026 OrcaRouter

Pour les fournisseurs

Vous exploitez une plateforme d'inférence ? Proposez vos modèles sur OrcaRouter.

Contactez-nous

Rejoignez notre communauté

DiscordEmailXGitHubYouTube