Una tarjeta de título principal para GLM-5.3 con una insignia de "INFORME DE FILTRACIÓN", el subtítulo "El buque insignia 'Epic Plus' de Z.ai — Lo que sabemos hasta ahora", y tres chips que dicen "Aún no lanzado", "Rastro de filtración: 3 ago" y "Sucesor de GLM-5.2".
Guides & Insights

GLM-5.3 se lanza: la filtración era real — el buque insignia post-entrenado de Z.ai para codificación y ciberdefensa

Autor

Rowan Sterling

Fecha de publicación

Últimos modelos · 20Ver todos los modelos
Benchmarks: Artificial Analysis · actualizado a diario
Volver a todas las publicaciones

The leak was real, the launch is here, and GLM-5.3 now has an independent score to argue about. Artificial Analysis' Intelligence Index — measured by the lab, not by Z.ai — puts GLM-5.3 at 60, tied with Kimi K3 for the top open-weights score on the board and 7 points clear of GLM-5.2's 53. The API went live this week at the same price as its predecessor, the open weights are confirmed for Friday, August 28, and the coding and cyber-defense claims Z.ai has been making since the August 14 announcement are starting to become testable. This page first tracked GLM-5.3 from its August 3 leak traces; this is the launch report, updated in place with what the launch, the API, and the first independent benchmark actually confirmed.

De la filtración al lanzamiento

The four traces that surfaced on August 3 — a "ZCode for GLM-5.3" harness page, an official docs page reachable for roughly an hour, a Bing index entry reading "GLM-5.3 Official Harness," and a commit adding a "glm-5.3" entry with JSON Schema support to Zhipu's official Java SDK — all pointed at a real, named release. Z.ai co-founder Tang Jie's "sooooooon" reply and the "epic-level plus" framing are now confirmed by an actual product rather than a rumor. The "roughly a week" timing signal this page tested — posted on X by @teortaxesTex after DeepSeek V4 Pro shipped on August 13 — held to the day: Z.ai formally announced GLM-5.3 on August 14 under the slogan "Built to Code. Ready for Cyber Defense." The follow-on came this week: on August 19 Z.ai said the GLM-5.3 API was live and open for calls, priced the same as GLM-5.2, with the model already wired into ZCode, AutoClaw, and the GLM Coding Plan.

Lo que la formación posterior realmente aportó

La afirmación arquitectónica es la parte que vale la pena precisar, porque el detalle más rotundo de la filtración —que los parámetros superan el billón— es incorrecto. GLM-5.3 no es un modelo más grande. Z.ai afirma que reutiliza exactamente la misma base de Mixture-of-Experts de 743B que GLM-5.2 (unos 40 mil millones de parámetros activos por token), mantiene la misma ventana de contexto de 1M de tokens y un máximo de salida de unos 128K, y consigue todas sus mejoras gracias a un post-entrenamiento ampliado: más entornos de tareas de horizonte largo, más tipos de entornos y ejecuciones de entrenamiento más largas, construido sobre el framework de contexto largo IndexShare, el RL asíncrono SAO y el framework de código abierto slime que ya produjo GLM-5.2. Ese enfoque de «sin reentrenamiento, solo post-entrenamiento» es propio de Z.ai y no ha sido auditado de forma independiente. El tamaño es el otro titular: con 743B de parámetros totales y solo unos 40B activos por token, GLM-5.3 es lo bastante ligero como para autoalojarse en un clúster modesto y lo bastante barato para servirse a gran escala — el posicionamiento de «más pequeño, más barato, abierto» que los comentarios del día del lanzamiento le atribuyeron, y un marcado contraste con los buques insignia cerrados de la frontera con los que se le compara.

Codificación y agentes: las cifras que Z.ai afirma

En cuanto a codificación, Z.ai informa —todos datos reportados por el proveedor y no reproducidos— que Terminal-Bench 3.0 sube de 4.6 a 28.3, lo que Z.ai califica como la mejor puntuación de pesos abiertos en ese banco de pruebas; DeepSWE v1.1 sube de 46.2 a 66.9; SWE-Marathon aproximadamente se duplica, pasando de 19.4 a 42.5; y Agents' Last Exam (CLI) sube de 23.8 a 28.5. En el banco de código interno de Z.ai, GLM-5.3 obtuvo un 31.4% con unos 50K tokens de salida por tarea a alto esfuerzo, frente al 29.5% de Claude Opus 4.8 con unos 120K tokens —con el argumento de Z.ai de que GLM-5.3 alcanza un resultado comparable mientras gasta muchos menos tokens de salida. Claude Fable 5 sigue liderando esa evaluación interna con un 39.5% al máximo esfuerzo, y Z.ai admite que GLM-5.3 aún va por detrás de GPT-5.6 Sol y Claude Fable 5 en varias evaluaciones de codificación más difíciles. Trata todos estos datos como cifras del proveedor hasta que un banco de pruebas independiente las reproduzca.

Ciberdefensa: la capacidad que nadie vio venir.

Las cifras de ciberseguridad son la verdadera noticia, y además provienen enteramente de los proveedores. En CyberGym, un benchmark de caja blanca para el descubrimiento y la validación de vulnerabilidades, Z.ai reporta GLM-5.3 con un 84.5%, frente al 77.2% de GLM-5.2 y por delante de Mythos 5 de Anthropic (83.8%) y de GPT-5.6 Sol (83.6%). En ExploitBench, que exige tanto análisis de causa raíz como un exploit funcional, GLM-5.3 se duplicó con creces, pasando del 24.4% al 54.4%, aunque Mythos 5 (78.0%) sigue a la cabeza. En ExploitGym, Z.ai reporta 105 tareas completadas en un plazo de 2 horas y 130 en 6 horas, frente a 29 y 39 para GLM-5.2, nuevamente por detrás de Mythos 5 (181 y 247). Z.ai enmarca la capacidad cibernética como una propiedad emergente del post-entrenamiento a escala—«la capacidad siguió acumulándose a medida que el entrenamiento escalaba», en palabras de la empresa—, más que como un objetivo deliberado.

Z.ai añade una afirmación del mundo real para acompañar a los benchmarks: en pruebas con equipos de seguridad, GLM-5.3 identificó 2.436 vulnerabilidades en 269 proyectos de código abierto, 1.097 de ellas calificadas como críticas o de alta gravedad, con el hallazgo más antiguo que data de 1981 y una "vida útil" media de 26,6 años. Esa es la cifra más llamativa del anuncio y también la menos comprobable de forma independiente. La empresa ha combinado esta capacidad con un Security Disclosure Ledger para la divulgación coordinada, un programa de "trusted access" que limita las funciones cibernéticas sensibles a usuarios verificados y una iniciativa "Open Source Shield" para auditar continuamente los proyectos de código abierto clave.

El modelo base GLM-5.3 que había que superar.

GLM-5.2 is the reference point the whole story hangs on. It shipped in June 2026 as a 743B Mixture-of-Experts model with roughly 40B active parameters per token, a 1M-token context window, a 128K max output, an MIT license, and open weights on Hugging Face. Independently, Artificial Analysis' Intelligence Index puts GLM-5.2 at 53 — the highest open-weights score on the index until this week. GLM-5.3 now clears it by 7 points: Artificial Analysis measures GLM-5.3 at 60 on the same index (v4.1.1), tying Kimi K3 for the top open-weights position and landing it in the frontier band alongside closed flagships like Claude Fable 5 and GPT-5.6 Sol. That is the first independent number attached to GLM-5.3, and it is consistent with the direction — if not every detail — of Z.ai's own claims. On long-horizon coding, the OrcaRouter harness measures 77.9 on Terminal-Bench 2.1, while Z.ai's best-reported GLM-5.2 figure is 82.7, which would be the first open-weight score above 80 but is vendor-reported and unreproduced. The list price is $1.40 per million input and $4.40 per million output tokens.

A two-column comparison scoreboard titled "GLM-5.3 vs GLM-5.2 — the scoreboard". Left column GLM-5.3 (rumored): Status "Leaked, not yet released", Size ">1T params (rumored)", Context "unconfirmed", Modality "text-first (rumored)", AA Index "~57-60 (projected)", License "unconfirmed". Right column GLM-5.2 (shipped): Status "Shipped June 2026", Size "753B MoE / 40B active", Context "1M tokens", Modality "text-only", AA Index "53", License "MIT". Footer reads "GLM-5.3 figures are unverified rumors; GLM-5.2 baseline per Artificial Analysis."

The scoreboard above is the leak-era projection this page published before launch — the ">1T params (rumored)" row, the unconfirmed context and license, the projected AA index. The launch corrected the biggest cell: GLM-5.3 reuses the same 743B base as GLM-5.2, so there is no parameter jump. The context window is confirmed at 1M, and the license stays unconfirmed because the open weights have not shipped yet. The projected index cell — this page's own guess of ~57–60 — was the rare projection that came in on the nose: the real number is 60, and the open question now is what happens when that score is reproduced against the actual weights.

A screenshot of the Artificial Analysis page for GLM-5.2 (max) showing an Intelligence Index of 53 (ranked #26), $1.40 per 1M input tokens and $4.40 per 1M output tokens, text input and text output, and a 1,000,000-token context window.

The capture above is the independent baseline GLM-5.3's claims are measured against. GLM-5.2 tops the open-weights leaderboard at an Artificial Analysis Intelligence Index of 53. The first test of whether GLM-5.3's post-training deltas move that number has now arrived: Artificial Analysis measures GLM-5.3 at 60 on the same index — tied with Kimi K3 for the open-weights lead, 7 points ahead of GLM-5.2, and reported by the lab as independently measured.

Precios y disponibilidad

GLM-5.3 is priced identically to GLM-5.2: $1.40 per million input and $4.40 per million output tokens (¥8 / ¥28 in the domestic listing), with cached-input reads at $0.26 / ¥2 per million. Z.ai announced the API was open on August 19, and it is reachable through Z.ai's own API, ZCode, AutoClaw, the GLM Coding Plan, and several partner gateways. One behavior change matters for API callers: requests now require "thinking" enabled across three effort levels — low, high, and max — with no off switch, a breaking change for existing integrations.

Same price does not mean same bill. GLM-5.3 runs roughly 20% more tokens per task than GLM-5.2 did on the same workloads, which a cost-per-task reading puts at about $0.68 against GLM-5.2's $0.44 — still under Kimi K3 (about $0.84) and GPT-5.6 Sol (about $1.23). That per-task math is a derived estimate from observed token usage, not a vendor figure, but it is the number that decides whether the flat $1.40 / $4.40 rate card actually saves you money.

Qué cambia el lanzamiento para ti

For API callers already on GLM-5.2, the practical step is a model-name change, not a project: GLM-5.2 is OpenAI-compatible and the integration carries over, with the thinking-effort caveat above. For self-hosters, the timeline is now a date rather than a guess: Zhipu promised the weights "two weeks after release" on August 14, which lands on Friday, August 28, and the open question is whether the license stays permissive. For anyone comparing models in the DeepSeek V4 Pro, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol, and Claude Fable 5 tier, GLM-5.3 is now a live, independently scored variable in that ranking instead of a rumor.

The launch-day argument around GLM-5.3 is that the coding frontier has converged: for most everyday tasks, the story goes, few users can reliably tell GPT-5.6 Sol, Claude Fable 5, Kimi K3, GLM-5.2, and Qwen3.8-Max apart. If that convergence is real, the deciding factors stop being raw capability and become price, openness, and switching cost — which is exactly the corner GLM-5.3 is staking out at $1.40 / $4.40 per million on a self-hostable 743B base with weights confirmed for August 28. Whether coding models are genuinely interchangeable is an opinion, not a benchmark; the prices, the parameter count, and the weight date are not.

On the routing side, GLM-5.3 went live on OrcaRouter on August 18, the same day Z.ai's API opened — at the first-party list price, $1.40 / $4.40 per million, passed through with zero markup. The screenshot below shows GLM-5.2's page, which is exactly the shape GLM-5.3 now has: same price, same 1M-token context, same 128K max output. Routing a slice of real traffic to GLM-5.3 with automatic failover to GLM-5.2 or another proven model is a configuration change, not a rewrite — same key, no second contract. If the new model regresses on your workload, the router falls back before a page turns, and you get a quality signal on your own traffic instead of a vendor's slide. For a model whose flagship claims are still mostly vendor-reported, that is the low-risk way to find out for yourself.

The OrcaRouter model page for z-ai/glm-5.2 showing the model id, Tools, JSON and Reasoning capability chips, a 1,000,000-token context window, a 128,000-token max output, text input and text output, $1.40 per 1M input tokens and $4.40 per 1M output tokens, and a p50 time-to-first-token of 5.95 seconds.

Qué ver a continuación

• The weights, on Friday, August 28, and the license line on the model card — permissive MIT like GLM-5.2, or something narrower. Zhipu's cyber-safety hardening is the stated reason for the two-week delay, and the "trusted access" program suggests some functions will be gated regardless.

• Whether the cyber claims hold up outside Z.ai's own harness. The 2,436-vulnerability real-world claim and the CyberGym lead are the numbers independent labs will probe first; the AA Intelligence Index measures general capability, not security.

• Where the index lands once the weights are out. The 60 is scored against the served API; the self-hosted version, with a license attached, is the one teams will actually redeploy.

• DeepSeek V4 Flash's announced price increase, which sets the pricing envelope GLM-5.3 is being judged against, and GPT-5.6 Sol's one-point lead at 61.

Preguntas frecuentes

¿Es GLM-5.3 un modelo más grande que GLM-5.2?

No. Z.ai afirma que GLM-5.3 usa la misma base Mixture-of-Experts de 743 mil millones de parámetros que GLM-5.2, con la misma ventana de contexto de 1M tokens y aproximadamente 40 mil millones de parámetros activos por token. Todas las mejoras reportadas provienen de un post-entrenamiento a mayor escala, no de un aumento de parámetros, lo que corrige directamente el rumor de la era de las filtraciones sobre una base de más de un billón. Esa afirmación sobre la arquitectura es propia de Z.ai y no ha sido auditada de forma independiente.

¿Cuándo estarán disponibles los pesos abiertos de GLM-5.3?

Friday, August 28. Zhipu promised the weights "two weeks after release" when it announced GLM-5.3 on August 14, and said "next Friday" when the API went live on August 19 — both readings land on the same date. The license has not been confirmed, and Z.ai has said sensitive cyber functions will be restricted to a verified-user "trusted access" program.

¿Cómo debo tratar los números de referencia?

Split the list. The coding jumps — Terminal-Bench 3.0 at 28.3, DeepSWE v1.1 at 66.9, SWE-Marathon at 42.5 — and the cyber results — CyberGym 84.5%, ExploitBench 54.4% — all come from Z.ai's own announcement and remain vendor-reported until an independent harness reproduces them. The Artificial Analysis Intelligence Index of 60 is the first independent measurement, and it is the number to weigh against everything Z.ai claims.

The leak was real, and the launch confirmed the name, the framing, and the timing — while correcting the one specific the rumor mill got loudest about. GLM-5.3 is the same base, post-trained hard, and now it carries an independent score to hold its vendor claims against: 60 on the Artificial Analysis Intelligence Index, tied with Kimi K3, seven ahead of GLM-5.2. The coding and cyber numbers are still Z.ai's own, the weights land on August 28, and the license line and a genuinely independent probe of the security claims are what's left to settle. Until then, the low-risk way to form your own view is a slice of real traffic and a failover to something proven.

Comparados en este artículo1

Detectado en este artículo · Benchmarks: Artificial Analysis · actualizado a diario

© 2026 OrcaRouter

Para proveedores

¿Operas una plataforma de inferencia? Publica tus modelos en OrcaRouter.

Contáctanos

Únete a la comunidad

DiscordEmailXGitHubYouTube