Thẻ tiêu đề chính cho GLM-5.3 với huy hiệu "BÁO CÁO RÒ RỈ", phụ đề "Flagship 'Epic Plus' của Z.ai — Những gì chúng ta biết cho đến nay", và ba chip ghi "Chưa phát hành", "Dấu vết rò rỉ: Aug 3", và "Kế nhiệm của GLM-5.2".
Guides & Insights

GLM-5.3 Ra Mắt: Thông Tin Rò Rỉ Là Thật — Sản Phẩm Chủ Lực Được Hậu Huấn Luyện Về Lập Trình & Phòng Thủ Mạng Của Z.ai

Tác giả

Rowan Sterling

Ngày đăng

Mô hình mới nhất · 20Xem tất cả mô hình
Benchmark: Artificial Analysis · cập nhật hằng ngày
Quay lại tất cả bài viết

The leak was real, the launch is here, and GLM-5.3 now has an independent score to argue about. Artificial Analysis' Intelligence Index — measured by the lab, not by Z.ai — puts GLM-5.3 at 60, tied with Kimi K3 for the top open-weights score on the board and 7 points clear of GLM-5.2's 53. The API went live this week at the same price as its predecessor, the open weights are confirmed for Friday, August 28, and the coding and cyber-defense claims Z.ai has been making since the August 14 announcement are starting to become testable. This page first tracked GLM-5.3 from its August 3 leak traces; this is the launch report, updated in place with what the launch, the API, and the first independent benchmark actually confirmed.

Từ rò rỉ đến ra mắt

The four traces that surfaced on August 3 — a "ZCode for GLM-5.3" harness page, an official docs page reachable for roughly an hour, a Bing index entry reading "GLM-5.3 Official Harness," and a commit adding a "glm-5.3" entry with JSON Schema support to Zhipu's official Java SDK — all pointed at a real, named release. Z.ai co-founder Tang Jie's "sooooooon" reply and the "epic-level plus" framing are now confirmed by an actual product rather than a rumor. The "roughly a week" timing signal this page tested — posted on X by @teortaxesTex after DeepSeek V4 Pro shipped on August 13 — held to the day: Z.ai formally announced GLM-5.3 on August 14 under the slogan "Built to Code. Ready for Cyber Defense." The follow-on came this week: on August 19 Z.ai said the GLM-5.3 API was live and open for calls, priced the same as GLM-5.2, with the model already wired into ZCode, AutoClaw, and the GLM Coding Plan.

Post-training thực sự mang lại điều gì

Phần đáng chốt lại chính là tuyên bố về kiến trúc, vì chi tiết gây ồn ào nhất của tin rò rỉ — tham số vượt mốc một nghìn tỷ — là sai. GLM-5.3 không phải là mô hình lớn hơn. Z.ai cho biết họ tái sử dụng đúng cùng nền tảng Mixture-of-Experts 743B như GLM-5.2 (khoảng 40 tỷ tham số hoạt động cho mỗi token), giữ nguyên cửa sổ ngữ cảnh 1M token và đầu ra tối đa khoảng 128K, đồng thời đạt được toàn bộ cải thiện nhờ mở rộng quy mô hậu huấn luyện: nhiều môi trường tác vụ tầm xa hơn, nhiều loại môi trường hơn và các lượt huấn luyện dài hơn, được xây dựng trên nền khung ngữ cảnh dài IndexShare, khung RL không đồng bộ SAO và khung mã nguồn mở slime vốn đã tạo ra GLM-5.2. Cách đóng khung "không huấn luyện lại, chỉ hậu huấn luyện" đó là của riêng Z.ai và chưa được kiểm toán độc lập. Quy mô là điểm nhấn còn lại: với tổng cộng 743B tham số nhưng chỉ khoảng 40B hoạt động cho mỗi token, GLM-5.3 đủ nhẹ để tự lưu trữ trên một cụm máy khiêm tốn và đủ rẻ để phục vụ ở quy mô lớn — đúng với định vị "nhỏ hơn, rẻ hơn, mở" mà các bài bình luận ngày ra mắt gán cho nó, và trái ngược hẳn với những mô hình tiên phong đóng kín đang được đem ra so sánh.

Lập trình và agent: những con số mà Z.ai đang tuyên bố

{{1}}Về mảng mã hóa, Z.ai báo cáo — tất cả đều do nhà cung cấp công bố và chưa được tái lập — {{2}}Terminal-Bench 3.0 tăng từ 4.6 lên 28.3{{/2}}, {{3}}mà Z.ai gọi là điểm số open-weights cao nhất trên bộ đánh giá đó{{/3}}; {{4}}DeepSWE v1.1 tăng từ 46.2 lên 66.9{{/4}}; {{5}}SWE-Marathon tăng xấp xỉ gấp đôi từ 19.4 lên 42.5{{/5}}; {{6}}và Agents' Last Exam (CLI) tăng từ 23.8 lên 28.5{{/6}}. {{7}}Trên bộ đánh giá mã nội bộ của Z.ai, GLM-5.3 đạt 31.4% với khoảng 50K token đầu ra mỗi tác vụ ở mức nỗ lực cao, trong khi Claude Opus 4.8 đạt 29.5% với khoảng 120K token{{/7}} — {{8}}quan điểm của Z.ai là GLM-5.3 đạt kết quả tương đương trong khi tiêu tốn ít token đầu ra hơn nhiều{{/8}}. {{9}}Claude Fable 5 vẫn dẫn đầu bộ đánh giá nội bộ đó với 39.5% ở mức nỗ lực tối đa{{/9}}, {{10}}và Z.ai thừa nhận GLM-5.3 vẫn xếp sau GPT-5.6 Sol và Claude Fable 5 trong một số bài đánh giá mã hóa khó hơn{{/10}}. {{11}}Hãy coi tất cả những con số này là số liệu của nhà cung cấp cho đến khi một bộ đánh giá độc lập tái lập được chúng.{{/11}}

Phòng thủ mạng: năng lực không ai ngờ tới

Các con số về an ninh mạng mới là tin tức thực sự, và chúng cũng hoàn toàn do nhà cung cấp tự báo cáo. Trên CyberGym, một chuẩn đánh giá phát hiện và xác thực lỗ hổng white-box, Z.ai báo cáo GLM-5.3 đạt 84,5%, tăng từ 77,2% của GLM-5.2 và vượt qua Mythos 5 (83,8%) của Anthropic cùng GPT-5.6 Sol (83,6%). Trên ExploitBench, chuẩn đánh giá đòi hỏi cả phân tích nguyên nhân gốc lẫn một exploit hoạt động được, GLM-5.3 đã tăng hơn gấp đôi từ 24,4% lên 54,4%, dù Mythos 5 (78,0%) vẫn dẫn trước. Trên ExploitGym, Z.ai báo cáo 105 tác vụ hoàn thành trong thời lượng 2 giờ và 130 tác vụ trong 6 giờ, so với 29 và 39 của GLM-5.2 — vẫn xếp sau Mythos 5 (181 và 247). Z.ai xem năng lực an ninh mạng này là một thuộc tính phát sinh từ quá trình hậu huấn luyện được mở rộng quy mô — "năng lực cứ tiếp tục tích lũy khi quy mô huấn luyện tăng lên," theo lời công ty — chứ không phải là một mục tiêu có chủ đích.

Z.ai đã bổ sung một tuyên bố thực tế để đi kèm với các điểm chuẩn: trong quá trình thử nghiệm với các nhóm bảo mật, GLM-5.3 đã xác định 2.436 lỗ hổng trên 269 dự án mã nguồn mở, trong đó 1.097 lỗ hổng được xếp hạng mức độ nghiêm trọng cao hoặc tới hạn, với phát hiện cũ nhất có từ năm 1981 và "vòng đời" trung bình là 26,6 năm. Đó là con số nổi bật nhất trong thông báo và cũng là con số khó kiểm chứng độc lập nhất. Công ty đã đi kèm khả năng này với Security Disclosure Ledger để tiết lộ có phối hợp, một chương trình "truy cập đáng tin cậy" giới hạn các chức năng an ninh mạng nhạy cảm đối với người dùng đã xác minh, và sáng kiến "Open Source Shield" để liên tục kiểm toán các dự án mã nguồn mở quan trọng.

Đường cơ sở GLM-5.3 cần phải vượt qua

GLM-5.2 is the reference point the whole story hangs on. It shipped in June 2026 as a 743B Mixture-of-Experts model with roughly 40B active parameters per token, a 1M-token context window, a 128K max output, an MIT license, and open weights on Hugging Face. Independently, Artificial Analysis' Intelligence Index puts GLM-5.2 at 53 — the highest open-weights score on the index until this week. GLM-5.3 now clears it by 7 points: Artificial Analysis measures GLM-5.3 at 60 on the same index (v4.1.1), tying Kimi K3 for the top open-weights position and landing it in the frontier band alongside closed flagships like Claude Fable 5 and GPT-5.6 Sol. That is the first independent number attached to GLM-5.3, and it is consistent with the direction — if not every detail — of Z.ai's own claims. On long-horizon coding, the OrcaRouter harness measures 77.9 on Terminal-Bench 2.1, while Z.ai's best-reported GLM-5.2 figure is 82.7, which would be the first open-weight score above 80 but is vendor-reported and unreproduced. The list price is $1.40 per million input and $4.40 per million output tokens.

A two-column comparison scoreboard titled "GLM-5.3 vs GLM-5.2 — the scoreboard". Left column GLM-5.3 (rumored): Status "Leaked, not yet released", Size ">1T params (rumored)", Context "unconfirmed", Modality "text-first (rumored)", AA Index "~57-60 (projected)", License "unconfirmed". Right column GLM-5.2 (shipped): Status "Shipped June 2026", Size "753B MoE / 40B active", Context "1M tokens", Modality "text-only", AA Index "53", License "MIT". Footer reads "GLM-5.3 figures are unverified rumors; GLM-5.2 baseline per Artificial Analysis."

The scoreboard above is the leak-era projection this page published before launch — the ">1T params (rumored)" row, the unconfirmed context and license, the projected AA index. The launch corrected the biggest cell: GLM-5.3 reuses the same 743B base as GLM-5.2, so there is no parameter jump. The context window is confirmed at 1M, and the license stays unconfirmed because the open weights have not shipped yet. The projected index cell — this page's own guess of ~57–60 — was the rare projection that came in on the nose: the real number is 60, and the open question now is what happens when that score is reproduced against the actual weights.

A screenshot of the Artificial Analysis page for GLM-5.2 (max) showing an Intelligence Index of 53 (ranked #26), $1.40 per 1M input tokens and $4.40 per 1M output tokens, text input and text output, and a 1,000,000-token context window.

The capture above is the independent baseline GLM-5.3's claims are measured against. GLM-5.2 tops the open-weights leaderboard at an Artificial Analysis Intelligence Index of 53. The first test of whether GLM-5.3's post-training deltas move that number has now arrived: Artificial Analysis measures GLM-5.3 at 60 on the same index — tied with Kimi K3 for the open-weights lead, 7 points ahead of GLM-5.2, and reported by the lab as independently measured.

Giá cả và tình trạng sẵn có

GLM-5.3 is priced identically to GLM-5.2: $1.40 per million input and $4.40 per million output tokens (¥8 / ¥28 in the domestic listing), with cached-input reads at $0.26 / ¥2 per million. Z.ai announced the API was open on August 19, and it is reachable through Z.ai's own API, ZCode, AutoClaw, the GLM Coding Plan, and several partner gateways. One behavior change matters for API callers: requests now require "thinking" enabled across three effort levels — low, high, and max — with no off switch, a breaking change for existing integrations.

Same price does not mean same bill. GLM-5.3 runs roughly 20% more tokens per task than GLM-5.2 did on the same workloads, which a cost-per-task reading puts at about $0.68 against GLM-5.2's $0.44 — still under Kimi K3 (about $0.84) and GPT-5.6 Sol (about $1.23). That per-task math is a derived estimate from observed token usage, not a vendor figure, but it is the number that decides whether the flat $1.40 / $4.40 rate card actually saves you money.

Việc ra mắt thay đổi điều gì cho bạn

For API callers already on GLM-5.2, the practical step is a model-name change, not a project: GLM-5.2 is OpenAI-compatible and the integration carries over, with the thinking-effort caveat above. For self-hosters, the timeline is now a date rather than a guess: Zhipu promised the weights "two weeks after release" on August 14, which lands on Friday, August 28, and the open question is whether the license stays permissive. For anyone comparing models in the DeepSeek V4 Pro, Qwen3.8-Max, Kimi K3, GPT-5.6 Sol, and Claude Fable 5 tier, GLM-5.3 is now a live, independently scored variable in that ranking instead of a rumor.

The launch-day argument around GLM-5.3 is that the coding frontier has converged: for most everyday tasks, the story goes, few users can reliably tell GPT-5.6 Sol, Claude Fable 5, Kimi K3, GLM-5.2, and Qwen3.8-Max apart. If that convergence is real, the deciding factors stop being raw capability and become price, openness, and switching cost — which is exactly the corner GLM-5.3 is staking out at $1.40 / $4.40 per million on a self-hostable 743B base with weights confirmed for August 28. Whether coding models are genuinely interchangeable is an opinion, not a benchmark; the prices, the parameter count, and the weight date are not.

On the routing side, GLM-5.3 went live on OrcaRouter on August 18, the same day Z.ai's API opened — at the first-party list price, $1.40 / $4.40 per million, passed through with zero markup. The screenshot below shows GLM-5.2's page, which is exactly the shape GLM-5.3 now has: same price, same 1M-token context, same 128K max output. Routing a slice of real traffic to GLM-5.3 with automatic failover to GLM-5.2 or another proven model is a configuration change, not a rewrite — same key, no second contract. If the new model regresses on your workload, the router falls back before a page turns, and you get a quality signal on your own traffic instead of a vendor's slide. For a model whose flagship claims are still mostly vendor-reported, that is the low-risk way to find out for yourself.

The OrcaRouter model page for z-ai/glm-5.2 showing the model id, Tools, JSON and Reasoning capability chips, a 1,000,000-token context window, a 128,000-token max output, text input and text output, $1.40 per 1M input tokens and $4.40 per 1M output tokens, and a p50 time-to-first-token of 5.95 seconds.

Xem gì tiếp theo

• The weights, on Friday, August 28, and the license line on the model card — permissive MIT like GLM-5.2, or something narrower. Zhipu's cyber-safety hardening is the stated reason for the two-week delay, and the "trusted access" program suggests some functions will be gated regardless.

• Whether the cyber claims hold up outside Z.ai's own harness. The 2,436-vulnerability real-world claim and the CyberGym lead are the numbers independent labs will probe first; the AA Intelligence Index measures general capability, not security.

• Where the index lands once the weights are out. The 60 is scored against the served API; the self-hosted version, with a license attached, is the one teams will actually redeploy.

• DeepSeek V4 Flash's announced price increase, which sets the pricing envelope GLM-5.3 is being judged against, and GPT-5.6 Sol's one-point lead at 61.

Câu hỏi thường gặp

GLM-5.3 có phải là mô hình lớn hơn GLM-5.2 không?

Không. Z.ai cho biết GLM-5.3 sử dụng cùng mô hình cơ sở Mixture-of-Experts 743 tỷ tham số như GLM-5.2, với cùng cửa sổ ngữ cảnh 1 triệu token và khoảng 40 tỷ tham số hoạt động trên mỗi token. Tất cả mức cải thiện được báo cáo đều đến từ việc mở rộng quy mô hậu huấn luyện, chứ không phải từ việc tăng tham số — điều này trực tiếp đính chính tin đồn từ thời kỳ rò rỉ về một mô hình cơ sở trên một nghìn tỷ tham số. Tuyên bố về kiến trúc đó là của riêng Z.ai và chưa được thẩm định độc lập.

Khi nào trọng số mở của GLM-5.3 sẽ có sẵn?

Friday, August 28. Zhipu promised the weights "two weeks after release" when it announced GLM-5.3 on August 14, and said "next Friday" when the API went live on August 19 — both readings land on the same date. The license has not been confirmed, and Z.ai has said sensitive cyber functions will be restricted to a verified-user "trusted access" program.

Tôi nên xử lý các con số benchmark như thế nào?

Split the list. The coding jumps — Terminal-Bench 3.0 at 28.3, DeepSWE v1.1 at 66.9, SWE-Marathon at 42.5 — and the cyber results — CyberGym 84.5%, ExploitBench 54.4% — all come from Z.ai's own announcement and remain vendor-reported until an independent harness reproduces them. The Artificial Analysis Intelligence Index of 60 is the first independent measurement, and it is the number to weigh against everything Z.ai claims.

The leak was real, and the launch confirmed the name, the framing, and the timing — while correcting the one specific the rumor mill got loudest about. GLM-5.3 is the same base, post-trained hard, and now it carries an independent score to hold its vendor claims against: 60 on the Artificial Analysis Intelligence Index, tied with Kimi K3, seven ahead of GLM-5.2. The coding and cyber numbers are still Z.ai's own, the weights land on August 28, and the license line and a genuinely independent probe of the security claims are what's left to settle. Until then, the low-risk way to form your own view is a slice of real traffic and a failover to something proven.

So sánh trong bài viết này1

Phát hiện từ bài viết này · Benchmark: Artificial Analysis · cập nhật hằng ngày

© 2026 OrcaRouter

Dành cho nhà cung cấp

Bạn vận hành nền tảng suy luận? Đưa mô hình của bạn lên OrcaRouter.

Liên hệ với chúng tôi

Tham gia cộng đồng

DiscordEmailXGitHubYouTube