
GPT-6.1 Sol 對決 GLM-5.3:權重已發布,帳單卻紋風不動
- openai新OpenAI: GPT-6.1 Sol2026-09-2952智能
- anthropic新Anthropic: Claude Sonnet 5.52026-09-2856智能
- typesafe新TypeSafe: Jev 1.132026-09-24$0.04 / $0.00 每百萬 tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238智能
- OpenAIOpenAI: GPT-6 Sol2026-09-2248智能
- AnthropicAnthropic: Claude Opus 5.52026-09-2258智能
- xAIGrok 4.72026-09-2146智能
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 每百萬 tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 每百萬 tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- OpenAIOpenAI: GPT-6 Astra2026-09-0453智能77程式
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241智能76程式
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245智能76程式
- AnthropicAnthropic: Claude Fable 5.12026-09-0153智能82程式
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 每百萬 tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百萬 tokens · 301 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72程式
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 每百萬 tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845智能75程式
- obsidianQwen3.8 27B2026-08-1534智能68程式
OpenAI released GPT-6.1 Sol on 29 September 2026 at its DevDay keynote, and Z.ai released GLM-5.3 six weeks earlier, on 18 August 2026, then held the checkpoint back for a week. That hold is over: zai-org/GLM-5.3 has been on Hugging Face since 25 August, a BF16/FP8 checkpoint at roughly 753 billion parameters of which about 40 billion are active per token, and it has since passed 1.4 million downloads. So the question most comparisons have been asking since August — is this an open model or a rental? — has an answer, and the answer did not change the arithmetic. On the independent board, GPT-6.1 Sol scores 51.8 and costs $0.72 per completed index task; GLM-5.3 scores 44.8 and costs $2.01. You can now own the cheaper-by-the-token model and still pay nearly three times more per finished job.
每個模型實際上是什麼
GPT-6.1 Sol is OpenAI's mid-tier GPT-6 model, one rung below the GPT-6 Astra flagship and above the budget GPT-6 Luna. It carries a 1,050,000-token context window with a 128,000-token output ceiling and an April 30, 2026 knowledge cutoff, and it accepts text, images and files as input. Reasoning runs from low to max with a default of medium; the none and minimal values that GPT-6 Sol accepted now return an error, which OpenAI's own migration guidance addresses by telling developers to substitute low. It costs $2.00 per million input tokens and $10.00 per million output, with cached input at $0.10.
GLM-5.3 是 Z.ai 目前針對軟體工程與長時程代理式工作的旗艦模型,建構於與 GLM-5.2 相同的混合專家基礎之上,其效能提升來自後訓練,而非更大的模型。它是文字輸入、文字輸出——無論價格為何,都不具備視覺能力——擁有 1M token 的上下文視窗、128,000 token 的輸出上限,且思考模式永久開啟。它沒有非推理模式,這對延遲敏感的呼叫而言是一項硬性限制。Z.ai 標價為每百萬 token 輸入 $1.40、輸出 $4.40,快取輸入為 $0.26;我們自家的目錄則列為 $1.26 與 $3.96,快取讀取為 $0.234,因此供應商的價目表是較保守的數字,也是值得據以爭論的依據。
價目表,以及反轉它的那一行
每個維度一行,每一行兩側都要顯示:
• Input — GPT-6.1 Sol $2.00 per 1M vs GLM-5.3 $1.40 per 1M; GLM is 30% cheaper
• Output — GPT-6.1 Sol $10.00 per 1M vs GLM-5.3 $4.40 per 1M; GLM is 56% cheaper
• Cached input — GPT-6.1 Sol $0.10 per 1M vs GLM-5.3 $0.26 per 1M; the ordering flips, Sol is 2.6× cheaper
• Long-context clause — GPT-6.1 Sol reprices the whole request above 272,000 input tokens at 2× input and cache and 1.5× output vs GLM-5.3 no equivalent clause on the published card
• Context and output — 1,050,000 in / 128,000 out vs 1M in / 128,000 out; a 5% difference, immaterial
• Modality — text, image and file in, text out vs text only
• Reasoning floor — GPT-6.1 Sol low, medium (default), high, xhigh, max vs GLM-5.3 always on, no floor to drop to
• Weights — closed, API only vs open, downloadable, FP8
• Independent score, Artificial Analysis v4.3.2 at max effort — 51.8 vs 44.8
• Independent cost per index task — $0.72 vs $2.01
• Output tokens for the index run — 67 million vs 210 million
img src="2.png" alt="一張生成的雙欄比較計分板,標題為「GPT-6.1 Sol vs GLM-5.3 - the scoreboard」。左欄「GPT-6.1 Sol」:各列寫著「Intelligence Index:51.8」、「每項指數任務成本:$0.72」、「輸入價格:每 1M $2.00」、「輸出價格:每 1M $10.00」、「指數測試執行的輸出 token 數:67M」、「權重:閉源」。右欄「GLM-5.3」:各列寫著「Intelligence Index:44.8」、「每項指數任務成本:$2.01」、「輸入價格:每 1M $1.40」、「輸出價格:每 1M $4.40」、「指數測試執行的輸出 token 數:210M」、「權重:於 Hugging Face 開源」。頁尾寫著「Independent figures per Artificial Analysis v4.3.2 at max effort; vendor list prices.」。OrcaRouter 標誌位於右下角。" /> 最後這一組數據,正是這場對決分出勝負的關鍵所在,而且這絕不是什麼細微的差距。

2.1 億個 token 的問題
Artificial Analysis 對排行榜上的每個模型都執行同一套十項評估,並公布每個分數背後的 token 數量。GLM-5.3 在完成該次測試時產生了大約 2.1 億個輸出 token。GPT-6.1 Sol 則產生了 6,700 萬個。對照排行榜上各模型所屬的同級比較類別,GLM-5.3 的 2.1 億高於該類別約 1.4 億的中位數,而 Sol 的 6,700 萬則遠低於其所屬旗艦類別約 8,100 萬的中位數。由於輸出是每張費率表中昂貴的一側,GLM-5.3 在輸出項目上主打的 56% 折扣經不起實際測量的檢驗:它以 0.44 倍的價格產生 3.1 倍的 token,而 3.1 × 0.44 等於 1.37 —— 儘管費率表實質上便宜得多,在這項工作負載上 GLM-5.3 才是較昂貴的模型。
該榜單自身的每項任務成本數字,印證了方向與大致規模。$2.01 對上 $0.72 是 2.8 倍,比單憑 token 比率所暗示的差距更大,因為 GLM-5.3 在每項任務的輸入與快取計費項目上也付出更多。值得精確說明這證明了什麼、又沒有證明什麼:該指數測試跑的是十項固定評估,而在這些評估上的冗長程度,並不等同於在你自家流量上的冗長程度。但對買方而言,會被計費的就是 token,而在任何要求模型產出長答案的任務集上,這種差異的形態都相同。
部分抵銷來自快取輸入這一項。GLM-5.3 的快取讀取是 $0.26,對比 $1.40 的基準;而 GPT-6.1 Sol 的是 $0.10,對比 $2.00 的基準——在一個對任何具有穩定前綴的工作負載都具支配地位的費率上,Sol 以 2.6 倍勝出。如果你的流量是大量固定語料,而答案只有一段,GLM-5.3 顯然是更便宜的選擇,你應該選它。如果你的流量看起來像這兩款模型當初鎖定打造的流量——代理迴圈、儲存庫規模的編輯、長時間生成的輸出——那麼相較於 3 倍的 token 比率,快取輸入的優勢會縮小到如同雜訊。
廠商編號最終落在哪裡
Z.ai 的發布數據很強勁,而且一如既往地未經重現。Terminal-Bench 2.1 為 88.2,DeepSWE v1.1 為 66.9,CyberGym 為 84.5,ExploitBench 為 54.4,AutomationBench 為 48.2,GDPval-AA v2 為 1769 Elo。唯一值得多看兩次的數字是 Terminal-Bench 3.0 的 28.3——這是後繼基準,也是對舊版基準 88.2 的修正。一個模型可以在其後訓練所鎖定的基準上表現卓越,卻在同一基準的下一代上表現平庸。
OpenAI's launch numbers for GPT-6.1 Sol are framed entirely around cost per completed task rather than raw score: DeepSWE v1.1 beating GPT-6 Sol's best by 6.4 percentage points at a lower reasoning effort, AutomationBench 1.0.6 at 2.2 points above Claude Opus 5.5 at medium effort, OSWorld 2.0 up seven points on GPT-6 Sol at maximum effort, Terminal-Bench Science 0.1 more than doubling GPT-6 Sol at less than half the cost per task. Note that these are OpenAI's harness at OpenAI's settings on OpenAI's selected benchmark set, and that none of them has been reproduced independently.
The overlap between the two vendors' published sets is thin, and the one direct comparison available is on the same benchmark: Z.ai reports 66.9 on DeepSWE v1.1, OpenAI reports 68.8 for GPT-6 Sol — the model 6.1 replaces — and a 6.4-point improvement for 6.1 Sol over its own predecessor. Two points apart on the vendor-reported comparison, and roughly six apart on OpenAI's own delta. That is close to a controlled comparison, and it says what the independent board says: near parity on capability, a wide gap on what finishing the job costs.
一個 API,還是兩份合約
Both models are on OrcaRouter, which is the unusual part of this pairing. GPT-6.1 Sol landed in the catalogue at OpenAI's own list price on release day — $2.00 and $10.00 per million tokens, cached input $0.10, with the 272,000-token step to $4.00 and $15.00 passed through exactly as the vendor lists it — and GLM-5.3 sits alongside it at $1.26 and $3.96 with a $0.234 cache read. Nothing is added on top of either: the platform passes provider list price through unchanged, which is why a vendor price cut shows up on our side the same day rather than after a reseller renegotiates.

That matters more for this specific pair than for most. The decision between them is not "which is better" — it is a threshold on your own output length, and it will move the first time Z.ai sharpens its rate card or OpenAI changes the 6.1 tier's tier boundary. Holding both behind one key, with automatic failover between them and a routing rule you change in configuration rather than in a deployment, is what lets you test that threshold on real traffic instead of arguing about it. If Sol's cache line wins your workload, route on it; if your answer lengths stay short and GLM-5.3's flat card with no context cliff wins the segment, route on that instead — the same key, no second contract, no second bill.

如果今天非得選一個,你會選哪個?
• 長時間生成的輸出、代理式迴圈、輸出密集型工作——GPT-6.1 Sol,而且一旦把 token 數量計入,差距接近三倍。這正是價目表會騙人、而實測不會的地方。
• 固定前綴、簡短回答——GLM-5.3。它的快取輸入與輸入費率確實更好,而它的冗長程度根本沒機會造成傷害。
• 任何輸入超過 272,000 個 token 的請求——GLM-5.3,因為 GPT-6.1 Sol 會在那個門檻上把整筆請求重新計價,而 GLM-5.3 的價目表沒有對應的級距。
• 任何以圖像或文件作為視覺輸入的內容——GPT-6.1 Sol。GLM-5.3 僅支援文字,而這差距毫不接近。
• 在自己的邊界內進行推論,或為供應商條款可能的變更避險——GLM-5.3,而且現在下載已經存在,不再只是承諾。要為硬體編列預算:該檢查點是個大型 FP8 產物,自行託管是資本決策,而非 API 決策。
• 對延遲敏感的呼叫——兩者都不能不加留意。GLM-5.3 無法關閉推理;GPT-6.1 Sol 的下限就在 低,而它在獨立排行榜上的首個 token 時間實測為 332 秒,對照排行榜中位數 3.7。
本文中的比較3
根據本文內容識別 · 基準測試:Artificial Analysis · 每日更新
