
DeepSeek V4 Pro 對決 Qwen3.8 Max:三倍價差,九月暫緩後重新評測
- deepseek新DeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- openai新OpenAI: GPT-6 Astra2026-09-0453智能77程式
- google新Google: Gemini 3.8 Flash2026-09-0241智能76程式
- qwen新Qwen: Qwen3.8 Max (0902)2026-09-0240智能72程式
- anthropic新Anthropic: Claude Fable 5.12026-09-0153智能82程式
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百萬 tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72程式
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 每百萬 tokens
- z-aiZ.ai: GLM 5.32026-08-1845智能75程式
- obsidianQwen3.8 27B2026-08-1534智能68程式
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236智能69程式
- grokSpaceXAI: Grok 4.62026-08-1244智能77程式
- metaMeta: Muse Spark 1.22026-08-0540智能72程式
- qwenQwen: Qwen3.8 Max2026-08-0340智能72程式
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135智能69程式
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 每百萬 tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451智能78程式
- googleGoogle: Gemini 3.6 Flash2026-07-2134智能69程式
DeepSeek V4 Pro and Qwen3.8 Max are the two most important open-or-cheap flagships of this generation, and the September events on the DeepSeek side changed the matchup more than either vendor's marketing has acknowledged. DeepSeek V4 Pro shipped on August 13, 2026, ten days after Alibaba's Qwen3.8 Max went generally available on August 3, and the pairing has always read as a one-sided contest: Qwen3.8 Max is the 2.4-trillion-parameter all-rounder at $2 per million input and $6 per million output tokens, while DeepSeek V4 Pro is the 1.6-trillion-parameter reasoning specialist at a fraction of that price, MIT-licensed, with public weights. The reason this comparison is worth re-running today is not the launch-week numbers but what happened a month later — DeepSeek spent the week of September 8 trying to retire V4 Pro, then reversed on September 11 and committed to keeping it online at unchanged billing. That reversal is the single most important fact about this matchup right now, because it decides whether the cheap half is a safe bet or a migration trap.
Independent measurements in this piece come from Artificial Analysis, captured today. Alibaba's and DeepSeek's own benchmark tables are labeled vendor-reported throughout — neither company's launch numbers have been fully reproduced by a third party, and DeepSeek's are now four weeks old against a moving scoreboard.
九月的暫緩才是重點,而不是八月的發布。
9月8日,DeepSeek 宣布,9月14日之後,所有對 deepseek-v4-pro 的請求都會被路由到 V4.1 Flash,並按 Flash 的價格計費,直到 V4.1 Pro 推出為止。任何認真比較 Qwen3.8 Max 與 DeepSeek V4 Pro 的人,都不得不開始設想這個便宜選項會消失。後來 DeepSeek 延後了截止期限,並在9月11日完全逆轉,表示9月14日之後將繼續為 DeepSeek V4 Pro 提供 API 服務,計費不變,若日後有任何改變會另行通知。這項逆轉當天由中國官媒與科技媒體報導。
這對這場對決的影響,是移除了那道一直在暗中左右決策的不對稱。在九月的第一週期間,「我該以 $0.66 輸出模型還是 $6 輸出模型為基礎來打造?」這個問題的誠實答案,一直受到生存風險的污染:價格領導者有個關閉日期。截至九月十五日,這種污染已經消失。便宜的選項具備持久性,而這項比較終於關乎它一直以來本應關乎的事情——智慧、價格與速度。

價格:同樣的遊戲,只需一小部分的輸出成本
在目前的獨立排行榜上,兩者相差四分,而它們之間的價格差距是開放權重層級中最大的。Qwen3.8 Max 在 Alibaba 的定價表中,每百萬輸入 token 為 $2.00,每百萬輸出 token 為 $6.00,並享有 88% 的快取折扣。DeepSeek V4 Pro 在離峰時段每百萬輸入為 $0.66、每百萬輸出為 $1.98,尖峰時段則加倍至 $1.32/$3.96,並享有 97% 的快取折扣。
• 價格 — DeepSeek V4 Pro 每 100 萬 $0.66/$1.98(離峰;尖峰 $1.32/$3.96),相較之下 Qwen3.8 Max 為均一價 $2.00/$6.00
• 智慧指數 — V4 Pro 36(#7/113)對比 Qwen3.8 Max 40(#30/200)
• 輸出速度 — V4 Pro 約 81 tokens/秒,對比 Qwen3.8 Max 約 41 tokens/秒
• 每項任務成本 — V4 Pro $0.67 對比 Qwen3.8 Max $2.67(以 AA 指數任務計)
• 上下文 — 兩者皆為 1M tokens
• 授權 — V4 Pro 的 MIT 開放權重,對比 Qwen3.8 Max 的專有權重、商業授權
The output-price ratio is the whole argument: Qwen3.8 Max costs three times as much per input token and roughly three times per output token at V4 Pro's peak rate — and the gap on output widens to six times when V4 Pro is billed at its off-peak rate, which covers most of the week outside a narrow window. On a reasoning-heavy workload, where output tokens dominate the bill, that gap is the difference between a model you can leave running and one you watch in the dashboard. The cache lines widen it further: V4 Pro's 97% discount against Qwen's 88% means a long-context task with a reused prefix costs a fraction of its list price on the DeepSeek side.
智能:相差四分,而差距才是真正的產品
The independent scoreboard makes this closer than the price ratio suggests. On the current Artificial Analysis Intelligence Index, Qwen3.8 Max scores 40 against DeepSeek V4 Pro's 36 — a four-point gap that shows up in reasoning-heavy work but disappears in token-heavy work. Qwen3.8 Max's edge is real and consistent: it ranks #30 of 200 overall, just inside the frontier, while V4 Pro ranks #7 of 113 in its open-weights class. On Alibaba's own benchmark table, Qwen3.8 Max claims a Terminal-Bench 2.1 of 86.6, GPQA Diamond 92.6, and a FrontierSWE 73.5 that nearly doubled its predecessor's 40.7 — all vendor-reported, none reproduced independently. DeepSeek's model card claims a Codeforces rating of 3,348 and a Terminal-Bench 2.1 of 87.9 at maximum reasoning effort — likewise vendor-reported and unverified by a third party.
The honest framing is that the two companies are arguing past each other. Qwen3.8 Max's vendor table is built around enterprise and scientific work — the 0902 refresh Alibaba shipped on September 2 was further post-trained on Coding and Cowork, which Qwen says strengthens exactly that profile. DeepSeek V4 Pro's own claims center on deep reasoning and agentic tool use, where its four-week-old Codeforces and Terminal-Bench numbers still look strong. Neither vendor's table has been reproduced by an independent evaluator, and on the one scoreboard that is independent, the four-point gap is the whole difference between the models.
在速度上,Qwen3.8 Max 正悄悄失去優勢。
規格表沒揭露的那個面向,是輸出速度。Artificial Analysis 測得 DeepSeek V4 Pro 約為每秒 81 個 token,而 Qwen3.8 Max 約為 41——回應速度相差兩倍。在聊天類工作負載中,那是使用者明顯感受得到的停頓;在需要大量循序呼叫的 agent 迴圈裡,則會累積成實實在在的實際耗時。價格比已經對 V4 Pro 有利,而速度差距更讓「每個答案的實際成本」差距,比單看 token 計算所顯示的還要更大。
那個數字裡藏著一個取捨,那就是 Qwen3.8 Max 的冗長程度。Artificial Analysis 指出該模型「明顯緩慢且非常冗長」,而其每項任務成本——2.67 美元,相較於 V4 Pro 的 0.67 美元——既反映了較高的費率,也反映了額外的輸出 token。一個每次回答寫得更多的推理模型,即使在乘以價目表之前,每項任務的成本就已經更高。如果你的工作負載對延遲敏感,或者你的 agent 迴圈會進行數十次呼叫,那就是你該自己跑評估的那個數字。
權重:MIT 與商業授權的比較
DeepSeek V4 Pro is MIT-licensed with public weights — you can download them, serve them yourself, fine-tune them, and keep your modifications closed. Qwen3.8 Max is the first Max-class Qwen ever to ship open weights (the 2.4T A95B checkpoint landed on August 12), but under a commercial license with scale-tier conditions, and Alibaba has not published the full inference stack the way DeepSeek has. For a company comparing the two, the deployment story is the tie-breaker: V4 Pro can be self-hosted as infrastructure; Qwen3.8 Max is primarily a service you rent. The weights difference matters most if your concern is vendor lock-in or if you want to escape the API price entirely.
該由誰來選擇哪一個
如果你的工作負載偏重推理、以輸出 token 為主,或對延遲敏感,那麼在帳單上會出現的幾乎每一個面向,DeepSeek V4 Pro 都是答案:輸出成本便宜約三到六倍、速度快約兩倍、快取折扣更深,還有 MIT 授權的權重可作為退路。那四分的智慧落差確實存在,但幅度有限,而且在實際生產流量下,這點差距通常禁不起價格差異的考驗。
如果你的工作負載屬於企業型態——複雜的多步驟分析、深度的邏輯推導、要求嚴苛的代理式工作,或任何邊際答案品質值得付出三到六倍 token 成本的任务——那麼 Qwen3.8 Max 在獨立指數上多出的四分,以及其更強的企業基準測試聲稱,就是值得支付較高價格的理由。0902 更新進一步強化了這個定位。而如果你需要圖像或影片輸入,選擇更是高下立判:Qwen3.8 Max 接受文字、圖像和影片,而 DeepSeek V4 Pro 僅支援文字。
兩個模型都能透過 OrcaRouter 以單一金鑰呼叫——DeepSeek V4 Pro 和 Qwen3.8 Max 都經由該平台路由——而且由於 OrcaRouter 以無加價的方式直接傳遞供應商定價,本文中的數字就是你實際支付的數字,供應商降價也會當天生效。路由 DSL 讓你能從同一套整合中,把 token 用量大或對延遲敏感的流量導向 DeepSeek V4 Pro,並把多模態或最高風險的推理任務導向 Qwen3.8 Max——而這正是大多數團隊最終使用這套組合的方式:兩者並用,各展所長。


裁決
九月的那次暫緩,讓一直困擾這場對決的問題塵埃落定:便宜的那一半不再是隨時可能退場的風險。DeepSeek V4 Pro 是性價比的答案——輸出成本便宜三到六倍、速度快一倍、採用 MIT 授權,而且如今已承諾持續保持上線。Qwen3.8 Max 是能力的答案——四項獨立指標各高出四分、更強的企業級訴求、多模態輸入,以及一份讓它進不了你技術堆疊的商業授權。當你的評測證明那四分確實重要時,就選那四分;當你的帳單才是真正拍板定案的評測時,就選價格。
將耗用大量 token 或對延遲敏感的流量導向 DeepSeek V4 Pro,並把多模態或最關鍵的推理交給 Qwen3.8 Max,兩者都透過同一個整合——而這正是大多數團隊最終使用這套組合的方式:讓兩者各自發揮所長。
本文中的比較1
根據本文內容識別 · 基準測試:Artificial Analysis · 每日更新
