
DeepSeek V4 Pro 対 Qwen3.8 Max:3倍の価格差、9月の猶予措置後の再実行
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040知能
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453知能77コーディング
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241知能76コーディング
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240知能72コーディング
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153知能82コーディング
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 100万トークンあたり
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642知能72コーディング
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 100万トークンあたり
- z-aiZ.ai: GLM 5.32026-08-1845知能75コーディング
- obsidianQwen3.8 27B2026-08-1534知能68コーディング
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236知能69コーディング
- grokSpaceXAI: Grok 4.62026-08-1244知能77コーディング
- metaMeta: Muse Spark 1.22026-08-0540知能72コーディング
- qwenQwen: Qwen3.8 Max2026-08-0340知能72コーディング
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135知能69コーディング
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 100万トークンあたり
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451知能78コーディング
- googleGoogle: Gemini 3.6 Flash2026-07-2134知能69コーディング
DeepSeek V4 Pro and Qwen3.8 Max are the two most important open-or-cheap flagships of this generation, and the September events on the DeepSeek side changed the matchup more than either vendor's marketing has acknowledged. DeepSeek V4 Pro shipped on August 13, 2026, ten days after Alibaba's Qwen3.8 Max went generally available on August 3, and the pairing has always read as a one-sided contest: Qwen3.8 Max is the 2.4-trillion-parameter all-rounder at $2 per million input and $6 per million output tokens, while DeepSeek V4 Pro is the 1.6-trillion-parameter reasoning specialist at a fraction of that price, MIT-licensed, with public weights. The reason this comparison is worth re-running today is not the launch-week numbers but what happened a month later — DeepSeek spent the week of September 8 trying to retire V4 Pro, then reversed on September 11 and committed to keeping it online at unchanged billing. That reversal is the single most important fact about this matchup right now, because it decides whether the cheap half is a safe bet or a migration trap.
Independent measurements in this piece come from Artificial Analysis, captured today. Alibaba's and DeepSeek's own benchmark tables are labeled vendor-reported throughout — neither company's launch numbers have been fully reproduced by a third party, and DeepSeek's are now four weeks old against a moving scoreboard.
9月の猶予が本題であり、8月のローンチではない
9月8日、DeepSeekは、9月14日以降、deepseek-v4-proへのすべてのリクエストをV4.1 Flashにルーティングし、V4.1 Proが提供されるまでFlashの料金で請求すると発表した。Qwen3.8 MaxとDeepSeek V4 Proを真剣な比較として扱っている人なら誰でも、安価な選択肢が消えることを想定して計画を始めなければならなかった。その後、DeepSeekは打ち切りを延期し、9月11日には完全に方針を覆して、9月14日以降もDeepSeek V4 ProのAPIサービスを請求内容を変えずに提供し続け、それが変更される場合は事前に通知すると述べた。この方針転換は同日、中国の国営メディアとテック系プレスによって報じられた。
これがこの対決にもたらすのは、これまで静かに意思決定を左右していた非対称性の除去だ。9月の第1週を通じて、「$0.66の出力モデルと$6の出力モデルのどちらを基盤に構築すべきか」という問いへの正直な答えは、存続リスクによって汚染されていた。価格リーダーにはサービス終了日があったからだ。9月15日時点で、その汚染は消えた。安価な選択肢は持続可能であり、比較はついに、本来常にそうあるべきだったもの——知能、価格、速度——についてのものになった。

価格:同じ結果を、出力コストのほんの一部で
現在の独立系スコアボードでは両者は4ポイント差につけており、その価格差はオープンウェイト部門で最大だ。Qwen3.8 MaxはAlibabaの定価で入力100万トークンあたり2.00ドル、出力100万トークンあたり6.00ドルで、キャッシュ割引は88%。DeepSeek V4 Proはオフピーク時で入力100万トークンあたり0.66ドル、出力100万トークンあたり1.98ドル、ピーク時には2倍の1.32ドル/3.96ドルになり、キャッシュ割引は97%だ。
• 価格 — DeepSeek V4 Pro は100万あたり $0.66/$1.98(オフピーク時、ピーク時は $1.32/$3.96)に対し、Qwen3.8 Max は $2.00/$6.00 の一律料金
• インテリジェンス指数 — V4 Pro 36(#7/113)vs Qwen3.8 Max 40(#30/200)
• 出力速度 — V4 Pro 約81トークン/秒 対 Qwen3.8 Max 約41トークン/秒
• タスクあたりのコスト — V4 Pro $0.67 対 Qwen3.8 Max $2.67(AA インデックスのタスクあたり)
• コンテキスト — 両方とも 1M トークン
• ライセンス — V4 Pro の MIT オープンウェイト 対 Qwen3.8 Max の独自ウェイト、商用ライセンス
The output-price ratio is the whole argument: Qwen3.8 Max costs three times as much per input token and roughly three times per output token at V4 Pro's peak rate — and the gap on output widens to six times when V4 Pro is billed at its off-peak rate, which covers most of the week outside a narrow window. On a reasoning-heavy workload, where output tokens dominate the bill, that gap is the difference between a model you can leave running and one you watch in the dashboard. The cache lines widen it further: V4 Pro's 97% discount against Qwen's 88% means a long-context task with a reused prefix costs a fraction of its list price on the DeepSeek side.
知能:4ポイント差、そしてその差こそが真の製品
The independent scoreboard makes this closer than the price ratio suggests. On the current Artificial Analysis Intelligence Index, Qwen3.8 Max scores 40 against DeepSeek V4 Pro's 36 — a four-point gap that shows up in reasoning-heavy work but disappears in token-heavy work. Qwen3.8 Max's edge is real and consistent: it ranks #30 of 200 overall, just inside the frontier, while V4 Pro ranks #7 of 113 in its open-weights class. On Alibaba's own benchmark table, Qwen3.8 Max claims a Terminal-Bench 2.1 of 86.6, GPQA Diamond 92.6, and a FrontierSWE 73.5 that nearly doubled its predecessor's 40.7 — all vendor-reported, none reproduced independently. DeepSeek's model card claims a Codeforces rating of 3,348 and a Terminal-Bench 2.1 of 87.9 at maximum reasoning effort — likewise vendor-reported and unverified by a third party.
The honest framing is that the two companies are arguing past each other. Qwen3.8 Max's vendor table is built around enterprise and scientific work — the 0902 refresh Alibaba shipped on September 2 was further post-trained on Coding and Cowork, which Qwen says strengthens exactly that profile. DeepSeek V4 Pro's own claims center on deep reasoning and agentic tool use, where its four-week-old Codeforces and Terminal-Bench numbers still look strong. Neither vendor's table has been reproduced by an independent evaluator, and on the one scoreboard that is independent, the four-point gap is the whole difference between the models.
スピードの面で、Qwen3.8 Maxは静かに後れを取っている
スペックシートが隠しているのは出力速度という側面です。Artificial AnalysisはDeepSeek V4 Proを約81トークン/秒と測定しており、Qwen3.8 Maxの約41トークン/秒と対比されます。これは回答が返ってくる速さにおける2倍の差です。チャットワークロードではユーザーが気づく一時停止であり、逐次呼び出しの多いエージェントループでは実際のウォールクロックタイムに累積します。価格比はすでにV4 Proに有利であり、速度差によって、実効的な回答あたりコストの差はトークンの計算が示唆する以上に広がります。
その数字にはトレードオフが潜んでいる。すなわち Qwen3.8 Max の冗長さだ。Artificial Analysis はこのモデルを「著しく遅く、非常に冗長」と指摘しており、タスクあたりのコスト——V4 Pro の $0.67 に対して $2.67——は、高い単価と余分な出力トークンの両方を反映している。回答ごとにより多くを書く推論モデルは、単価を掛けるまでもなくタスクあたりのコストが高くつく。ワークロードがレイテンシに敏感だったり、エージェントループが何十回も呼び出しを行うなら、自分の評価を回すべきなのはまさにその数字だ。
重み: MIT 対 商用ライセンス
DeepSeek V4 Pro is MIT-licensed with public weights — you can download them, serve them yourself, fine-tune them, and keep your modifications closed. Qwen3.8 Max is the first Max-class Qwen ever to ship open weights (the 2.4T A95B checkpoint landed on August 12), but under a commercial license with scale-tier conditions, and Alibaba has not published the full inference stack the way DeepSeek has. For a company comparing the two, the deployment story is the tie-breaker: V4 Pro can be self-hosted as infrastructure; Qwen3.8 Max is primarily a service you rent. The weights difference matters most if your concern is vendor lock-in or if you want to escape the API price entirely.
誰がどれを選ぶべきか
ワークロードが推論重視、出力トークン支配的、またはレイテンシに敏感なら、請求に現れるほぼすべての軸でDeepSeek V4 Proが答えです。出力ではおよそ3~6分の1のコスト、約2倍の速度、より大きなキャッシュ割引、そして逃げ道としてMITライセンスの重みがあります。4ポイントの知能差は実在しますが狭く、本番トラフィックでは価格差に直面するとほとんど生き残れません。
ワークロードがエンタープライズ向けの性質を帯びているなら——複雑な多段階分析、深い論理的導出、要求の厳しいエージェント作業、あるいは限界的な回答品質がトークンコストの3〜6倍の価値を持つあらゆるタスク——Qwen3.8 Max が独立系インデックスで追加の4ポイントを獲得していることと、より強力なエンタープライズベンチマークの主張が、割増料金を支払う理由です。0902のリフレッシュは、そのプロファイルをさらに際立たせます。また、画像や動画の入力が必要なら、選択の余地はまったくありません。Qwen3.8 Max はテキスト、画像、動画を受け付けますが、DeepSeek V4 Pro はテキスト専用です。
両モデルはOrcaRouterを通じて単一のキーで呼び出せます — DeepSeek V4 ProとQwen3.8 Maxはどちらもこのプラットフォーム経由でルーティングされます — そしてOrcaRouterはプロバイダーの定価をマークアップなしでそのまま通すため、本記事に載っている数字は実際に支払う金額そのままであり、ベンダーの値下げも当日に反映されます。ルーティングDSLを使えば、トークン消費の大きいトラフィックやレイテンシに敏感なトラフィックをDeepSeek V4 Proへ、マルチモーダルなタスクや最も重要な推論タスクをQwen3.8 Maxへと、同じ統合から振り分けられます — これこそが、ほとんどのチームがこの組み合わせを最終的にこう使う理由です。つまり、それぞれが最も得意とする処理のために両方を使うのです。


評決
9月の猶予が、この対決を覆い尽くしていた問いに決着をつけた:安価な側はもはや離脱リスクではない。DeepSeek V4 Proはコストパフォーマンスの答えだ——出力で3〜6倍安く、2倍高速、MITライセンスで、しかも今やオンラインに留まることが約束されている。Qwen3.8 Maxは能力の答えだ——独立したインデックスで4ポイント高く、より強力なエンタープライズ向けの主張、マルチモーダル入力、そしてあなたのスタックから締め出す商用ライセンス。あなたの評価がその4ポイントの重要性を証明するなら4ポイントを選べ。請求書こそが実際に決める評価であるなら、価格を選べ。
トークン消費の大きいトラフィックやレイテンシに敏感なトラフィックは DeepSeek V4 Pro へ、マルチモーダルな処理や最も重要な推論はQwen3.8 Maxへと、同じ統合から振り分ける——これこそ、ほとんどのチームがこの組み合わせをこう使う理由です。どちらも、それぞれが最も得意とする用途のために。
この記事で比較したモデル1
この記事から検出 · ベンチマーク:Artificial Analysis · 毎日更新
