
Tencent HY4 Preview vs Qwen3.8-Max: The August Open-Weight Shootout
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
Tencent HY4 Preview vs Qwen3.8-Max is the collision this month was building toward. Alibaba's Qwen3.8-Max went general-availability on August 3 and became the first Max-class Qwen with open weights on August 12; Tencent HY4 Preview landed this morning, 25 days later, with a launch card that names Qwen3.8-Max directly — Tencent reports HY4 Preview at 74.1 on Toolathlon-Verified, ahead of Qwen3.8-Max on that benchmark. Both are open-weight Chinese flagships, both are priced to move, and only one of them has independent scores.
The two models are separated by more than scale — Qwen3.8-Max is 2.4 trillion parameters against HY4 Preview's 770 billion — but on the agentic benchmarks that matter to a buyer they are aiming at the same target: long-horizon work where a model reads, calls tools, and iterates without a human in the loop. That is exactly the territory where the August open-weight generation is trying to prove it can replace closed frontier models, and Tencent's first move was to take a shot at the month's biggest open release.
The scoreboard

• Scale — HY4 Preview 770B total / 49B active vs Qwen3.8-Max 2.4T total / ~95B active
• Context — HY4 Preview >1M tokens vs Qwen3.8-Max 1M tokens on the API (262K native in open weights, extendable to ~1.01M)
• Price per 1M — HY4 Preview ¥6 in / ¥18 out (≈$0.85 / $2.50) vs Qwen3.8-Max $2.00 in / $6.00 out, cache reads $0.25
• Terminal-Bench 2.1 — HY4 Preview 85.4 (vendor-reported) vs Qwen3.8-Max 86.6
• Modality — both text-only in their open-weight forms; Qwen3.8-Max API adds image and video input, HY4 Preview preview does not
• Independent score — HY4 Preview none vs Qwen3.8-Max AA Intelligence Index 58, Agentic Index 58
The two cards tell the same story from opposite sides. On Terminal-Bench 2.1, Qwen3.8-Max's 86.6 and HY4 Preview's 85.4 are separated by a point and a half — one point in Qwen's favor, and the HY4 figure is unaudited. On price, HY4 Preview is cheaper on both input and output per token. On verification, Qwen3.8-Max has an independent Intelligence Index of 58 and an Agentic Index of 58; HY4 Preview has nothing but its own press. The spread is the classic pattern of a challenger arriving with better prices and unproven claims against a verified incumbent.
Reading the Toolathlon claim
Toolathlon-Verified measures agentic tool-use reliability — how consistently a model drives a tool through a full task with errors, recovery and multi-step calls. Tencent's 74.1 is aimed directly at the strongest part of Qwen3.8-Max's profile, because Qwen's Agentic Index of 58 and its OSWorld-Verified 86.1 are the numbers that made Alibaba's agentic pitch credible. Qwen3.8-Max also posts 82.8 on IFBench (instruction following) and 93.0 on PaperBench (research-grade agentic work), both topping models like GPT-5.6 Sol on the vendor's own tables. So Tencent chose the one benchmark where it claims a direct win over the month's biggest open model — and the claim is currently as verifiable as any other single vendor row: not at all.
Verified vs vendor-reported
This is the axis where the matchup is least fair, and the asymmetry is worth being explicit about. Qwen3.8-Max's Intelligence Index of 58 comes from Artificial Analysis, its Agentic Index of 58 comes from Artificial Analysis, and the same independent harnesses that gave it those scores also measured its weaknesses: a 40% hallucination rate on AA-Omniscience, up from 23% on the previous generation, and an enormous 64 agentic turns per task — the model works hard, and you pay for every turn. Tencent HY4 Preview has none of that instrumentation. Its strengths are claims, and its failure modes — Tencent itself flags long thinking and over-self-verification — are admissions, not measurements.
For a team choosing between the two, the verification gap is not abstract. Qwen3.8-Max's 64-turns-per-task behavior is a real, measured cost driver: at $2/$6 it is still the most expensive agent per completed task in some third-party comparisons, and teams have reported the turn count eroding its price advantage. HY4 Preview's equivalent behavior is unknown — it could be cheaper per task than its already-low per-token price suggests, or its self-verification habit could eat the difference. Nobody can tell you which yet, because nobody outside Tencent has run it.
Price and the open-weights practicalities
Per token, HY4 Preview is cheaper on both sides of the ledger: ¥6/¥18 against Qwen3.8-Max's $2.00/$6.00, and cache hits at ¥0.3 against $0.25. That makes HY4 Preview the cheapest open-weight flagship on either input or output, full stop. But "open weights" comes with two different sets of fine print. Qwen3.8-Max's open-weights release — the Qwen3.8-2.4T-A95B — is text-only with thinking forced on, native context of 262K rather than the API's 1M, and a custom license with scale-tier conditions: display-model-name requirements above 100M monthly users, and a separate license for large Model-as-a-Service businesses. Tencent HY4 Preview's weights are on HuggingFace, ModelScope and GitHub as of this morning, but its exact license terms and whether the weights match the served API have not been stated with that kind of clarity. Both models are downloadable; neither is drop-in-trivial.
The bottom line
Qwen3.8-Max is the safer pick in every way that verification buys safety: independently scored on intelligence and agentic work, known failure modes, a settled license, and three weeks of production feedback. It is also the more expensive pick per token, and its measured turn-heavy behavior is a known cost risk. Tencent HY4 Preview is the value bet: cheaper on both input and output, comparable Terminal-Bench claim, and a direct Toolathlon-Verified claim against Qwen — all of it unverified on day one. If you are putting an open-weight model into production this week, Qwen3.8-Max is the defensible choice and it is live on OrcaRouter at Alibaba's list price — $2.00/$6.00, passed through with no markup. HY4 Preview is the evaluation project: run it through Tencent's own API against your own Toolathlon-style tasks, keep the production path on the verified model, and when the independent scores land — or when HY4 Preview reaches a router and its ¥6/¥18 rate becomes a same-day pass-through — the swap is a routing rule, not a rewrite.


What to watch
• The first independent run of HY4 Preview on an agentic harness. The Toolathlon-Verified 74.1 is the number Tencent wants to win; a third-party Agentic Index would be the confirmation.
• HY4 Preview's measured turn count. Qwen3.8-Max's 64 turns per task is the cautionary tale; whether HY4 Preview is cheaper per completed task is the entire economic argument.
• The license terms on the weights. They will determine whether "open" means "usable at your scale" for HY4 Preview, as the Qwen3.8-Max License conditions did for Alibaba's model.
• Whether either vendor cuts price again. In a month where two open-weight flagships shipped with aggressive pricing, the pass-through on a router makes any further cut visible the same day.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
