A hero title card for the comparison 'Tencent HY4 Preview vs Qwen3.8-Max'. A central banner reads 'The August open-weight shootout', annotated 'Two open-weight flagships, 25 days apart'. Left card 'Tencent HY4 Preview' lists '770B / 49B active, launched Aug 28', '¥6 / ¥18 per 1M', 'No independent scores'. Right card 'Qwen3.8-Max' lists '2.4T / ~95B active, launched Aug 3', '$2.00 / $6.00 per 1M', 'AA Index 58'. A footer reads 'HY4 Preview figures vendor-reported; Qwen3.8-Max scores per Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Tencent HY4 Preview vs Qwen3.8-Max: The August Open-Weight Shootout

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Tencent HY4 Preview vs Qwen3.8-Max is the collision this month was building toward. Ali​baba's Qwen3.8-Max went general-availability on August 3 and became the first Max-class Qw​en with open weights on August 12; Tencent HY4 Preview landed this morning, 25 days later, with a launch card that names Qwen3.8-Max directly — Tencent reports HY4 Preview at 74.1 on Toolathlon-Verified, ahead of Qwen3.8-Max on that benchmark. Both are open-weight Chinese flagships, both are priced to move, and only one of them has independent scores.

The two models are separated by more than scale — Qwen3.8-Max is 2.4 trillion parameters against HY4 Preview's 770 billion — but on the agentic benchmarks that matter to a buyer they are aiming at the same target: long-horizon work where a model reads, calls tools, and iterates without a human in the loop. That is exactly the territory where the August open-weight generation is trying to prove it can replace closed frontier models, and Tencent's first move was to take a shot at the month's biggest open release.

The scoreboard

A two-column scoreboard titled 'Tencent HY4 Preview vs Qwen3.8-Max — the scoreboard'. Left column 'Tencent HY4 Preview': 'Total / active: 770B / 49B', 'Context: >1M tokens', 'Terminal-Bench 2.1: 85.4 (vendor-reported)', 'Toolathlon-Verified: 74.1 (vendor-reported)', 'Price: ¥6 / ¥18 per 1M', 'Independent score: none'. Right column 'Qwen3.8-Max': 'Total / active: 2.4T / ~95B', 'Context: 1M tokens', 'Terminal-Bench 2.1: 86.6', 'AA Agentic Index: 58', 'Price: $2.00 / $6.00 per 1M', 'Independent score: AA Intelligence 58'. Footer reads 'HY4 Preview figures are Tencent-reported and unreproduced; Qwen3.8-Max scores from Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.

• Scale — HY4 Preview 770B total / 49B active vs Qwen3.8-Max 2.4T total / ~95B active

• Context — HY4 Preview >1M tokens vs Qwen3.8-Max 1M tokens on the API (262K native in open weights, extendable to ~1.01M)

• Price per 1M — HY4 Preview ¥6 in / ¥18 out (≈$0.85 / $2.50) vs Qwen3.8-Max $2.00 in / $6.00 out, cache reads $0.25

• Terminal-Bench 2.1 — HY4 Preview 85.4 (vendor-reported) vs Qwen3.8-Max 86.6

• Modality — both text-only in their open-weight forms; Qwen3.8-Max API adds image and video input, HY4 Preview preview does not

• Independent score — HY4 Preview none vs Qwen3.8-Max AA Intelligence Index 58, Agentic Index 58

The two cards tell the same story from opposite sides. On Terminal-Bench 2.1, Qwen3.8-Max's 86.6 and HY4 Preview's 85.4 are separated by a point and a half — one point in Qw​en's favor, and the HY4 figure is unaudited. On price, HY4 Preview is cheaper on both input and output per token. On verification, Qwen3.8-Max has an independent Intelligence Index of 58 and an Agentic Index of 58; HY4 Preview has nothing but its own press. The spread is the classic pattern of a challenger arriving with better prices and unproven claims against a verified incumbent.

Reading the Toolathlon claim

Toolathlon-Verified measures agentic tool-use reliability — how consistently a model drives a tool through a full task with errors, recovery and multi-step calls. Tencent's 74.1 is aimed directly at the strongest part of Qwen3.8-Max's profile, because Qw​en's Agentic Index of 58 and its OSWorld-Verified 86.1 are the numbers that made Ali​baba's agentic pitch credible. Qwen3.8-Max also posts 82.8 on IFBench (instruction following) and 93.0 on PaperBench (research-grade agentic work), both topping models like GPT-5.6 Sol on the vendor's own tables. So Tencent chose the one benchmark where it claims a direct win over the month's biggest open model — and the claim is currently as verifiable as any other single vendor row: not at all.

Verified vs vendor-reported

This is the axis where the matchup is least fair, and the asymmetry is worth being explicit about. Qwen3.8-Max's Intelligence Index of 58 comes from Artificial Analysis, its Agentic Index of 58 comes from Artificial Analysis, and the same independent harnesses that gave it those scores also measured its weaknesses: a 40% hallucination rate on AA-Omniscience, up from 23% on the previous generation, and an enormous 64 agentic turns per task — the model works hard, and you pay for every turn. Tencent HY4 Preview has none of that instrumentation. Its strengths are claims, and its failure modes — Tencent itself flags long thinking and over-self-verification — are admissions, not measurements.

For a team choosing between the two, the verification gap is not abstract. Qwen3.8-Max's 64-turns-per-task behavior is a real, measured cost driver: at $2/$6 it is still the most expensive agent per completed task in some third-party comparisons, and teams have reported the turn count eroding its price advantage. HY4 Preview's equivalent behavior is unknown — it could be cheaper per task than its already-low per-token price suggests, or its self-verification habit could eat the difference. Nobody can tell you which yet, because nobody outside Tencent has run it.

Price and the open-weights practicalities

Per token, HY4 Preview is cheaper on both sides of the ledger: ¥6/¥18 against Qwen3.8-Max's $2.00/$6.00, and cache hits at ¥0.3 against $0.25. That makes HY4 Preview the cheapest open-weight flagship on either input or output, full stop. But "open weights" comes with two different sets of fine print. Qwen3.8-Max's open-weights release — the Qwen3.8-2.4T-A95B — is text-only with thinking forced on, native context of 262K rather than the API's 1M, and a custom license with scale-tier conditions: display-model-name requirements above 100M monthly users, and a separate license for large Model-as-a-Service businesses. Tencent HY4 Preview's weights are on HuggingFace, ModelScope and GitHub as of this morning, but its exact license terms and whether the weights match the served API have not been stated with that kind of clarity. Both models are downloadable; neither is drop-in-trivial.

The bottom line

Qwen3.8-Max is the safer pick in every way that verification buys safety: independently scored on intelligence and agentic work, known failure modes, a settled license, and three weeks of production feedback. It is also the more expensive pick per token, and its measured turn-heavy behavior is a known cost risk. Tencent HY4 Preview is the value bet: cheaper on both input and output, comparable Terminal-Bench claim, and a direct Toolathlon-Verified claim against Qw​en — all of it unverified on day one. If you are putting an open-weight model into production this week, Qwen3.8-Max is the defensible choice and it is live on OrcaRouter at Ali​baba's list price — $2.00/$6.00, passed through with no markup. HY4 Preview is the evaluation project: run it through Tencent's own API against your own Toolathlon-style tasks, keep the production path on the verified model, and when the independent scores land — or when HY4 Preview reaches a router and its ¥6/¥18 rate becomes a same-day pass-through — the swap is a routing rule, not a rewrite.

A screenshot of the Artificial Analysis Intelligence Index leaderboard (captured August 28, 2026) showing Claude Opus 5 (max) and Claude Opus 5 (xhigh) at the top of the ranking, followed by Claude Fable 5 (with fallback) and GPT-5.6 Sol (max). Qwen3.8-Max is indexed further down and outside this capture. Tencent HY4 Preview does not appear: it launched today and has no independent index.A screenshot of the OrcaRouter model page for Qwen3.8-Max (qwen/qwen3.8-max) showing text, image and video input chips, a 1M-token context window, $2.00 per 1M input tokens and $6.00 per 1M output tokens, released August 3 2026.

What to watch

• The first independent run of HY4 Preview on an agentic harness. The Toolathlon-Verified 74.1 is the number Tencent wants to win; a third-party Agentic Index would be the confirmation.

• HY4 Preview's measured turn count. Qwen3.8-Max's 64 turns per task is the cautionary tale; whether HY4 Preview is cheaper per completed task is the entire economic argument.

• The license terms on the weights. They will determine whether "open" means "usable at your scale" for HY4 Preview, as the Qwen3.8-Max License conditions did for Ali​baba's model.

• Whether either vendor cuts price again. In a month where two open-weight flagships shipped with aggressive pricing, the pass-through on a router makes any further cut visible the same day.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube