
Tencent HY4 Preview vs GPT-5.6 Sol: The Tool-Calling Claim That Puts Open Weights Against the Frontier
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
Tencent HY4 Preview is the first open-weight model in months whose launch card explicitly claims a win over GPT-5.6 Sol. The number in question is Toolathlon-Verified, a benchmark for agentic tool-calling, where Tencent reports HY4 Preview at 74.1 — ahead of Qwen3.8-Max and, per Tencent's own comparison, ahead of GPT-5.6 Sol. On every other dimension, the two models are not really peers: GPT-5.6 Sol is OpenAI's closed, flagship, independently-measured frontier model, and Tencent HY4 Preview is a 770-billion-parameter open-weight preview that shipped this morning with a vendor-reported scorecard and no third-party verification.
So the interesting question is not "which model is smarter" — it is whether an open-weight challenger at roughly one-sixth of Sol's input price can credibly claim a slice of the frontier, and what a team should do with a claim that is currently unverifiable.
The claim that makes this a real matchup
Toolathlon-Verified measures how reliably a model drives a tool across a full agentic task — reading results, deciding the next call, recovering from errors — rather than how well it writes a single function. It matters because the largest share of real API spend on frontier models goes to agent loops, not to one-shot completions. Tencent's 74.1 on that benchmark is the specific claim that put HY4 Preview in GPT-5.6 Sol's conversation, and it is worth sitting with for a moment: it is the kind of number that, if it holds up under independent replication, makes the open-weight vs closed-frontier cost conversation real. If it does not hold up, it becomes another vendor scorecard that evaporates under a third-party harness.
Two ways to buy intelligence

• Parameters — HY4 Preview 770B total / 49B active, open weights vs GPT-5.6 Sol, parameter count undisclosed, closed API
• Context — HY4 Preview >1M tokens vs GPT-5.6 Sol 1.05M tokens, 128K max output
• Price per 1M — HY4 Preview ¥6 in / ¥18 out (≈$0.85 / $2.50) vs GPT-5.6 Sol $4.00 / $20.00 at the base tier, rising to $8.00 / $30.00 for large-context requests, cache reads $0.40
• Input modalities — HY4 Preview text-only in this preview vs GPT-5.6 Sol text and image input
• Reasoning modes — HY4 Preview no published equivalent vs GPT-5.6 Sol's Ultra mode (up to 16 parallel subagents) and Max reasoning effort
• Independent verification — HY4 Preview none yet vs GPT-5.6 Sol present on every major leaderboard, with a 1.05M-context ranking in the top tier of long-context composite tests
The price column is where the whole matchup lives. At the base tier, GPT-5.6 Sol costs $4.00 for a million input tokens and $20.00 for a million output tokens. Tencent HY4 Preview costs about $0.85 and $2.50 at current exchange rates. For a workload that moves a lot of output tokens — which agentic coding does — the open-weight model is priced at roughly an eighth of the closed one, and that gap survives even the caveat that Sol's numbers come with a lot more verification behind them.
Where GPT-5.6 Sol still holds the edge
Verification is the first edge. GPT-5.6 Sol has been general-availability since July 9 and independently measured since: 88.8 on Terminal-Bench 2.1 at base reasoning, 91.9 in Ultra mode, 53.6 on Agents' Last Exam (a record at release), 94.6 on GPQA Diamond, and a 64.6 SWE-bench Pro. OpenAI also reports a 54% token-efficiency improvement on agentic programming tasks over its own previous flagship. Those are numbers you can find reproduced outside OpenAI's own documentation, which is the exact property Tencent HY4 Preview lacks today.
Capability is the second edge, and it is structural rather than marginal. GPT-5.6 Sol accepts image input; HY4 Preview is text-only in this preview, with Tencent acknowledging vision will have to wait for the full release. GPT-5.6 Sol ships Ultra mode — a multi-agent configuration that coordinates parallel subagents at three to four times the token cost, useful for exactly the long-horizon engineering tasks where the two models would actually compete. And Sol's computer-use and programmatic tool-calling paths are documented, exercised at production scale, and load-tested. Tencent's Toolathlon-Verified claim, if it replicates, narrows the tool-use gap; it does not close the multimodal or multi-agent gaps, which HY4 Preview does not attempt.
Where HY4 Preview is genuinely interesting
Set the benchmark theater aside, and three real things remain. First, the price: at ¥6/¥18 — roughly $0.85/$2.50 — Tencent is undercutting every Western frontier model on output tokens by a factor of four to eight, and undercutting its own Chinese open-weight peers on input. Second, the open weights: a 49B-active MoE is deployable on a single node in a way that Sol's undisclosed architecture will never be, and the weights are already on HuggingFace, ModelScope and GitHub. Third, the context and the engineering trajectory: DeepSWE 64.3 and a 1M-plus context window are aimed at the same long-horizon agentic workloads Sol's Ultra mode targets, at a fraction of the cost.
The honest caveat is that none of this has survived contact with an independent benchmark. The self-optimization story Tencent tells — a 31.8% end-to-end throughput gain from the model iterating on its own inference stack — is internally measured. The blind-test win over GLM 5.3 and Kimi K3 is internally run. Every interesting thing about HY4 Preview is, today, a claim.
The decision
If you are building a production system today, GPT-5.6 Sol is the defensible choice, and it is reachable through OrcaRouter at OpenAI's own tiered list price — $4.00/$20.00 at the base tier, $8.00/$30.00 above it, passed through with zero markup, so any vendor price change is live on our side the same day. If your workflow is agentic and tool-heavy, the tiered structure means you should be checking your actual input sizes against the $272K-token threshold rather than assuming a flat rate.
If you are willing to run a shadow evaluation, HY4 Preview is cheap enough to be worth a real one: route your tool-calling workload through Sol today, run the same prompts against HY4 Preview through Tencent's own API, and compare on your own Toolathlon-style tasks. And if the independent scores land where Tencent says they will, the switching path is short — the open-weight model drops into a router and your failover rule swaps endpoints, with the vendor's ¥6/¥18 rate passed through untouched. That is the honest structure of this matchup: a verified frontier model you can build on today, facing an unverified open challenger that is cheap enough to test and positioned to make the price conversation about the entire open-weight class.


What to watch
• The first independent replication of Toolathlon-Verified. If a third-party harness confirms HY4 Preview in the 70s, the "open weights can't do agentic tool use" argument collapses.
• Whether Tencent ships vision before the full Hy4 release. Text-only limits HY4 Preview to a strict subset of the workloads GPT-5.6 Sol handles.
• The actual token bills from HY4 Preview's long-thinking tendency. At ¥18 per million output, verbose self-verification eats the price advantage faster than on a $20 model.
• Whether Sol's tiered pricing sees another cut before Hy4 full launches. The pass-through on a router means a drop is visible the same day.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
