A hero title card for the comparison 'Tencent HY4 Preview vs GPT-5.6 Sol'. A central highlight reads 'Toolathlon-Verified 74.1 - the open-weight claim against the frontier', annotated 'Tencent reports its model ahead of GPT-5.6 Sol on agentic tool-calling'. Left card 'Tencent HY4 Preview' lists 'Open weights, 770B / 49B active', '¥6 / ¥18 per 1M', 'No independent scores'. Right card 'GPT-5.6 Sol' lists 'Closed API, OpenAI flagship', '$4 / $20 base per 1M', 'Terminal-Bench 2.1: 88.8'. A footer reads 'HY4 Preview figures vendor-reported; Sol figures widely reproduced.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Tencent HY4 Preview vs GPT-5.6 Sol: The Tool-Calling Claim That Puts Open Weights Against the Frontier

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Tencent HY4 Preview is the first open-weight model in months whose launch card explicitly claims a win over GPT-5.6 Sol. The number in question is Toolathlon-Verified, a benchmark for agentic tool-calling, where Tencent reports HY4 Preview at 74.1 — ahead of Qwen3.8-Max and, per Tencent's own comparison, ahead of GPT-5.6 Sol. On every other dimension, the two models are not really peers: GPT-5.6 Sol is Ope​nAI's closed, flagship, independently-measured frontier model, and Tencent HY4 Preview is a 770-billion-parameter open-weight preview that shipped this morning with a vendor-reported scorecard and no third-party verification.

So the interesting question is not "which model is smarter" — it is whether an open-weight challenger at roughly one-sixth of Sol's input price can credibly claim a slice of the frontier, and what a team should do with a claim that is currently unverifiable.

The claim that makes this a real matchup

Toolathlon-Verified measures how reliably a model drives a tool across a full agentic task — reading results, deciding the next call, recovering from errors — rather than how well it writes a single function. It matters because the largest share of real API spend on frontier models goes to agent loops, not to one-shot completions. Tencent's 74.1 on that benchmark is the specific claim that put HY4 Preview in GPT-5.6 Sol's conversation, and it is worth sitting with for a moment: it is the kind of number that, if it holds up under independent replication, makes the open-weight vs closed-frontier cost conversation real. If it does not hold up, it becomes another vendor scorecard that evaporates under a third-party harness.

Two ways to buy intelligence

A two-column scoreboard titled 'Tencent HY4 Preview vs GPT-5.6 Sol — the scoreboard'. Left column 'Tencent HY4 Preview': 'Params: 770B / 49B active, open weights', 'Context: >1M tokens', 'Toolathlon-Verified: 74.1 (vendor-reported)', 'APEX-Agents: 37.1', 'Price: ¥6 / ¥18 per 1M', 'Independent score: none'. Right column 'GPT-5.6 Sol': 'Params: undisclosed, closed API', 'Context: 1.05M tokens, 128K max output', 'Terminal-Bench 2.1: 88.8 (91.9 Ultra)', 'Agents' Last Exam: 53.6', 'Price: $4 / $20 base per 1M', 'Independent score: on all major leaderboards'. Footer reads 'HY4 Preview figures are Tencent-reported and unreproduced; GPT-5.6 Sol figures widely reproduced.' The OrcaRouter logo is composited in the bottom-right corner.

• Parameters — HY4 Preview 770B total / 49B active, open weights vs GPT-5.6 Sol, parameter count undisclosed, closed API

• Context — HY4 Preview >1M tokens vs GPT-5.6 Sol 1.05M tokens, 128K max output

• Price per 1M — HY4 Preview ¥6 in / ¥18 out (≈$0.85 / $2.50) vs GPT-5.6 Sol $4.00 / $20.00 at the base tier, rising to $8.00 / $30.00 for large-context requests, cache reads $0.40

• Input modalities — HY4 Preview text-only in this preview vs GPT-5.6 Sol text and image input

• Reasoning modes — HY4 Preview no published equivalent vs GPT-5.6 Sol's Ultra mode (up to 16 parallel subagents) and Max reasoning effort

• Independent verification — HY4 Preview none yet vs GPT-5.6 Sol present on every major leaderboard, with a 1.05M-context ranking in the top tier of long-context composite tests

The price column is where the whole matchup lives. At the base tier, GPT-5.6 Sol costs $4.00 for a million input tokens and $20.00 for a million output tokens. Tencent HY4 Preview costs about $0.85 and $2.50 at current exchange rates. For a workload that moves a lot of output tokens — which agentic coding does — the open-weight model is priced at roughly an eighth of the closed one, and that gap survives even the caveat that Sol's numbers come with a lot more verification behind them.

Where GPT-5.6 Sol still holds the edge

Verification is the first edge. GPT-5.6 Sol has been general-availability since July 9 and independently measured since: 88.8 on Terminal-Bench 2.1 at base reasoning, 91.9 in Ultra mode, 53.6 on Agents' Last Exam (a record at release), 94.6 on GPQA Diamond, and a 64.6 SWE-bench Pro. Ope​nAI also reports a 54% token-efficiency improvement on agentic programming tasks over its own previous flagship. Those are numbers you can find reproduced outside Ope​nAI's own documentation, which is the exact property Tencent HY4 Preview lacks today.

Capability is the second edge, and it is structural rather than marginal. GPT-5.6 Sol accepts image input; HY4 Preview is text-only in this preview, with Tencent acknowledging vision will have to wait for the full release. GPT-5.6 Sol ships Ultra mode — a multi-agent configuration that coordinates parallel subagents at three to four times the token cost, useful for exactly the long-horizon engineering tasks where the two models would actually compete. And Sol's computer-use and programmatic tool-calling paths are documented, exercised at production scale, and load-tested. Tencent's Toolathlon-Verified claim, if it replicates, narrows the tool-use gap; it does not close the multimodal or multi-agent gaps, which HY4 Preview does not attempt.

Where HY4 Preview is genuinely interesting

Set the benchmark theater aside, and three real things remain. First, the price: at ¥6/¥18 — roughly $0.85/$2.50 — Tencent is undercutting every Western frontier model on output tokens by a factor of four to eight, and undercutting its own Chinese open-weight peers on input. Second, the open weights: a 49B-active MoE is deployable on a single node in a way that Sol's undisclosed architecture will never be, and the weights are already on HuggingFace, ModelScope and GitHub. Third, the context and the engineering trajectory: DeepSWE 64.3 and a 1M-plus context window are aimed at the same long-horizon agentic workloads Sol's Ultra mode targets, at a fraction of the cost.

The honest caveat is that none of this has survived contact with an independent benchmark. The self-optimization story Tencent tells — a 31.8% end-to-end throughput gain from the model iterating on its own inference stack — is internally measured. The blind-test win over GLM 5.3 and Kimi K3 is internally run. Every interesting thing about HY4 Preview is, today, a claim.

The decision

If you are building a production system today, GPT-5.6 Sol is the defensible choice, and it is reachable through OrcaRouter at Ope​nAI's own tiered list price — $4.00/$20.00 at the base tier, $8.00/$30.00 above it, passed through with zero markup, so any vendor price change is live on our side the same day. If your workflow is agentic and tool-heavy, the tiered structure means you should be checking your actual input sizes against the $272K-token threshold rather than assuming a flat rate.

If you are willing to run a shadow evaluation, HY4 Preview is cheap enough to be worth a real one: route your tool-calling workload through Sol today, run the same prompts against HY4 Preview through Tencent's own API, and compare on your own Toolathlon-style tasks. And if the independent scores land where Tencent says they will, the switching path is short — the open-weight model drops into a router and your failover rule swaps endpoints, with the vendor's ¥6/¥18 rate passed through untouched. That is the honest structure of this matchup: a verified frontier model you can build on today, facing an unverified open challenger that is cheap enough to test and positioned to make the price conversation about the entire open-weight class.

A screenshot of the Artificial Analysis Intelligence Index leaderboard (captured August 28, 2026) showing Claude Opus 5 (max) and Claude Opus 5 (xhigh) at the top of the ranking, followed by Claude Fable 5 (with fallback), with GPT-5.6 Sol (max) visible in the near-top rows — the only model in this matchup with an independent index. Tencent HY4 Preview does not appear: it launched today and has no independent index.A screenshot of the OrcaRouter model page for GPT-5.6 Sol (openai/gpt-5.6-sol) showing a 1.05M-token context window, 128K max output, base pricing of $4.00 per 1M input and $20.00 per 1M output tokens, and a July 9 2026 release date.

What to watch

• The first independent replication of Toolathlon-Verified. If a third-party harness confirms HY4 Preview in the 70s, the "open weights can't do agentic tool use" argument collapses.

• Whether Tencent ships vision before the full Hy4 release. Text-only limits HY4 Preview to a strict subset of the workloads GPT-5.6 Sol handles.

• The actual token bills from HY4 Preview's long-thinking tendency. At ¥18 per million output, verbose self-verification eats the price advantage faster than on a $20 model.

• Whether Sol's tiered pricing sees another cut before Hy4 full launches. The pass-through on a router means a drop is visible the same day.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube