
Claude Sonnet 5.5 vs Tencent HY4 Preview: Both Labs Are Selling the Token Bill
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 931 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 192 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1177 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 70 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 107 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Two launches, six weeks apart, making the same promise in different words. The vendor's claim is that Claude Sonnet 5.5, shipped September 28, 2026, costs "up to 30% less per task" than Claude Sonnet 5 at unchanged per-token rates. The second claim is that the September 7 optimisation to the hosted build of Tencent HY4 Preview, first released and open-sourced on August 28, 2026, cuts reasoning turns and per-task token consumption at the same task quality. Neither vendor raised or lowered a price to make that claim. Both are telling you that the tokens are the product now, and the interesting part of this matchup is that only one of the two claims can be checked against a published per-task figure — because only one of these two models has an independent measurement to check it against.
That asymmetry is not a knock on Tencent. HY4 Preview has no Artificial Analysis model page, no independent Intelligence Index score, and no cost-per-task figure from any third party. What it has is a vendor benchmark table — Terminal-Bench 2.1 at 85.4, DeepSWE 64.3, an internal blind evaluation across 203 engineering tasks where it averaged 2.99 against GLM-5.3 at 2.92 and Kimi K3 at 2.94 — and a public crowd-voted debut on the WebDev Arena leaderboard, where Tencent reports it landing around fifth overall and third among open-weight models. All of those are real signals. None of them is a per-task cost, which means the exact number both vendors are competing on is, for HY4 Preview, unavailable.
What each side actually shipped
Tencent HY4 Preview is a 770-billion-parameter mixture of experts with 49 billion active per token, 78 layers, 256 routed experts plus one shared expert, a native MTP layer for speculative decoding, and a context window that clears a million tokens. It is text-only by design — Tencent built it for coding agents, complex tool use and productivity work, not for images — and it defaults to heavy reasoning, with three configurable levels (high, low, none) and a vendor-recommended sampling setting of temperature 0.9 and top_p 1.0. The weights are Apache 2.0 and downloadable from Hugging Face, GitHub, ModelScope and GitCode. Tencent originally flagged two rough edges in the preview, spending longer than necessary thinking and over-verifying its own work, and the September 7 update targets exactly those.
Claude Sonnet 5.5 is Anthropic's second 5.5-generation model and the one it positions as the best combination of speed and intelligence in the lineup. One million tokens of context, 128,000 output tokens synchronously and 300,000 on the Message Batches API, text, image and file input, adaptive thinking at selectable effort, a June 2026 knowledge cutoff, zero data retention available from launch, and a retirement floor of September 28, 2027. The rate card did not move from Claude Sonnet 5: $2.00 per million input, $10.00 per million output, $0.20 for cache reads and $2.50 for cache writes. Closed weights, Anthropic API plus Bedrock, Google Cloud and Microsoft Foundry.
Line the two against each other and the positions are almost symmetrical:
• Price — Claude Sonnet 5.5 $2.00 in / $10.00 out per million vs Tencent HY4 Preview $0.83 in / $2.50 out on our catalogue, from a vendor list of ¥6 / ¥18
• Cache reads — Claude Sonnet 5.5 $0.20 per million vs Tencent HY4 Preview $0.042 per million on our catalogue
• Weights — Claude Sonnet 5.5 closed vs Tencent HY4 Preview Apache 2.0, downloadable
• Context — Claude Sonnet 5.5 1,000,000 tokens vs Tencent HY4 Preview 1,048,576 tokens, effectively identical
• Maximum output — Claude Sonnet 5.5 128,000 synchronous, 300,000 on Batches vs Tencent HY4 Preview 64,000 on our catalogue
• Input modality — Claude Sonnet 5.5 text, image and file vs Tencent HY4 Preview text only
• Independent scores — Claude Sonnet 5.5 Intelligence Index 56 at $7.60 per task, revision v4.3.2 vs Tencent HY4 Preview none published
• Measured output speed — Claude Sonnet 5.5 138.7 tokens/sec vs Tencent HY4 Preview 22.3 tokens/sec on our own traffic

The claim both labs are making, and why one of them is checkable
Anthropic's "30% less per task" and Tencent's "fewer reasoning turns, lower token consumption" are the same engineering project described from two directions: make the model spend fewer tokens arriving at the same answer. Anthropic is explicit that the rates did not change, which means any per-task saving has to come from token counts — and the independent board lets you see the shape of the problem rather than the size of the fix. Claude Sonnet 5.5 generates 410 million output tokens across the Intelligence Index evaluation, against 117 million for MiniMax M3 and 77 million for Grok 4.5. It is a verbose model, measured, on a board that publishes the number. Whatever the 30% improvement over Claude Sonnet 5 amounts to, it is being applied to a very large base.
For HY4 Preview the equivalent figure does not exist publicly. Tencent's own optimisation note is the only account of it, and it states the result without a number: fewer turns, lower input and output tokens, same quality, validated by benchmark metrics and human evaluation in a dual pass. That is a vendor claim in the strict sense — stated, plausible, unreproduced — and it is labelled that way here. Adopting a preview model on the strength of a vendor's token-economy claim, when the token economy is the thing you cannot measure from outside, is precisely the bet a preview release is asking you to make.
One further wrinkle is worth knowing before anyone plans a self-hosted deployment around that optimisation. Tencent's September 7 change applies to the build Tencent serves on its own hosted endpoints. The open-weight snapshot was not refreshed — the Hugging Face repository's last modification date is still August 28 — so a self-hosted copy of the weights runs the original, longer-reasoning behaviour that the preview notes flagged. The optimised version and the downloadable version are not the same artifact.
Where the open weights do pay off is the same place they always do: a deployment that cannot call an API. A 770B-parameter model with 49B active is a datacenter decision rather than a workstation one, so this is not the laptop-class story a 30B model tells — but for a buyer who needs the model inside their own perimeter, Apache 2.0 at that scale is an option Claude Sonnet 5.5 does not offer at any price.
Trying a preview without betting a production path on it
Preview models are the case where routing stops being a cost optimisation and becomes a risk control. The reason to put a preview build behind a router rather than calling it directly is that a preview is defined by the possibility of changing under you, and the two things you want from the layer in front of it are a percentage split and something to catch the failures.
That is the shape OrcaRouter provides: one OpenAI-compatible endpoint covering 200-plus models, provider list price passed through with no markup added, automatic failover between upstream providers, and a routing DSL for composing calls out of several models. Tencent HY4 Preview is routable here today as tencent/hy4-preview at $0.83 input and $2.50 output per million tokens with $0.042 cache reads across its 1,048,576-token window — the same build, at the provider's rate, reachable from the key you already use for everything else in the catalogue. Claude Sonnet 5 is routable here too, at Anthropic's $2.00 and $10.00, and has been since June 30; it makes an obvious failover partner while a day-old model accumulates production evidence.

Claude Sonnet 5.5 itself is not in our catalogue. It shipped yesterday, the route to it is Anthropic's own API, and this page is not going to suggest otherwise — the point of the routing advice is that the safe way to run a brand-new tier in production is to give it a slice of traffic with something proven behind it, and that only works when both ends of the arrangement live somewhere you can reach from one place. HY4 Preview is carrying 372 thousand tokens of traffic a week here, which is what a two-day-old preview looks like before anyone has committed to it; Claude Sonnet 5 has been carrying 7.9 million a week since June 30, and that is the difference between a candidate and a floor.

Where the money actually goes
On rates alone, HY4 Preview is four times cheaper on output than Claude Sonnet 5.5 and nearly five times cheaper on cache reads, and it matches the window. On the measured side, Claude Sonnet 5.5 generates output six times faster on our own traffic — 138.7 tokens per second against 22.3 — which is the kind of gap that turns a four-fold rate advantage into a much smaller wall-clock advantage, and occasionally into a loss, once queueing is taken into account.
Between those two facts the decision rests on four questions that only the reader can answer. Does the work need an image or a file in the prompt? HY4 Preview is text-only and that ends the discussion. Does it need to complete a multi-step software task, where HY4 Preview's Terminal-Bench 2.1 score of 85.4 is Tencent's own number and unreproduced? Then the preview status is the risk you are accepting, and the September 7 optimisation is worth re-testing rather than trusting. Does the workload have to run inside your own perimeter? Then Apache 2.0 weights are the deciding fact. And does the composite capability level matter more than the rate? Then Claude Sonnet 5.5's independently measured 56 is the number to weigh, at $7.60 per Index task, with the caveat that no equivalent figure exists on the Tencent side to compare it to.
What both launches really demonstrate is that the frontier has moved on from rate cards. Two labs, six weeks apart, shipped models and led with per-task token consumption rather than price per million, and the practical consequence for a buyer is that the useful number is no longer on either vendor's pricing page. It is the token count your own workload produces, measured on both models, and it is the one figure in this article neither vendor can give you.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
