
GPT-6 Sol vs Tencent HY4 Preview: The Cheaper Model Is Carrying Six Hundred Times the Traffic
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Comparisons between a frontier model and a cheap one usually end at the price table, and this one has a better ending available. Tencent HY4 Preview and GPT-6 Sol are both routed by OrcaRouter, which means the same telemetry covers both, and over the same rolling seven-day window the HY4 Preview model moved roughly 2.4 billion tokens against GPT-6 Sol's roughly 3.9 million. That is not a benchmark and it is not a verdict; it is a rolling window on one platform's traffic and it will read differently next week. But it is a real signal about what people building things actually reach for when a strong open-weight model costs a quarter of the closed one, and it is the kind of number a spec sheet cannot produce.
The two models are also a study in what a benchmark can and cannot tell you. GPT-6 Sol, released 22 September 2026 as the middle tier of OpenAI's GPT-6 generation, has an independent Intelligence Index score. HY4 Preview, released and open-sourced by Tencent on 28 August 2026, has no Artificial Analysis model page, no third-party index, and no independent per-task cost figure at all. Everything published about how well it works is Tencent's own measurement. That asymmetry does not make HY4 Preview the weaker model. It makes it the less audited one, and those are different claims.
The specs, including the two that decide most deployments
• Output price — GPT-6 Sol $10.00 per million vs Tencent HY4 Preview $2.50 per million
• Input price — GPT-6 Sol $2.00 per million vs Tencent HY4 Preview $0.83 per million
• Cached input — GPT-6 Sol $0.20 per million vs Tencent HY4 Preview $0.04 per million
• Weights — GPT-6 Sol closed, API only vs HY4 Preview open-sourced under Apache 2.0
• Input modalities — GPT-6 Sol text, image and file vs HY4 Preview text only
• Context window — GPT-6 Sol 1,050,000 tokens vs HY4 Preview 1,048,576 tokens
• Max output — GPT-6 Sol 128,000 tokens vs HY4 Preview 64,000 tokens
• Long-request step — GPT-6 Sol reprices the whole request above 272,000 input tokens to $4.00 / $15.00 vs HY4 Preview no step on its card
• Architecture — GPT-6 Sol undisclosed vs HY4 Preview 770B total parameters, 49B active, 78 layers, 256 routed experts plus one shared, native multi-token-prediction layer
• Independent score — GPT-6 Sol 47.6 Intelligence Index vs HY4 Preview none published
• Measured output speed — GPT-6 Sol roughly 244 tokens per second vs HY4 Preview roughly 54, both rolling seven-day figures
The two lines that decide most real deployments are the ones furthest down. The first is that HY4 Preview takes no image or file input, which removes it from any pipeline that has to look at something. The second is the output ceiling: 64,000 tokens is half of GPT-6 Sol's 128,000, and on a model whose stated purpose is coding agents and long tool-use chains, that is the constraint a team hits first — not quality, but the point at which a generated file no longer fits in one response.
Against those, a quarter of the output price and a fifth of the cached-input price buy a great deal of retrying. And HY4 Preview's architecture is documented in a way GPT-6 Sol's is not: Tencent published the parameter count, the layer count, the expert layout and the speculative-decoding layer, which means the serving characteristics are at least partly predictable from the paper rather than observable only through the API.

What HY4 Preview's benchmark table actually is
Tencent's published numbers are strong and they are all vendor-run. On Terminal-Bench 2.1 the model posts 85.4; on DeepSWE, 64.3; and in an internal blind evaluation across 203 engineering tasks it averaged 2.99 against GLM-5.3 at 2.92 and Kimi K3 at 2.94 — a lead of seven hundredths of a point on a ten-point scale in a study Tencent designed. It also debuted on the WebDev Arena leaderboard, where Tencent reports it landing around fifth overall and third among open-weight models. The arena result is crowd-voted rather than vendor-scored and is the most independent evidence in the set; the rest is a claim, and a claim is still worth reading as long as it is labelled one.
The gap that matters is not that these numbers might be generous — vendor harnesses usually are, in both directions. It is that there is no second measurement to compare them against. A model with an Artificial Analysis page can be checked; a model without one cannot, and the difference shows up months later when someone discovers that a headline number was measured on a harness nobody else can reproduce. Nothing here suggests that has happened to HY4 Preview. It just means an adoption decision resting on the 85.4 is resting on one source.
The practical translation is that HY4 Preview should be evaluated on your own traffic rather than on the table, and the cost structure makes that cheap to do — at a quarter of GPT-6 Sol's output rate, a week of shadow traffic against your real prompts costs less than a single API tier review. That is the correct form of due diligence for an under-measured model, and it is the form almost nobody actually does.
The open-weight question, which is the real dividing line
HY4 Preview is open-sourced under Apache 2.0 with a permissively licensed FP8 build alongside the base checkpoint. GPT-6 Sol is closed and has no licence file because there is nothing to license. Everything else in this comparison is a trade; this one is a category difference.
What that buys you is not mainly a cheaper bill — at $2.50 per million output the hosted API is already below most teams' cost of running a 770-billion-parameter model, and if cost were the whole argument the permissive licence would be beside the point. What it buys is a deployment that does not depend on anyone's endpoint: weights that cannot be rate-limited, a checkpoint that can be quantised to fit the hardware you have, a model that can be fine-tuned, and traffic that can stay inside a network boundary. GPT-6 Sol's answer to all four is that you should not need any of them, which is a real argument for a real set of teams and an irrelevant one for the teams that do.
The modality split and the licence line point the same way, which is worth noticing. A model that reads only text and ships under Apache 2.0 is built to be a component — something you put inside a pipeline you control, feeding it text you have already reconciled. A model that reads text, images and files and is reachable only over an API is built to be the entry point, taking the messy input directly. Those are different jobs, and a comparison that ranks them on one index is answering a question neither vendor asked.

Serving them together, and what the traffic number implies
Both models are behind one OrcaRouter key with 200+ others, and the routing shape this pairing suggests is the inverse of the obvious one. Put HY4 Preview first — it is the cheaper model by a factor of four on output, it is the one with the documented architecture, and on our routes it is the one carrying the load — and keep GPT-6 Sol behind it for the requests that need an image read, or that need to emit more than 64,000 tokens, or that fail a quality check. The composition is expressed in the routing DSL rather than in application code, so the boundary between the two is a rule you can edit without a deploy.
The speed figures deserve a caveat and then a conclusion. HY4 Preview is measured at roughly 54 output tokens per second in our rolling window against GPT-6 Sol's roughly 244, and both are shared-infrastructure observations rather than a controlled benchmark — the Tencent figure in particular reflects a model with a very large active parameter count, and 49 billion active per token is a lot of arithmetic per output token. So the cheaper model is the slower one by a wide margin, which is exactly the trade a batch pipeline is happy to make and an interactive one is not. Tencent's hosted build also received a performance optimisation on 7 September 2026 that the vendor says reduces reasoning turns and per-task token consumption at the same task quality — another vendor-stated figure, but one that predicts a lower bill rather than a faster stream.
Failover is what makes the pairing safe in either order. A 770-billion-parameter open-weight model is served by more than one infrastructure provider, and a route that can fall back means a degraded host is a retry rather than an outage. Our catalogue passes provider list prices through with no markup, which also means the currency-quoted vendor rate is not something you have to wait on: Tencent lists HY4 Preview at 6 yuan input and 18 yuan output per million, and the dollar figure our page shows is that rate converted, not a reseller's spread.

What this pair is actually for
Take Tencent HY4 Preview for volume, for text-only pipelines you control, for anything that needs to run on your own hardware, and for any workload where a wrong answer can be caught and retried — the price makes retrying the cheapest quality lever you have. Take GPT-6 Sol when the input is an image or a file, when one response has to carry more than 64,000 tokens, or when the answer is going to a person who will act on it without checking. Take both if your traffic contains some of each, which it almost certainly does.
Two closing checks. HY4 Preview is still labelled a preview and remains the newest Tencent model in our catalogue, so the name is accurate rather than stale — but a preview can change underneath you, and a permissively licensed checkpoint can be superseded by its own final release without warning. And GPT-6 Sol has already been replaced by GPT-6.1 Sol at identical short-context rates with cached input halved from $0.20 to $0.10, with OpenAI's own model page now pointing readers there. If the reason you are weighing Sol rather than the Tencent model is the eight-day-old integration, that stands. If it is capability, the successor is what you want to price.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
