
GPT-6 Astra vs Tencent HY4 Preview: One Model Has an Audit Trail and the Other Has a Benchmark Table
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 56 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 348 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Tencent HY4 Preview and GPT-6 Astra are the two cheapest routes to near-frontier agentic coding available today if you read the vendors' own numbers, and they are the two least comparable models in that tier if you read anyone else's. Tencent HY4 Preview was released and open-sourced on 28 August 2026 under Apache 2.0, at $0.834 per million input tokens and $2.501 per million output tokens, with a 770-billion-parameter mixture-of-experts architecture, a context window that clears a million tokens, and an output ceiling of 64,000 tokens. GPT-6 Astra has been generally available since 3 September 2026 at $10.00 and $50.00, with a 1,050,000-token window and 128,000 maximum output. Astra's strengths are backed by eight published third-party evaluations; HY4 Preview's are backed by Tencent's own terminal, tool-calling and blind-evaluation runs, and by nothing else. That asymmetry does not mean the cheaper model is weaker. It means one of these two comparisons can be checked and the other has to be believed, and the practical decisions differ accordingly.
What the Apache 2.0 release actually gives you
The licence is the part of this matchup that a price table cannot express. Tencent HY4 Preview is not a hosted teaser with open-weight branding — the weights are downloadable from Hugging Face, GitHub, ModelScope and GitCode, which means the model survives any decision Tencent makes about its own pricing, its own endpoint, or its own continued interest in serving it.
• Architecture — 770B total parameters, 49B active per token, 78 layers, 256 routed experts plus one shared expert, and a native multi-token-prediction layer for speculative decodingbr /> • Serving characteristics — 54.6 output tokens per second and a 6.26-second median time to first token, measured across OrcaRouter's routes over a rolling seven-day windowbr /> • Context and ceiling — a 1,048,576-token window against a 64,000-token maximum outputbr /> • Input surface — text only; no image, no file, no audiobr /> • Reasoning control — three levels, high, low and none, with high as the default, plus a documented sampling recommendation of temperature 0.9 and top_p 1.0br /> • Reliability — a 2.8% error rate over the same seven-day window, against 10.3% on the Astra route
That last pair of numbers is the most actionable thing on this page, and it is the kind of figure no vendor publishes about its own model. Astra's independent benchmark scores are higher and its route has been failing roughly one request in ten over the past week, while a model with no independent scores at all has been failing one in thirty-five. Neither figure is a statement about the models themselves — error rates measure the route, the rate limits and the upstream, not the weights — but they are the numbers a production system actually experiences.

Where Tencent's numbers can be checked against someone else's
This is, as far as published evidence goes, a genuinely unusual opportunity: Astra and HY4 Preview both have a Terminal-Bench 2.1 figure, and only one of them is vendor-reported.
• Terminal-Bench 2.1, driving a real terminal — GPT-6 Astra 88.4, independently evaluated; Tencent HY4 Preview 85.4, vendor-reportedbr /> • Tool-calling — Astra 41.4 on τ²-Banking, independently evaluated; HY4 Preview 74.1 on Toolathlon-Verified, vendor-reported, on a different testbr /> • Software engineering — HY4 Preview 64.3 on DeepSWE, up from 28.0 on its predecessor Hy3, vendor-reportedbr /> • Agent tasks — HY4 Preview 37.1 pass@1 on APEX-Agents, vendor-reportedbr /> • Human preference — a 2.99 out of 4.00 average across 203 engineering tasks scored by 163 experts, vendor-run, against GLM-5.3 at 2.92 and Kimi K3 at 2.94br /> • Independent index — GPT-6 Astra 54.7 Intelligence Index and 96.1 GPQA Diamond, third-party; no equivalent index exists for HY4 Preview
The Terminal-Bench row is the one worth dwelling on. A 3-point gap between an independent 88.4 and a vendor-reported 85.4 is well inside the range that a different harness, a different scaffold or a different sampling temperature can produce, which means the honest reading is not "Astra is three points better" but "the cheapest open-weight model in this tier is claiming parity, and the claim is at least plausible." Tencent's own framing is careful on this point: it describes DeepSWE as "gradually narrowing the gap" rather than closing it, and its one clean parity claim — the Terminal-Bench tie with a top closed model from a different vendor — is exactly the claim most in need of a second measurement.
There is also a change worth knowing about, because it moves the economics rather than the scores. On 7 September 2026 Tencent pushed an optimisation into the live preview build that it says cuts reasoning turns and per-task token consumption on complex work. Because the model bills per token, fewer tokens per finished task lowers the cost of the task even though the per-token rate has not moved. The size of that saving is Tencent's own measurement and nobody outside Tencent has reproduced it — but the mechanism is real and it is the reason a preview released in August can be materially cheaper to operate in October than its launch write-up suggested.
The 64,000-token ceiling is the spec that decides deployments
Halfway down the spec sheet is the constraint that bites first, and it is not quality. HY4 Preview's maximum output is 64,000 tokens against Astra's 128,000, on a model whose stated purpose is coding agents and long tool-use chains. A generated file, a multi-file patch, a long structured report — these are the outputs that hit a ceiling, and on an agent run the failure is not a slightly worse answer, it is a truncated one that the surrounding pipeline has to detect and handle.
Both models take roughly a million tokens of input, so the document-reading case is genuinely a tie. The difference is entirely on the generation side, and it is a factor of two rather than a rounding error. Add the input-surface difference — Astra accepts images and files, HY4 Preview is text-only — and the picture that emerges is a clean division of labour rather than a ranking: the open-weight model for text-in, text-out agent loops where the deliverable is a tool call or a diff, and the closed model for anything that has to look at something or write something long.
What 20× the output rate buys, line by line
• Input price — Tencent HY4 Preview $0.834 per 1M vs GPT-6 Astra $10.00, about 12×br /> • Output price — HY4 Preview $2.501 per 1M vs Astra $50.00, about 20×br /> • Cached input — HY4 Preview $0.042 per 1M vs Astra $1.00, about 24×br /> • Long-request step — Astra reprices the whole request above 272,000 input tokens to $20.00 and $75.00; HY4 Preview's card has no equivalent stepbr /> • Effective input rate on a 400,000-token prompt — HY4 Preview unchanged at $0.834 vs Astra $20.00, about 24×br /> • Cost per million tokens of ordinary agent loop, 4:1 input to output — HY4 Preview roughly $1.17 vs Astra roughly $18.00br /> • Cost of a 100-million-token month at the same ratio — HY4 Preview roughly $117 vs Astra roughly $1,800
The cached-input line is the one that scales worst for the expensive side, because agent loops re-send the same system prompt and conversation prefix on every turn. At $0.042 against $1.00, a long loop's cheapest tokens are 24 times cheaper on the Tencent model, which is precisely the workload where a small per-token difference compounds fastest. Per-task cost, the metric that would settle this cleanly, is not published by Tencent at all — so the only way to get the number that matters is to run both models through your own task set and read the usage object.

What a router adds here, with the error rates in hand
Both models are in the OrcaRouter catalogue — tencent/hy4-preview and openai/gpt-6-astra — behind one OpenAI-compatible endpoint, each at the provider's own list rate passed through with 0% markup. That last detail matters more on this pair than on most, because the Tencent model's list rate is published in yuan: the dollar figures above are the converted rate you actually see, not a reseller spread, and if Tencent moves the price, the move is live here the same day rather than at the next contract renewal.
The routing DSL is where the two models stop being alternatives. The natural rule for this pair is escalation on output: send the loop to the cheap model, and when a response comes back truncated against the 64,000-token ceiling or fails your schema check, retry that step against openai/gpt-6-astra, which has twice the output room. Automatic failover is the other half of the argument, and it is not theoretical when one of the two routes has been erroring on one request in ten over the past week — a long agent run pinned to a single model dies with its accumulated context, and neither of these models is cheap enough to re-run from the top by accident.

What would change this comparison
Three things, in order of how much they would move it. An independent evaluation of HY4 Preview — the model has been public for six weeks, has no third-party index, no reproduced Terminal-Bench number and no per-task cost figure, and the single vendor-run blind evaluation is the only human-preference data in existence. A published cost per completed task for both models, which would replace every ratio in this article with one number. And a Tencent decision on third-party inference, because Apache 2.0 weights can be served by anyone, and the moment competing hosts stand up HY4 Preview the price on the cheap side of this comparison becomes a market rather than a list.
Until then, the usable summary is narrower than either vendor's launch post and more honest than a scoreboard: GPT-6 Astra is the model whose capability claims have been checked, with twice the output ceiling and a route that has been failing one request in ten. Tencent HY4 Preview is the model whose capability claims have not been checked, with a published architecture, an open licence, a third of the price and a much lower error rate over the same window. If your work is text-in, text-out and fits inside 64,000 generated tokens, the cheaper model is not a compromise — it is the right answer, and the audit trail is the only thing you are giving up.
The natural rule for this pair is escalation on output: send the loop to the cheap model, and when a response comes back truncated against the 64,000-token ceiling or fails your schema check, retry that step against openai/gpt-6-astra , which has twice the output room.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
