A hero title card for the comparison 'Tencent HY4 Preview vs Kimi K3'. A central scoreboard-style card reads 'Internal blind test, 4-point scale' with 'Tencent HY4 Preview 2.99' on the left and 'Kimi K3 2.94' on the right. Left card 'Tencent HY4 Preview' lists '770B / 49B active, open weights', '¥6 / ¥18 per 1M', 'Text-only preview'. Right card 'Kimi K3' lists '2.8T / 104B active, open weights', '$3.00 / $15.00 per 1M', 'Native vision, AA Index 57'. A footer reads 'Blind test run by Tencent with 163 engineers across 203 tasks; results vendor-reported.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Tencent HY4 Preview vs Kimi K3: The Blind Test Tencent Chose to Publish

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Tencent HY4 Preview vs Kimi K3 is the one matchup in this model's launch that Tencent itself chose to run. In its internal blind test — 163 engineers, 203 engineering tasks, scored on a four-point scale — HY4 Preview scored 2.99/4.00 against Kimi K3's 2.94, and Tencent's headline called it a win. The fine print of that win is worth reading slowly: HY4 Preview beat Kimi K3 on 51.2% of tasks, tied on 7.9%, and lost on 40.9%. A 10-point net win-rate is real, but it is a narrow edge against a model with a third of the total parameter count and less than half the active parameters — and the whole test was run by Tencent.

Kimi K3 is Moonshot's 2.8-trillion-parameter open-weight flagship, the largest open model ever released, and the reference point every Chinese lab has been measuring against since July 16. Tencent chose it as the opponent because it is the open-weight standard — which makes this article's job unusually concrete: read the one head-to-head the challenger volunteered, and decide what it is worth.

The benchmark Tencent chose to publish

The blind test was scored by Tencent's own engineers on Tencent's own engineering tasks, which is the first thing to discount. A company-curated task set is exactly the environment where a new model gets its best possible showing. Even granting that, the shape of the result is informative: 51.2% wins against 40.9% losses is a modest margin, and the 7.9% tie rate means the models are closely matched on most tasks. On the hard tasks, the ones where a frontier model separates itself, the gap is where it always is — and Tencent's headline, "slightly ahead of Kimi K3," is honestly worded, which is not what every lab publishes.

Scale isn't destiny

The parameter comparison makes the result either impressive or suspicious, depending on your priors. Kimi K3 is 2.8 trillion total parameters with 104 billion active — 896 experts, 16 active per token — and it was trained to reason always-on with exposed chain-of-thought. HY4 Preview is 770 billion total with 49 billion active, a third of the total scale and less than half the active compute. A smaller model scoring a narrow win in a blind test over a larger one is the direction of travel the whole open-weight field has been trending in — Qwen3.8-Max, at 2.4T, and GLM-5.2, at 753B, have been trading blows on agentic benchmarks for months — so the result is plausible rather than outlandish. It is also, crucially, unverified by anyone outside Tencent.

The scoreboard

A two-column scoreboard titled 'Tencent HY4 Preview vs Kimi K3 — the scoreboard'. Left column 'Tencent HY4 Preview': 'Total / active: 770B / 49B', 'Context: >1M tokens', 'Blind test vs K3: 2.99 vs 2.94 (vendor-run)', 'APEX-Agents: 37.1', 'Price: ¥6 / ¥18 per 1M', 'Vision: not in preview'. Right column 'Kimi K3': 'Total / active: 2.8T / 104B', 'Context: 1M tokens', 'AA Intelligence Index: 57', 'FrontierSWE: 81.2 (vendor-reported)', 'Price: $3.00 / $15.00 per 1M', 'Vision: native, image + video'. Footer reads 'HY4 Preview figures are Tencent-reported; Kimi K3 Index 57 from Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.

• Scale — HY4 Preview 770B total / 49B active vs Kimi K3 2.8T total / 104B active

• Context — HY4 Preview >1M tokens vs Kimi K3 1M tokens

• Price per 1M — HY4 Preview ¥6 in / ¥18 out (≈$0.85 / $2.50) vs Kimi K3 $3.00 in / $15.00 out, cache hits $0.30

• Modality — HY4 Preview text-only vs Kimi K3 native image and video input

• Reasoning — HY4 Preview no published mode vs Kimi K3 always-on reasoning with streamed chain-of-thought

• Independent score — HY4 Preview none vs Kimi K3 AA Intelligence Index 57, top-three globally at release

On the rows where both sides report a number, the models cluster. Tencent reports HY4 Preview at 37.1 on APEX-Agents, essentially level with Kimi K3's 37.2 — a near-tie the blind-test scoreboard would predict. HY4 Preview's Toolathlon-Verified 74.1 and DeepSWE 64.3 have no Kimi K3 equivalent in my hands, and Kimi K3's FrontierSWE 81.2, ProgramBench 77.8 and BrowseComp 91.2 have no HY4 Preview equivalent. The scoreboard is honest that the two cards barely overlap, which is itself a warning about treating either vendor's rows as a full description of the model.

Vision is the biggest real difference

For a developer choosing between the two today, the structural gap is not intelligence — it is input modality. Kimi K3 is natively multimodal: it takes image and video input through its MoonViT-V2 encoder and is among the best open models at vision tasks, ranking second globally on the vision arena. Tencent HY4 Preview is text-only in this preview, and Tencent has said vision must wait for the full release. Any workload that needs screenshots, UI understanding, or visual grounding is Kimi K3's by default, regardless of what the blind test says about engineering text. If your work is purely code, documents and terminal output, the gap does not matter; the moment a task involves looking at something, it decides the matchup.

The cost question

Per token, HY4 Preview is dramatically cheaper: roughly $0.85/$2.50 against Kimi K3's $3.00/$15.00, with cache hits at ¥0.3 (about four cents) against $0.30. But the two models have opposite token behaviors that narrow the gap. Kimi K3 reasons always-on and streams its chain-of-thought, which inflates output tokens on every request; reviewers have noted the model is slower — about 62 tokens per second — and token-heavy for agent loops. HY4 Preview, per Tencent's own release note, "tends toward overly long thinking and excessive self-verification" on complex tasks. Both burn extra output tokens; HY4 Preview just charges less for them. A fair comparison is cost-per-completed-task, and nobody outside Tencent has measured HY4 Preview's yet.

The verdict

Tencent HY4 Preview is the cheaper, smaller, unverified challenger that claims a narrow blind-test edge over the open-weight standard; Kimi K3 is the larger, verified, vision-capable incumbent that costs four to five times as much per token and has an independent Intelligence Index of 57. Neither model's benchmark card is fully independent — Kimi K3's headline scores are Moonshot-reported too, though its Index 57 is not. If your workload is multimodal, Kimi K3 is the only choice in this matchup. If it is text and engineering, HY4 Preview is the value bet with a real, if vendor-run, head-to-head to its name, and the cheap path is to route Kimi K3 on the production key — it is live on OrcaRouter at Moonshot's $3.00/$15.00 list price, passed through with no markup — while you run HY4 Preview through Tencent's own API on shadow workloads. When HY4 Preview reaches a router, the same key and the same failover rules will carry both, and Tencent's ¥6/¥18 rate will be passed through unchanged.

A screenshot of the Artificial Analysis Intelligence Index leaderboard (captured August 28, 2026) showing Claude Opus 5 (max) and Claude Opus 5 (xhigh) at the top of the ranking, followed by Claude Fable 5 (with fallback) and GPT-5.6 Sol (max). Kimi K3 is indexed further down and outside this capture. Tencent HY4 Preview does not appear: it launched today and has no independent index.A screenshot of the OrcaRouter model page for Kimi K3 (kimi/kimi-k3) showing the capability chips, a 1M-token context window, $3.00 per 1M input tokens and $15.00 per 1M output tokens, released July 15 2026 by MoonshotAI.

What to watch

• An independent rerun of the blind-test tasks. The 51.2% win rate is the single most testable claim in this launch, and a third-party harness would settle it in a week.

• Whether HY4 Preview ships vision before the full Hy4 release. That is what would convert this from a text-only value matchup into a genuine flagship fight.

• Cost-per-completed-task on standard agentic harnesses. Both models are token-heavy; whoever is cheaper per finished task wins the production argument.

• Kimi K3's response to the price pressure. If Moonshot cuts the $15 output rate, the "cheap challenger vs expensive standard" frame flips.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube