Qwen 3.8 vs Kimi K3: The Two Chinese Open-Weight Giants, Head to Head
Guides & Insights

Qwen 3.8 vs Kimi K3: The Two Chinese Open-Weight Giants, Head to Head

Author

jinhao song

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Of all the models Qwen 3.8 Max gets compared against, Kimi K3 from Moonshot is the most natural rival: both are enormous Chinese Mixture-of-Experts models, both are open-weight-class, both are famously verbose, thorough, and slow. They are, in effect, the two biggest giants of the Chinese open-weight world — and unusually, we don't have to guess how they stack up, because someone ran them against each other on the same task. That makes this the richest comparison in the series: not vendor claims versus audited numbers, but one genuine head-to-head.

Before the details, a note for builders — the honest way to settle a rivalry this close is to run both on your own workload and watch the tradeoff between thoroughness and speed. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint, so you can pit Qwen 3.8 against Kimi K3 on the same prompts without wiring up two SDKs.

TL;DR verdict. In the one direct head-to-head we have (Trilogy AI, an architecture-understanding task), Kimi K3 scored 83/100 and Qwen 3.8 Max 80/100 — a narrow 3-point edge to Kimi. But Qwen was the more *thorough* one: 354 repo citations vs 274, 22 gateway requests vs 53, and zero failed tool calls vs two, at the cost of higher latency and more tokens. Both cleared 90% cache hit and reached the same core architectural decision. Today Kimi K3 is the safer open-weight pick — it has an audited score (57.1, #3 on Artificial Analysis), the Arena frontend crown, and weights that are essentially here. Qwen 3.8 is the one to watch: bigger in ambition, more exhaustive in exploration, but unscored on public boards with weights only "coming soon."

Key takeaways

•  Qwen 3.8 Max is a 2.4T-parameter MoE preview; Kimi K3 is a 2.8T MoE with 1M context and native vision — meaning Qwen is roughly 400B parameters smaller, a reminder that bigger param count doesn't equal better.

•  In the only direct test (Trilogy AI), Kimi K3 edged Qwen 80 to 83 overall — but Qwen made more repo citations (354 vs 274), fewer gateway requests (22 vs 53), and had zero failed tool calls (vs two).

•  Kimi K3 has an audited AA Intelligence Index of 57.1 (#3) — behind Fable 5 and GPT-5.6 Sol, ahead of everything else. Qwen 3.8 has no independent benchmark at all.

•  Kimi K3 also holds the LMArena frontend-code #1 spot (1679); Qwen 3.8 has not been scored on public leaderboards.

•  Openness is near-parity but timing differs: Kimi K3's weights are open/promised and essentially ready; Qwen 3.8's are "coming soon" with no date, license, or repo yet.

Qwen 3.8 figures are vendor claims or early, uncontrolled hands-on impressions — not independently audited. Kimi K3's index and price figures are attributed to Artificial Analysis and may differ from Moonshot's own reporting. The Trilogy AI head-to-head is a single task, not a full benchmark suite. Qwen 3.8's preview pricing is a temporary discount; verify rates before budgeting.

The specs and price, side by side

Here is the hard data, each competitor figure attributed to its source.

•  Maker / status — Qwen 3.8 Max: Alibaba; preview (July 19, 2026); Kimi K3: Moonshot; open-weight release

•  Architecture — Qwen 3.8 Max: ~2.4T total, sparse MoE (active params undisclosed); Kimi K3: 2.8T MoE, native vision (source: Artificial Analysis)

•  Context window — Qwen 3.8 Max: Reported ~1M tokens (not formally confirmed); Kimi K3: 1M tokens (source: Artificial Analysis)

•  AA Intelligence Index — Qwen 3.8 Max: None yet — unscored; Kimi K3: 57.1 — #3 (Artificial Analysis)

•  Arena / leaderboard — Qwen 3.8 Max: Not scored publicly; Kimi K3: Frontend-code #1, 1679 (LMArena)

•  Pricing (in / out) — Qwen 3.8 Max: Preview only (~90% off; per-token API unpublished); Kimi K3: $3 / $15 per 1M (AA blended ~$2.31)

•  Open weights — Qwen 3.8 Max: Promised "soon" (not released); Kimi K3: Open-weight; weights ready/near

•  Speed — Qwen 3.8 Max: Slowest model tested (hands-on); Kimi K3: Verbose and slow (~34s TTFT, per AA)

Two things frame the whole comparison. First, the empty "AA Intelligence Index" cell for Qwen 3.8 is the story: Kimi K3's 57.1 is a neutral referee's verdict that places it #3 in the world, behind only Fable 5 and GPT-5.6 Sol; Qwen 3.8's standing rests on vendor framing and a handful of hands-on runs. Second, notice the size inversion — Qwen 3.8 (2.4T) is actually about 400B parameters smaller than Kimi K3 (2.8T). Alibaba's model is the one people describe as "bigger in ambition," but on raw parameter count Kimi is the larger giant, and it's the one with the audited score to back it up. Both are famously verbose and slow, so latency is a wash; the real separators are proof and weight availability, and today both favor Kimi.

The one real head-to-head we have

This is the rare matchup where we don't have to infer. Trilogy AI ran both models on a single architecture-understanding task — reading and reasoning about a real codebase — and published the side-by-side. Kimi K3 won on the headline score, but the underlying behavior tells a more interesting story.

•  Overall score — Qwen 3.8 Max: 80 / 100; Kimi K3: 83 / 100

•  Repo citations — Qwen 3.8 Max: 354; Kimi K3: 274

•  Gateway requests — Qwen 3.8 Max: 22; Kimi K3: 53

•  Failed tool calls — Qwen 3.8 Max: 0; Kimi K3: 2

•  Tool-use rating — Qwen 3.8 Max: 9 / 10; Kimi K3: 8 / 10

•  Cache hit rate — Qwen 3.8 Max: >90%; Kimi K3: >90%

Read the columns, not just the top row. Kimi K3 took the 3-point win, but Qwen 3.8 was the more meticulous worker: it cited 80 more places in the repo (354 vs 274), reached its conclusion in far fewer gateway requests (22 vs 53), made zero failed tool calls to Kimi's two, and edged Kimi on the tool-use rating (9 vs 8). The tradeoff is exactly what you'd expect from Qwen's character elsewhere — it explores exhaustively, which costs higher latency and more total tokens. Crucially, both models reached the same core architectural decision, so this isn't one being right and one wrong; it's Kimi getting to a slightly higher-graded answer faster, and Qwen getting there with more evidence and cleaner tool discipline. If your work rewards exhaustive, well-cited exploration, Qwen's profile is attractive; if you want the proven, faster-to-the-answer option, Kimi took the round.

FAQ

Is Qwen 3.8 better than Kimi K3?

Not on the evidence that exists today. In the one direct head-to-head (Trilogy AI), Kimi K3 scored 83/100 to Qwen's 80/100, and Kimi has an audited Artificial Analysis Index of 57.1 (#3) plus the Arena frontend-code crown, while Qwen 3.8 has no independent score at all. Qwen's edge was thoroughness — more repo citations (354 vs 274), fewer gateway requests, and zero failed tool calls — but that's a style advantage, not a proven capability lead.

Qwen 3.8 is bigger — does that make it better?

No, and it's actually the other way around on size: Qwen 3.8 is ~2.4T parameters versus Kimi K3's ~2.8T, so Kimi is the larger of the two by roughly 400B. Bigger param count doesn't equal better anyway — Kimi is both larger and higher-scored, which is a clean reminder that architecture, training, and tuning matter more than raw parameter totals.

How does the pricing compare?

Kimi K3 is a published, durable $3 / $15 per 1M tokens (Artificial Analysis blends it to about $2.31). Qwen 3.8's preview is heavily discounted (~90% off) but has no published per-token API price and the discount is temporary, so a stable price comparison isn't possible yet. For budgeting today, Kimi is the one with a real number.

Which is the better open-weight choice?

Both are open-weight-class, but timing differs. Kimi K3's weights are open and essentially ready; Qwen 3.8's are promised "coming soon" with no date, license, or repo published. If open weights are the reason you're choosing one, Kimi is the pick you can act on now.

Which should I use today?

Kimi K3 is the safer default — it's scored (#3), weights are ready, and it took the head-to-head. Reach for Qwen 3.8 when you value exhaustive, well-cited exploration and clean tool use over speed, and treat it as the one to watch for when its weights land and an independent score arrives.

Bottom line

Qwen 3.8 vs Kimi K3 is the closest thing this series has to a fair fight: two Chinese open-weight giants, both verbose, both thorough, both slow. Kimi K3 wins the round today because it wins the things that decide adoption — an audited #3 ranking, the Arena frontend crown, a published price, and weights that are ready — and it edged Qwen 80 to 83 on the one real head-to-head. But that head-to-head also shows why Qwen 3.8 is worth watching: it was the more meticulous model, citing more of the codebase, using fewer requests, and making zero failed tool calls. If exhaustive exploration is what your work rewards, and if Alibaba ships the weights and an independent score lands near the top, this rivalry could flip. Until then, Kimi K3 is the safer open-weight pick, and Qwen 3.8 is the challenger to test on your own prompts.


© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube