
Atria Dawn vs Kimi K3: An Open-Weights Agent That Never Finishes, Against a 44-Point Reasoning Ceiling
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Before the benchmarks, sit with the difference in what these two models are for. Atria Dawn Preview is the Shanghai AI Laboratory's new open-weights agentic model — announced "now live" on September 14, built on the 744B MoE GLM-5.2 foundation, MIT-licensed, 256K context, text-only, downloadable in BF16 and FP8, and served at api.atria-asi.ai with no published price. Kimi K3 is Moonshot AI's 2.8-trillion-parameter flagship — released July 16, priced at $3.00/$15.00 per million tokens, 1M context, multimodal input, open weights under a Modified MIT license since July 27, and an independent Artificial Analysis Intelligence Index of 44 that makes it the second-highest-scoring open-weights model on the board. Both are open. Both are agentic. Both come from Chinese labs with international API arms. And they could hardly be aimed at more different jobs: Kimi K3 is a reasoning ceiling you can rent or download; Atria Dawn Preview is a loop that keeps working until something is verifiable.
The sourcing labels are the skeleton of this comparison. Kimi K3's index score, pricing, and context are independent facts — AA Index 44, $3/$15, 1M window, Modified MIT weights. Atria Dawn Preview's benchmark rows are vendor-reported from its model card, unreproduced by any neutral lab, and its international API has no published price. Where the card's grid pits the two against a common field, Atria Dawn Preview reports leading on discovery and tool-use suites (AutomationBench 53.8, BrowseComp 92.5, DeepSearchQA 96.0, BFCL v4 77.0, CyberGym 86.5) while trailing Kimi K3's lineage on coding rows (SWE-bench Pro 59.6, Terminal-Bench 2.1 78.3, JobBench 50.3). None of the Atria numbers is independently verified.
Two kinds of open weights
The superficial similarity — both open, both Chinese, both agentic — hides a genuine fork. Kimi K3 is Moonshot's attempt to prove that an open-weights model can hold the frontier on raw intelligence: 2.8T total parameters, a 1M context, native image and audio input, a Modified MIT license, and an independent index score of 44 that only GLM 5.3 tops among open models. You can rent it at $3/$15 or download it and serve it yourself, and either way you are buying a reasoning engine.
Atria Dawn Preview is not trying to be a reasoning engine in that sense. It is a 744B-total MoE — roughly a quarter of Kimi K3's size — tuned for a loop rather than for the index: analyse, design, use tools, write and run code, read the experimental result, recover from failure, iterate. The lab's four pillars (Discovery, Creation, Delivery, Cybersecurity) all reduce to that loop. It has a 256K context (a quarter of Kimi K3's 1M), is text-only where Kimi K3 takes images and audio, and publishes its strengths on long-horizon suites rather than single-shot reasoning. On paper it looks like a strictly worse model; on its own terms it is a different product aimed at a different failure.

• Price — Atria Dawn Preview unpublished (international API) vs Kimi K3 $3.00/$15.00 per 1M
• Context — Atria 256K text-only vs Kimi K3 1M, image+audio input
• Size — Atria 744B total MoE vs Kimi K3 2.8T total
• Independent record — none for Atria vs Kimi K3 AA Index 44, #2 open-weights
Why the 44-point wall matters less than the loop
If your workload is a hard reasoning question, this matchup is not close: Kimi K3's independently measured 44 is the highest open-weights ceiling on the board, and it is a 13-point-plus independent gap over anything Atria Dawn Preview can claim today. The premium is real and Kimi K3's $3/$15 is what it costs. But the index is mostly short, single-shot reasoning — the exact regime where a loop-first agent is trying to change the rules. Atria Dawn Preview's reported strengths are BrowseComp, DeepSearchQA, AutomationBench, CyberGym: long-horizon suites where the model gets to read, act, and correct itself over many turns. The vendor-reported split is the signature of a model tuned for iterative execution rather than for the leaderboard — and it is also unverified, which is why the first neutral score will matter more than any of this.
For a team already running Kimi K3, Atria Dawn Preview is not a replacement — its coding rows lose, its context is a quarter, its record is unproven. For a team whose bottleneck is the loop itself — the model giving up, asking, losing the thread on a long experiment — the newcomer is aimed exactly there, and its MIT weights mean you can find out on your own hardware without paying Moonshot's rate. Both sit behind one key on a routing platform like OrcaRouter, where Kimi K3's pass-through $3/$15 is billed with zero markup and automatic failover, and where Atria Dawn Preview will be routable the moment the Shanghai lab publishes a rate — so a head-to-head eval against your real workload costs setup time, not a second contract.


The verdict
Do not mistake a 44-point independent ceiling for the whole story, and do not mistake a vendor's grid for a verdict — one column is verifiable today, the other is a self-report. If you need proven open-weights reasoning at the frontier, Kimi K3 is the model; its 1M window, multimodality, and neutral scores are exactly why teams pay $3/$15. If your problem is a loop that never finishes, Atria Dawn Preview is a legitimate experiment — open, free to run, aimed precisely at the failure you are hitting — worth a controlled eval behind failover, and nothing more until a neutral lab actually runs it.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
