
Atria Dawn vs GLM 5.2: 동일한 기반, 정반대 방향으로 밀어붙이다
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040지능
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453지능77코딩
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241지능76코딩
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240지능72코딩
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153지능82코딩
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 100만 토큰당
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642지능72코딩
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 100만 토큰당
- z-aiZ.ai: GLM 5.32026-08-1845지능75코딩
- obsidianQwen3.8 27B2026-08-1534지능68코딩
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236지능69코딩
- grokSpaceXAI: Grok 4.62026-08-1244지능77코딩
- metaMeta: Muse Spark 1.22026-08-0540지능72코딩
- qwenQwen: Qwen3.8 Max2026-08-0340지능72코딩
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135지능69코딩
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 100만 토큰당
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451지능78코딩
- googleGoogle: Gemini 3.6 Flash2026-07-2134지능69코딩
이 대결에 관해 가장 유용한 단일 사실은 더 작은 모델의 자체 스펙 시트 안에 숨어 있습니다: Atria Dawn Preview는 9월 14일 라이브로 발표된 상하이 AI 연구소의 새 오픈 웨이트 에이전트형 모델로, 744B 파라미터 MoE GLM 5.2 기반 위에 구축되었습니다 — 바로 Z.ai가 6월에 출시하고 백만 토큰당 $1.40/$4.40에 오픈 웨이트 플래그십으로 판매해 온 그 모델입니다. 따라서 이는 신생 주자가 기존 강자를 앞지르려 애쓰는 사례가 아니라, 동일한 실리콘을 두 개의 서로 다른 연구소가 두 방향으로 후속 학습시킨 뒤 두 개의 서로 다른 제품으로 가격을 매긴 사례입니다. 어떤 벤치마크 수치보다도, 이 사실만으로 이미 이 비교가 실제로 무엇에 관한 것인지 알 수 있습니다: GLM 5.2의 추론 능력 중 얼마만큼을 에이전트 루프와 맞바꿀 의향이 있는지 — 그리고 둘 다 유지하는 데 얼마를 지불할 것인지.
여기서의 소싱은 중요한 지점에서 한쪽으로 치우쳐 있습니다. GLM 5.2의 독립적인 기록은 빈약합니다. GLM 5.2의 오픈 웨이트 상한은 8월 14일에 이를 대체한 GLM 5.3에 의해 설정되었고, 한때 GLM 5.2를 53으로 측정했던 AA Intelligence Index는 그 이후 재채점되고 명칭이 변경되었습니다. Atria Dawn Preview에 유리한 모든 벤치마크는 자사의 모델 카드에서 비롯된 벤더 보고 수치로, 어떤 중립 연구소도 재현하지 않았으며, api.atria-asi.ai에 있는 국제 API는 아직 공개된 가격이 없습니다. 모델 카드가 동일한 그리드에서 두 모델을 나란히 비교한 부분에서는 Atria Dawn Preview가 탐색과 도구 사용(BrowseComp 92.5, DeepSearchQA 96.0, BFCL v4 77.0)에서 앞서고, 순수 코딩 항목에서는 GLM 5.2의 계보에 뒤처집니다. 이 중 외부 연구소가 아직 확인한 것은 하나도 없습니다.
같은 베이스, 두 가지 후속 학습
Both models are 744B-total MoEs with the same 8-expert-per-token activation pattern, because they literally start from the same weights. The difference is what each lab did after that. Z.ai took GLM 5.2 and tuned it as a general open-weights flagship: text-only, 1M context, 128K output, MIT license, with a 40B-active class profile that made it one of the cheapest frontier-grade models to serve — the thing it is still remembered for is being the highest-scoring open-weights model on the independent index before GLM 5.3 arrived a week after Intern-S2.

The Shanghai AI Laboratory took that same base and tuned it into a research-loop agent. Atria Dawn Preview's four capability pillars — Discovery, Creation, Delivery, Cybersecurity — all describe the same loop: analyse a problem, design a solution, use tools, write and run code, read the experimental result, recover from failure, iterate. The lab is explicit that the result is text-only (image and PDF input get a 400), that the context window is 256K rather than 1M, and that the weights are MIT-licensed and downloadable today in BF16 (353 shards) or FP8 (177 shards). The base is GLM 5.2's; the product is not.
• Base — both 744B-total MoE on the GLM-5.2 foundation, 8 experts active per token
• Context — Atria Dawn Preview 256K text-only vs GLM 5.2 1M, 128K output
• License — both MIT; Atria BF16 + FP8 shards, GLM 5.2 the familiar 1.5TB-class download
• Price — Atria unpublished (international API) vs GLM 5.2 $1.40/$4.40 per 1M
What the vendor-reported grid actually shows
The model card compares Atria Dawn Preview against a field that includes GLM 5.3 rather than GLM 5.2 itself, which is a small editorial choice that flatters the open-weights column. Against that field, Atria Dawn Preview's reported strengths are discovery and tool use: AutomationBench 53.8, BrowseComp 92.5, DeepSearchQA 96.0, WideSearch 81.9, BFCL v4 77.0, CyberGym 86.5. Its reported weaknesses are the coding rows where a tuned generalist usually wins: SWE-bench Pro 59.6, Terminal-Bench 2.1 78.3, JobBench 50.3. Nothing here is independently verified, and every row is the lab's own harness on the lab's own prompts — but the split is consistent enough to read as a real design choice rather than noise. GLM 5.2, for its part, is the model whose June-era independent ceiling (AA Index 53, the then-highest open-weights score) was superseded by GLM 5.3 the month before Atria shipped.
Which one you should route to
If you need a million-token open-weights generalist for long-context coding and knowledge work, GLM 5.2 at $1.40/$4.40 is a proven, MIT-licensed workhorse that has been self-hosted and production-served for a quarter — and on a gateway like OrcaRouter, one API key and zero per-request markup later, it is available the moment you need it, with automatic failover if Z.ai's endpoint ever wobbles. If your job is an open-ended research loop — a task that needs the model to keep reading, running code, and iterating until something is verifiable — Atria Dawn Preview's tuning is aimed exactly there, but you are trading the 1M window, a known price, and an independent record for a 256K preview with no published rate and no neutral scores.


솔직한 결론은 갈림길이다. 이미 긴 컨텍스트 개방형 작업에 GLM 5.2를 의존하는 팀이라면, 프로덕션 경로를 Atria Dawn Preview로 전환하기 전에 독립 평가자들이 Atria Dawn Preview에 대해 보고하는 내용을 지켜보며 그대로 머물러야 한다. 병목이 루프 자체인 팀 — 모델이 실행하는 대신 멈춰서 묻는 문제 — 에게는 새로운 모델을 시도할 진짜 이유가 있다. 이상적으로는 GLM 5.2로의 페일오버 경로를 뒤에 두어, 프리뷰가 삐끗하더라도 실험에 비용이 들지 않게 하는 것이다. 같은 기반, 두 제품이며, 대부분의 독자에게 결정적인 열은 벤치마크 그리드가 아니라 그 그리드가 결코 보여주지 않는 컨텍스트 윈도와 가격이다.
이 글에서 비교한 모델1
이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트
