
Muse Spark 1.2 vs GLM 5.2: The Closed Coder vs the Open Generalist
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
Two 1,000,000-token context models, priced within fifteen cents of each other per million tokens, both aimed at coding-heavy agent work — and they share almost nothing else. Muse Spark 1.2, which Meta shipped on August 5, 2026 alongside a co-trained terminal agent called Muse Code, is closed-weight, cheaper per token, scores 57 on the current Artificial Analysis Intelligence Index, and cannot reply without reasoning first. GLM 5.2, Z.ai's flagship open-weights model from mid-June, is MIT-licensed, self-hostable, roughly five times faster to the first token, and still carries four and a half times more real traffic through our gateway — 213 million tokens over the last seven days against Muse Spark 1.2's 46 million. The higher-scoring, cheaper-per-token model is the one with less traffic. The open one is the one people actually route.
The honest version of that fact is what this article is about. Scoring and pricing decide less than how a model fits your pipeline, and these two make opposite choices on every axis that determines fit. Where independent numbers exist, they come from Artificial Analysis; where they come from Meta or Z.ai, they are labeled vendor-reported; the latency and traffic figures are OrcaRouter's own seven-day production telemetry.
The spec contrast, one dimension at a time
• Price — Muse Spark 1.2 at $1.25 in / $4.25 out per million tokens, $0.15 on a cache hit; GLM 5.2 at $1.40 in / $4.40 out, $0.26 cached. Meta is cheaper on every row.
• Context — both 1,048,576 tokens. Max output: Meta documents a 131,072-token ceiling for Muse Spark 1.2; GLM 5.2 declares 128,000.
• Reasoning — reasoning is mandatory on Muse Spark 1.2 across five effort levels (minimal to xhigh, default medium); GLM 5.2 runs hybrid reasoning controlled by reasoning_effort at high or max, with deep reasoning on by default.
• Inputs — Muse Spark 1.2 advertises text, image, video, audio and PDF input; GLM 5.2 is text-in, text-out.
• Weights — Muse Spark 1.2 is closed, with no repository and no self-host path; GLM 5.2 ships MIT-licensed open weights, a 753B-parameter Mixture-of-Experts with roughly 40B active per token.
• Age — Muse Spark 1.2 is three days old as this is written; GLM 5.2 has been in production since June 13, with eight weeks of field experience and a whole ecosystem of hosts and tools built around it.
The index gap is four points. The version trap is real.
Artificial Analysis's current Intelligence Index (v4.1.1) scores Muse Spark 1.2 at 57, ranked #12 of 185, and GLM 5.2 at 53 — the highest open-weights score on the board, but four points behind. Most coverage still quotes 54 for Muse Spark 1.2 because that is what the launch-day evaluation reported; the v4.1.1 rescore added 2.7 points to it, the largest increase of any model in that update. If you see "54 vs 51" or "54 vs 53" anywhere, check which index version each number is on before treating the gap as real.
Worth saying what the four-point gap does and does not mean. It means that on the nine-benchmark composite, Meta's closed model lands ahead of Z.ai's open one at comparable effort. It does not mean Muse Spark 1.2 beats GLM 5.2 at everything, because the two models reach their scores by different routes. GLM 5.2 is a math and reasoning outlier — AIME 2026 at 99.2 — but it is famously verbose and its weak spot is deep repository-scale work. Muse Spark 1.2 is the opposite shape: strong on terminal coding and multi-file tasks, and dramatically cheaper per task to evaluate, because it burns far fewer tokens while thinking.

Where the benchmarks actually diverge
On Terminal-Bench 2.1 the two are closer than the index gap suggests, and the sourcing matters more than usual. On the same Artificial Analysis harness, Muse Spark 1.2 scores 80.1 and GLM 5.2 scores 77.9. Meta claims 82.9 for Muse Spark 1.2 on its own harness; Z.ai's best-reported figure for GLM 5.2 is 82.7, which made it the first open-weight model above 80 on that benchmark. Same benchmark, four different numbers — harness pairing moves these scores by several points in both directions, so treat any single-terminal-bench claim as a range.
The divergence to pay attention to is DeepSWE 1.1, the benchmark that stresses deep, multi-file, repository-scale reasoning over a long agent loop. Meta reports 59.3 for Muse Spark 1.2 — a vendor figure, no third party has reproduced it yet. GLM 5.2's published DeepSWE is 46.2, and its predecessor GLM-5.1 was at 18.0, so this is the dimension Z.ai has improved most and still trails on. If your workload is "change fifty files coherently," that gap is the whole story; if your workload is terminal commands and scaffolding, it barely registers.
One more finding that upends the obvious framing. The independent Vals AI suite, which runs agent workloads per domain, puts Muse Spark 1.2 fifth overall at 71.88% — and first at Finance Agent, first at TaxEval, and first on Harvey's Legal Agent benchmark, against 44, 136 and 31 models respectively — while Terminal-Bench 2.1 is only its fourteenth-best domain out of fifty. Meta sold this model as a coder, and an independent evaluator found it is strongest at finance, tax and legal document agents. If your team writes contracts or reconciles ledgers, that is the single most useful number in this article.
The open-weights question is the real fork
Everything above is about capability. The decision is mostly about ownership, and here the two could not be further apart.
GLM 5.2 is MIT-licensed open weights. That changes the negotiation, not just the bill: you can self-host it if a workload is sensitive or the volume is high; you can move it between providers without re-integrating, because the weights are yours; and any host that routes it has to compete on price and latency, which is why the open-weight line gets cheaper over time. The 753B MoE is not a laptop model — most teams will still rent it rather than run it — but the option to leave, or to run the same weights in-house for a fixed cost, puts a floor under what anyone can charge you.
Muse Spark 1.2 buys the opposite bundle. The weights are closed and the only first-party route is Meta's Model API, which we now route as well. In exchange you get a co-trained product: Muse Code, the terminal agent that shipped with the model, was tuned on rejection-sampled harness trajectories and runs persistent, restart-safe subagents on a replay-exact event-log runtime, with bundled /plan, /grill and /goal skills. That is a genuinely different thing from "a good model you wire into your own harness" — it is an agent that arrives working. The trade-off is that it only works Meta's way, in Meta's tool, on macOS or Linux, with no Windows client, no MCP and no IDE plugins yet.

Price per token, price per task
Per token, Muse Spark 1.2 is cheaper on input and output, and it stays cheaper per task because it is also the less verbose model. Artificial Analysis measured GLM 5.2 burning 141 million output tokens, 95% of them reasoning tokens, just to run the Intelligence Index at launch — the most verbose leading model on the board. Muse Spark 1.2 burned 95 million on the same exercise at $0.40 per task. More tokens at a higher rate means GLM 5.2's per-task cost compounds in a way its list price hides, and this is the correct way to compare reasoning models: price the completed task, not the million-token rate.
There is a third price that only one of these models has. Muse Spark 1.2 ships a contributor tier at $0.10 in / $0.20 out — an order of magnitude under the standard tier, and under GLM 5.2's list price by a similar margin. The cost of admission is permission for Meta to train future models on your traffic, plus a 60-requests-per-minute cap that rules out production fan-out. Treat it as a legal review with a discount attached, not a cheaper SKU; for a team that qualifies and works under the cap, it makes the price comparison stop being close.
Latency: the two models invert
If the index gap favors Meta, the latency profile is where GLM 5.2 wins decisively on feel. Across the last seven days of real traffic on OrcaRouter, Muse Spark 1.2 shows a p50 time to first token of 8.13 seconds and a p95 of 10.00 seconds, then blazes at 576 output tokens per second with a zero error rate. GLM 5.2 shows a p50 of 1.55 seconds — more than five times faster to the first token — but trickles at 78.5 output tokens per second with a 0.16% error rate. In the lab at maximum effort the gap widens: Artificial Analysis measured Muse Spark 1.2's time to first token at 26 seconds at xhigh, its most expensive reasoning setting.
So one model makes you wait and then delivers fast; the other starts immediately and takes its time. For anything a human is watching — a chat turn, an interactive terminal session — GLM 5.2 feels better, and that is why it dominates the interactive traffic through our gateway. For a long generation, a big patch or a document draft, Muse Spark 1.2 will finish first once it starts, because its output rate is seven times GLM 5.2's. Measure the wrong number and you will pick the wrong model for the wrong reason.
Trying both without a second contract
Both models are in the OrcaRouter catalog at provider list price — Muse Spark 1.2 at $1.25 and $4.25, GLM 5.2 at $1.40 and $4.40 — because OrcaRouter passes provider rates through at 0% markup. That makes the pricing sections above the invoice, not the sticker, and it means switching between the two is a one-line change in the model string against the same key rather than a new integration. If a vendor cuts a price, the cut lands on our side the same day.
The routing layer earns its keep here precisely because the two models are complements. Muse Spark 1.2's weakness is a slow first token and a high price at scale; GLM 5.2's weakness is verbosity and deep-repo reasoning. A routing rule that sends the planning pass to Muse Spark 1.2 and the parallel worker fan-out to GLM 5.2 — or fails over to GLM 5.2 when Muse Code's endpoint is slow — costs nothing extra to build, because both sit behind one endpoint. And for a model three days old, automatic failover is how you try it without betting a production path on it.

Who should pick which
Choose Muse Spark 1.2 when the unit of work is a large, coherent artifact and someone can wait a few seconds for it to start: multi-file refactors, whole-repository generation, long agent loops, and — per Vals AI — finance, tax and legal document work. The DeepSWE lead, the high output rate and the contributor tier all point the same way. You are accepting closed weights, Meta's tooling and a slower first token, and you should read Meta's 82.9 and 59.3 as claims, not facts.
Choose GLM 5.2 when ownership and feel matter more than the composite: open weights you can self-host or move, a fast first token for interactive work, an ecosystem of harnesses and tools that already works with it, and the strongest open-weight general capability available. Accept the verbosity, the 46.2 DeepSWE, and a list price that is higher per token even before the token-count difference.
Most teams will find the honest answer is neither one alone. The open one is the workhorse you route most traffic through; the closed one is the specialist you point at the deep tasks and the sensitive documents. Four index points apart on a composite, a world apart on who owns the weights — that is the choice, and it is a lot clearer than the scores.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
