
Muse Spark 1.2 vs Muse Spark 1.1: Should You Take the Upgrade?
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- qwenNEWQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2033 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
- z-aiZ.ai: GLM 5.22026-06-1653Intelligence69Coding60Math
Meta replaced Muse Spark 1.1 with Muse Spark 1.2 on August 5, 2026 without changing the price, the context window, the modality list, the supported parameters or the reasoning dial. The migration therefore looks free — same model string family, same billing, nothing to renegotiate. It is not free. Independent measurement puts the new version's time to first token at 26.12 seconds against the old version's 2.90, and its sustained output at 165 tokens per second against 213.5. You are buying three points of measured intelligence with roughly nine times the wait before anything comes back.
Whether that is worth doing depends almost entirely on how long your tasks run. Here is what actually differs, what stayed identical, and what nobody can tell you yet.
The decision in four lines
• Intelligence Index: 51 → 54, independently measured by Artificial Analysis at each version's xhigh setting.
• Time to first token: 2.90s → 26.12s, same source, same conditions.
• Cost to run the identical benchmark suite: $548.07 → $637.85, on identical per-token pricing. The extra 16% is pure token burn.
• Everything a procurement form asks about: unchanged.

What is genuinely identical
This is worth enumerating because it is the reason a migration looks trivial on paper, and because a difference that does not exist is one you do not need to test for.
• Price — $1.25 per million input tokens, $4.25 per million output, $0.15 per million on a cache hit, $2.50 per thousand built-in web searches. Every one of those figures is the same on both versions.
• Context — 1,048,576 tokens on both, with no long-context surcharge tier on either.
• Modalities — text, images, video, audio and PDF in; text out. Both listings advertise the same five input types.
• Reasoning dial — mandatory on both, five levels (minimal, low, medium, high, xhigh), default medium on both. There is no setting on either version that turns thinking off.
• Parameters — tools, tool_choice, structured_outputs, response_format, reasoning_effort, include_reasoning, max_tokens, temperature, top_p, top_k, repetition_penalty. Identical lists.
• Serving — one provider, Meta itself, on both. No third-party hosts, no quantization variants.
• Blended price — Artificial Analysis puts both at $0.78 per million tokens on its blended weighting, which is what you would expect from identical rate cards.
If you were hoping the upgrade came with more context, a completion-length guarantee, or a cheaper cache, it did not. This is a weights-and-post-training release with a rate card copy-pasted from the previous one.
What actually moved
Artificial Analysis is currently the only independent party to have evaluated both versions, and it evaluated both at their highest reasoning effort, so the comparison is like-for-like.
• Intelligence Index — 51 for 1.1 (rank #22 of 185) to 54 for 1.2 (rank #13). Nine places on a board where the class median is 32, so the gain buys real position, not just a decimal.
• Time to first token — 2.90s to 26.12s. This is the headline regression and it is not marginal.
• Output speed — 213.5 to 165.0 tokens per second. Still ranked #16 of 185 and still described by Artificial Analysis as "notably fast," but 23% slower than the model it replaces.
• Verbosity — 94M output tokens to get through the index, versus 95M. Effectively unchanged, and both about 45% above the 65M median.
• Total eval spend — $548.07 to $637.85. Since verbosity barely moved and prices did not move at all, that 16% increase is coming from somewhere other than the visible output: longer internal reasoning traces, more retries, or more tool turns per task.
The pattern is coherent. Meta did not make the model cheaper, faster or bigger. It made the model deliberate longer before committing, and the extra deliberation shows up as latency, as slightly slower streaming, and as a bill 16% higher for identical work. Three index points is what that bought.
How bad the latency really is
The 26.12-second figure will get quoted out of context, so treat it correctly: it is measured at xhigh, the most expensive of five effort levels, on a benchmark suite designed to be hard. The API defaults to medium.
For a sense of what the default actually feels like, Muse Spark 1.1 on OrcaRouter — serving real callers at whatever effort they request — shows a p50 time to first token of 1.84 seconds and a p95 of 6.00 seconds across a rolling week. That is production data at production effort levels, and it is more than an order of magnitude below the benchmark figure. Version 1.2 has no production telemetry yet.

So the honest reading is not "1.2 takes 26 seconds to respond." It is: at the setting where 1.2 earns its three extra points, the first token costs you 26 seconds, and on the same setting the old model cost you three. If you were already running at xhigh — which is exactly the configuration where the intelligence gain exists — that is the number you inherit. If you run at medium or below, you will see something much shorter, and you will also see less of the improvement.
That coupling is the real trap in this upgrade. You cannot take the intelligence without the deliberation, because the deliberation is where the intelligence comes from.
The repositioning behind the numbers
Meta published no launch post for 1.2 — the Muse Spark announcement page still describes 1.1 — but it rewrote the product description, and the rewrite explains the trade.
Muse Spark 1.1 was sold as a multimodal reasoning model with an unusually broad input surface, aimed at "multimodal agents, deep research over mixed media, and long-context analysis": long reports, media libraries, recordings and video alongside text.
Muse Spark 1.2's description drops nearly all of that and names software engineering instead — multi-file refactors, extended debugging sessions, whole-repository generation — then adds an explicit multi-agent claim: the model can act as the primary agent that gathers context, plans and delegates, or as a subagent executing in parallel beneath one. It also claims prompt caching can cut effective cost 60–80% depending on how much context repeats between calls. That last figure is Meta's estimate rather than a measured result, though the $0.15 cache rate makes it arithmetically credible.
Read the latency regression against that repositioning and it stops looking like a mistake. Nobody refactoring a dozen files cares about 26 seconds of planning. Meta chose a customer, and it is not the one 1.1 was sold to.
What 1.1 still has that 1.2 does not
Four weeks of existence is not enough to accumulate a track record, and in three concrete ways the older model is currently the better-documented product.
• An arena record. Muse Spark 1.1 has a substantial Design Arena history: 1334 Elo at rank 5 in game development, 1334 at rank 7 in UI components, 1320 at rank 3 in ASCII art, 1318 in data visualisation, 1309 in general code categories, plus a full agents-arena ladder covering web apps, HTML slides and Godot game development. Muse Spark 1.2 has no entries at all. On the axis Meta is now marketing hardest, there is zero head-to-head data for the new model.
• Sub-index scores. Artificial Analysis publishes a coding index of 71.3 and an agentic index of 37.5 for 1.1. There is no published counterpart for 1.2. Since the agentic score was 1.1's weakest dimension by a wide margin, and agentic capability is exactly what 1.2 claims to improve, this is the number that would settle the upgrade question — and it does not exist yet.
• Production telemetry. 1.1 has a week of real p50 and p95 latency behind it and 4.4M tokens of live traffic on OrcaRouter. 1.2 has neither; its endpoint currently reports no latency or throughput statistics at all because volume is too low to sample.
• Somewhere to rent it besides Meta. Muse Spark 1.1 is in the OrcaRouter catalog at $1.25 and $4.25 — Meta's own list price, passed through at 0% markup. Muse Spark 1.2 is not there yet, so today it is a direct Meta integration with a single provider and no failover path.

None of this means 1.2 is worse. It means the evidence for 1.2 consists of one composite score from one evaluator, while the evidence for 1.1 is a month of arenas, sub-indices and live traffic. That asymmetry should affect how you stage the migration, not whether you attempt it.
A migration checklist that is actually short
Because the rate card and parameter surface are identical, the swap is a model-string change. The work is in verifying the three things that did move.
• Measure your own time to first token at your own effort setting. Do not accept either the 2.90s or the 26.12s figure. Run your real prompts at whatever reasoning_effort you actually pass and compare directly. This is the one number that can kill the migration outright.
• Check your timeout and streaming configuration. A 9x increase in first-token latency at high effort will breach client timeouts that were tuned against 1.1, and will make any non-streaming integration look like a hang.
• Recompute cost from measured token counts, not from the rate card. Prices are identical, so any cost change is behavioural. Artificial Analysis saw 16% more spend for the same work; whether you see that depends on your task mix.
• Verify cache-key stability first. With a stable 60,000-token prefix, a typical agent step falls from about 8.8 cents to about 2.2 cents on either version. That 75% swing is larger than anything the version change does, and it is entirely under your control.
• Re-test audio and video paths explicitly. Both listings advertise the same five input modalities, but Artificial Analysis's specification card for 1.2 records only text and image as evaluated. Meta rewrote the pitch away from mixed-media work. Nobody independent has checked whether the capability held.
• Keep 1.1 reachable as a fallback. Running both behind one key means the rollback is a config change rather than an incident, and it lets you A/B the two on live traffic instead of arguing about a benchmark.
Who moves now, who waits
Move now if you run long-horizon coding agents at high reasoning effort. The tasks already take minutes, the latency penalty disappears into that, the intelligence gain is measured by a third party rather than claimed by Meta, and three index points is a meaningful jump at this tier — nine places on a 185-model board.
Wait if anything you serve is interactive. There is no configuration in which 1.2 is more responsive than 1.1, and mandatory reasoning means you cannot buy responsiveness back by disabling thinking. Version 1.1 remains the faster model at every effort level and costs exactly the same.
Wait if your workload is mixed-media analysis — the long-report, video-and-audio use case 1.1 was explicitly built and marketed for. Meta stopped advertising it and the only independent evaluation of 1.2 covers text and images. The capability may well be intact; there is simply no evidence either way, and 1.1 has a month of it.
Wait if you need the arena data. A model sold on whole-repository generation with no Design Arena entries and no published agentic sub-index is asking you to take Meta's word for the thing it changed. Give it three or four weeks and that will no longer be true — at which point this decision gets much easier to make on evidence rather than on a single composite number.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
