
Muse Spark 1.2 vs GPT-5.6 Luna: Three Index Points for 4.6x the Token Price
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
Artificial Analysis ran the same nine-benchmark suite through both of these models and kept the receipts. Getting Muse Spark 1.2 through it cost $637.85. Getting GPT-5.6 Luna through it cost $174.06. For that 3.7x difference in spend, Meta's model came out three points ahead — 54 against 51 on the Intelligence Index. Whether that is a good trade is the entire question, and the answer depends far less on the benchmark than on two things almost nobody writes about: how your agent framework handles prompt caching, and whether your workload ever needs to hear or watch anything.
Every figure below is either an independent Artificial Analysis measurement, a live API listing, or OrcaRouter's own production telemetry — the last of which is worth flagging up front because it contradicts the benchmark numbers in an instructive way. Nothing here is a vendor claim dressed up as a result; where the vendor is the only source, the sentence says so.
The scoreboard is closer than the price tag suggests
Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index at its xhigh reasoning setting, ranking #13 of 185 models against a class median of 32. GPT-5.6 Luna scores 51 at max effort. Three points on a composite of GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR.
Three points is real but small, and the sub-scores that exist for the previous generation suggest where it comes from. Artificial Analysis's per-dimension figures put Muse Spark 1.1 at a coding index of 71.3 and an agentic index of 37.5; GPT-5.6 Luna sits at 71.4 and 45.6. Coding was a dead heat, and Luna was eight points ahead on agentic work — the exact dimension Meta rewrote 1.2's product description to claim. The composite moved three points in Meta's favour. Nobody has yet published a per-dimension breakdown for 1.2, so whether the agentic gap actually closed is unverified.

On blended token price the gap is not subtle. Artificial Analysis's blended rate puts Muse Spark 1.2 at $0.78 per million tokens and GPT-5.6 Luna at $0.17. List prices: $1.25 in and $4.25 out for Muse Spark 1.2, $0.20 in and $1.20 out for Luna.
Priced per unit of work rather than per token, a single coding-agent step reading 60,000 tokens of context and writing 3,000 tokens of output costs about 8.8 cents on Muse Spark 1.2 and about 1.6 cents on GPT-5.6 Luna. Across a thousand steps a day that is $88 against $16 — a difference of roughly $26,000 a year for a single agent loop.
Caching is where the gap either closes or doubles
Both models sell aggressive cached-input rates, and the ratio between them is not the same as the ratio between their list prices. Muse Spark 1.2 reads cached input at $0.15 per million, an 88% discount off its $1.25 rate. GPT-5.6 Luna reads cached input at $0.01 per million, a 95% discount off $0.20.
Rerun the same agent step with the 60,000-token prefix served warm: Muse Spark 1.2 drops from 8.8 cents to about 2.2 cents. Luna drops from 1.6 cents to about 0.4 cents. The absolute gap narrows from 7.2 cents to 1.8 cents per step, but the multiple stays roughly the same — caching does not rescue the more expensive model, it just makes both cheaper by similar factors. What it does change is which line item dominates: cached, Muse Spark 1.2's cost is 59% output tokens rather than 85% input tokens, so the model's verbosity starts to matter more than its context size.
That is not a hypothetical concern here. Muse Spark 1.2 generated 95M output tokens getting through the Intelligence Index, against a 65M median. Artificial Analysis's own summary calls it "somewhat verbose." Luna burned 130M, which sounds worse until you multiply by $1.20 instead of $4.25.
Meta's product copy claims prompt caching can cut effective cost 60–80% depending on how much context repeats. That is a vendor estimate, not a measurement, but the arithmetic above lands inside the range.
Latency: the axis where the benchmark inverts the truth
Read only the Artificial Analysis latency figures and you would conclude that Muse Spark 1.2 is the responsive one. Time to first token: 26.12 seconds for Muse Spark 1.2, 126.99 seconds for GPT-5.6 Luna. Nearly five times faster.
Both of those are measured at each model's most expensive reasoning setting — xhigh for Muse Spark, max for Luna. Neither is what you will see. On OrcaRouter, across a week of real traffic at whatever effort callers actually request, GPT-5.6 Luna shows a p50 time to first token of 1.84 seconds and a p95 of 9.68 seconds. Muse Spark 1.1 shows an identical 1.84-second p50 with a tighter 6.00-second p95. There is no production telemetry for 1.2 yet because it is too new.

The structural difference is more useful than either number. GPT-5.6 Luna's reasoning is optional — the effort dial runs from none through low, medium, high, xhigh to max, and you can genuinely switch thinking off for classification, routing or extraction. Muse Spark 1.2's reasoning is mandatory. The dial bottoms out at minimal, and there is no setting that gives you the weights without paying for deliberation. For anything high-volume and latency-sensitive, that is a design constraint, not a tuning parameter.
Sustained throughput is close: 165.0 tokens per second for Muse Spark 1.2 against 177.9 for Luna, both well above the class median of 71.
The pricing footnote everyone will trip over
GPT-5.6 Luna's price is currently listed differently in different places, and if you are building a cost model you should know before you quote a number. Artificial Analysis and OrcaRouter's catalog both show $0.20 per million input and $1.20 per million output — the rate that followed the July GPT-5.6 price cut, and the rate Artificial Analysis used to compute the $174.06 eval bill above. Some marketplace listings currently show $0.10 and $0.60 with a separate higher tier that kicks in above 272,000 prompt tokens. The safe move is to check the price on the endpoint you are actually calling rather than the one in the comparison article.
The tiering detail is real regardless of which base rate applies, and it is genuinely under-covered: Luna's price steps up once a single request exceeds roughly 272,000 prompt tokens. Input roughly doubles, output rises by half, and cached reads scale with it. If your reason for looking at Luna is its 1.05M-token context window, note that you leave the headline rate at about a quarter of the way in. Muse Spark 1.2 has a flat rate across its full 1,048,576 tokens with no long-context surcharge — for genuinely long single requests, the effective gap between the two narrows considerably.
What Luna cannot do at any price
The modality difference is the cleanest reason to pick one over the other, and it does not appear on any leaderboard.
• Inputs — Muse Spark 1.2 accepts text, images, video, audio and PDF documents per Meta's listing. GPT-5.6 Luna accepts text, images and files. No audio, no video. If your pipeline touches recordings or footage, Luna is not a candidate at all.
• Maximum output — Luna declares a 128,000-token completion ceiling. Muse Spark 1.2's endpoint declares no maximum completion length, which is unusual enough that you should test it rather than plan around it.
• Knowledge cutoff — Luna's is published as February 16, 2026. Meta publishes no cutoff for Muse Spark 1.2, so budget for retrieval either way.
• Serving — Luna is generally available worldwide through multiple routes. Muse Spark 1.2 is served by exactly one provider: Meta. One endpoint, one region policy, one thing to go down.
• Moderation — both are moderated endpoints.
One honest caveat on the multimodal advantage: Artificial Analysis's technical-specification card for Muse Spark 1.2 lists text and image input as what it evaluated, not the full five-modality surface Meta advertises. The API accepts more than anyone independent has tested.

Adoption, since it happens to be measurable
Benchmarks tell you what a model can do; traffic tells you what people trust it with. Across a rolling week on OrcaRouter, GPT-5.6 Luna moved 4,679.4 million tokens. Muse Spark 1.1 — same 1M context, comparable index score, four weeks older in the catalog — moved 4.4 million. That is a factor of roughly a thousand.
Some of that is distribution: Luna is a default in more tooling and has been GA worldwide for longer, while the Muse Spark family launched US-only. But a thousand-to-one gap on a router where both are one model string away from each other is not purely a distribution story. It is the practical reason Luna is the safer default and Muse Spark 1.2 is the thing you evaluate deliberately.
Worth noting what you can and cannot rent today: GPT-5.6 Luna is in the OrcaRouter catalog at $0.20 and $1.20, the same numbers on the provider's own price list, because we run 0% markup and pass provider list price straight through. Muse Spark 1.1 is there at $1.25 and $4.25 on the same basis. Muse Spark 1.2 is not in the catalog yet — it is served by Meta directly. So the cheapest way to run this comparison honestly today is Luna against Muse Spark 1.1 on one key, changing one string between runs, and then re-testing against 1.2 when it lands.
Pick one
Take GPT-5.6 Luna if the workload is high-volume and text-shaped: chat, classification, extraction, routing, lightweight tool use. It is a quarter of the blended price, it lets you turn reasoning off entirely, its 1.84-second p50 is a real measured production number rather than a benchmark artefact, and it is the model with a thousand times the live traffic behind it. Three index points do not survive that arithmetic.
Take Muse Spark 1.2 if the workload is long-horizon agentic coding — multi-file refactors, extended debugging, whole-repository work — or if it involves audio or video input, where Luna simply cannot compete. It is also the better choice for genuinely enormous single requests, since its rate is flat across the full million tokens while Luna's steps up past 272K.
Take both if you are being honest about how agent systems actually get built. The pattern that keeps winning is a cheap model handling the routine majority of calls and an expensive one reserved for the steps where quality pays for itself. Both models sit behind one OpenAI-compatible endpoint, so that split is a routing rule rather than a second vendor relationship — and if the expensive arm goes down mid-run, automatic failover means your evaluation does not quietly poison itself with half a dataset.
The thing to watch over the next month is the per-dimension breakdown. Muse Spark 1.2's whole pitch is agentic capability, and on the previous generation that was the dimension where Luna led by eight points. Until someone publishes an agentic index for 1.2, the three-point composite gain is the only evidence there is that Meta closed it.
