
MiniMax M3 vs Muse Spark 1.2: A 4x Price Gap Against a Benchmark Sweep
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
MiniMax M3 costs $0.30 per million input tokens on our catalogue. Muse Spark 1.2 costs $1.25. That is a 4.2x gap on input and a 3.5x gap on output, and it is the single most useful fact about this pairing, because on the same catalogue pages Muse Spark 1.2 leads MiniMax M3 on every benchmark row the two models share. One is the cheap open-weight option with a vendor-guaranteed 512K of usable context and a 512K output ceiling. The other is a Meta reasoning checkpoint that is now one generation behind its own family and still outscores it almost everywhere. The rest of this page is about which of those two facts actually decides a workload — and the answer is not the same for every workload.
What each of these models actually is
Both shipped this year and neither is new. That matters, because a comparison page written as if either were launching dates itself within a month.

MiniMax M3 is MiniMax's own release, announced on the vendor's blog on 2026-06-01 under the headline "MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — All in One Model." The post opens flatly: "MiniMax M3 is officially released today." MiniMax's three claims for it are that it reaches frontier-level performance on coding and agentic work, that it uses MSA (MiniMax Sparse Attention), a new attention architecture the company proposed, and that it is natively multimodal — it accepts image and video input and, in MiniMax's words, "can operate a desktop computer." MiniMax adds that these three capabilities are table stakes for closed-source frontier models, and that M3 is "the first and only open-weight model to bring all three together." Our catalogue lists the model as available since 2026-05-31, a day earlier than the blog date.
Muse Spark 1.2 is Meta's. Our catalogue dates it to 2026-08-05. It is not a new tier — it is an updated checkpoint over Muse Spark 1.1 with slightly higher performance, served on the same Standard tier at identical pricing, which makes it a drop-in upgrade rather than a separate product. Its distinguishing feature is the input surface: text, images, video, audio and PDF. Output is text.
One piece of context belongs here rather than buried later. Meta moved the Muse Spark family on to a 1.3 checkpoint on 2026-09-02, which makes 1.2 the previous generation of its own line. That happened three weeks ago, so it is not news and this page does not treat it as news — but anyone choosing between these two today should know they are comparing a current MiniMax model against a superseded Meta one.
Price per million tokens, and what a workload actually costs

The published rates, taken from each model's page on our own catalogue, which passes provider list price through with no markup added:
• Input — MiniMax M3 $0.30 per 1M tokens vs Muse Spark 1.2 $1.25 per 1M tokens
• Output — MiniMax M3 $1.20 per 1M vs Muse Spark 1.2 $4.25 per 1M
• Cache read — MiniMax M3 $0.060 per 1M vs Muse Spark 1.2 $0.150 per 1M
• Ratio — Muse Spark 1.2 is 4.2x on input, 3.5x on output, 2.5x on cache reads
Those abstract ratios become concrete in the cost calculator on each page. Set both to 10 million tokens a month at a 70% input share — a reasonable shape for a summarisation or extraction job, and the default the estimator ships with — and MiniMax M3 comes out at $5.70 a month. Muse Spark 1.2 comes out at $11.30 a month. Turn prompt caching on and the same workload lands at roughly $4.86 versus roughly $9.63. The calculator labels both as estimates based on list price, which is the right caveat: real token counts depend on the provider's tokenizer.
For a cheap workload the gap is coffee money. The point is the shape of the curve, not the absolute: a workload of a hundred million output-heavy tokens a month crosses from three figures to four on the Muse Spark side alone. If cost is the constraint that made you look at this pairing, that is the number to model on your own traffic.
Because we pass provider list price straight through with no markup, a vendor price cut shows up on our side the same day it is announced, and both of these models are routed here — one key covers both, so pricing them against each other does not require a second contract or a code change.
Context window, and what actually fits in it
This is the section where the two look most alike on a spec sheet and least alike in use.
• Context window — MiniMax M3 1,048,576 tokens vs Muse Spark 1.2 1,048,576 tokens
• Maximum output — MiniMax M3 512,000 tokens vs Muse Spark 1.2 131,072 tokens
• Input surface — MiniMax M3 text, image and video vs Muse Spark 1.2 text, image, video, audio and PDF
• Output surface — both emit text only
• Structured parameters — MiniMax M3 exposes tools and tool_choice; Muse Spark 1.2 exposes reasoning_effort and structured_outputs
Both windows are the same million tokens, so the "who has more context" question has no winner. The maximum-output row is where they separate: a MiniMax M3 response can run to 512,000 tokens, roughly four times the Muse Spark 1.2 ceiling of 131,072. If your work is long-form generation — a full-file rewrite, a long report, a transcript-scale output — that difference is structural, not incremental. If your work is short answers over a long input, it is irrelevant.
The input-surface row cuts the other way. Muse Spark 1.2 accepts audio and PDF directly; MiniMax M3's documented input surface stops at image and video. A pipeline that ingests call recordings or scanned documents has a different amount of preprocessing to do on each of these two.
On the MiniMax side, the context claim carries an architecture note worth labelling clearly as the vendor's own. MiniMax attributes M3's long-context behaviour to MSA (MiniMax Sparse Attention), which our catalogue describes as supporting up to 1M tokens of context "with a guaranteed minimum of 512K." The guaranteed-minimum framing is MiniMax's; there is no independent measurement of it we can point to, so treat the 512K floor as a vendor claim rather than a tested figure.
Latency and throughput on our own gateway
These are the figures we can speak to most directly, because they are our own numbers. They cover the seven days to 2026-09-23 and describe traffic on our gateway, not a vendor benchmark.
• p50 time to first token — MiniMax M3 4.28 s vs Muse Spark 1.2 4.82 s
• p95 time to first token — MiniMax M3 10.00 s vs Muse Spark 1.2 10.00 s
• Output speed — MiniMax M3 289 tok/s vs Muse Spark 1.2 507 tok/s
• Error rate — MiniMax M3 0.12% vs Muse Spark 1.2 0.15%
• Traffic, last 7 days — MiniMax M3 153.8M tokens vs Muse Spark 1.2 79.6M tokens
Read this honestly and the picture is split. On time to first token the two are close — 4.28 against 4.82 seconds at the median, and the p95 is identical to the hundredth at 10.00 seconds, which is a rounding artifact of both sitting near the same ceiling rather than a real tie. On output speed they are not close at all: Muse Spark 1.2 streams at roughly 1.75x MiniMax M3's rate, which changes how a long generation feels to a user watching it, even when total cost is higher. Error rates are within a rounding error of each other and both are low.
The traffic row is worth reading as a signal rather than a scoreboard: over the same seven days, MiniMax M3 moved about twice the tokens Muse Spark 1.2 did on our gateway. That is what the market is choosing, not what the benchmarks say, and on this pairing the two disagree.
Routing both behind one endpoint is also the cheap way to find out which one your workload prefers, because failover means a provider-side degradation on either model falls through to a working path instead of an error.
Tool calling and long-horizon agentic work
This is where the two are furthest apart on measured results, and where the caveats matter most.
• Terminal-Bench v2.1 — MiniMax M3 52.4 vs Muse Spark 1.2 80.1
• tau_banking (tool use) — MiniMax M3 12.3 vs Muse Spark 1.2 34.8
• MCP Atlas — MiniMax M3 74.2 (vendor-reported) vs Muse Spark 1.2 not measured
• BrowseComp — MiniMax M3 83.5 (vendor-reported) vs Muse Spark 1.2 not measured
• OSWorld-Verified — MiniMax M3 70.0 (vendor-reported) vs Muse Spark 1.2 not measured
The first two rows come from the benchmark panel on our own catalogue pages and are the only agentic rows both models have filled in. They are not close: Muse Spark 1.2 more than doubles MiniMax M3 on tau_banking and leads Terminal-Bench v2.1 by nearly 28 points. If your workload is a long-horizon agent with a tool surface, that is the largest single gap on this page.
The next three rows are MiniMax's own numbers, published in the launch post. They are not comparable to the first two — different harnesses, no independent audit — but they are the only published figures for MiniMax M3 on those tasks, so leaving them out would understate the model. Read them as vendor-reported and nothing more.
One divergence deserves its own paragraph, because it is the kind of thing a comparison page usually smooths over. MiniMax's launch post renders Terminal-Bench 2.1 at 66.0 for M3, inside a chart image with no accompanying numeric table. The independent Terminal-Bench v2.1 row on our catalogue page shows 52.4 for the same model. The task names are close enough that readers will meet them side by side, and the two figures differ by more than 13 points. We are not going to declare one of them wrong: the likely explanation is a different harness or a different run configuration, and the honest statement is that the vendor's own number and the audited number disagree, with the vendor's being the higher one.
The parameter lists tell a smaller, more practical story. MiniMax M3 exposes tools and tool_choice, so tool selection is steerable from the request. Muse Spark 1.2 exposes reasoning_effort, so the depth of reasoning is steerable instead. Neither exposes the other's dial. If your application's control problem is "make it call the right function," that points one way; if it is "make it think longer on the hard cases," it points the other.
Coding
Coding is MiniMax M3's headline claim and the place where the measured gap is most awkward for it.
• AA Coding Index — MiniMax M3 38.7 vs Muse Spark 1.2 72.2
• SciCode — MiniMax M3 43.1 vs Muse Spark 1.2 57.4
• SWE-Bench Pro — MiniMax M3 59.0 (vendor-reported) vs Muse Spark 1.2 not measured
• KernelBench Hard — MiniMax M3 28.8 (vendor-reported) vs Muse Spark 1.2 not measured
MiniMax's positioning for M3 is coding-first — the launch headline puts "Frontier Coding" ahead of the million-token context and the multimodality. Its own reported SWE-Bench Pro figure of 59.0 is a strong number on the vendor's harness. On the independent Coding Index and SciCode rows that both models have filled in, though, Muse Spark 1.2 leads by a wide margin, and the Coding Index gap of 33 points is the second-largest on this page after tau_banking.
There is no honest way to present MiniMax M3 as the better coding model on the evidence available. What can be said is that the vendor's own coding numbers are high and the independent ones are not, in the same pattern as the Terminal-Bench divergence above. Anyone whose decision turns on coding should weight the audited rows more heavily than the launch-post charts, and should test on their own repository rather than on either.
Where they are genuinely close
Three rows refuse to produce a winner, and saying so is more useful than manufacturing one.
• GPQA Diamond — MiniMax M3 89.9 vs Muse Spark 1.2 90.4, a 0.5-point gap that is a tie for any practical purpose
• p95 time to first token — both 10.00 s over the last seven days on our gateway
• Context window and output modality — identical at 1,048,576 input tokens and text-only output
On graduate-level science reasoning the two are effectively indistinguishable, and that is the one reasoning row where MiniMax M3's much cheaper model holds its own against the more expensive one. It is worth remembering when reading the aggregate indices: the gap between these models is not uniform across capability, it is concentrated in agentic and coding work, and it is roughly zero on knowledge-heavy multiple choice.

Pick MiniMax M3 if, pick Muse Spark 1.2 if
The two facts from the opening paragraph resolve differently depending on what you are building, so here is the split without hedging.
• Pick MiniMax M3 if cost per token is the binding constraint, if you need long-form output past 131,072 tokens, if you need open weights, if you want video input without a separate pipeline, or if you want reasoning-heavy knowledge work at roughly a quarter of the input price — GPQA Diamond is a dead heat and that is the row where the cheap model earns its place.
• Pick Muse Spark 1.2 if your workload is a long-horizon agent with tools, if you need the highest audited coding scores, if you want an explicit reasoning_effort dial, if you need to ingest audio or PDF directly, or if you need fast streaming output — 507 tok/s against 289 tok/s is a visible difference, not a benchmark one.
• If you do not yet know which, the deciding evidence is not on this page. Set both models against a sample of your own traffic and measure it; the price calculator on each model page will tell you the cost side in a minute, and the seven-day latency figures above are the closest thing to a preview of the performance side.
Both are on one API with no markup over provider list price, so the trial is a model-id change rather than a migration, and automatic failover means a bad day on either provider degrades into a slower answer rather than an outage.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
