
Muse Spark 1.3 Max Reasoning Is Live: The Setting Behind Meta's Top Scores Now Ships to Everyone
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 220 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 113 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 215 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Muse Spark 1.3 — the frontier reasoning model Meta Superintelligence Labs released on September 2 — just gained the mode its own scoreboard was built on. On September 4, after two days of additional safety testing, Meta opened the model's most compute-intensive "max" reasoning setting to developers on the Meta Model API and in Muse Code, its terminal coding agent. Chief AI officer Alexandr Wang announced the rollout on X and the company's @AIatMeta account confirmed it; what was a configuration limited to Meta and a partner preview at launch is now a reasoning level any developer can select.
That sequence matters because several of the numbers that dominated Muse Spark 1.3's launch coverage — 75.4 on DeepSWE v1.1, 66.9 on OSWorld 2.0, 1,754 on GDPval-AA v2 — came from the max configuration, while the model developers could actually call shipped with reasoning effort capped at "xhigh." Meta was explicit at launch that max was still in safety testing, but the split meant the headline scoreboard and the callable model were two different things for two days. Thursday's release collapses that gap, and it does so without touching the per-token price. The benchmark figures in that list are Meta's own and remain unreproduced by an independent harness; everything vendor-reported in this piece is labeled as such.
What was held back — and what changed on September 4
Muse Spark 1.3's reasoning ladder has been the same shape since the family's paid API opened: reasoning is always on, and the model's effort parameter on the Meta Model API takes minimal, low, medium, high, and xhigh, with the model spending more of its output budget thinking as the level rises. Max sits above xhigh — more compute, more reasoning tokens, and in Meta's telling, meaningfully stronger on long-horizon agentic work. At the September 2 launch Meta shipped every level up to xhigh but kept max out of the public API pending what it called additional safety testing, sharing it only with a limited set of partners — enough for an independent evaluation to run, not enough for a production workload.
On September 4 Wang said that testing was complete. Max is now a selectable setting in Muse Code — the install script is at dev.meta.ai/install.sh — and on the muse-spark-1.3 model id through the Meta Model API, where the effort parameter now accepts max. Wang described the two-day stagger as Meta's way of gating access "at the configuration level," and told Axios Meta has not had to pause its broader work over safety concerns. The window was short, but it is a real precedent: Meta is now willing to hold a specific reasoning mode behind a safety review while shipping the rest of a checkpoint immediately.
One practical caveat for anyone calling the model through a gateway rather than Meta's own endpoint: several third-party documentation pages and aggregator listings were written on launch day and still enumerate reasoning effort only up to xhigh. The max value is a Meta-side rollout, so check that the route you call through has updated its parameter surface before assuming max is reachable — a proxy that validated its accepted values on September 2 may reject "max" until its maintainers catch up.
What max actually buys — and where it does not help
Meta's Thursday materials include an explicit max-versus-xhigh split on the benchmarks it runs, and this is the most useful thing it has published about the model yet. All of it is vendor-reported, produced inside Meta's own Muse Code harness, and Meta says its comparisons against Claude Opus 5 and GPT-5.6 Sol were "best-effort" runs at those rivals' own maximum settings that "may not reflect provider-optimized performance." With that caveat, the split tells a clear directional story:
• OSWorld 2.0 (computer-use agent tasks) — 66.9 at max vs 57.2 at xhigh
• GDPval-AA v2 (knowledge-work agent Elo) — 1,754 at max vs 1,709 at xhigh
• JobBench (long-horizon professional tasks) — 64.9 at max vs 61.2 at xhigh
• DeepSearchQA — a tie at 89.4
• Terminal-Bench 2.1 (terminal-agent coding) — 89.2 at xhigh vs 88.8 at max, the one suite where the configuration that shipped on September 2 wins

The pattern is worth reading carefully. Max helps most where the task is a long agentic loop — operating a computer, working a knowledge-work brief, holding a schedule of subtasks the way JobBench does. On a single-turn terminal benchmark it added nothing and cost a little: Terminal-Bench is the reminder that "more reasoning" is not a monotonic upgrade, and it is the reason to A/B max against xhigh on your own workload rather than assume the higher setting is better everywhere. Meta also flags the DeepSWE v1.1 result of 75.4 — the day-one headline, up from the 59.3 Meta reported for Muse Spark 1.2 six weeks earlier — as a max-configuration number. That is the figure independent evaluators will re-run first.
The independent read is starting to catch up
What was missing at launch was any independent measurement, and that gap is closing on a schedule that happens to match the max rollout. Artificial Analysis' Coding Agent Index — which scores each model inside its native coding-agent harness and blends DeepSWE, Terminal-Bench, and SWE-Atlas tasks — placed Muse Spark 1.3 (max) in the Muse Code harness at 68 this week, second on the board behind Claude Opus 5 (xhigh) in Claude Code and ahead of Claude Fable 5 (max) at 67 and GPT-5.6 Sol (max) at 65. The same board scores the shipping xhigh configuration at 64 in the same Muse Code harness, which is the cleanest like-for-like read yet of what max adds — roughly four index points — on an independent task set. That board row was published while max was still in partner preview; the configuration it describes is the one that became generally available on September 4.
Artificial Analysis has also updated its own model card for the max configuration since the rollout. Through the preview period the card listed no provider and no price; the live card now lists Meta at the standard $1.25 per million input and $4.25 per million output tokens — the same list price as Muse Spark 1.2 — while adding a verbosity caveat: AA's evaluation run consumed roughly 130 million tokens against a 79-million median, cost about $1,335 in total, and lands near $0.95 per task on AA's per-task estimate at about 190 output tokens per second.
One number from earlier coverage deserves a version check. Artificial Analysis refreshed its Intelligence Index to v4.2 this week — the same week max shipped — re-scoring the field and shifting the medians. On the current card the max configuration reads 53 against a comparable-model median of 29; the 62 that circulated in launch-week coverage was measured on the previous version of the index, and the two are not comparable. Meta's own xhigh-versus-max split above is independent of AA's index and unaffected by the refresh.

What max costs — the bill is in the tokens, not the price list
Meta publishes no separate price for the max configuration, and no dedicated latency analysis either. The per-token price is identical to Muse Spark 1.2 and to every other effort level on Muse Spark 1.3: $1.25 per million input tokens, $4.25 per million output, and $0.15 for cached input on the standard tier. Reasoning tokens are billed as output tokens and count against the output budget, so choosing max is not a line-item price increase — it is a promise to spend more of that budget thinking before answering.
That makes the real cost question one of token volume, and the early data says max is spendy exactly where xhigh is lean. AA's instrumented run consumed about 130 million tokens on a single evaluation of the configuration. For a developer already running long agentic loops at xhigh, the A/B test that matters is not "does max pass more tasks" but "does max pass enough additional tasks to pay for a materially larger token bill on the same workload."
Meta's Contributor tier — $0.10 in and $0.20 out per million tokens in exchange for letting Meta train on your prompts and completions — applies to the same checkpoint, and Meta says a meaningful double-digit percentage of developers choose it. Contributor's terms make every submission training data and its rate limits are tight, so it is a poor default for production fan-out but a legitimate way to run expensive max experiments cheaply if the data trade is acceptable to you.
The pricing anchor for all of this is easy to check without taking Meta's word for it, because the checkpoint Muse Spark 1.3 is priced identically to is already listed on OrcaRouter at Meta's own list price: Muse Spark 1.2 is on OrcaRouter at $1.25 in and $4.25 out per million tokens with zero markup, so the "unchanged price" claim is checkable against a live page that shows exactly those numbers. Muse Spark 1.3 itself is not on OrcaRouter yet. When a provider we carry starts serving it with max, the price shown on our side will be Meta's list price passed straight through — we do not mark provider prices up — and any vendor cut goes live the same day Meta makes it. For a release like this one, where the interesting economics sit in token volume rather than headline price, that pass-through is what keeps the number you are actually billed equal to the number Meta publishes.

Who should switch to max
The decision splits cleanly:
• Switch for long-horizon agentic workloads — computer-use, knowledge-work briefs, multi-step tool loops in Muse Code — where Meta's own split and AA's coding-agent board both show the largest max gains. A/B it against xhigh on a sample of your real tasks first, because verbosity is real.
• Stay on xhigh for high-volume, latency-sensitive, or short single-turn calls, where max mainly adds tokens. Terminal-Bench is the proof that it does not always add capability.
• Watch three things from here: the open-weights release Zuckerberg has promised without a date, the consumer rollout of Muse Spark 1.3 to Instagram, Facebook, and Meta AI in the coming weeks, and whether Artificial Analysis' coding-agent board starts pricing the max row now that the configuration is generally available.
The two-day gate turned out to be short, but it made a point that outlasts the rollout itself: Meta's strongest Muse Spark 1.3 numbers were never a mirage — they were just waiting on a safety review. As of September 4, the model a developer can call and the model on the scoreboard are finally the same model. Whether max earns its tokens is now an empirical question any developer can answer, on Meta's API or on whatever route they already use to reach the Muse Spark family.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
