Muse Spark 1.2 vs Claude Haiku 4.5: Identical Blended Price, Opposite Ideas About Thinking
Guides & Insights

Muse Spark 1.2 vs Claude Haiku 4.5: Identical Blended Price, Opposite Ideas About Thinking

Author

Jim Song

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Artificial Analysis prices Muse Spark 1.2 at $0.78 per million tokens blended and Claude Haiku 4.5 at $0.77. One cent apart, from two labs that agreed on nothing else. Meta's model cannot stop reasoning; Claude Haiku 4.5 only reasons when you ask it to. Meta gives you 1,048,576 tokens of context; Claude Haiku 4.5 gives you 200,000. Meta shipped this one on August 5, 2026 as a coding model with a terminal agent attached; Claude Haiku 4.5 shipped in October 2025 as the fast, cheap worker a bigger model tells what to do. Same price per token, entirely different machines.

That coincidence is the useful place to start a comparison, because it removes the variable people usually decide on and forces the question to be about shape instead. What follows uses independent numbers where they exist — Artificial Analysis for index scores and lab latency, OrcaRouter's own seven-day production telemetry for real-traffic latency — and flags vendor-reported figures as such, including one benchmark headline of the vendor's that has no third-party counterpart.

The spec contrast, one dimension at a time

Price — Muse Spark 1.2 at $1.25 in / $4.25 out per million; Claude Haiku 4.5 at $1.00 in / $5.00 out. Meta is cheaper on input, Claude Haiku 4.5 on output.

Cache — $0.15 per million on a Muse Spark 1.2 cache hit, an 88% discount; $0.10 per million on a Claude Haiku 4.5 cache read, a 90% discount, with cache writes billed at $1.25.

Context — 1,048,576 tokens vs 200,000. A 5.2× difference, and the single largest gap between them.

Max output — Claude Haiku 4.5 declares a 64,000-token ceiling; the Muse Spark 1.2 endpoint declares no maximum completion length at all, which is unusual enough to test rather than assume.

Reasoning — mandatory on Muse Spark 1.2, with five effort levels from minimal to xhigh and a default of medium; optional on Claude Haiku 4.5, which is a non-reasoning model until you switch extended thinking on and hand it a thinking budget.

Inputs — Meta advertises text, images, video, audio and PDF; Claude Haiku 4.5 takes text and images. Both return text.

Weights and providers — both closed. Muse Spark 1.2 is served only by Meta's own Model API; Claude Haiku 4.5 is available from the vendor and through the usual cloud resellers and aggregators.

Age — Muse Spark 1.2 is days old. Claude Haiku 4.5 has been in production since October 15, 2025, with a July 2025 knowledge cutoff and nine months of accumulated field experience behind it.

The index gap is real. The number everyone will quote is not.

Artificial Analysis scores Muse Spark 1.2 at 54 on its Intelligence Index, rank #13 of 185. Its current entry for Claude Haiku 4.5 is 24. Put those side by side and you have a headline; look at what the entries actually are and the headline falls apart.

The 54 is Muse Spark 1.2 evaluated at xhigh, its most expensive reasoning setting. The 24 is Claude Haiku 4.5 evaluated as a non-reasoning model — thinking off — and Artificial Analysis flags it as an estimate rather than a completed run. On an earlier version of the same index, before the v4.1 reweighting, Artificial Analysis put Claude Haiku 4.5 at 55 with reasoning enabled and 42 without it. Those older figures cannot be compared to 54 either, because index versions rescale and a v3-era score is not a v4.1 score.

So the honest statement is narrower than the leaderboard makes it look: Muse Spark 1.2 at maximum effort scores 54 on the current index; there is no published v4.1 score for Claude Haiku 4.5 with extended thinking on, so no like-for-like comparison of these two at full effort exists yet. Anyone showing you 54 versus 24 is comparing a model at full deliberation against a model told not to deliberate.

Where the vendor has published a number, it is a specific one: Claude Haiku 4.5 scores 73.3% on SWE-bench Verified, averaged over 50 trials with a 128K thinking budget. That is a vendor figure, and the thinking budget in it is enormous for a model marketed on speed — but it is the same evaluation the industry quotes for frontier models, and it puts Haiku in a bracket its price does not suggest. Meta's counterpart claims for Muse Spark 1.2 are Terminal-Bench 2.1 at 82.9% and DeepSWE v1.1 at 59.3%, also vendor-run, on Meta's own harness with each competitor paired to its own agent product. Two vendors, two benchmarks, no overlap. There is currently no single evaluation on which both of these models have a published score from the same evaluator.

Four latency numbers, two different stories

Latency is where the "same price" framing stops being a curiosity and starts being a decision, and it is also where measurement source matters most.

In the benchmark harness, Artificial Analysis records time to first token at 26.12 seconds for Muse Spark 1.2 at xhigh, and 1.00 second for Claude Haiku 4.5. A 26× gap, and if that were the whole picture the comparison would be over.

In production it inverts. Across seven days of real traffic on OrcaRouter — callers using whatever effort they actually want — Claude Haiku 4.5 shows a p50 time to first token of 3.84 seconds and a p95 of 10.00 seconds, at 214 tokens per second and a 1.0% error rate. Muse Spark 1.1, the version of Meta's model that is in the catalog, shows a p50 of 1.84 seconds and a p95 of 6.00 seconds, at 519 tokens per second with no errors recorded. On live traffic, the Meta model is the faster one.

Both sets of numbers are true and they measure different things. The lab figure is each model at maximum deliberation with a cold cache, which is a fair test of the ceiling and a poor test of a chat box. The production figure includes cache hits, short prompts, low effort levels and everything else real applications do, which is a fair test of the median and no test at all of the worst case. The trap is quoting one and reasoning about the other. If you are sizing a user-facing timeout, the p95 is your number. If you are asking what a deep agentic step costs in wall-clock, the benchmark figure is closer to the truth — and the honest caveat is that no such production figure exists for Muse Spark 1.2 itself yet, because it is too new and not widely routed.

Both labs sold you a subagent. They meant different things.

The vendor's pitch for Claude Haiku 4.5 was explicit about hierarchy: Sonnet-class models plan and orchestrate, Haiku instances execute in parallel underneath. Cheap, fast, disposable, many at once. The product design follows — a 200K window is plenty for a scoped subtask and deliberately not enough for a whole repository, thinking is off by default because a worker that deliberates is a worker that costs orchestrator money, and the vendor's own marketing names real-time chat, customer service and pair programming as the target.

Meta's pitch for Muse Spark 1.2 claims both ends of the same hierarchy. Meta's copy says the model can be the primary agent that gathers context, plans and delegates, or a subagent executing in parallel beneath one. Muse Code, the terminal agent that shipped alongside it, coordinates multiple persistent subagents per task in isolated worktrees, on a replay-exact event-log runtime. The million-token window and the always-on reasoning are what a planner needs; they are exactly what you would not choose for a fan-out worker, because every parallel instance is paying for deliberation you did not ask for.

Read as architecture rather than marketing, that is the real difference: the vendor built a worker and expects you to bring an orchestrator; Meta built an orchestrator and claims it can also work. If you already run a hierarchy, the interesting experiment is not choosing between them — it is Muse Spark 1.2 planning and Claude Haiku 4.5 executing, which is a shape neither vendor will sell you as a package.

What one real step costs on each

Per-million rates decide nothing on reasoning models, because the model chooses how many tokens to spend. Price the step instead.

Take a scoped code task: 60,000 tokens of repository context in, 3,000 tokens of patch out. Cold, Claude Haiku 4.5 costs $0.060 of input plus $0.015 of output — 7.5 cents. Muse Spark 1.2 costs $0.075 plus $0.013 — 8.8 cents. Close enough to be noise, exactly as the blended rate promised.

Now add the thing the blended rate hides. Muse Spark 1.2 cannot turn reasoning off, and at its default medium setting a step like this plausibly emits a few thousand reasoning tokens before the patch — they bill as output. Add 3,000 and the same step is $0.075 + $0.026 = 10.1 cents, about 35% above Haiku on identical inputs. Claude Haiku 4.5 with thinking off pays nothing for deliberation and stays at 7.5 cents; with a large thinking budget it will exceed Muse Spark 1.2, because Claude Haiku 4.5 charges $5.00 per million output against Meta's $4.25.

Caching flips the emphasis again. Hold the repository context stable so it hits cache and Haiku's input drops to $0.006 and Muse Spark 1.2's to $0.009 — the input side stops mattering entirely and the whole comparison becomes a fight over output tokens, which is a fight about how much each model thinks. On a workload that repeats context, cache-key stability is worth more than model choice, and that is true of both of these models.

The 1M window is the one place the arithmetic is not close. Anything that genuinely needs more than 200,000 tokens in a single call cannot be run on Claude Haiku 4.5 at any price. You either chunk and orchestrate — which is what the vendor intends — or you use a model with the room.

Trying both without a second contract

Claude Haiku 4.5 is in the OrcaRouter catalog at $1.00 and $5.00 — the vendor's list price, because OrcaRouter runs 0% markup and passes provider pricing through untouched, which also means a vendor price change is live on our side the same day rather than at the next contract cycle. The production latency figures above come from that same routing layer.

Muse Spark 1.2 is not in the catalog. Meta's Model API is currently the only place to call it, so a genuine head-to-head today means one integration through us and one direct with Meta. Muse Spark 1.1 is hosted, at the identical $1.25 and $4.25 that Meta charges for 1.2, which makes it a reasonable stand-in for the family's cost and latency shape but not for the coding gains 1.2 is being sold on. When 1.2 becomes routable, the comparison becomes a one-line model-string change against the same key — and for the orchestrator-plus-worker experiment described above, the routing DSL is how you wire a planner and a fan-out worker into a single call rather than building the handoff yourself.

Pick by the shape of the work, not the score

Choose Claude Haiku 4.5 when the work is bounded and the user is waiting. Support agents, classification, extraction, chat, pair-programming completions, and the parallel-worker tier of an agent hierarchy. It is a known quantity with nine months of field history, thinking is optional so you only pay for deliberation when you want it, and 200K is enough for almost any scoped subtask. Its published SWE-bench Verified result says it is far more capable than the price implies, provided you are willing to spend the thinking budget that result used.

Choose Muse Spark 1.2 when the unit of work is a repository rather than a request. Multi-file refactors, extended debugging, whole-project generation, long-horizon agent loops where nobody is watching a spinner. The million-token window removes a chunking problem Haiku cannot solve at any price, and the mandatory reasoning that ruins it for chat is the right default when a wrong plan costs more than a slow one.

What should decide it is not the index number. Ask two questions instead: does a single call ever need to see more than 200,000 tokens, and is there a human waiting on the first token? Yes to the first points at Meta. Yes to the second points at the vendor. Yes to both means you want two models, which is the honest answer more often than either vendor's marketing admits.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube