Muse Spark 1.2 vs Claude Fable 5
Guides & Insights

Muse Spark 1.2 vs Claude Fable 5: What Six Index Points Actually Cost

Author

Jim Song

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Artificial Analysis publishes something most benchmark tables leave out: what it spent. Running its nine-evaluation Intelligence Index cost $637.85 on Muse Spark 1.2 and $5,630.52 on Claude Fable 5. Identical benchmark suite, identical harness, one bill 8.8 times the other. Fable 5 scored 60 to Muse Spark 1.2's 54.

Divide the difference and you get the number this comparison is really about: on that workload, the last six points of intelligence cost roughly $832 per point. Whether you should pay it is not a benchmark question. It is a question about how much of your traffic actually needs those points, and the answer for most teams is a small and identifiable fraction.

The marginal index point has a price, and it is steep

Both figures come from the same independent evaluator running the same composite — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Neither is a vendor claim. Fable 5 sits at rank #3 of 185 models; Muse Spark 1.2 at #13, against a class median of 32. Both are well above average; only one is at the frontier.

Muse Spark 1.2 vs Claude Fable 5

The list prices explain most of the gap and understate the rest:

Input — Muse Spark 1.2 $1.25 per million, Claude Fable 5 $10.00. Eight times.

Output — $4.25 against $50.00. Nearly twelve times.

Cached input — $0.15 against $1.00. Roughly seven times, and Fable 5 additionally charges $12.50 per million to write the cache, or $20.00 for a one-hour TTL, where Muse Spark 1.2 publishes no separate cache-write fee at all.

Blended rate — Artificial Analysis puts Muse Spark 1.2 at $0.78 per million tokens and Fable 5 at $7.70. Just under ten times.

Fable 5 was the most expensive model Artificial Analysis had ever benchmarked when it launched, and it has held that title. It is worth being precise about why: the model burned fewer output tokens than Muse Spark 1.2 getting through the index — 87M against 95M — so the entire cost difference is rate, not verbosity. Fable 5 is not wasteful. It is simply priced as a frontier product.

Translated into an agent step reading 60,000 tokens of repository context and writing 3,000 tokens of patch: about 8.8 cents on Muse Spark 1.2, about 75 cents on Claude Fable 5. A thousand steps a day is $88 against $750. Over a year, $32,000 against $274,000.

Where Claude Fable 5 is not replaceable

The cost case would be one-sided if the six points were the whole difference. They are not, and the place the gap widens is precisely the place Meta has repositioned Muse Spark 1.2 to compete.

Fable 5 holds first place across a broad slice of the Design Arena head-to-head ladders: 1399 Elo and rank 1 in game development, 1367 and rank 1 in ASCII art, 1353 and rank 1 in SVG generation, plus rank 1 in the agents arena for Godot game development (1344), Android native (1318), agentic game development (1300) and HTML slides (1260), with second place in web apps and full-stack work. Muse Spark 1.2 has no Design Arena entries whatsoever — its predecessor accumulated a respectable mid-table record (1334 in game development at rank 5, 1309 in general code categories at rank 9), but the new version has not been rated once.

The per-dimension indices point the same way. Artificial Analysis scores Claude Fable 5 at a coding index of 76.5 and an agentic index of 52.8. Muse Spark 1.1 sat at 71.3 and 37.5 — five points behind on coding and fifteen behind on agentic work. No per-dimension figures have been published for 1.2 yet, so the honest statement is that we know the composite closed by three points and we do not know how much of that landed on the agentic axis.

Muse Spark 1.2 vs Claude Fable 5

Fable 5's own product positioning claims the lead widens with task length and complexity. That is a vendor claim and unverified as stated, but it is consistent with the shape of the independent data: Fable 5's biggest margins are in the agents arena, where a model has to hold a plan together across many turns, rather than in single-shot generation.

Where Muse Spark 1.2 is simply the right answer

Three things make it the correct pick regardless of the six points.

Audio and video input. Meta's listing advertises text, images, video, audio and PDF in. Claude Fable 5 accepts text, images and files — no audio, no video. For any pipeline touching recordings or footage this is not a trade-off, it is a hard filter. One honest caveat: Artificial Analysis's specification card for Muse Spark 1.2 records only text and image as evaluated, so the wider surface is advertised rather than independently verified.

Speed. 165.0 tokens per second against 71.7. Muse Spark 1.2 streams output well over twice as fast, and ranks #16 of 185 on speed against a class median of 71 — Fable 5 sits at the median.

Cost at volume. Everything in the previous section.

Add a fourth for anyone whose requests are genuinely enormous: both models advertise roughly a million tokens of context — 1,048,576 for Muse Spark 1.2, 1,000,000 for Fable 5 — but Muse Spark 1.2 charges a flat rate across all of it, while Fable 5's cache-write pricing means holding a large stable prefix costs real money up front before it starts saving you any.

Neither of these is a fast model

This is the section most head-to-heads skip, and it changes the architecture rather than the invoice.

At maximum reasoning effort, Artificial Analysis measures time to first token at 26.12 seconds for Muse Spark 1.2 and 141.26 seconds for Claude Fable 5, with an end-to-end response time for Fable 5 of 148.24 seconds. Neither belongs in front of a user at those settings.

Production numbers are far kinder and worth knowing. On OrcaRouter, Claude Fable 5 shows a p50 time to first token of 7.04 seconds and a p95 of 10.00 seconds across a rolling week — that is real traffic at whatever effort callers request, not a benchmark at maximum. Muse Spark 1.1, for reference, runs a 1.84-second p50. Muse Spark 1.2 has no production telemetry yet.

Muse Spark 1.2 vs Claude Fable 5

The structural point: both models have mandatory reasoning. Fable 5's effort dial runs low through max and defaults to high; Muse Spark 1.2's runs minimal through xhigh and defaults to medium. Neither lets you switch thinking off. If you need sub-second responses, this entire comparison is the wrong shelf — that is what a Luna- or Flash-Lite-class model is for.

The two-model stack, priced out

Framing this as a winner-takes-all choice is the mistake, because the traffic is not homogeneous. Almost every production agent has a long tail of routine steps — read a file, summarise a diff, format a patch, decide which tool to call next — and a short head of genuinely hard ones where a wrong answer costs an hour.

Take the same thousand agent steps a day. All on Claude Fable 5: about $750. All on Muse Spark 1.2: about $88. Split 900 routine to the cheap model and 100 hard ones to the expensive one: about $79 plus $75, or $154 a day. That is a fifth of the all-Fable bill while keeping frontier capability on the tenth of calls that actually need it — and it is roughly $220,000 a year of difference on one agent loop.

The reason that split is worth building rather than just describing is that it used to require two vendor relationships, two SDKs and two sets of credentials. It does not now. Claude Fable 5 is in the OrcaRouter catalog at $10.00 and $50.00 — vendor list price, passed through at 0% markup — and Muse Spark 1.1 is there at $1.25 and $4.25 on the same basis, behind the same OpenAI-compatible endpoint. The routing DSL lets you express "cheap model first, escalate on low confidence or on a failed test run" as a single call rather than as orchestration code you maintain, and model fusion lets you send the genuinely ambiguous steps to a panel of models at once and compare. Muse Spark 1.2 itself is not in the catalog yet — Meta serves it directly, as its only provider — so today the closest thing you can build to this stack uses 1.1 on the cheap arm.

That single-provider detail deserves its own line in a risk assessment. Claude Fable 5 has multiple serving routes and automatic failover between them. Muse Spark 1.2 has exactly one endpoint, run by Meta. If it goes down, there is no fallback that is the same model.

A cost model, not a verdict

The useful output of this comparison is not "which model wins." It is a threshold you can compute for your own workload.

Estimate the fraction of your calls where a Muse Spark 1.2-class answer fails and a Fable 5-class answer succeeds. Multiply by what a failure costs — engineer minutes, a bad merge, a retry loop, a customer. Compare that to 66 cents per step, which is what upgrading a single call from Muse Spark 1.2 to Claude Fable 5 costs on the 60K-in / 3K-out example above. If the expected loss from a wrong answer exceeds 66 cents, route that step to Fable 5. If it does not, do not.

For most teams that calculation puts Fable 5 on code review, final patch generation, architectural planning and anything touching production data, and puts Muse Spark 1.2 on everything else — file reading, summarisation, test triage, tool selection, the high-volume middle of an agent loop where the cost is real and the stakes per call are not.

One thing to keep watching: the Design Arena ladders and the per-dimension indices for Muse Spark 1.2, neither of which exists yet. Fable 5's strongest evidence is precisely in the head-to-head arenas where Meta's new model has no record at all. When that record appears, this threshold moves — possibly by a lot.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube