Muse Spark 1.2 对比 Claude Haiku 4.5:混合价格相同,对思考的理念截然相反
Guides & Insights

Muse Spark 1.2 对比 Claude Haiku 4.5:混合价格相同,对思考的理念截然相反

作者

Jim Song

发布日期

最新模型 · 20查看全部模型
基准测试:Artificial Analysis · 每日更新
返回全部文章

Artificial Analysis prices Muse Spark 1.2 at $0.78 per million tokens blended and Claude Haiku 4.5 at $0.77. One cent apart, from two labs that agreed on nothing else. Meta's model cannot stop reasoning; Claude Haiku 4.5 only reasons when you ask it to. Meta gives you 1,048,576 tokens of context; Claude Haiku 4.5 gives you 200,000. Meta shipped this one on August 5, 2026 as a coding model with a terminal agent attached; Claude Haiku 4.5 shipped in October 2025 as the fast, cheap worker a bigger model tells what to do. Same price per token, entirely different machines.

That coincidence is the useful place to start a comparison, because it removes the variable people usually decide on and forces the question to be about shape instead. What follows uses independent numbers where they exist — Artificial Analysis for index scores and lab latency, OrcaRouter's own seven-day production telemetry for real-traffic latency — and flags vendor-reported figures as such, including one benchmark headline of the vendor's that has no third-party counterpart.

规格对比,一次一个维度

Price — Muse Spark 1.2 at $1.25 in / $4.25 out per million; Claude Haiku 4.5 at $1.00 in / $5.00 out. Meta is cheaper on input, Claude Haiku 4.5 on output.

Cache — Muse Spark 1.2 缓存命中:每百万次 $0.15(88% 折扣);Claude Haiku 4.5 缓存读取:每百万次 $0.10(90% 折扣);缓存写入按 $1.25 计费。

上下文—— 1,048,576 tokens 对比 200,000。差距为 5.2 倍,是两者之间最大的差距。

最大输出 — Claude Haiku 4.5 声明了 64,000 token 的上限;而 Muse Spark 1.2 端点则完全没有声明最大完成长度,这种情况很不寻常,值得测试而非臆断。

推理——在 Muse Spark 1.2 上为强制功能,具有五个努力级别(从 minimal 到 xhigh),默认级别为 medium;在 Claude Haiku 4.5 上为可选功能,该模型本身不具备推理能力,除非你开启扩展思维并为其设置思维预算。

Inputs — Meta advertises text, images, video, audio and PDF; Claude Haiku 4.5 takes text and images. Both return text.

Weights and providers — both closed. Muse Spark 1.2 is served only by Meta's own Model API; Claude Haiku 4.5 is available from the vendor and through the usual cloud resellers and aggregators.

发布时间——Muse Spark 1.2 才发布几天。Claude Haiku 4.5 自 2025 年 10 月 15 日起就已投入生产,知识截止日期为 2025 年 7 月,背后积累了九个月的实战经验。

指数差距是真实存在的。但每个人都会引用的那个数字却不是。

Artificial Analysis 在 Intelligence Index 智能指数上给 Muse Spark 1.2 打出 54 分,在 185 个模型中排名第 13。其对 Claude Haiku 4.5 的当前评分是 24。把这两个数字放在一起,就是一个头条新闻;但仔细看看这些评分实际代表什么,这个头条就站不住脚了。

54 是 Muse Spark 1.2 在 xhigh 下的评估结果,这是它最昂贵的推理设置。24 是 Claude Haiku 4.5 作为非推理模型(关闭思考)评估的结果,并且 Artificial Analysis 将其标记为估计值而非已完成运行。在同一个指数的更早版本中,也就是 v4.1 重新加权之前,Artificial Analysis 将启用推理的 Claude Haiku 4.5 评为 55,而关闭推理评为 42。那些旧数字也不能与 54 进行比较,因为指数版本会重新缩放,v3 时代的分数并不等同于 v4.1 的分数。

因此,诚实的说法比排行榜所显示的更为有限:Muse Spark 1.2 在最大努力下于当前指数上得分为 54;目前尚无 Claude Haiku 4.5 开启扩展思维的已发布 v4.1 分数,因此这两者在全力状态下尚不存在同口径对比。 任何向你展示 54 对 24 的人,都是在拿一个全力深思的模型与一个被要求不要深思的模型作比较。

Where the vendor has published a number, it is a specific one: Claude Haiku 4.5 scores 73.3% on SWE-bench Verified, averaged over 50 trials with a 128K thinking budget. That is a vendor figure, and the thinking budget in it is enormous for a model marketed on speed — but it is the same evaluation the industry quotes for frontier models, and it puts Haiku in a bracket its price does not suggest. Meta's counterpart claims for Muse Spark 1.2 are Terminal-Bench 2.1 at 82.9% and DeepSWE v1.1 at 59.3%, also vendor-run, on Meta's own harness with each competitor paired to its own agent product. Two vendors, two benchmarks, no overlap. There is currently no single evaluation on which both of these models have a published score from the same evaluator.

四个延迟数字,两个不同的故事

延迟是“同价”框架从好奇变为决策的关键所在,也是测量来源最为重要的环节。

在基准测试工具中,Artificial Analysis 记录到 Muse Spark 1.2 在 xhigh 设置下的首 token 时间为 26.12 秒,而 Claude Haiku 4.5 为 1.00 秒。26 倍的差距,如果仅凭这一点,比较就已经结束了。

在生产环境中,情况发生了逆转。在OrcaRouter上七天的真实流量中——调用方使用他们实际想要的任何努力程度——Claude Haiku 4.5的p50首token延迟为3.84秒,p95为10.00秒,每秒输出214个token,错误率为1.0%。目录中所收录的Meta模型版本Muse Spark 1.1的p50为1.84秒,p95为6.00秒,每秒输出519个token,且未记录到任何错误。在线上流量中,Meta模型反而更快。

两组数字都是真实的,但衡量的对象不同。实验室数据是每个模型在冷缓存(cold cache)下进行最大程度深思熟虑的结果,适合测试性能上限,却不适合衡量聊天场景。生产环境数据包含缓存命中、短提示、低推理强度以及真实应用会做的所有其他事情,适合衡量中位数,但完全无法反映最坏情况。陷阱在于引用其中一组数字,却用另一组的逻辑来推理。如果你在确定面向用户的超时时间,p95 才是你要用的数字。如果你想知道深度智能体步骤在墙钟时间上的成本,基准测试数字更接近真相——但诚实的告诫是,Muse Spark 1.2 本身还没有这样的生产环境数据,因为它太新了,尚未被广泛路由。

两家实验室都向你推销了一个子代理,但它们所指的并不相同。

The vendor's pitch for Claude Haiku 4.5 was explicit about hierarchy: Sonnet-class models plan and orchestrate, Haiku instances execute in parallel underneath. Cheap, fast, disposable, many at once. The product design follows — a 200K window is plenty for a scoped subtask and deliberately not enough for a whole repository, thinking is off by default because a worker that deliberates is a worker that costs orchestrator money, and the vendor's own marketing names real-time chat, customer service and pair programming as the target.

Meta对Muse Spark 1.2的宣传声称其兼顾同一层级的两端。Meta的文案称,该模型既可以作为收集上下文、制定计划并委派任务的主代理,也可以作为在主代理之下并行执行的子代理。与其一同发布的终端代理Muse Code,在可精确重放的事件日志运行时上,于隔离的工作树中为每个任务协调多个持久子代理。百万token的窗口和始终开启的推理能力是规划者所需要的;而它们恰恰是你不会为扇出型工作节点所选择的,因为每个并行实例都在为你不想要的深思熟虑付费。

Read as architecture rather than marketing, that is the real difference: the vendor built a worker and expects you to bring an orchestrator; Meta built an orchestrator and claims it can also work. If you already run a hierarchy, the interesting experiment is not choosing between them — it is Muse Spark 1.2 planning and Claude Haiku 4.5 executing, which is a shape neither vendor will sell you as a package.

每一步实际花费多少

对于推理模型,按每百万令牌计费说明不了什么,因为模型自己决定消耗多少令牌。应该按步骤定价。

以一个限定范围的代码任务为例:输入 60,000 个仓库上下文 token,输出 3,000 个补丁 token。客观来说,Claude Haiku 4.5 的输入成本为 $0.060,输出成本为 $0.015 — 7.5 美分。Muse Spark 1.2 的成本为 $0.075 加上 $0.013 — 8.8 美分。差距小到可以忽略,正如混合费率所承诺的那样。

Now add the thing the blended rate hides. Muse Spark 1.2 cannot turn reasoning off, and at its default medium setting a step like this plausibly emits a few thousand reasoning tokens before the patch — they bill as output. Add 3,000 and the same step is $0.075 + $0.026 = 10.1 cents, about 35% above Haiku on identical inputs. Claude Haiku 4.5 with thinking off pays nothing for deliberation and stays at 7.5 cents; with a large thinking budget it will exceed Muse Spark 1.2, because Claude Haiku 4.5 charges $5.00 per million output against Meta's $4.25.

缓存再次改变了重点。保持仓库上下文稳定以命中缓存后,Haiku 的输入成本降至 $0.006,Muse Spark 1.2 的降至 $0.009——输入侧完全不再重要,整个比较变成了一场围绕输出 token 的较量,拼的是每个模型的思考量。在上下文重复的工作负载上,缓存键的稳定性比模型选择更有价值,这两个模型皆是如此。

The 1M window is the one place the arithmetic is not close. Anything that genuinely needs more than 200,000 tokens in a single call cannot be run on Claude Haiku 4.5 at any price. You either chunk and orchestrate — which is what the vendor intends — or you use a model with the room.

尝试两者而无需第二份合同。

Claude Haiku 4.5 is in the OrcaRouter catalog at $1.00 and $5.00 — the vendor's list price, because OrcaRouter runs 0% markup and passes provider pricing through untouched, which also means a vendor price change is live on our side the same day rather than at the next contract cycle. The production latency figures above come from that same routing layer.

Muse Spark 1.2不在目录中。Meta的Model API目前是唯一可以调用它的地方,所以如今要进行真正的正面比较,意味着一个集成通过我们,另一个直接与Meta。Muse Spark 1.1是托管的,价格与Meta对1.2收取的$1.25和$4.25完全相同,这使它成为该系列成本和延迟形态的合理替代,但并不能替代1.2所宣传的编码增益。当1.2变得可路由时,比较就变成了对同一个密钥的一行模型字符串更改——而对于上述编排器加工作器的实验,路由DSL是您将规划器和扇出工作器连接成单个调用、而不是自己构建交接的方式。

按作品的形态来挑选,而不是按分数。

选择 Claude Haiku 4.5,当工作范围明确且用户等待时。支持代理、分类、提取、聊天、结对编程补全,以及代理层次结构中的并行工人层级。它是一款已知产品,拥有九个月的现场历史;思考是可选的,因此你只需在需要时支付推理费用;200K 对于几乎任何有范围的子任务来说都足够了。其已发布的 SWE-bench Verified 结果表明,它远比价格所暗示的更强大,前提是你愿意花费该结果所使用的思考预算。

选择Muse Spark 1.2当工作单元是仓库而非请求时。多文件重构、长时间调试、整个项目生成、长时间跨度的智能体循环——这些场景下没有人盯着加载指示器。百万词元的上下文窗口消除了一个 Haiku 无论付出何种代价都无法解决的分块问题;而那种会毁掉聊天体验的强制性推理,在错误计划比缓慢计划代价更高时,恰恰是正确的默认设置。

What should decide it is not the index number. Ask two questions instead: does a single call ever need to see more than 200,000 tokens, and is there a human waiting on the first token? Yes to the first points at Meta. Yes to the second points at the vendor. Yes to both means you want two models, which is the honest answer more often than either vendor's marketing admits.

本文中的对比1

根据本文内容识别 · 基准测试:Artificial Analysis · 每日更新

© 2026 OrcaRouter

推理服务商

运营推理平台?让您的模型上线 OrcaRouter。

联系我们

加入我们的社区

DiscordEmailXGitHubYouTube