
Muse Spark 1.2 vs Gemini 3.5 Flash-Lite: Two Labs, Two Definitions of Subagent
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxNEWMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2041 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
Within a fortnight of each other, two labs shipped a model and described it with the same word. Per its own model card, Gemini 3.5 Flash-Lite is "suited for subagents that execute focused tasks within complex, multi-agent workflows." Meta says Muse Spark 1.2 can serve "as either a primary agent for planning and delegation or as a subagent executing in parallel." Same vocabulary, and it took Artificial Analysis running both through its Intelligence Index to show how differently they mean it: Flash-Lite finished the suite in 43 million output tokens, Muse Spark 1.2 needed 95 million. One of these models thinks briefly and often; the other thinks hard and long. Both call the result a subagent.
That token ratio is the most decision-relevant number in the comparison, because a subagent tier is a thing you fan out — twenty at a time, a hundred at a time — and fan-out multiplies whatever per-call habits the model has. What follows works through what each lab actually built, what the fan-out arithmetic looks like at real prices, and where the cheap one turns out to be the expensive choice.
What each lab means by the word
The Flash-Lite version is a role definition by size. Gemini 3.5 Flash-Lite is the cheapest tier in its family, and that positioning is notable for tying thinking levels directly to multi-step subagent workloads — the first time a tier in that family has been described by agent role rather than by size or speed. Reasoning is mandatory but the default effort is minimal, the dial tops out at high, and the model is built to be spawned in bulk and answer fast. The vendor reports 54.2% on SWE-Bench Pro against 49.6% for Gemini 3 Flash, and 74.0% on OSWorld-Verified against 65.1% — both vendor-reported figures, not independently reproduced, and both framed as agentic gains rather than raw-intelligence ones.
Meta's version is a role definition by topology. Muse Spark 1.2's product copy describes a model that can gather context, make a plan, and delegate execution across parallel subagents — or be one of those subagents itself. Its reasoning dial runs minimal through xhigh with a default of medium, and Meta names the target workloads explicitly: multi-file refactors, extended debugging sessions, whole-repository generation. This is a model designed to be the orchestrator that can also do the work, which is a different product from a model designed to be spawned twenty at a time.
Neither lab published a benchmark for the specific claim. Meta shipped 1.2 on August 5 with no launch post at all — its Muse Spark 1.1 announcement page still describes the previous version.
The gap the index actually measures

Independent scores, from Artificial Analysis running the same nine-evaluation composite on both:
• Intelligence Index — Muse Spark 1.2 (xhigh) 54, rank #13 of 185. Gemini 3.5 Flash-Lite 36, rank #20 of 163. Eighteen points, against a class median of 32.
• Output speed — 165.0 tokens per second against 343.6. Flash-Lite streams more than twice as fast.
• Time to first token — 26.12 seconds against 10.30, both at maximum effort. Artificial Analysis notes Flash-Lite's 10.30s is itself at the high end for reasoning models in its price tier.
• Verbosity — 95M output tokens against 43M for the same suite, with a 65M median across all models. Muse Spark 1.2 is 46% above median; Flash-Lite is 34% below.
• Cost to run the index — $637.85 against $153.08.
• Per-dimension — Artificial Analysis's sub-indices put Gemini 3.5 Flash-Lite at 49.3 coding and 26.8 agentic. There is no published breakdown for Muse Spark 1.2 yet; its predecessor scored 71.3 and 37.5.
Eighteen index points is not a rounding difference. These are not two candidates for the same slot.
The fan-out arithmetic
Per-token price is the wrong lens for a subagent tier, because the whole point of the tier is that you run many of them. Work an example at list prices — $0.30 in and $2.50 out for Gemini 3.5 Flash-Lite, $1.25 in and $4.25 out for Muse Spark 1.2.
An orchestrator dispatches twenty subagents. Each reads 25,000 tokens of scoped context and writes 1,500 tokens back.
• Gemini 3.5 Flash-Lite — 20 × ($0.0075 + $0.00375) = about 22.5 cents per fan-out.
• Muse Spark 1.2 — 20 × ($0.03125 + $0.006375) = about 75 cents per fan-out.
Three and a third times, which sounds manageable. Now apply the measured verbosity difference. If Muse Spark 1.2 writes proportionally more than Flash-Lite the way it did on the index — roughly 2.2 times as many output tokens for identical work — those 1,500-token replies become about 3,300, and the fan-out rises to about 91 cents. Four times, not three. At a hundred fan-outs an hour in a busy agent system that is the difference between $22 and $91 an hour, or roughly $600,000 a year of difference at continuous load.
Two pricing details make the real gap wider still in one direction and narrower in the other:
• Flash-Lite bills internal reasoning at the output rate. Its listing prices internal reasoning tokens at $2.50 per million — the same as visible output. Thinking is not free, it is merely invisible on the response. If a subagent deliberates 2,000 tokens before answering, add half a cent per call, or 10 cents across the fan-out.
• Caching favours Muse Spark 1.2 in ratio terms. Cached input reads at $0.15 per million against a $1.25 list — an 88% discount. Flash-Lite reads cached input at $0.03 against $0.30, a 90% discount, but also charges about $0.083 per million to write the cache, where Muse Spark 1.2 publishes no separate cache-write fee. For a fan-out where every subagent shares the same large prefix, the gap closes meaningfully.

Production latency is where Flash-Lite pulls furthest ahead, and benchmark numbers hide it. On OrcaRouter, across a rolling week of real traffic, Gemini 3.5 Flash-Lite shows a p50 time to first token of 816 milliseconds and a p95 of 4.57 seconds. Muse Spark 1.1 — the closest available proxy, since 1.2 has no telemetry yet — shows a 1.84-second p50 and a 6.00-second p95. When twenty subagents run in parallel, the slowest one sets the wall clock, so p95 is the number that governs your fan-out latency, not p50.
Where the cheap one is the wrong call
Four constraints on Gemini 3.5 Flash-Lite that do not show up in a price comparison and will show up in an incident review.
• A 65,536-token output ceiling. Half of what most models in this class allow, and a hard wall for a subagent asked to generate a large file or return a long structured document. Muse Spark 1.2's endpoint declares no maximum completion length at all — which is unusual enough to test rather than trust, but it is not a documented cap.
• It is an unmoderated endpoint. The Flash-Lite listing carries no moderation flag; Meta's for Muse Spark 1.2 does. That is a genuine advantage for some workloads and a compliance problem for others, and it should be a deliberate choice rather than a discovery.
• Expensive built-in search. Flash-Lite's web-search tool is priced at roughly $14 per thousand calls against $2.50 for Muse Spark 1.2. If your subagents research rather than just read, the cheap model becomes the expensive one — five and a half times over — and search volume in a fan-out architecture scales with the fan-out.
• Its price went up, not down. Gemini 3.5 Flash-Lite lists at $0.30 and $2.50 where the previous Gemini 3.1 Flash-Lite was $0.25 and $1.50. Output is 67% more expensive than the tier it replaces. If you sized a subagent budget on the older model, resize it.

One more asymmetry worth stating plainly: Muse Spark 1.2 is served by exactly one provider — Meta — with no third-party hosts and no fallback. Gemini 3.5 Flash-Lite has the ordinary first-party multi-region serving footprint. For an orchestrator that is a single point of failure for your entire agent tree; for a worker tier it is a single point of failure for the fan-out.
The pairing most teams will actually land on
Read against each other, these two models are not competitors so much as the two halves of one architecture, and Meta's own product copy says as much when it describes the primary-agent role separately from the subagent role.
The shape that follows from the numbers: Muse Spark 1.2 as the orchestrator, Gemini 3.5 Flash-Lite as the workers. One expensive, deliberate call that reads the repository, holds the plan and decides what to delegate — where 26 seconds of thinking and 95M-token-scale verbosity buy you a plan worth executing — followed by twenty cheap, fast, narrow calls that each do one scoped thing and return in under a second at p50. You pay frontier-adjacent rates once per task and budget rates per unit of work, which is exactly the cost curve a fan-out architecture wants.
Invert it and the economics collapse. Twenty Muse Spark 1.2 subagents at 91 cents a fan-out is four times the bill for work that an 18-point-weaker model finishes in a third of the time, and a Flash-Lite orchestrator at index 36 will produce plans the workers cannot rescue.
Building that split across two vendors used to be the awkward part. Gemini 3.5 Flash-Lite is in the OrcaRouter catalog at $0.30 and $2.50 — vendor list price, passed through at 0% markup, so a price change lands on our side the same day rather than after a contract cycle — and Muse Spark 1.1 is there at $1.25 and $4.25 on the same basis, both behind one OpenAI-compatible endpoint. That means the orchestrator-and-workers pattern is two model strings in one codebase rather than two SDKs, two keys and two invoices, and automatic failover covers the case where one vendor has a bad afternoon in the middle of a fan-out. To be exact about what is available: Muse Spark 1.2 itself is not in the catalog yet — Meta serves it directly as its sole provider — so today the closest version of this stack runs 1.1 in the orchestrator slot.
Choose by role, not by score
If you are picking a worker tier — document parsing, data extraction, scoped file edits, classification, the narrow repetitive middle of an agent tree — Gemini 3.5 Flash-Lite is the better instrument, and the 18-point index deficit is mostly irrelevant because you are not asking it to reason. Watch the output ceiling and the search pricing.
If you are picking the model that decides what the workers do, Muse Spark 1.2 is the stronger candidate of the two, with the caveats that it has no independent per-dimension agentic score, no arena record of any kind, no production telemetry, and one provider. Those are the gaps to close before it holds a production plan.
And if you were hoping one model could do both jobs, the honest answer from the measurements is that neither does. Flash-Lite cannot plan at index 36; Muse Spark 1.2 cannot be spawned twenty at a time at 91 cents a fan-out. The word "subagent" appearing in both product descriptions is a marketing coincidence, not a sign they belong in the same slot.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
