
GPT-6 Sol vs Muse Spark 1.2: One of Them Has Ears
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Put GPT-6 Sol and Muse Spark 1.2 side by side and the price columns are merely different, while the modality columns are not close at all. Muse Spark 1.2 accepts text, image, video, file and audio. GPT-6 Sol accepts text, image and file. Audio and video are the difference, and they are not a rounding detail: they separate a model that can be handed a meeting recording from one that has to be handed a transcript of it. GPT-6 Sol, the middle tier of that generation, shipped 22 September 2026; Muse Spark 1.2, Meta's current checkpoint in the Muse Spark family, shipped 5 August 2026 as a drop-in update to Muse Spark 1.1 on the same Standard tier at identical pricing.
That last point matters for how this page should be read. Muse Spark 1.2 is not a new generation and should not be written up as one — Meta's own description is an updated checkpoint with slightly higher performance, served at unchanged rates. So the interesting question is not what 1.2 adds over 1.1. It is what the Muse Spark family has that the GPT-6 family does not, and whether you need it.
The two spec sheets, on the dimensions that differ
• Input modalities — GPT-6 Sol text, image and file vs Muse Spark 1.2 text, image, video, file and audio
• Output — both text only
• Input price — GPT-6 Sol $2.00 per million vs Muse Spark 1.2 $1.25 per million
• Output price — GPT-6 Sol $10.00 per million vs Muse Spark 1.2 $4.25 per million
• Cached input — GPT-6 Sol $0.20 per million vs Muse Spark 1.2 $0.15 per million
• Context window — 1,050,000 tokens vs 1,048,576 tokens, effectively the same
• Long-request step — GPT-6 Sol reprices the whole request above 272,000 input tokens to $4.00 / $15.00 vs Muse Spark 1.2 no step on its card
• GPQA Diamond — GPT-6 Sol 96.1 vs Muse Spark 1.2 90.4
• Humanity's Last Exam — GPT-6 Sol 47.9 vs Muse Spark 1.2 45.5
• Long-context recall — GPT-6 Sol 83.7 vs Muse Spark 1.2 79.0
• tau-banking — GPT-6 Sol not published on this revision vs Muse Spark 1.2 34.8
• Weights — both closed, API only
The long-request line is doing quiet work in that list. Muse Spark 1.2 has no long-context repricing clause on its published card, so its $1.25 / $4.25 holds across the whole window. GPT-6 Sol's headline does not: past 272,000 input tokens the entire request moves to $4.00 / $15.00. On a full-context prompt the output comparison is therefore not ten dollars against $4.25, it is fifteen against $4.25 — three and a half times, not two and a third.
And the benchmark gap is narrower than the price gap suggests. GPQA Diamond, 96.1 against 90.4, is roughly the spread you would expect from two different reasoning-effort settings rather than from a capability difference, and on Humanity's Last Exam the two are 47.9 and 45.5, which is 2.4 points on a hundred-point scale. Muse Spark 1.2 is the cheaper model and the more broadly capable reader; GPT-6 Sol is ahead on the reasoning figures but not by the margin the price implies. Neither column dominates, which is what makes the choice worth making deliberately.

What audio and video input actually change
A model that accepts audio can be handed a recording instead of a transcript. That sounds like a convenience and is really a change in what is possible: transcription is a lossy step that discards tone, overlap, laughter, and the difference between a question and a statement read from text alone. A model that hears the file directly sees all of it. The same argument applies to video, where frame-by-frame description loses timing and a model that ingests the clip does not.
GPT-6 Sol's answer is a pipeline rather than a modality — transcribe or caption first, then pass text. That is a perfectly workable architecture and for many workloads it is the better one, because a transcript is inspectable, cacheable and cheap to store. It is worse exactly when the audio or the motion is the signal: call-centre quality review, media logging, anything where a speaker's hesitation carries the meaning.
Meta's position on this is worth reading carefully, because the audio tag is not decoration. Muse Spark 1.2 is listed with audio among its capabilities alongside vision, tools, JSON and reasoning, and Meta has shipped the family as its multimodal line, while the smaller Muse Glimmer was distilled from a Muse Spark teacher. If you are choosing between these two on input surface, the choice is not between a slightly broader spec sheet and a slightly narrower one. It is between having the modality and constructing a pipeline to approximate it.
Where the two genuinely separate
The clearest separation is agentic tool use against a simulated service, and it is the one place the numbers are far enough apart to act on. Muse Spark 1.2 records 34.8 on tau-banking and just 7.1 on the newer Terminal-Bench 4.0 revision — both third-party measurements, and both describing the same weakness from two directions. A model that reads a meeting well and then miscounts a customer's balance when asked to call an API is not a bad model; it is a model with a clearly bounded job.
Run the same check on GPT-6 Sol and the picture is consistent but thinner. Its long-context recall sits at 83.7, SciCode at 57.6 and its Terminal-Bench 4.0 at 43.9, all third-party. Its coding and terminal-agent figures on the v2.1 revision are simply absent, and the absence of a number is not a low number — anyone who fills that cell with an estimate is inventing.
One asymmetry worth naming before the numbers get compared too eagerly. Muse Spark 1.2 carries a full vendor benchmark suite from Meta alongside the third-party set, and Artificial Analysis's own launch analysis put it at 54 on the Intelligence Index at its xhigh reasoning setting, three points above Muse Spark 1.1. GPT-6 Sol lands at 47.6 on the same index in our catalogue. Those two figures are not subtractable: they come from different configurations measured at different times, and an index revised between them is not the same instrument. The honest comparative claims are the single-benchmark ones above, where both models carry a figure from the same evaluator. When a model has more published numbers it will always look better documented than one that has fewer, and on this pair much of that difference is about who reports rather than what the models can do.

Serving two models that behave differently under load
Both are on OrcaRouter, one key, and the pair is a natural routed split rather than an either-or. Send the multimodal traffic — the recordings, the clips, the document bundles with scanned attachments — to Muse Spark 1.2 at $1.25 / $4.25 with no long-context cliff. Send the reasoning-heavy and agentic traffic to GPT-6 Sol and accept the higher rate for the requests where the answer has to survive scrutiny. Through one endpoint that split is a routing rule, not two integrations.
One number on the serving side cuts against the price logic. Muse Spark 1.2 is measured at roughly 420 output tokens per second in our rolling seven-day window against GPT-6 Sol's roughly 244 — Meta's model is both cheaper and faster on our routes. Latency runs the other way on first token: a p50 time-to-first-token around 5.2 seconds for Muse Spark 1.2 against roughly 6.8 for GPT-6 Sol in the same window, so the cheaper model also starts answering sooner. Both are rolling measurements on shared infrastructure rather than a controlled benchmark, and both move between reads.
Automatic failover earns its place on a mixed pipeline like this one. When a request can arrive as audio and leave as text through one route, or as a mandatory transcription step through another, the failure mode you care about is a provider that accepts the request and then stalls on the media upload — a second configured route turns that into a retry rather than a job someone has to notice and re-run. Our catalogue passes provider list prices through unchanged, so if Meta moves the $4.25 line, or adds a long-context clause to the Muse Spark card the way other vendors already have, the change reaches you the day it happens rather than after a reseller notices.

The decision, put plainly
If your input is a recording or a clip and the timing or tone is part of the meaning, use Muse Spark 1.2. There is no configuration of GPT-6 Sol that reaches that capability through the API, and no price at which the substitute pipeline becomes equivalent. If your input is text and images and the output will be acted on, GPT-6 Sol's reasoning figures are ahead and its tool-use profile is the safer one — Muse Spark 1.2's tau-banking and Terminal-Bench 4.0 scores are the loudest signal on either sheet, and they point away from agentic work.
And note what Muse Spark 1.2 is not: a new generation. It is Meta's current checkpoint of an existing family, a drop-in replacement for 1.1 at unchanged pricing, which makes the upgrade decision free and the comparison against GPT-6 Sol a separate question. On the other side, GPT-6 Sol has already been superseded by GPT-6.1 Sol at identical short-context rates with cached input halved from $0.20 to $0.10, and OpenAI's own model page now points readers to it. If the reason you are weighing Sol rather than Muse Spark is the eight-day-old integration, that stands. If it is capability, check the successor first.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
