
Eleven v4 vs Cartesia Sonic 3.6: A 43-Point Elo Gap Against a 39% Cheaper Voice
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
ElevenLabs Eleven v4 sits 43 Elo points above Cartesia Sonic 3.6 on Artificial Analysis' Provider Voice Arena — 1,319 to 1,276 — and it wins the Controlled Voice board too, second against Sonic 3.6's fourth. It loses on price by a wide margin in the other direction: $80.00 per million characters against $49.00, and about nine times the per-character rate on the vendors' own rate cards during the launch window. This is the rare matchup where the scoreboard is clean and the decision still isn't, because the two models are not really competing for the same job.

What each model is, in one paragraph each
ElevenLabs Eleven v4 shipped on September 28 as the company's flagship speech synthesis model, alongside Eleven v4 Turbo. The documentation gives it 90-plus languages, a 10,000-character input cap per request, multi-speaker dialogue with stable speaker identity, improved IPA phoneme support, and inline directives that let you write delivery into the script. Voice cloning runs from ten seconds of audio for instant clones, with a professional tier above that.
Cartesia Sonic 3.6 is the current release of a model line Cartesia positions on speed and naturalness. Its own product page calls it the fastest and most natural text-to-speech model, claims sub-90 ms latency, natively multilingual operation across 44 languages, voice cloning from ten seconds of audio, custom pronunciation dictionaries, and a compliance set spanning HIPAA, SOC 2 Type 2, GDPR and PCI. Both vendors publish latency figures for their own models; neither set is independently verified, and they are not measured the same way.
The scoreboard, dimension by dimension
• Blind preference, each vendor's own voices — Eleven v4 rank 1, Elo 1,319 vs Sonic 3.6 rank 2, Elo 1,276. A 43-point gap.
• Blind preference, matched cloned voices — Eleven v4 rank 2, Elo 1,157 vs Sonic 3.6 rank 4, Elo 1,138. A 19-point gap, narrower than the first.
• Arena price normalisation, per million characters — Eleven v4 $80.00 vs Sonic 3.6 $49.00. Sonic is 39% cheaper on the board's common denominator.
• Vendor rate card, per thousand characters — Eleven v4 $0.022 promotional, $0.08 after October 12 vs Sonic 3.6 $0.049. Sonic is 39% cheaper once the promotion ends.
• Language coverage, vendor-documented — Eleven v4 90-plus vs Sonic 3.6 44. Roughly double.
• Input cap per request — Eleven v4 10,000 characters vs Sonic 3.6 not stated on its product page.
• Latency, vendor-stated — Eleven v4 Turbo about 150 ms median to first speech vs Sonic 3.6 sub-90 ms. Again, different tests.
• Voice cloning — both from ten seconds of audio, both with a higher professional tier.
• Enumeration controls — Eleven v4 documents IPA phoneme input and a professional clone tier; Sonic 3.6 documents custom pronunciation dictionaries.
• Compliance posture — Sonic 3.6 states HIPAA, SOC 2 Type 2, GDPR and PCI. ElevenLabs states its compliance separately from the model page and does not put a matching list next to v4.

Two of those rows are the ones that decide real projects. The Elo gap is 43 points on the vendor-voice board and 19 on the cloned-voice board — the Sonic line closes by more than half when the speakers are held constant. That is the signature of a model whose tuning is competitive and whose default voice cast is not, and it tells you Sonic's problem in this matchup is presentation rather than synthesis quality.
Why the price gap is smaller than it looks
The headline rate comparison is misleading in both directions. Eleven v4's $0.022 per thousand characters is a promotional figure with an expiry date of October 12; the number that survives the window is $0.08, which is more than 60% above Sonic 3.6's rate card. But the arena normalisation does not care about the promotion: it prices the models at their list rates and lands on $80.00 versus $49.00, which is the comparison that will still be true in November.
There is a second wrinkle. Cartesia does not publish a per-character rate on its product page, so the $0.049 figure comes from the arena normalisation rather than the vendor, and the two meters may not tally exactly — normalisation makes assumptions about how a provider bills. Treat the ordering as reliable and the precise percentages as approximate.
A worked month at five million characters: Eleven v4 on the promotional rate costs $110, and $400 after October 12. Sonic 3.6 costs $245. So for the next two weeks Eleven v4 is less than half Sonic's cost, and after that it is 63% more. Anyone writing a procurement comparison today and shipping it next month will get this backwards unless they say which rate they used.
The segment where Sonic 3.6 is the correct answer
There are production voice workloads where the arena boards do not measure the thing that matters, and Sonic 3.6 is built for them.
Real-time conversational agents are the clearest case. Sub-90 ms latency is not a marketing number in a phone agent; it is the difference between an interruption feeling handled and feeling like a dropped call, and ElevenLabs publishes no equivalent figure for Eleven v4 itself — only for the Turbo tier at about 150 ms median to first speech. If your agent's socket budget is tight, the 43 Elo points may be worth less to you than the milliseconds, and your own barge-in testing will settle it faster than any board.
Regulated deployment is the second. Cartesia states HIPAA, SOC 2 Type 2, GDPR and PCI on the product page. A compliance list is not an Elo and cannot be substituted by one; if a healthcare or payments reviewer is in your release path, that row of the scoreboard outweighs the other nine.

Deterministic pronunciation is the third, and it is subtler. Eleven v4's documented IPA phoneme input is a strength, but Sonic 3.6's custom pronunciation dictionaries are the workflow that teams with large product-name vocabularies tend to already have built. Porting a dictionary is real work; the model that consumes the artifact you own has a head start.
Where Eleven v4 is simply the better model
Having said all that, the boards measure something and Eleven v4 won both. On Provider Voice it leads Sonic 3.6 by 43 points with 1,674 samples behind the estimate, and the previous flagship — ElevenLabs Eleven v3 at 1,169 — sits 107 points behind Sonic 3.6, which means ElevenLabs closed a real gap in one generation rather than defending a lead.
The language count is the second unambiguous win. 90-plus against 44 is roughly double, and for a product shipping beyond English and a handful of European languages, that is a coverage question rather than a quality one — Sonic 3.6 may not have a voice at all for some of your locales. And the multi-speaker dialogue handling plus inline delivery directives are a genuine capability difference: writing a laugh into the script is a workflow Sonic 3.6 does not document, and reproducing that effect on a model without it means post-processing or a second take, which costs more than the Elo gap suggests.
Running either one, and the part we can't help with
OrcaRouter hosts neither of these models. ElevenLabs and Cartesia are not in our catalogue, so nothing in this comparison can be called through us, and the honest answer to "how do I get it on OrcaRouter" is that you cannot — you go to the vendor. What we do handle is the case where a speech stack ends up spanning several providers rather than one: a single API key across 200-plus models, with provider list prices passed through at 0% markup so that a vendor price change is live on our side the same day it takes effect. For a decision that hinges on whether you are paying $110 or $400 for the same month's characters, that timing is not a small thing.
Where a voice path is customer-facing and cannot drop, automatic failover moves traffic to a healthy upstream when one degrades. That is the pattern that makes a two-week promotional window worth exploiting: you can point production at the cheaper model for a fortnight without rewriting the integration, and move back when the rate reverts.
Who should switch, and who should wait
Switch to Eleven v4 if your product is judged on how it sounds, if you ship in more than a couple of dozen languages, or if your users hear whichever voice the vendor chose for them — that last group is exactly who the Provider Voice board represents and the gap there is the largest of any row.
Stay on Sonic 3.6 if you are running real-time conversation under a latency budget, if a compliance reviewer signs off your stack, or if you have a pronunciation dictionary you cannot afford to rebuild. On those three rows Sonic is not behind; it is the only option with a published answer.
Wait, in either case, if your decision depends on price. Two weeks from now Eleven v4's rate card changes by a factor of nearly four and Sonic 3.6's does not, and a comparison written today will read as wrong by the middle of October. The Elo ordering will still hold. The money will not.
