
Eleven v4 vs the Rest of ElevenLabs: Should You Migrate Off Eleven v3?
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
If you are already on ElevenLabs Eleven v3, the answer is yes for anything a listener hears and no for the two things ElevenLabs kept unchanged. ElevenLabs Eleven v4 gains 150 Elo points over its own predecessor on Artificial Analysis' Provider Voice Arena — 1,319 against 1,169 — and 84 points on the Controlled Voice board, where the same eight cloned voices are used for everyone. It also carries a character limit double v3's and roughly double the language coverage. What it does not change is your contract, your voice clones, or the tier structure you are invoiced against, which is why this migration is cheaper than most and still worth doing in a staged way rather than all at once.

The intra-family scoreboard
This is not a vendor-versus-vendor comparison, so the price normalisation on the arena boards behaves differently than usual: both models bill per character, on the same rate card, from the same account. The strengths and weaknesses below are measured against siblings rather than competitors.
• Provider Voice Arena — ElevenLabs Eleven v4 rank 1, Elo 1,319 vs ElevenLabs Eleven v3 rank 18, Elo 1,169. A 150-point gap.
• Controlled Voice Arena, matched cloned voices — Eleven v4 rank 2, Elo 1,157 vs Eleven v3 rank 7, Elo 1,073. An 84-point gap.
• Arena price normalisation — Eleven v4 $80.00 per million characters vs Eleven v3 $100.00. The new flagship is the cheaper of the two on the board.
• Vendor rate card — Eleven v4 $0.022 per thousand characters promotional, $0.08 after October 12; Eleven v3 $0.08. Identical at list, half price during the window.
• Input cap per request — Eleven v4 10,000 characters vs Eleven v3 5,000. Exactly double.
• Language coverage — Eleven v4 90-plus vs Eleven v3 70-plus.
• Delivery directives — Eleven v4 documents inline audio tags and IPA phoneme input; Eleven v3 documents neither on its model page.
• Latency, vendor-stated — Eleven v3 Conversational about 280 ms; Eleven v4 Turbo about 150 ms median to first speech. Different tiers, different tests.
• Voice clones — unchanged. Existing instant and professional clones carry across.
• Plan tiers — unchanged. Free 10,000 characters, Starter $6, Creator $22, Pro $99, Scale $299, Business $990.
• Migration path — Eleven v4 and Eleven v4 Turbo are both live in the API, agent and creative surfaces; the older Turbo identifiers are deprecated.
Read the price row twice. The unusual fact about this launch is that the new model is cheaper than the one it replaces, on the same board, at list, with no promotion applied — $80 per million against $100. The promotional rate widens that rather than creating it. A migration that raises quality by 150 Elo and lowers cost per character is not the trade most teams expect to be making.
Migration checklist, in the order that matters
Move the audio-quality path first, because that is where the gain is and where a regression would be obvious to your users before it is obvious to you.
• Audition the voices you actually ship. Provider Voice is an average over an arena voice set; your brand voice is not in it. The 150-point gap is a reason to audition, not a substitute for it, and the cloned-voice board's 84-point gap is the closer estimate of what a clone-based product will feel.
• Re-check your character budgets. The 10,000-character cap doubles v3's, so scripts you previously split may now fit in one request. That changes request counts, retry logic and the seam quality of stitched output — and it is the change most likely to break code that assumed a 5,000 cap.
• Test the seam. ElevenLabs specifically documents more dependable request stitching in v4. If you were splitting long-form content at v3, re-run the same content through v4 in one call and compare, rather than assuming the split is still necessary.
• Re-test your pronunciation cases with the new phoneme support. Lexicons built for v3 should be re-validated rather than trusted; improved IPA handling can mean a previously-correct override now interacts differently with the model's own guess.
• Leave your clone library alone. Clones carry; re-uploading is not needed and re-cloning risks losing a voice your users recognise.
• Keep v3 Conversational where you use it, for now. It is a conversational tier rather than a synthesis model, it is still documented at about 280 ms, and it has no published v4 equivalent. If your agent depends on it, v4 is not a like-for-like replacement for that slot — v4 Turbo is the lower-latency v4 label, and its latency tier is quoted differently.

What actually gets better, and what does not
The gains cluster in three places. Emotional range and pacing are the first — that is what the Provider Voice board is measuring, and it is also what the vendor's own announcement emphasises, along with a claim that blind comparisons preferred v4 in roughly three-quarters of cases. Treat that figure as vendor-stated and unreproduced; the arena boards are the independent half of the same story, and they agree on direction.

Multi-speaker dialogue with stable speaker identity is the second. If you have shipped two-hander content on v3 and spent effort fixing speaker bleed or drift across a long script, that is a workload where the improvement is structural rather than cosmetic, and it is one the Elo number only partially captures.
Inline delivery directives are the third — writing a sigh or a whisper directly into the script. On v3 that meant a second take, an upstream prompt, or post-processing. On v4 Turbo it is a line in the text you already have.
What does not get better is the part teams often assume improves with a new flagship: latency. ElevenLabs publishes no latency figure for Eleven v4 itself, only for v4 Turbo at roughly 100 ms median inference and about 150 ms median time to first speech, measured in September 2026. Eleven v3's conversational tier is quoted around 280 ms but is a different tier doing a different job. If your architecture is built around a hard latency budget, the honest position is that v4 Turbo is the candidate and you should measure it yourself, because the tier names do not map cleanly and the vendor numbers are not like-for-like.
The two-week window, and what happens after
The promotional pricing runs until October 12. Until then Eleven v4 is $0.022 per thousand characters and Eleven v4 Turbo is $0.011, against stated list rates of $0.08 and $0.04 respectively. After the window, v4 returns to exactly what v3 costs today — $0.08 — and v4 Turbo returns to half of v3's rate.
A worked month at five million characters: v3 costs $400. v4 during the window costs $110 and afterwards costs $400 again. v4 Turbo costs $55 during the window and $200 afterwards. So the interesting migration is arguably the Turbo one, because it is the only option on this rate card that stays cheaper than v3 after the promotion ends, and the audio-tag support that v3 lacks lives there.
The practical move is to do the quality work now while the price is a bonus, and to make the pricing decision on the post-October-12 numbers. Any comparison that quotes $22 per million for Eleven v4 without a deadline attached is describing a rate that expires in a fortnight.
Staging the switch without betting the release on it
OrcaRouter does not host ElevenLabs, so there is nothing in this migration we can move for you — the change happens on your ElevenLabs account. What we do handle is the layer around a stack that is not single-vendor. If your speech path is one of several models behind one endpoint, we serve 200-plus models through a single API key with provider list prices passed through at 0% markup, which means a vendor repricing — like this one — lands on our side the same day it lands on theirs, without a second contract or an integration change.
Where the voice path is customer-facing and cannot drop, automatic failover routes around a degraded upstream instead of failing the call. That is the shape a migration like this should take: run v4 behind the same endpoint as the incumbent, shift a slice of traffic, and let the routing layer hold the fallback while you find out whether the 150 Elo points show up in your own audio.
The decision, in one paragraph
Migrate if a listener hears your output and you are still on v3 — the quality gap is the largest single-generation jump ElevenLabs has published on these boards, the character cap doubles, the language list doubles, and the list price is lower. Do the migration in stages, start with your highest-traffic voice rather than your most complex script, and plan your budget on the October 12 numbers rather than this week's. Stay on v3 only if your integration cannot absorb a different input cap, or if v3 Conversational is load-bearing in an agent and you have not measured v4 Turbo against your own latency budget yet. Those are the two cases where patience costs less than a rushed switch.
