A hero card for 'Eleven v4 Tops the Voice Arena': a leaderboard with ElevenLabs Eleven v4 at rank 1 and Elo 1,319, Cartesia Sonic 3.6 at rank 2 and Elo 1,276, Google Gemini 3.8 Flash TTS at rank 3 and Elo 1,267, ElevenLabs Eleven v3 at rank 18 and Elo 1,169, and a ribbon reading '72% off until Oct 12'. Footer: 'Provider Voice Arena, per Artificial Analysis, Sept 2026.' The OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

Eleven v4 Tops the Voice Arena: What ElevenLabs' New Flagship Actually Won

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

ElevenLabs Eleven v4 arrived on September 28 and went straight to the top of Artificial Analysis' Provider Voice Arena with an Elo of 1,319 — 43 points clear of Cartesia Sonic 3.6 at 1,276 and 52 clear of Google Gemini 3.8 Flash TTS at 1,267. It is also, at $80.00 per million characters on that same board, the most expensive model in the top ten. And on the second arena — the one built from cloned voices rather than each vendor's own voice set — it does not win at all: it finishes second, behind a model that costs a quarter as much. Both of those are true, and the useful part of this launch is working out which one your use case lives under.

What actually shipped on September 28

ElevenLabs put two models into general availability, not one. ElevenLabs Eleven v4 is the flagship synthesis model — the company calls it its most emotive, and its documentation sets the input cap at 10,000 characters per request and language coverage at 90-plus. ElevenLabs Eleven v4 Turbo is the low-latency sibling: same 90-plus languages, a 10,000-character cap, and support for the inline audio tags that let you write a delivery direction into the script itself — a laugh, a whisper, a buzz on the line. Both are live across the company's agent, creative and API surfaces on day one, which is a faster rollout than the v3 generation got.

The announcement page is where the feature list lives and, notably, not the pricing. ElevenLabs describes a new architecture, better handling of tone, pacing and character, more reliable multi-speaker dialogue with stable speaker identity, improved IPA phoneme input, and instant voice clones built from ten seconds of audio alongside the existing professional clone tier. The one number on it is a claim about listeners: the vendor says blind comparisons preferred v4 in roughly three-quarters of cases. That is a vendor figure, unreproduced, and it should be read next to the board numbers below rather than instead of them.

Where the pricing actually is — the API pricing page — is the more consequential document, and it is covered further down. Short version: this launch is not a price rise.

Two boards, two different answers

Artificial Analysis runs two speech arenas and they are not variations on one another. Provider Voice has listeners compare each vendor's own voices. Controlled Voice clones the same eight voices — four US, four UK — for every entrant, so the speaker is held constant and only the model varies. A model with a great default voice cast can win the first and lose the second. Eleven v4 does exactly that.

• Provider Voice, Eleven v4 — rank 1, Elo 1,319, $80.00 per million characters, 1,674 samples, eight arena voices, September 2026

• Provider Voice, Cartesia Sonic 3.6 — rank 2, Elo 1,276, $49.00

• Provider Voice, Google Gemini 3.8 Flash TTS — rank 3, Elo 1,267, $16.50

• Provider Voice, the previous flagship ElevenLabs Eleven v3 — rank 18, Elo 1,169, $100.00

• Controlled Voice, Alibaba Qwen-Audio-3.1-TTS-Plus — rank 1, Elo 1,178

• Controlled Voice, Eleven v4 — rank 2, Elo 1,157, $80.00

• Controlled Voice, Cartesia Sonic 3.6 — rank 4, Elo 1,138

• Controlled Voice, ElevenLabs Eleven v3 — rank 7, Elo 1,073

• Controlled Voice, Google Gemini 3.8 Flash TTS — rank 13, Elo 1,048

• Best open-weights entry on Provider Voice — Breeze TTS 2, Elo 1,206, $34.00

A single-column scoreboard card for ElevenLabs Eleven v4 with six rows: 'Provider Voice rank: 1 (Elo 1,319)', 'Controlled Voice rank: 2 (Elo 1,157)', 'Board price: $80.00 per 1M characters', 'Rate card: $0.022 per 1K characters until Oct 12', 'Languages and cap: 90-plus, 10,000 characters', 'Latency: not stated for v4 itself'. Footer: 'Elo and board prices per Artificial Analysis; rate card and specs vendor-documented.'

The lineage jump is the headline for anyone already on ElevenLabs: against its own predecessor on the same board, Eleven v4 gains 150 Elo points under Provider Voice and 84 under Controlled Voice. That is a generational move, not a tuning pass, and the improvement is bigger on the board where the vendor gets to choose the voices — which is the honest way to say its new default cast is a real part of the gain.

A screenshot of the Artificial Analysis Provider Voice Arena leaderboard, captured in English, showing ElevenLabs Eleven v4 first at Elo 1,319 with $80.0 per 1M chars, 1,674 samples and a -19/19 confidence interval, Cartesia Sonic 3.6 second at Elo 1,276 and $49.0, Google Gemini 3.8 Flash TTS third at Elo 1,267 and $16.5, and ElevenLabs Eleven v3 eighteenth at Elo 1,169 and $100.0, under columns for samples, arena voices, release month and API pricing.

The pronunciation board we can't put a number on

Artificial Analysis does run a Pronunciation Robustness section, and it is worth describing properly because the ranking summary floating around describes Eleven v4 as topping it. The board measures the share of highlighted spans that reviewers judged pronounced correctly, with up to three reviewers per clip, and the prompts are sent exactly as written — no number expansion, no text normalisation. It is broken into sub-categories, and the point of the design is that a model which silently rewrites "Dr." or "$4.50" before speaking scores badly on purpose.

What the board does not publish, in the rendering we could read, is a citable numeric score per model. So there is no figure to quote for Eleven v4 there, and the Elo of 1,319 belongs to arena preference — how listeners liked the voice — not to pronunciation accuracy. If you see the two conflated, that is the conflation. The defensible claim from this launch is the one with numbers attached: first on Provider Voice at 1,319, second on Controlled Voice at 1,157.

The launch discount is the real news for your bill

ElevenLabs repriced the new models at the same time as it released them, and the discount is large enough to change a purchasing decision on its own. On the API pricing page, Eleven v4 lists at $0.022 per thousand characters against a stated list of $0.08 — a 72% reduction — and Eleven v4 Turbo lists at $0.011 against $0.04. Both carry the same window: the promotion runs until October 12. After that, the numbers revert, and the reverted v4 price is double what ElevenLabs' own v3 costs today at $0.08. Plan from the post-promotion number, not the promotional one.

The repricing also moves v4's position against the field on a per-character basis, which the arena normalisation already shows at $80.00 per million. Sonic 3.6 is $49.00 there, Breeze TTS 2 is $34.00, Qwen-Audio-3.0-TTS-Plus is $19.30, and Gemini 3.8 Flash TTS is $16.50. On the promotional rate, Eleven v4 at $22 per million undercuts Sonic 3.6 outright. On the list rate it is the priciest voice on the board by a wide margin. Anyone building a budget for Q1 should be looking at the second number.

A screenshot of ElevenLabs' API pricing page, captured in English, showing the Eleven v4 row at $0.022 per 1K characters with a struck-through $0.08 and a '72% off until Oct 12' badge, and the Eleven v4 Turbo row at $0.011 with a struck-through $0.04, alongside the v3 row at $0.08 and the Eleven v3 Conversational row at $0.04.

A worked month makes the spread concrete. At five million characters — a mid-size product reading roughly a hundred hours of speech — Eleven v4 during the window costs $110. At the reverted $0.08 it costs $400. Sonic 3.6 costs $245. Gemini 3.8 Flash TTS costs $82.50. The same month on ElevenLabs' own v3 costs $400 as well, which is the odd fact underneath this launch: for the next two weeks the new flagship is cheaper than the model it replaces, and the day the window closes it stops being cheaper and becomes merely better.

The latency claim is vendor-stated and tier-dependent

ElevenLabs publishes latency figures per tier, not per model family, and they are measurements the company took rather than anything independent. Eleven v4 Turbo is quoted at roughly 100 ms median inference and about 150 ms median time to first speech, measured in September 2026. Eleven v3 Conversational is quoted at about 280 ms, and the older Flash and Turbo tiers at about 75 ms. The documentation does not publish an equivalent figure for Eleven v4 itself, which is the tier you would want if you are building a real-time agent and reading a spec sheet rather than testing.

Cartesia claims sub-90 ms latency for Sonic and describes it as ranked first for naturalness, with voice cloning from ten seconds of audio, custom pronunciation dictionaries and a compliance set covering HIPAA, SOC 2 Type 2, GDPR and PCI. Those are the vendor's own claims on its own product page — the same category of evidence as ElevenLabs' latency tier numbers, and not interchangeable with the arena Elo. Where the two vendors are directly comparable is only on the boards.

Where this leaves you, and how to try it

If you are picking a voice for a product where the default cast is what your users hear, the Provider Voice result is the one that matters and Eleven v4 just took it, with a discount attached for two weeks. If you are cloning a specific person's voice and every candidate model is being asked to reproduce the same eight speakers, the Controlled Voice board is the one that matters and Eleven v4 is second — 21 Elo behind the leader and, at list price, more than four times the cost per character. Neither result is wrong; they are answers to different questions, and the second one is the question most production voice-cloning work is actually asking.

OrcaRouter does not host ElevenLabs, Cartesia or any of the other vendors on these boards, so there is nothing here to call through us and we will not imply otherwise. What we do offer is the surrounding plumbing if your speech stack ends up spanning providers: one API key across 200-plus models with provider list prices passed through at 0% markup, so a vendor price cut — including a promotional window like this one — is live on our side the same day rather than at the next invoice cycle. Where a speech path is production-critical, automatic failover swaps to a healthy upstream when one degrades, which is the practical way to test a brand-new endpoint without betting a release on it.

The date to diarise is October 12. That is when Eleven v4 stops being the cheapest good voice on the board and becomes the most expensive one, and any comparison written this week that quotes $22 per million without saying so is a comparison you should re-run in three weeks.