A generated hero card for ElevenLabs Eleven v4 vs SpeechifyAI Simba 3.2 with two score cards: Eleven v4 at Elo 1,319, rank 1 and $80.00 per 1M characters; Simba 3.2 at Elo 1,240, rank 7 and $6.58 per 1M characters; captioned '79 Elo points against roughly twelve times the price'. Footer: 'Provider Voice Arena, per Artificial Analysis, September 2026.' The OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

Eleven v4 vs Simba 3.2: The 79-Point Gap That Costs Twelve Times as Much

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On Artificial Analysis' Provider Voice Arena, ElevenLabs Eleven v4 is first at an Elo of 1,319 and SpeechifyAI Simba 3.2 is seventh at 1,240 — 79 points apart, on a board where the top ten span barely 113 points. The prices are nowhere near as close: $80.00 per million characters for ElevenLabs' new flagship, released September 28, 2026, against $6.58 for Simba 3.2. That is a gap of roughly twelve times, and it is the largest price-to-quality spread among any two models in the top ten. Whether that spread is a bargain or a trap depends entirely on which of the two you are replacing.

The board, read properly

The figures worth carrying into a decision are not the two headline numbers alone. Rank on a preference board is a moving target; the interval around each estimate is what tells you how much of the ordering is signal.

• ElevenLabs Eleven v4 — rank 1, Elo 1,319, ±19, 1,674 appearances, eight arena voices, $80.00 per million characters, released September 28, 2026

• SpeechifyAI Simba 3.2 — rank 7, Elo 1,240, ±14, 2,535 appearances, eight arena voices, $6.58 per million characters, released July 1, 2026

• SpeechifyAI Simba 3.0 — Elo 1,123, $6.58 per million characters, released February 19, 2026

A generated scoreboard comparing ElevenLabs Eleven v4 and SpeechifyAI Simba 3.2 across six dimensions: Provider Voice rank and Elo (1 / 1,319 against 7 / 1,240), interval and samples (plus or minus 19 on 1,674 against plus or minus 14 on 2,535), board price ($80.00 against $6.58 per 1M characters), rate card ($0.022 per 1K characters until October 12 against $6.58 per 1M unchanged), arena voices (8 against 8) and release date (September 28, 2026 against July 1, 2026). Footer: 'Ranks, Elo, intervals and board prices per Artificial Analysis, Sept 2026.' The OrcaRouter logo sits in the bottom-right corner.

Two things fall out of that. First, the Eleventh v4 interval runs roughly 1,300 to 1,338 and Simba 3.2's runs about 1,226 to 1,254; the ranges are separated by around 46 points of clear air, so the ordering is not a coin flip. Second, and more usefully, Simba 3.2 improved exactly 100 Elo points over Simba 3.0 at an identical board price (1,240 against 1,140). SpeechifyAI shipped a generational gain and did not raise the rate, which is the opposite of what ElevenLabs did with this launch.

A screenshot of Artificial Analysis' Provider Voice Arena leaderboard. ElevenLabs Eleven v4 is first at Elo 1,319 with an interval of plus or minus 19, 1,674 samples, eight arena voices and $80.0 per 1M characters. Cartesia Sonic 3.6 is second at 1,276, Google Gemini 3.8 Flash TTS third at 1,267 and $16.5, Google Gemini 3.8 Flash-Lite TTS sixth at 1,241 and $11.0, and SpeechifyAI Simba 3.2 seventh at Elo 1,240 with an interval of plus or minus 14, 2,535 samples, eight arena voices, released Jul 2026 at $6.6 per 1M characters.

What twelve times the price buys, concretely

The Elo number is the abstract part. Here is the part you can put in a budget line. At five million characters a month — roughly 100 hours of finished speech — the comparison reads like this:

• Eleven v4 during the launch window, at $0.022 per 1,000 characters — $110 a month

• Eleven v4 at list, $0.08 per 1,000 characters after October 12 — $400 a month

• Simba 3.2, at the board's normalised $6.58 per million — about $33 a month

• ElevenLabs Eleven v3, the previous flagship, at $0.08 per 1,000 characters — $400 a month

For two weeks the gap is a factor of three. After October 12 it is a factor of twelve, and it stays there, because the $80.00 per million on the board is the list rate rather than the promotion. Any comparison written this week that quotes the promotional rate as if it were the standing price will be wrong by the middle of October, and wrong in the direction that makes ElevenLabs look cheap.

Where Simba 3.2 is the correct answer

Seventh on the board is not a bad model. It is above Google's Gemini 3.1 Flash TTS, above ElevenLabs' own v3 Conversational at 1,197, and above Cartesia's previous-generation Sonic 3.5 at 1,185. Three workloads make it the obvious pick.

A screenshot of Artificial Analysis' Simba 3.2 model page, titled Simba 3.2 Quality Elo, Speed & Price Analysis, credited to SpeechifyAI with an Elo of 1,239.94. The Provider Voice Arena Preference Elo bar chart shows Simba 3.2 highlighted in blue at 1,240 in the middle of the row, with Eleven v4 at 1,319 at the left, and the model selector reads '27 of 96 models'.

Long-form volume is the first. Audiobook back catalogues, podcast archives and accessibility reads are billed per character and judged over hours, not seconds, and at $6.58 per million the marginal cost of a hundred more hours is trivial in a way it never is at $80.00. The 79-point preference gap is a difference a listener would notice in a side-by-side; over six hours of narration, most listeners notice the bill less.

Default-cast workloads are the second, and this is where the board's construction matters. Provider Voice measures each vendor using its own voices, so Simba 3.2's rank was earned with eight arena voices carrying its own character. If your product ships whichever voice the vendor chose, you are buying exactly what the board measured, and you are buying it at a seventh of the price of the top-ranked alternative.

Already-integrated Simba deployments are the third. If you are on Simba 3.0, the upgrade to 3.2 is 100 Elo points for zero change in rate and, presumably, no change in integration. That is the single best price-performance move in this entire comparison, and it involves neither ElevenLabs nor us.

Where the extra twelve times is not optional

The gap that Simba 3.2 does not close is the one the arena cannot score. Two capability differences separate these models.

Script direction is the clearest. Eleven v4 documents inline audio tags that let you write a performance into the text — [laughs], [whispers], [said angrily in French accent] — plus natural-language direction prompts and improved International Phonetic Alphabet support for custom pronunciations. Reproducing that on a model without it means a second take, a human editor, or a different voice. Those cost more per hour than the character delta does.

Language coverage is the second. Eleven v4 documents 90-plus languages against a 10,000-character input cap. SpeechifyAI does not publish an equivalent count for Simba 3.2 that we could verify, so treat the comparison as unproven in ElevenLabs' favour rather than proven: if your product ships in two dozen locales, confirm coverage on your own list before you assume either model has a voice for them.

A third difference is worth naming because it is invisible on the boards: Eleven v4's instant cloning runs from ten seconds of audio. If your workflow is built on cloning specific people rather than using a vendor cast, the sample-length bar matters more than 79 Elo points, and it is a per-model fact you have to check rather than a general property of speech APIs.

The routing question, answered honestly

Neither SpeechifyAI Simba 3.2 nor any ElevenLabs model is in the OrcaRouter catalogue. There is nothing here to call through us, and the right place to reach both is the vendor's own API — SpeechifyAI's for Simba, ElevenLabs' for Eleven v4. We would rather say that plainly than dress up a comparison as a sales pitch.

The place a single key does earn its keep is the fortnight after a launch, when the sane engineering answer is "run both on our own scripts and decide from that". Doing it across two vendors normally means two accounts, two billing relationships and two request shapes before you have a single comparison point. One API key across 200-plus models with provider list prices passed through at 0% markup removes that setup cost, and automatic failover means the model you are trialling can sit behind a production path without becoming a single point of failure for it.

Who should switch, and who should not move at all

Switch to Eleven v4 if voice quality is the product: if users hear a default voice you did not choose, if you need performance direction written into the script, or if your cloning workflow is measured in seconds of source audio. The 79-point gap is real and the capability delta behind it is larger than the number suggests. Budget for $0.08 per 1,000 characters, not $0.022.

Stay on Simba 3.2 if you are producing volume: long-form narration, archival content, anything where the per-character rate compounds over hours of audio. You are giving up the top of the board and keeping roughly eleven-twelfths of the money. And if you are on Simba 3.0, upgrade to 3.2 first and re-measure — 100 Elo points for nothing is a better return than anything either side of this comparison is offering.

What you should not do is decide from this article alone. The honest limitation of everything above is that the arena measures blind preference on vendor voices at one point in time, and the ElevenLabs rate card changes on October 12. Re-run the arithmetic after that date, with your own monthly character count, before you move a production path.