A generated hero card for ElevenLabs Eleven v4 vs Google Gemini 3.1 Flash TTS with two score cards: Eleven v4 at Elo 1,319 and $80.00 per 1M characters, and Gemini 3.1 Flash TTS at Elo 1,203 and $18.31 per 1M characters, captioned '116 Elo points, 4.4 times the board price'. Footer: 'Provider Voice Arena, per Artificial Analysis, September 2026.' The OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

Eleven v4 vs Gemini 3.1 Flash TTS: A 116-Point Gap, and the Cheaper Google Model

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Shopping for Goog​le's voice today, you will land on the wrong one. Goog​le Gemini 3.1 Flash TTS is listed at rank 11 on Artificial Analysis' Provider Voice Arena with an Elo of 1,203 across 3,512 appearances, at $18.31 per million characters. ElevenLabs Eleven v4, released September 28, 2026, sits at rank 1 with an Elo of 1,319 across 1,674 appearances, at $80.00 per million. That looks like a 116-point gap against a cheaper challenger — until you notice that Goog​le's own Gemini 3.8 Flash TTS sits third on the same board at 1,267 and a lower normalised price of $16.49. The real contest is not the one the headline matchup implies, and the reason is a release date.

One of these is eighteen months newer than it looks

Gemini 3.1 Flash TTS carries an April 15, 2026 release date on the board. Eleven v4 carries September 28, 2026. That is five and a half months, which is not a generation in speech synthesis but is long enough that Google has already shipped a successor: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS both went to general availability on September 23, 2026, five days before Eleven v4, and both are ranked above 3.1 on the same board. Gemini 3.8 Flash-Lite TTS lands sixth at 1,241 and $11.03 per million.

A generated scoreboard comparing ElevenLabs Eleven v4 and Google Gemini 3.1 Flash TTS across six dimensions: Provider Voice rank and Elo (1 / 1,319 against 11 / 1,203), samples (1,674 with an interval of plus or minus 19 against 3,512 with plus or minus 12), board price ($80.00 against $18.31 per 1M characters), rate card ($0.022 per 1K characters until October 12 against billed per token), script direction (inline audio tags against a voice set and prompting) and release date (September 28, 2026 against April 15, 2026). Footer: 'Ranks, Elo, intervals and board prices per Artificial Analysis; rate meters and capabilities vendor-documented.' The OrcaRouter logo sits in the bottom-right corner.

So the honest comparison has three points, not two, and the ordering by price is the interesting part:

• Rank 1, ElevenLabs Eleven v4 — Elo 1,319, $80.00 per million characters, 1,674 appearances, ±19 interval

• Rank 3, Google Gemini 3.8 Flash TTS — Elo 1,267, $16.49 per million characters, released September 23, 2026

• Rank 6, Google Gemini 3.8 Flash-Lite TTS — Elo 1,241, $11.03 per million characters, released September 23, 2026

• Rank 11, Google Gemini 3.1 Flash TTS — Elo 1,203, $18.31 per million characters, released April 15, 2026, 3,512 appearances, ±12 interval

Note what the intervals do to that ordering. Gemini 3.1 Flash TTS estimates 1,203 with a ±12 interval, so its true value sits somewhere around 1,191 to 1,215. Eleven v4 estimates 1,319 with a ±19 interval — roughly 1,300 to 1,338. The two ranges do not touch, and the gap between them is over a hundred points wide. That is as clean as a preference board gets.

A screenshot of Artificial Analysis' Provider Voice Arena leaderboard, top of the table. ElevenLabs Eleven v4 is rank 1 at Elo 1,319 with an interval of plus or minus 19, 1,674 samples, eight arena voices, released Sept 2026 at $80.0 per 1M characters. Cartesia Sonic 3.6 is second at 1,276. Google Gemini 3.8 Flash TTS is third at 1,267 with $16.5 per 1M characters. Google Gemini 3.8 Flash-Lite TTS is sixth at 1,241 with $11.0. SpeechifyAI Simba 3.2 is seventh at 1,240 with $6.6. Google Gemini 3.1 Flash TTS is eleventh at 1,203 with an interval of plus or minus 12, 3,512 samples, seven arena voices, released Apr 2026 at $18.3 per 1M characters.

Why the sample counts matter more than the ranks

Gemini 3.1 Flash TTS has 3,512 appearances behind its estimate; Eleven v4 has 1,674, because it is one day old on this board. A newer model's estimate is carrying more uncertainty per sample, and the ±19 interval reflects that. If you are the kind of buyer who waits for a number to settle, the rule of thumb is that the ordering above the interval width is the part you can act on. Here the ranking survives its error bars comfortably in both directions — ahead of 3.8 Flash TTS and behind nobody in this comparison.

The arena also counts arena voices: eight for Eleven v4, seven for Gemini 3.1 Flash TTS. That is the number of distinct vendor voices used in the blind tests, and it is a rough proxy for how much of a default cast a vendor brings. Eleven v4's rank 1 was earned with eight voices, which means the win is not a single lucky narrator.

The capability differences that the Elo does not price

The boards measure preference, not feature coverage. Four differences decide real projects and none of them appear in an Elo.

• Script direction — Eleven v4 documents inline audio tags ([laughs], [whispers], [said angrily in French accent]) and natural-language delivery prompts; Google's 3.1 generation documents voice choice and prompting but not an equivalent tag syntax for performance direction.

• Voice creation — the Gemini 3.8 generation added voice design from a written description and voice replication from a 30-second sample with watermarking and provenance metadata on the output. Gemini 3.1 Flash TTS predates that and ships a prebuilt voice set.

• Cloning from real audio — Eleven v4's Instant Voice Cloning runs from ten seconds of audio, with a professional tier above it. Google's replication path in the 3.8 generation asks for 30 seconds and adds consent verification. Different bars, different consent machinery.

• Input cap per request — Eleven v4 documents 10,000 characters. Google bills its TTS models in tokens rather than characters, so there is no like-for-like per-request character cap to compare, and the two meters should not be placed side by side as if they were.

That last point is the one that trips up procurement spreadsheets. ElevenLabs quotes a price per 1,000 characters. Google quotes a price per million tokens, in and out. Artificial Analysis normalises both onto a per-million-character axis — that is where the $80.00 and $18.31 come from — but the normalisation is the board's conversion, not either vendor's invoice. Use it to rank, not to forecast.

A screenshot of Google's Gemini API documentation page for text-to-speech generation, headed by a banner reading 'Gemini 3.8 Flash is now available'. The page covers single-speaker and multi-speaker TTS, states that the TTS capability differs from the Live API in being designed for exact text recitation with fine-grained control, names the gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts models, notes that TTS models accept text-only input and produce audio-only output, and shows a Python example that attaches a speech_metadata annotation with the style 'cheerful and friendly' and a voice of 'Kore' in generation_config.speech_config.

What the Eleven v4 rate card does to this comparison

For two weeks the price comparison above is wrong in ElevenLabs' favour, and then it is wrong against it. On ElevenLabs' API pricing page, Eleven v4 is listed at $0.022 per 1,000 characters against a stated list of $0.08, labelled 72% off until October 12. Eleven v4 Turbo is listed at $0.011 against $0.04. After October 12 the promotional number disappears and the $80.00-per-million board figure becomes the rate you actually pay.

Run the arithmetic on a five-million-character month. Eleven v4 costs $110 during the window and $400 after it. Gemini 3.1 Flash TTS, at the board's $18.31 normalisation, comes to about $92 for the same month — cheaper than the promotion and roughly four times cheaper than the list rate. A team that benchmarks in September and ships in November will have made the decision against a figure that no longer exists.

Where routing fits, and where it does not

Neither of these models is in the OrcaRouter catalogue. There is no ElevenLabs endpoint and no Google TTS endpoint here to call, and the honest answer to "how do I reach them through you" is that you do not: Eleven v4 is reached through ElevenLabs' own API, and Gemini 3.1 Flash TTS through Google's. Saying anything else would be inventing an integration.

What a single key is good for in a comparison like this is the decision itself. If you are weighing a $400-a-month endpoint against a $92-a-month one, the answer depends on your own scripts, your own languages and your own listeners — and getting that answer means calling both from the same code path with the same request shape, rather than standing up two vendor integrations to run one test. OrcaRouter covers that layer: 200-plus models behind one key with provider list prices passed through at 0% markup, so a vendor's price change is reflected the day it takes effect, plus automatic failover if one upstream degrades mid-evaluation.

Who should pick which

Pick Eleven v4 when the voice is the product. It leads the board by 116 points with a cleaner interval than the model below it, it is the only one of the two with a documented inline tag syntax for performance direction, and its cloning bar is ten seconds rather than thirty. The premium is real — four times the normalised price at list — and it buys the top of the board.

Pick Gemini 3.1 Flash TTS only if you are already committed to it. It is a five-month-old preview-generation model with a lower score and a higher normalised price than Google's own September release, and the only reason to keep it is inertia in an existing integration. If you have that inertia, migrating inside the Google stack to 3.8 Flash TTS buys 64 Elo points and a lower normalised price without changing vendors — which is a rare case where the upgrade is also the cheaper option.

For everyone else the live question is Eleven v4 against Gemini 3.8 Flash TTS, and the answer is a genuine trade rather than a ranking: 52 Elo points and the tag syntax against roughly a fifth of the cost per character. Run both on the same scripts after October 12, when the ElevenLabs promotion is gone and the price gap is the one you will actually be paying.