
Eleven v4 vs Gemini 3.1 Flash TTS: A 116-Point Gap, and the Cheaper Google Model
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Shopping for Google's voice today, you will land on the wrong one. Google Gemini 3.1 Flash TTS is listed at rank 11 on Artificial Analysis' Provider Voice Arena with an Elo of 1,203 across 3,512 appearances, at $18.31 per million characters. ElevenLabs Eleven v4, released September 28, 2026, sits at rank 1 with an Elo of 1,319 across 1,674 appearances, at $80.00 per million. That looks like a 116-point gap against a cheaper challenger — until you notice that Google's own Gemini 3.8 Flash TTS sits third on the same board at 1,267 and a lower normalised price of $16.49. The real contest is not the one the headline matchup implies, and the reason is a release date.
One of these is eighteen months newer than it looks
Gemini 3.1 Flash TTS carries an April 15, 2026 release date on the board. Eleven v4 carries September 28, 2026. That is five and a half months, which is not a generation in speech synthesis but is long enough that Google has already shipped a successor: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS both went to general availability on September 23, 2026, five days before Eleven v4, and both are ranked above 3.1 on the same board. Gemini 3.8 Flash-Lite TTS lands sixth at 1,241 and $11.03 per million.

So the honest comparison has three points, not two, and the ordering by price is the interesting part:
• Rank 1, ElevenLabs Eleven v4 — Elo 1,319, $80.00 per million characters, 1,674 appearances, ±19 interval
• Rank 3, Google Gemini 3.8 Flash TTS — Elo 1,267, $16.49 per million characters, released September 23, 2026
• Rank 6, Google Gemini 3.8 Flash-Lite TTS — Elo 1,241, $11.03 per million characters, released September 23, 2026
• Rank 11, Google Gemini 3.1 Flash TTS — Elo 1,203, $18.31 per million characters, released April 15, 2026, 3,512 appearances, ±12 interval
Note what the intervals do to that ordering. Gemini 3.1 Flash TTS estimates 1,203 with a ±12 interval, so its true value sits somewhere around 1,191 to 1,215. Eleven v4 estimates 1,319 with a ±19 interval — roughly 1,300 to 1,338. The two ranges do not touch, and the gap between them is over a hundred points wide. That is as clean as a preference board gets.

Why the sample counts matter more than the ranks
Gemini 3.1 Flash TTS has 3,512 appearances behind its estimate; Eleven v4 has 1,674, because it is one day old on this board. A newer model's estimate is carrying more uncertainty per sample, and the ±19 interval reflects that. If you are the kind of buyer who waits for a number to settle, the rule of thumb is that the ordering above the interval width is the part you can act on. Here the ranking survives its error bars comfortably in both directions — ahead of 3.8 Flash TTS and behind nobody in this comparison.
The arena also counts arena voices: eight for Eleven v4, seven for Gemini 3.1 Flash TTS. That is the number of distinct vendor voices used in the blind tests, and it is a rough proxy for how much of a default cast a vendor brings. Eleven v4's rank 1 was earned with eight voices, which means the win is not a single lucky narrator.
The capability differences that the Elo does not price
The boards measure preference, not feature coverage. Four differences decide real projects and none of them appear in an Elo.
• Script direction — Eleven v4 documents inline audio tags ([laughs], [whispers], [said angrily in French accent]) and natural-language delivery prompts; Google's 3.1 generation documents voice choice and prompting but not an equivalent tag syntax for performance direction.
• Voice creation — the Gemini 3.8 generation added voice design from a written description and voice replication from a 30-second sample with watermarking and provenance metadata on the output. Gemini 3.1 Flash TTS predates that and ships a prebuilt voice set.
• Cloning from real audio — Eleven v4's Instant Voice Cloning runs from ten seconds of audio, with a professional tier above it. Google's replication path in the 3.8 generation asks for 30 seconds and adds consent verification. Different bars, different consent machinery.
• Input cap per request — Eleven v4 documents 10,000 characters. Google bills its TTS models in tokens rather than characters, so there is no like-for-like per-request character cap to compare, and the two meters should not be placed side by side as if they were.
That last point is the one that trips up procurement spreadsheets. ElevenLabs quotes a price per 1,000 characters. Google quotes a price per million tokens, in and out. Artificial Analysis normalises both onto a per-million-character axis — that is where the $80.00 and $18.31 come from — but the normalisation is the board's conversion, not either vendor's invoice. Use it to rank, not to forecast.

What the Eleven v4 rate card does to this comparison
For two weeks the price comparison above is wrong in ElevenLabs' favour, and then it is wrong against it. On ElevenLabs' API pricing page, Eleven v4 is listed at $0.022 per 1,000 characters against a stated list of $0.08, labelled 72% off until October 12. Eleven v4 Turbo is listed at $0.011 against $0.04. After October 12 the promotional number disappears and the $80.00-per-million board figure becomes the rate you actually pay.
Run the arithmetic on a five-million-character month. Eleven v4 costs $110 during the window and $400 after it. Gemini 3.1 Flash TTS, at the board's $18.31 normalisation, comes to about $92 for the same month — cheaper than the promotion and roughly four times cheaper than the list rate. A team that benchmarks in September and ships in November will have made the decision against a figure that no longer exists.
Where routing fits, and where it does not
Neither of these models is in the OrcaRouter catalogue. There is no ElevenLabs endpoint and no Google TTS endpoint here to call, and the honest answer to "how do I reach them through you" is that you do not: Eleven v4 is reached through ElevenLabs' own API, and Gemini 3.1 Flash TTS through Google's. Saying anything else would be inventing an integration.
What a single key is good for in a comparison like this is the decision itself. If you are weighing a $400-a-month endpoint against a $92-a-month one, the answer depends on your own scripts, your own languages and your own listeners — and getting that answer means calling both from the same code path with the same request shape, rather than standing up two vendor integrations to run one test. OrcaRouter covers that layer: 200-plus models behind one key with provider list prices passed through at 0% markup, so a vendor's price change is reflected the day it takes effect, plus automatic failover if one upstream degrades mid-evaluation.
Who should pick which
Pick Eleven v4 when the voice is the product. It leads the board by 116 points with a cleaner interval than the model below it, it is the only one of the two with a documented inline tag syntax for performance direction, and its cloning bar is ten seconds rather than thirty. The premium is real — four times the normalised price at list — and it buys the top of the board.
Pick Gemini 3.1 Flash TTS only if you are already committed to it. It is a five-month-old preview-generation model with a lower score and a higher normalised price than Google's own September release, and the only reason to keep it is inertia in an existing integration. If you have that inertia, migrating inside the Google stack to 3.8 Flash TTS buys 64 Elo points and a lower normalised price without changing vendors — which is a rare case where the upgrade is also the cheaper option.
For everyone else the live question is Eleven v4 against Gemini 3.8 Flash TTS, and the answer is a genuine trade rather than a ranking: 52 Elo points and the tag syntax against roughly a fifth of the cost per character. Run both on the same scripts after October 12, when the ElevenLabs promotion is gone and the price gap is the one you will actually be paying.
