
Gemini 3.8 TTS vs ElevenLabs: What Three Times the Price Buys You
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
At the same volume, ElevenLabs' flagship synthesis model costs about three times what the vendor's new one does. ElevenLabs Eleven v3 bills $0.10 per thousand characters — $100 per million — and Gemini 3.8 Flash TTS bills $9.00 per million audio tokens, which Artificial Analysis normalises to $33.00 per million characters. That is a 3.0x gap on the meter, and it is the single biggest fact in this comparison. The interesting question is not whether the gap is real; it is what the extra sixty-seven dollars per million characters is actually buying, because the answer is not "better speech".

Three meters, not one
The first thing to get straight is that these two products do not measure the same thing, and the published rate cards hide that.
• ElevenLabs bills by character. Eleven v3 is $0.10 per 1,000 characters, capped at 5,000 characters per request. Eleven v3 Conversational is $0.05 per 1,000, and the Flash and Turbo models are also $0.05 per 1,000 but with a 40,000-character request limit and a claimed ~75 ms latency against v3 Conversational's ~280 ms.
• Google bills by token, and audio tokens are not text tokens. Gemini 3.8 Flash TTS charges $0.50 per million text tokens in and $9.00 per million audio tokens out, metered at 25 audio tokens per second of generated speech. There is no per-request character cap published.

• Artificial Analysis bills neither; it normalises. The $100.00 and $33.00 per million characters on its Provider Voice Arena leaderboard are a derived common denominator so the board can rank on price at all. Useful for a ratio, not for a quote.
So the honest version of the 3.0x figure is: on the only like-for-like basis available, Google is about a third of the price. What that means in your own spreadsheet depends on whether your workload is dominated by short strings or long ones — a per-second audio meter and a per-character text meter diverge sharply as utterances get longer, because silence, pauses and prosodic padding all cost audio tokens and no characters.
A worked month
Take a million characters of text, roughly the length of a mid-sized novel, converted to speech once.
• ElevenLabs Eleven v3 — $100.00 for the synthesis, plus whatever plan tier you need to run it at your concurrency. The published plans run Starter at $6, Creator at $22, Pro at $99, Scale at $299 and Business at $990, and the character allowances are bundled into those rather than billed separately.
• Gemini 3.8 Flash TTS — $33.00 equivalent, or roughly $16.50 if the job is batchable, because the Batch and Flex tier halves the audio rate to $4.50 per million tokens. Priority service raises it to $16.20 per million, or about $59.40 equivalent, which is still well under ElevenLabs.
• Gemini 3.8 Flash-Lite TTS — $6.00 per million audio tokens, about $22.10 equivalent on the same normalisation, at the cost of 101 languages instead of 130.
Run the same million characters through Google's promotional pricing and the gap widens further, because both new Gemini TTS models are at half rate until December 31, 2026 and then double. Anyone building a business case on the September numbers is building it on a discount that has a published expiry date.

The one place the comparison flips is very short utterances. A per-character meter charges almost nothing for a three-word confirmation; a per-second audio meter charges for the second the confirmation takes to say, including the silence around it. If your product is a voice interface built out of one-sentence responses, the arithmetic moves toward ElevenLabs. If it is narration, audiobooks, long-form localisation or anything where the text is long and the audio is longer, it moves hard toward Google.
The quality question, honestly
On the Artificial Analysis Provider Voice Arena — blind pairwise votes across each provider's own native voices — the two models are not close, and Google is ahead.
• Gemini 3.8 Flash TTS — rank 2, Elo 1,260, a 17-point interval, 1,999 samples, 8 arena voices, released September 2026.
• ElevenLabs Eleven v3 — rank 17, Elo 1,167, an 11-point interval, 4,345 samples, 8 arena voices, listed as February 2026 on the board.
Ninety-three Elo points separate them, and unlike the gap between the top three on that board, this one is real: Eleven v3's interval is 1,156 to 1,178 and Gemini's is 1,243 to 1,277, with no overlap. Eleven v3's sample count is more than double Gemini's, so the estimate is not thin either.
That is a genuinely awkward result for the price argument, and it cuts against the obvious conclusion. ElevenLabs charges three times as much and places seventeen places lower on blind preference. Whatever the extra sixty-seven dollars buys, it is not a better-sounding voice in a head-to-head.
What ElevenLabs is actually selling
ElevenLabs is not primarily selling a synthesis endpoint, and pricing it as though it were is where most comparisons go wrong. What the money buys:
• A voice library. ElevenLabs' shared voice library is in the hundreds of voices and is a genuine product surface — the thing people come to ElevenLabs for is often a specific voice they heard somewhere, not the model behind it. Gemini 3.8 Flash TTS ships 30 prebuilt studio voices, a voice-design path that generates one from a natural-language description, and a replication path that clones from a 30-second sample with consent verification, a SynthID watermark and C2PA credentials attached. The default catalogue is smaller; the ability to create a specific voice is comparable.
• Latency tiers as separate products. Flash and Turbo at ~75 ms and a 40,000-character limit are a different product from v3 at ~280 ms and 5,000 characters, and both are cheaper than v3. If your application needs sub-100 ms, ElevenLabs sells that explicitly and Google's launch material does not make a comparable latency claim for either new TTS model.
• An ecosystem. Dubbing, voice agents, conversational endpoints, a plan structure with concurrency attached — the same reason Cartesia sells telephony: it is a stack, not a model.
If you are buying a stack, the 3.0x is the price of not assembling it. If you are buying a synthesis call, you are paying it for a voice library and a latency tier.
What Google is actually selling
The counter-case is narrower and stronger than "cheaper".
• Language coverage — 130 languages on Flash TTS and 101 on Flash-Lite TTS, auto-detected from the text, against the 70-plus languages ElevenLabs documents for Eleven v3. For a localisation pass, that difference is not a feature comparison, it is a coverage gap.
• Direction — line-by-line delivery control, two-speaker scene staging inside a single call, and non-verbal cues written into the script as <laughs>, <sigh>, <gasp>, |mhm| and |yeah|. Google claims the 3.8 generation improved long-form consistency and dual-speaker control over Gemini 3.1 Flash TTS, which is exactly what those features exist for. Vendor claim, unreproduced.
• Voice state model — 200 stateful custom voices per project with a one-year TTL, or stateless voices addressed by a voicekey_... handle that expire in seven days. That is an infrastructure decision you can plan around.
• Output control — WAV 24 kHz mono 16-bit signed PCM on the unary endpoint, headerless PCM (audio/l16) on the streaming endpoint, and audio/mulaw and audio/alaw with a configurable sample rate.
Google's launch benchmarks are vendor-reported and should be read that way: first on the Hume AI Voice Design Benchmark at 71.4 and first on accent modelling at 60.8, first and second on the Hume Overall Quality Index for Flash and Flash-Lite respectively, and top blind-preference placings claimed in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. None of those runs is public. The arena Elo above is the only independently collected number in this article, and it happens to agree with the direction of the vendor claims.
The switch calculus
Switch if your workload is long-form, multi-language, or high-volume enough that 3.0x shows up as a line item, and if you can live with a smaller default voice catalogue and no published sub-100 ms tier. That covers narration, dubbing, accessibility, and most product-embedded speech.
Stay if you are buying the stack — the voice library, the latency tiers, the dubbing and agent products — or if your workload is dominated by very short utterances where a per-second audio meter punishes you. Also stay if you have already fine-tuned a voice identity on ElevenLabs that your users recognise, because voice continuity is a product feature and re-creating it on a new model is a project.
Either way the sane move is to run both for a month on real traffic before committing, and that is the argument for not wiring either one directly into your codebase. OrcaRouter puts 200+ models behind a single key at provider list price with no markup, so a price move from either lab reaches your bill the same day instead of at renewal, and automatic failover keeps a production speech path alive when a brand-new endpoint misbehaves. We do not host ElevenLabs — its synthesis and agent layers are its own product — and we do not host the Gemini 3.8 TTS endpoints, which come from Google's API. What we carry is the text layer underneath and the routing above it.
The number worth watching is Eleven v3's Elo on the next leaderboard revision. Ninety-three points is a large deficit for a model charging three times as much, and if the interval narrows without the gap closing, the price argument stops being a trade and becomes a verdict.
