
Gemini 3.8 Flash TTS Says Hello: Google Splits Its Speech Line in Two and Cuts the Price
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
The vendor shipped two speech models on September 23, 2026, and the announcement is written as though the headline were voice quality. It isn't. Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are the first text-to-speech models the vendor has put on the Gemini Developer API at a price below the preview model they replace — and for anyone who has been paying per character somewhere else, that is the part that changes a budget line. The audio output rate on Gemini 3.8 Flash TTS is $9.00 per million tokens through the end of 2026. Audio bills at 25 tokens per second, which works out at roughly 0.02 cents per second of generated speech. The model it supersedes, Gemini 3.1 Flash TTS Preview, was priced at $20.00 per million audio tokens.
The second half of the story is that there are two models now, not one. Google split the line: Flash TTS carries the full language set and the higher audio rate, Flash-Lite TTS carries a shorter language list and a cheaper one. That is a different shape from the single-preview-model arrangement that has been on the API since April, and it tells you what Google thinks the market looks like — a small number of applications that need 130 languages and voice cloning, and a much larger number that need a competent English voice for a support bot and nothing else.
What actually shipped on September 23
Both models are generally available on the Gemini API and in AI Studio as of the announcement, not gated behind a waitlist. The names on the API are gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
• Flash TTS — 130 languages, language auto-detected from the text, 30 prebuilt studio voices including Kore and Puck.
• Flash-Lite TTS — 101 languages, same voice-design surface, positioned as the high-volume tier.
• Both — voice design from a natural-language description, line-by-line delivery direction, two-speaker scene staging, and non-verbal cues written straight into the script.
• Output — WAV 24 kHz mono 16-bit signed PCM on the unary endpoint; headerless PCM (audio/l16) on the streaming endpoint; audio/mulaw and audio/alaw also accepted, with a configurable sample rate.
The rate card, per million tokens, is worth reading in full because the tiers differ by more than the headline number:
• Flash TTS Standard — $0.50 text in / $9.00 audio out, rising to $1.00 / $18.00 after December 31, 2026.
• Flash TTS Batch and Flex — $0.25 / $4.50.
• Flash TTS Priority — $0.90 / $16.20.
• Flash-Lite TTS Standard — $0.50 / $6.00, rising to $1.00 / $12.00 after December 31, 2026.
• Flash-Lite TTS Batch and Flex — $0.25 / $3.00.
• Flash-Lite TTS Priority — $0.90 / $10.80.
• Gemini 3.1 Flash TTS Preview, for comparison — $1.00 / $20.00, batch $0.50 / $10.00.
The promotional window is the thing to note. Google has priced both new models at half rate until the end of the year and published the post-promotion numbers in the same table. Anyone modelling cost per thousand minutes should use the January figure, not the September one.

Voice design is the genuinely new capability
Prebuilt voices are table stakes. What Google added in 3.8 is a way to describe a voice in words and get it, and a way to clone one from a short sample with a consent check attached.
Voice design takes a natural-language prompt — pacing, timbre, accent, the kind of thing you would otherwise spend an afternoon auditioning for — and returns a usable voice without a reference recording. Voice replication goes the other way: you supply a 30-second sample, the system verifies consent, and the resulting voice is marked with a SynthID watermark and C2PA credentials on the generated audio. Voice remixing, which lets you blend or modify an existing custom voice, is listed as coming soon rather than shipped.
Custom voices come in two flavours and they behave differently enough that the distinction matters in production. Stateful custom voices are stored against the project, capped at 200 per project, with a one-year time to live. Stateless voices are addressed by a voicekey_... handle and expire after seven days. If your application provisions a voice per customer, the cap and the TTL are the two numbers that decide whether you need a storage layer of your own.
The delivery controls are the other half of "new". You can direct a line the way you would direct an actor — pace, emphasis, emotional register — rather than only choosing a voice and hoping. Two-speaker scenes can be staged inside a single call. And non-verbal vocalisations are first-class script elements: <laughs>, <sigh> and <gasp> as inline cues, plus |mhm| and |yeah| for the small back-channel noises that make a synthetic conversation sound like a conversation. Long-form reading is claimed to hold speaker identity with minimal drift, which is the failure mode that shows up first when you generate a 40-minute audiobook chapter.
The benchmarks, and who ran them
Google's launch material leans on two outside evaluations and a set of arena placings. All of these are vendor-reported — Google is quoting them, and the underlying runs are not public — so treat them as claims rather than measurements.

• Hume AI Voice Design Benchmark — Gemini 3.8 Flash TTS placed first at 71.4, and first on accent modelling at 60.8.
• Hume Overall Quality Index — Flash TTS first, Flash-Lite TTS second.
• Blind preference in the Speech Arena — top placings claimed in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
• Versus Gemini 3.1 Flash TTS — Google claims gains in long-form consistency and dual-speaker screenplay control, which are exactly the two things the delivery-direction features are built for.
One documented discrepancy is worth flagging because it will confuse anyone comparing sources. The blog post describes a library of 2,000+ production-ready voices. The developer documentation describes "hundreds" of voices in the Extended Voice Library alongside the 30 prebuilt studio voices. Third-party coverage has quoted figures an order of magnitude higher again. These are not reconcilable from the outside — the honest reading is that the prebuilt set you get by default is 30, the Extended Voice Library is larger and vaguely described, and the big numbers refer to a catalogue that includes user-created voices. If a voice count is load-bearing for your decision, test the specific voice you need rather than trusting any of the three figures.
What it means if you are building on it
The practical change is that text-to-speech has moved from a specialist line item to something you can put in the same cost model as your language tokens. At 25 tokens per second, an hour of generated audio is 90,000 audio tokens, which at the promotional Flash-Lite rate is about 54 cents. That is cheap enough that the interesting question stops being "can we afford a synthetic voice" and becomes "which of the two tiers does this surface need".
The second practical change is that a brand-new TTS model is exactly the kind of thing you want to try without betting a production path on it. That is the case for putting a new model behind a router rather than wiring it directly: OrcaRouter exposes 200+ models behind a single key at provider list price with no markup, and its automatic failover means a new speech endpoint that wobbles under load degrades to a model you already trust instead of dropping the call. We do not host the Gemini 3.8 TTS endpoints — those come from Google's own API — but the text layer underneath the voice, and the fallback path when a synthesis call fails, are the parts a router is actually for.
What is still missing
Three things are absent from the launch and each of them is a real constraint, not a nitpick.
There is no published independent evaluation. Google's Hume numbers are the vendor quoting a third party's benchmark, which is better than nothing and worse than an audit. Until Artificial Analysis or an equivalent publishes its own run on the same models, every quality claim in this article should be read as a claim.
There is no open-weights option. Both models are closed and API-only, which puts them in a different category from the open speech models that have been landing in the same arena. If your deployment requires running the model yourself, this launch does not address you.
And the promotional pricing ends. The January 2027 rate for Flash TTS is double the September rate, and any business case built on the launch numbers will need redoing in the first quarter. That is not a criticism — publishing both numbers up front is more honest than most — but it is the number that belongs in the spreadsheet.
Google has not said whether Gemini 3.1 Flash TTS Preview will be deprecated, only that Flash-Lite TTS is positioned as its workhorse replacement. Anyone still on the preview model should plan a migration rather than wait for a shutdown notice.

