A generated hero card for 'Gemini 3.8 TTS vs Gemini 3.1 Flash TTS' on a white background with soft blue-and-cyan gradient accents. The left card 'Gemini 3.1 Flash TTS' lists Elo 1,199, 3,430 samples, 7 arena voices and April 2026; the right card 'Gemini 3.8 Flash TTS' lists Elo 1,260, 1,999 samples, 8 arena voices and September 2026. A footer line reads 'Elo per Artificial Analysis Provider Voice Arena, Sept 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Gemini 3.8 TTS vs Gemini 3.1 Flash TTS: The One Gap in This Family That Is Real

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Upgrading inside a model family is usually the easy decision, and this one has a wrinkle that makes it less easy than it looks. Gemini 3.1 Flash TTS — still listed on the vendor's API as a preview — sits at rank 10 on the Artificial Analysis Provider Voice Arena with an Elo of 1,199 and a 12-point interval. Gemini 3.8 Flash TTS, shipped September 23, 2026, sits at rank 2 with an Elo of 1,260 and a 17-point interval. That is a 61-point gain, and unlike most of the gaps on that board it survives the error bars: 3.1's range runs 1,187 to 1,211, 3.8's runs 1,243 to 1,277, and there is a thirty-two-point dead zone between them. The quality upgrade is real. The price story is where it gets strange.

A generated two-column scoreboard for Gemini 3.1 Flash TTS and Gemini 3.8 Flash TTS — 3.1 at Arena Elo 1,199, $20.00 per 1M audio tokens, 7 arena voices, 3,430 samples and an April 2026 release; 3.8 at Elo 1,260, $9.00 per 1M audio tokens, 8 arena voices, 1,999 samples and a September 2026 release — over the footer 'Prices per Google; Elo per Artificial Analysis, Sept 2026.'

Google's own rate card says the audio output rate fell from $20.00 per million tokens on 3.1 to $9.00 on 3.8 — a 55 percent cut. Artificial Analysis normalises the same two models to $18.30 and $33.00 per million characters respectively, which says the opposite. Both numbers are published; neither is wrong; they are measuring different things, and the reconciliation is the most useful part of this comparison for anyone planning a migration.

What the upgrade actually changed

The 3.8 generation is not a retune. Three surfaces changed materially.

• Voice count and voice creation — 3.1 Flash TTS shipped a prebuilt voice set. 3.8 Flash TTS ships 30 prebuilt studio voices including Kore and Puck, plus voice design from a natural-language description and voice replication from a 30-second sample with consent verification, a SynthID watermark and C2PA credentials on the output. Voice remixing is listed as coming soon rather than shipped.

• Delivery control — line-by-line direction, two-speaker scene staging inside a single call, and non-verbal cues written straight into the script as <laughs>, <sigh>, <gasp>, |mhm| and |yeah|. Google's launch material claims improved long-form consistency and dual-speaker screenplay control against 3.1 specifically. That is a vendor claim with no public run behind it.

• Language coverage — 130 languages on Flash TTS and 101 on Flash-Lite TTS, auto-detected from the text. 3.1 Flash TTS's published coverage is lower.

The structural change is the split. There is no single successor to 3.1 Flash TTS; there are two, and Google's documentation positions Flash-Lite TTS as the "workhorse replacement". That matters because it means a migration is a tier decision, not a version bump. The cheaper tier is not the newer version of what you have — it is a different point on the curve.

Why the price went down and the leaderboard says it went up

Start with what Google publishes, because that is what you will actually be billed.

Gemini 3.1 Flash TTS Preview — $1.00 per million text tokens in, $20.00 per million audio tokens out. Batch is $0.50 and $10.00.

• Gemini 3.8 Flash TTS Standard — $0.50 and $9.00 through December 31, 2026, then $1.00 and $18.00. Batch and Flex $0.25 and $4.50. Priority $0.90 and $16.20.

• Gemini 3.8 Flash-Lite TTS Standard — $0.50 and $6.00 through December 31, 2026, then $1.00 and $12.00. Batch and Flex $0.25 and $3.00. Priority $0.90 and $10.80.

A screenshot of Google's Gemini API text-to-speech documentation, captured in English, showing the 'Gemini 3.8 Flash is now available' banner, the single-speaker TTS section naming gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, and a Python sample that passes speech_metadata style annotations and the Kore voice.

On the audio-output line alone, 3.8 Flash TTS at $9.00 against 3.1 at $20.00 is a 55 percent cut, and 3.8 Flash-Lite TTS at $6.00 is a 70 percent cut. Even at the post-promotion $18.00, Flash TTS is still below the model it replaces. That is the number that lands on an invoice, and it is unambiguous.

Now the leaderboard. Artificial Analysis publishes a per-million-characters figure so that models billing in different units can be ranked side by side, and for this pair it gives $18.30 for 3.1 Flash TTS and $33.00 for 3.8 Flash TTS. Read alone, that says the upgrade nearly doubled in price.

The reconciliation is that a per-character rate is a conversion, not a price, and the conversion factor is not the same for the two models. Converting audio tokens to characters requires assuming both a speech rate and an audio-token ratio — Google meters audio at 25 tokens per second of output — and it requires choosing which service tier to normalise against. If the board normalises 3.8 at the post-promotion $18.00 rate while normalising 3.1 at its own $20.00, the two are close enough that any difference in the assumed seconds-per-character dominates the result. A per-character rate built on a generous speech-rate assumption flatters the model it is applied to.

The practical instruction is therefore: use Google's token rates for budgeting, and use the leaderboard's character rates only for the ratio between models that the board measured the same way. Do not mix them, and do not quote the $18.30-to-$33.00 move as a price increase — it is an artefact of a normalisation that the two models do not share.

The genuine migration cost is a different one, and it is the promotional window. Both new tiers are at half rate until December 31, 2026 and double on January 1. A migration planned on September's numbers and executed in November will re-price in the first quarter. The January 2027 rate for Flash TTS, $18.00 per million audio tokens, is still below 3.1's $20.00 — so the cut is real past the promo, just much smaller than it looks today.

The quality gain, and what is not behind it

The arena Elo is the one number here collected by someone other than the vendor, and it is the one that clears its error bars. Sixty-one points from 1,199 to 1,260, with 3,430 samples on 3.1 and 1,999 on 3.8, and 7 arena voices against 8. The sample count on the newer model is lower, which widens its interval, and even so the ranges do not touch.

A screenshot of the Artificial Analysis Provider Voice Arena leaderboard, captured in English, showing the top rows with Cartesia Sonic 3.6 first at Elo 1,273, Google's Gemini 3.8 Flash TTS second at 1,260, Alibaba's Qwen-Audio-3.0-TTS-Plus third at 1,259, and Google's Gemini 3.1 Flash TTS tenth at 1,199 across 3,430 samples.

What is not behind it: any independent benchmark other than the arena. Google's launch numbers — first on the Hume AI Voice Design Benchmark at 71.4, first on accent modelling at 60.8, first and second on the Hume Overall Quality Index for Flash and Flash-Lite, and top blind-preference placings claimed in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi — are the vendor quoting a third party's benchmark. The underlying runs are not public. Read them as claims, in the same direction as the arena result but not adding independent weight to it.

The deprecation question, and a migration checklist

Google has not announced a shutdown date for Gemini 3.1 Flash TTS Preview. It has said that Flash-Lite TTS is positioned as its workhorse replacement, which is the kind of language that precedes a deprecation notice rather than following one. If you are still on the preview model, the correct read is that you have a migration to plan, not a deadline to meet — and that waiting for the deadline is the wrong strategy for a model whose replacement is already 61 Elo points ahead.

• Check which tier your workload needs. Flash TTS for 130 languages and the full delivery surface; Flash-Lite TTS for 101 languages at $6.00 per million audio tokens. The language list is the dividing line, not the quality.

• Re-price against the January 2027 rate, not the promotional one. $18.00 per million audio tokens on Flash TTS, $12.00 on Flash-Lite.

• Decide whether the job is batchable. Batch and Flex halves the audio rate on both tiers — $4.50 and $3.00 per million tokens. Anything that is not interactive should not be on the Standard tier.

• Re-check anything that depended on the voice set. The default catalogue is 30 prebuilt voices; the voice you were using on 3.1 may need re-creating, either through voice design or through a 30-second replication sample.

• Re-check request shaping. There is no published per-request character cap on the 3.8 models, and audio is metered per second rather than per character, so a long utterance costs more than a per-character quote would suggest.

• Plan the custom-voice lifecycle. 200 stateful voices per project with a one-year time to live, or stateless voicekey_... handles that expire in seven days.

Running both while you decide

An in-family upgrade with a promotional expiry and a normalisation disagreement is exactly the case for not cutting over in one step. OrcaRouter routes the older model today — models/google/gemini-3.1-flash-tts-preview is on our model list — so you can stand the new model up beside it and compare on your own traffic before you move the production path. We do not host the Gemini 3.8 TTS endpoints; those come from Google's own API. What we provide is the single key across 200+ models at provider list price with no markup, so a vendor rate change reaches your bill the same day rather than at renewal, and automatic failover so that a brand-new model is something you can try without betting a customer-facing path on it.

The thing to watch is whether Google attaches a date to the 3.1 Flash TTS preview. That single line of documentation is worth more to a migration plan than any of the benchmarks in this article, and it has not been written yet.