A generated hero card for 'Gemini 3.8 TTS vs Cartesia Sonic 3.6' on a white background with soft blue-and-cyan gradient accents. The left card 'Cartesia Sonic 3.6' lists Elo 1,273, 1,757 samples, $49.0 per 1M chars and August 2026; the right card 'Gemini 3.8 Flash TTS' lists Elo 1,260, 1,999 samples, $33.0 per 1M chars and September 2026. A footer line reads 'Elo per Artificial Analysis Provider Voice Arena, Sept 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Gemini 3.8 TTS vs Cartesia Sonic 3.6: Thirteen Elo Points and Sixteen Dollars

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two closest text-to-speech models on the Artificial Analysis blind-vote leaderboard are one place apart and sixteen dollars apart, and those two gaps point in opposite directions. Cartesia Sonic 3.6 sits at number one with a Provider Voice Arena Elo of 1,273 across 1,757 samples. Gemini 3.8 Flash TTS sits at number two with 1,260 across 1,999 samples. The Elo gap is thirteen points — inside the published confidence interval for both — and the price gap runs the other way, with Gemini at $33.00 per million characters against Cartesia's $49.00. If you are choosing between them, the honest framing is that you are paying sixteen dollars per million characters for a quality difference the leaderboard cannot resolve.

That is the whole decision, and it is worth slowing down on because both models are new enough that almost nothing else has been measured. Cartesia Sonic 3.6 arrived in August 2026. Gemini 3.8 Flash TTS arrived on September 23, 2026. Neither has a published independent benchmark beyond the arena votes, and the arena votes put them in a statistical tie at the top.

What the leaderboard actually says

The Artificial Analysis Provider Voice Arena is the useful board here because it compares each provider's own native voices against every other provider's, in a blind pairwise vote. That is a fairer comparison than a controlled-voice board for the question "which product sounds better", because it lets each vendor ship the voices it actually ships.

• Rank 1 — Cartesia Sonic 3.6, Elo 1,273 with a 17-point interval, 1,757 samples, 8 arena voices, released August 2026, $49.0 per 1M characters.

• Rank 2 — Gemini 3.8 Flash TTS, Elo 1,260 with a 17-point interval, 1,999 samples, 8 arena voices, released September 2026, $33.0 per 1M characters.

• Rank 3 — Alibaba Qwen-Audio-3.0-TTS-Plus, Elo 1,259, 1,447 samples, 15 arena voices, $27.6 per 1M characters.

• Rank 4 to 9, for context — Inworld Realtime TTS-2 at 1,245, SpeechifyAI Simba 3.2 at 1,237, Google Gemini 3.8 Flash-Lite TTS at 1,235, VUI Labs Luna TTS at 1,230, Inworld Realtime TTS-2 Flash at 1,210, and BreezeBlue Breeze TTS 2 at 1,204.

Read the intervals and the top three collapse into one another. Sonic 3.6's 1,273 carries a range of seventeen points; Gemini's 1,260 carries the same. Their ranges overlap across almost their entire width. Qwen-Audio-3.0-TTS-Plus at 1,259 overlaps both. The board is telling you these three are indistinguishable by blind preference on the samples collected so far, and that the ordering between them is not stable evidence.

The sample counts are close enough not to explain the difference either — 1,757 for Cartesia against 1,999 for Gemini. Both are in the low thousands, which is enough to place a model in a band and not enough to separate neighbours in it.

A generated two-column scoreboard for Cartesia Sonic 3.6 and Gemini 3.8 Flash TTS — Sonic 3.6 at Arena Elo 1,273, $49.00 per 1M characters, 8 arena voices, 1,757 samples and hosted beta; Gemini 3.8 Flash TTS at Elo 1,260, $33.00 per 1M characters, 8 arena voices, 1,999 samples and generally available — over the footer 'Per Artificial Analysis Provider Voice Arena, Sept 2026.'

Where the sixteen dollars comes from

Cartesia's own pricing page does not publish a per-million-character rate for Sonic. It sells subscription tiers with bundled minutes — Free at $0 with 20,000 credits, roughly 27 minutes of Sonic 3.6; Pro at $5 with 100,000 credits, roughly 133 minutes; Startup at $49 with 1.25 million credits, roughly 1,667 minutes; Scale at $299 with 8 million credits, roughly 10,667 minutes — plus concurrency limits of 2, 3, 5 and 15 respectively, voice localisation at 225 credits per added accent, voice agent call duration at $0.06 per minute, and telephony at $0.014 per minute.

So the $49.00 per million characters is not Cartesia's list price in its own units. It is Artificial Analysis normalising Cartesia's credits into a common denominator so the board can rank on price at all. That is a legitimate thing for a comparison board to do and it is the only like-for-like figure available, but it means the Cartesia number is a derived rate while the Gemini number is a published one.

Google's rate is straightforward. Gemini 3.8 Flash TTS bills $0.50 per million tokens of text in and $9.00 per million tokens of audio out through December 31, 2026, then $1.00 and $18.00. Audio is metered at 25 tokens per second, so the audio-output rate converts to about 0.02 cents per second of speech. Flash-Lite TTS, the cheaper sibling, is $6.00 per million audio tokens on the same 25-token second, and Artificial Analysis has it at $22.10-per-1M-characters equivalent at rank 6 with an Elo of 1,235.

A screenshot of Google's Gemini API text-to-speech documentation, captured in English, showing the 'Gemini 3.8 Flash is now available' banner, the single-speaker TTS section naming gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, and a Python sample that passes speech_metadata style annotations and the Kore voice.

• Price per 1M characters — Cartesia Sonic 3.6 $49.00 vs Gemini 3.8 Flash TTS $33.00 vs Gemini 3.8 Flash-Lite TTS $22.10.

• Arena Elo — 1,273 vs 1,260 vs 1,235, all with 17-point intervals at the top two.

• Arena voices — 8 vs 8 vs 8.

• Samples — 1,757 vs 1,999 vs 2,000.

• Availability — Cartesia hosted beta vs Google's own Gemini Developer API, generally available.

• Weight release — closed for both; neither is downloadable.

The part the price comparison hides

Per-million-characters is a convenient unit and it is not how either product bills.

Cartesia's unit is a credit against a subscription, with concurrency caps attached to the tier. That means the marginal cost of your next thousand characters is zero until you hit the cap, and then it is the cost of the next tier up. For a workload with a predictable ceiling, that is a better deal than the per-character rate suggests — you are buying headroom, not consumption. For a workload that spikes, it is worse, because concurrency of 2 on the Free tier and 3 on Pro is a hard wall rather than a price.

Gemini's unit is a token, metered per call, with no concurrency ceiling published and three service tiers — Standard, Batch and Flex, and Priority — at different rates. Batch and Flex halves the audio rate to $4.50 per million tokens on Flash TTS; Priority raises it to $16.20. That is a much finer-grained cost model, and it rewards the thing Google's pricing usually rewards: knowing in advance whether a job is interactive or batch.

Which is cheaper therefore depends on whether your speech workload looks like a subscription or a meter. Cartesia's tiers are priced for products with steady, bounded volume. Gemini's are priced for consumption you can classify.

What Cartesia still has that Google does not

Two things, and both are about being a voice product rather than a voice model.

Cartesia sells telephony and voice-agent infrastructure directly: $0.014 per minute for telephony, $0.06 per minute for voice agent call duration, and voice localisation at 225 credits per added accent. If you are building a phone agent, that is a stack you can buy in one place. Gemini 3.8 Flash TTS is a synthesis endpoint — it returns audio and stops. You supply the transport, the turn detection and the telephony yourself, or you assemble them from other vendors.

A screenshot of Cartesia's pricing page, captured in English, showing the Free, Pro, Startup, Scale and Enterprise tiers at $0, $5, $49, $299 and custom per month, with monthly credit allowances of 20K, 100K, 1.25M and 8M credits and the features bundled into each tier.

Cartesia has also been on the board longer. Its August release date against Google's September means more of its sample count accumulated against a settled version, whereas a September launch is still collecting votes. A model's Elo in its first week is a noisier estimate than the same model's Elo after a month, and that cuts against reading the thirteen-point gap as evidence for Cartesia.

What Google has that Cartesia does not

Language coverage is the clearest one. Flash TTS handles 130 languages with automatic detection; Flash-Lite handles 101. Cartesia's localisation model is additive — you buy accents on top of a base voice at 225 credits each — which is a different approach and a different cost curve once you go past a handful of languages.

The delivery surface is the second. Gemini 3.8 Flash TTS lets you direct a line the way you would direct an actor, stage two-speaker scenes inside one call, and write non-verbal cues — <laughs>, <sigh>, <gasp>, |mhm|, |yeah| — directly into the script. It also offers voice design from a natural-language description and voice replication from a 30-second sample with consent verification, a SynthID watermark and C2PA credentials on the output. Voice remixing is listed as coming soon.

And it is generally available, not in beta. Cartesia Sonic 3.6 is a hosted beta, which is a statement about the vendor's own confidence in the endpoint under load.

Which one to actually call

If your workload is a phone agent with predictable volume and you want telephony, concurrency and synthesis from one vendor, Cartesia Sonic 3.6 is the coherent choice, and the sixteen-dollar gap is the price of not assembling three vendors yourself.

If your workload is synthesis inside a product you already operate — narration, an accessibility layer, a support bot's voice, a 130-language localisation pass — Gemini 3.8 Flash TTS is cheaper per character, wider in language coverage, finer in delivery control, and generally available. The Elo difference does not argue against it, because the Elo difference is not statistically real.

If you want the cheapest of the three at the top of the board and can accept a shorter language list, Gemini 3.8 Flash-Lite TTS is $6.00 per million audio tokens and sits six places down at 1,235 — a gap that is also inside the noise.

Whoever you pick, the migration path matters more than the pick, because both models are new and both vendors are still tuning. OrcaRouter is the hedge: one key across 200+ models at provider list price with no markup, so a price cut from either lab lands on your bill the same day rather than at renewal, automatic failover if a synthesis call fails, and a routing DSL if you want to send a batch job to one model and an interactive call to another. We do not host Cartesia Sonic 3.6 or the Gemini 3.8 TTS endpoints — Cartesia's synthesis is its own product and the Gemini speech models come from Google's API — but the text layer and the failover path are ours, and those are the parts that make a brand-new model safe to try.

The thing to watch is the next Artificial Analysis revision. If Gemini 3.8 Flash TTS closes the thirteen points once its sample count matures, the price gap stops being a trade and starts being a straightforward answer. If it doesn't, Cartesia keeps the quality crown and Google keeps the price one, and the choice stays exactly where it is now: a subscription with telephony attached, or a meter with 130 languages attached.