
Gemini 3.8 TTS vs Inworld Realtime TTS-2: Which Scoreboard You Believe Changes the Answer
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Put Gemini 3.8 Flash TTS and Inworld Realtime TTS-2 next to each other on the Artificial Analysis Provider Voice Arena and the answer looks obvious: rank 2 against rank 4, Elo 1,260 against 1,245, 8 arena voices each, and a normalised $16.50 per million characters against $20.80. The vendor wins on position and on price, and the 15-point gap is inside both rows' confidence intervals anyway.
Move to the other board Artificial Analysis publishes and the ordering inverts. Inworld Realtime TTS-2 holds the top row on the Controlled Voice Arena at an Elo of 1,123, with its Flash variant at 1,075. That is not a rounding difference — it is a different model winning, and the reason is that the two boards are not measuring the same thing. If you are choosing a voice for a product, knowing which of those two questions your use case is actually asking is the whole decision.

Two boards, two different questions
The Provider Voice Arena scores a model's own native voices — the ones the vendor ships and hosts. The Controlled Voice Arena holds the voice constant and scores how well each model performs when it is asked to speak in a voice that is not its own. Same judges, same kind of blind comparison, different variable.
That distinction maps onto two different jobs:
• Provider ranking answers "does this model sound good out of the box". It is the right board for a product that ships with a fixed set of narrator voices and never clones anything.
• Controlled ranking answers "does this model obey a voice specification". It is the right board for a product where the voice is the customer's — a cloned brand voice, a licensed actor, a per-tenant identity.
Google's models lead the first. Inworld's lead the second. Both statements are true at the same time and neither is a contradiction, which is exactly why a single "which TTS model is best" answer is usually a sign that the writer picked one board and did not look at the other.

What Inworld's number-one position implies in practice
A first-place Controlled result is a claim about fidelity to an instruction, and fidelity to an instruction is the property that decides whether cloning is usable at all.
Inworld's cloning path comes in two forms. Instant cloning works from a short sample — the vendor's figure is 5 to 15 seconds of reference audio — and a professional cloning tier sits in beta and wants on the order of ten minutes. Both are vendor-reported, and neither is independently verified by the arena; what the arena measures is the model's behaviour once it has a voice, not the quality of the cloning step that produced it.
The latency claims attached to the family are in the same category. The Flash variant's first-byte figure — 25 milliseconds, reported through a third-party measurement — and the p99 numbers for the full model are Inworld's own, and the arena has nothing to say about either. They are plausible for a realtime-oriented model and they are unverified, which is the correct way to hold them.
So the practical reading is: if your product clones voices, Inworld's Controlled result is the more relevant of the two scores, and it is the reason a lower Provider rank is not disqualifying. If your product does not clone, that advantage is invisible to you and the Provider board is the one to read.
The price ladder has three rungs, and the cheap one is a subscription
Comparing one price against one price gets this matchup wrong, because Inworld does not have a price — it has a ladder.
• On demand — Realtime TTS-2 at $25.00 per million characters, Flash at $15.00. This is the rate a pay-as-you-go caller sees, and it is the vendor's own published figure.
• Creator plan — $20.00 and $10.00 for the same two models, which is a 20% and a 33% discount for committing to a plan rather than metering. Vendor-reported, and the rung most worth re-checking on the live pricing page before you commit, since plan terms move faster than list rates.
• Normalised — the arena's per-million-characters conversion gives $20.80 for Realtime TTS-2 and $10.40 for Flash, which is the board's arithmetic rather than either vendor's quote.
Against that, Google's rates are token-based and promotional: $0.50 per million text tokens in and $9.00 per million audio tokens out through December 31, 2026, then $18.00; Flash-Lite TTS is $0.50 and $6.00, then $12.00. Normalised, that is $16.50 and $11.00.

Read the two together and the picture is more interesting than "Google is cheaper". On the numbers as published today, Flash TTS is the cheapest of the three models and Inworld's Flash variant is the second cheapest — the two Flash models are close enough that a small change in your audio-to-text ratio can flip which one bills lower. On January 1, 2027, when the Google rate doubles, the ordering settles and Inworld becomes the cheaper option at every rung of its ladder. Anyone comparing these two in September and committing in December is comparing the wrong numbers.
Where each one is the better answer
The two models are close enough on quality that the differentiators are structural, so the honest version is a list of conditions rather than a verdict.
Pick Gemini 3.8 Flash TTS when the voice is yours to choose. Voice design from a written description, two-speaker dialogue in a single call, turn-level styling, and the widest language coverage of the two by Google's own count — none of which Inworld matches on this comparison. It is also the cheaper model today, and its sibling Flash-Lite gives you a second price point inside the same API.
Pick Inworld Realtime TTS-2 when the voice is the customer's. A top Controlled Voice row, an instant cloning path measured in seconds of reference audio, a professional cloning tier for the cases that need more, and a realtime orientation backed by first-byte claims the vendor publishes — with the caveat that the claims are the vendor's. It is English-only by the vendor's own documentation, which is a hard constraint if you need anything else.
Neither of these models is hosted by OrcaRouter, and it is worth saying so rather than implying otherwise: the place to call Gemini 3.8 Flash TTS is Google's own API and the surfaces Google named on September 23, and the place to call Inworld Realtime TTS-2 is Inworld's own platform. What OrcaRouter is for is everything around that choice — 200-plus models behind a single key at provider list price with no markup, so a vendor rate change like the January step is live on our side the same day, automatic failover so an evaluated model does not have to be a production dependency, and a routing DSL for composing models into one call when no single voice carries the whole job.
The reason to keep both of these in the running is that the January rate change and the Controlled-versus-Provider distinction are two separate reasons the answer can move. A pipeline that can only call one of them cannot follow either.
