
Gemini 3.8 Live vs Qwen-Audio-3.1-TTS: Two Halves of a Voice Stack, Repriced Eight Days Apart
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 36 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 181 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1277 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Gemini 3.8 Live is Google's speech-to-speech model, shipped on September 15, 2026 as gemini-3.8-live and gemini-3.8-live-extended-thinking on the Live API. Qwen-Audio-3.1-TTS is Alibaba's text-to-speech model, announced at the Apsara Conference in Hangzhou on September 23, 2026, eight days later, with a price cut of roughly 70% on speech synthesis attached to it. That eight-day gap is the reason this comparison is worth reading now rather than in July. Nothing about the two products' roles changed — one listens and talks, the other reads text aloud — but both halves of a production voice pipeline were rebuilt or repriced inside a single fortnight, and the arithmetic that used to settle this matchup quietly moved.
Two halves, refreshed eight days apart
Start with the thing that makes this pairing unusual: the two models do not compete. A voice product has a mouth and it has an ear, and these are one of each. Gemini 3.8 Live closes the loop — audio and video in, audio out, interruptions handled, tool calls running underneath the conversation. Qwen-Audio-3.1-TTS takes a string and a reference voice and returns a performance. It is a better mouth than it is anything else, and it is not trying to be an ear.
What makes the September timing interesting is that the two announcements are aimed at opposite ends of the same budget line. Google's September 15 launch was a capability story with no price move attached; Alibaba's September 23 announcement was mostly a margin story, with new capability riding along. A team that built a hybrid pipeline in August has, in the space of eight days, been handed a reason to re-examine both legs — and only one of those legs got cheaper.
What the 3.1 release actually changed on the Qwen side
Alibaba did not ship one model on September 23. The 3.1 generation covers five endpoints — Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-ASR-Next, Qwen-Audio-3.1-TTS, Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-Realtime — and only one of them, the plain TTS tier, is a drop-in for anyone already calling the 3.0 TTS model.
On that tier, three things changed and one of them is a genuine migration cost. Delivery control moved from a fixed vocabulary of 86 inline tags to natural-language instruction: you describe the emotion, pace and style you want instead of inserting the right bracket at the right position. Cross-language timbre transfer arrived, letting one reference voice carry its identity across Mandarin, Cantonese, English and Japanese. And the model now tolerates noisy or echoed reference audio directly, without a separate denoise step. Language coverage is unchanged at a vendor-reported 16 languages plus 20 Chinese dialect regions, and — this is worth flagging because it gets smoothed over elsewhere — Alibaba's own material says the 3.1 tier adds seven languages, which does not reconcile with the 16 that published coverage of the July generation listed. Treat both counts as vendor-reported until you check the specific languages you need against Model Studio.
The genuinely new product in the release is Qwen-Audio-3.1-TTS-Next, which Alibaba calls an "Audiogen" model: one pass that produces voice, sound effects and background ambience together, for multi-speaker dialogue, podcast assembly and scene work. It lists at $0.848 per million input tokens and $1.696 per million output tokens on the Beijing node, capped at 3,000 characters of input, 240 seconds of generated audio for podcasts and 120 seconds otherwise, three requests per second, accepting up to three reference clips of 30 seconds each. It handles Chinese and English only. If your pipeline is a single narrator reading a script, that product is irrelevant to you; if it is two hosts and a music bed, it is the reason to look at this release at all.
The seam is the thing you are actually choosing
Because the two models do not do the same job, the decision is not a bake-off. It is a question about where you put the handoff, and how much you are willing to pay for the seam to be invisible.
There are two honest architectures. The first is one network: Gemini 3.8 Live handles the whole turn, hearing the caller, deciding what to say and saying it, and you accept the voice it gives you — which you cannot clone and cannot brand. The second is a cascade: recognition, then reasoning, then Qwen-Audio-3.1-TTS as the mouth, giving you a thousand voices including one recorded by your own talent, at the cost of latency across two or three vendor boundaries and the prosody that gets lost when a sentence is split across an API call.
Most teams end up with both, and that is not a failure of nerve. Use the realtime model where latency is felt — interruption, acknowledgement, the "mm-hm" that tells a caller you are still there — and the synthesis model where quality is heard: the long read-back of a policy, the confirmation number read at a deliberate pace, the branded voice that has to be identical on every single call. The hybrid is the correct answer for a large class of products. It is also where the cost model gets hard, because you are now metering two different units.

Priced per audio-hour, which is the only unit both fit
Google bills the Live API by the minute; Alibaba bills the Qwen TTS tiers by the character and by the token. To compare them you have to pick a normaliser, and the only defensible one for voice is the hour of finished audio.
• Metered in — Gemini 3.8 Live per minute of audio, $0.005 in and $0.018 out vs Qwen-Audio-3.1-TTS per million characters, with the 3.0 Plus tier tracking near $27.60 and the Flash tier near $15 before the cut, and the TTS Flash tier listing at 1.5 yuan per million input tokens and 12 yuan per million output tokens after it
• Metered out — Gemini 3.8 Live two-way conversation runs about $1.38 for an hour of wall-clock talk at list, before any caching vs Qwen-Audio-3.1-TTS a one-way stream, with the 3.1-Next tier at $0.848 per million input tokens and $1.696 per million output on the Beijing node
• Hears the caller — Gemini 3.8 Live yes, native audio and video input vs Qwen-Audio-3.1-TTS no, text in and audio out
• Voices — Gemini 3.8 Live one voice, not cloneable, not brandable vs Qwen-Audio-3.1-TTS a thousand-plus with zero-shot cloning from noisy reference audio
• Languages — Gemini 3.8 Live a vendor-reported 97, switched mid-conversation vs Qwen-Audio-3.1-TTS a vendor-reported 16 plus 20 Chinese dialect regions, with timbre transfer across Mandarin, Cantonese, English and Japanese
• Delivery control — Gemini 3.8 Live prompt and conversation context vs Qwen-Audio-3.1-TTS natural-language instruction, replacing the 3.0 generation's 86 inline tags
• Independent score — Gemini 3.8 Live 82.6 on the Artificial Analysis Speech to Speech Index for the Extended Thinking variant and 76.0 for the standard one vs Qwen-Audio-3.1-TTS none; the nearest Qwen number is 1,259 Elo for the previous-generation 3.0 Plus tier on the Provider Voice Arena, which is a scoring of the predecessor and not of this model
The percentage cut is quoted against rates that vary by region and tier, so the headline 70% is not a promise about your traffic, and Alibaba Cloud also moved realtime transcription on the Singapore node from duration-based to token-based metering — which is a less predictable basis, since silence and multi-turn context both inflate token consumption. Read the new rate card against your own audio.

What the 3.1 upgrade does not settle
Here is the part that is easy to get wrong in both directions. The 3.1 improvements do not make Qwen-Audio-3.1-TTS a candidate for the job Gemini 3.8 Live does — instruction-based control, timbre transfer and noisy-input tolerance all make it a better mouth, and none of them give it an ear. Equally, Gemini 3.8 Live's benchmark position does not tell you anything about whether your branded voice survives a sentence in Cantonese.
There is also a hole in the evidence that a price cut makes more tempting to ignore. Qwen-Audio-3.1-TTS has no independent score. The 1,259 Elo that appears next to the Qwen name on the Provider Voice Arena belongs to Qwen-Audio-3.0-TTS-Plus, the model being replaced, measured on 1,447 samples. The successor's own Arena row does not exist yet. If your sign-off process requires a third-party number — and for a brand voice it probably should — the newer model is the one without it, and the cheaper tier is the one you cannot yet defend. The single event that would settle this is Qwen-Audio-3.1-TTS appearing on a blind-listening board.
Where one key does and does not help
One consequence of the 3.1 release is easy to miss because it looks like a prompt-engineering detail rather than an infrastructure one. Replacing 86 hardcoded tags with free-form instruction turns your delivery directions into text artifacts. You now generate them, version them, and re-validate every one of them against the new model's interpretation — and the failure mode is drift, not an error, because two phrasings of "warm but not sentimental" will not land identically.
That is a text-inference problem wrapped around a speech pipeline, and it is where we are useful. OrcaRouter does not route Qwen-Audio-3.1-TTS or any Qwen-Audio model, and we do not route the Gemini 3.8 Live endpoints either — neither half of this stack is callable through us, and nothing here should be read as a claim that it is. What we do carry is the layer that writes and checks the direction text: qwen/qwen3.8-flash and google/gemini-3.8-flash among more than 200 models from a single key at provider list price with no markup, with a routing DSL that composes several models into one call and automatic failover when an upstream degrades. When you are about to re-validate ten thousand delivery prompts against a model that just changed how it reads them, generating the variants and scoring them for consistency is the part that should be one contract, not four.
What to measure before you choose
Three experiments answer more than any board. Take one reference voice and one paragraph of your own script through Qwen-Audio-3.1-TTS and listen for whether the instruction-based control gives you the same delivery twice — that is the migration risk, and it is cheaper to measure than to plan around. Take a real call recording with your own background noise through Gemini 3.8 Live and check whether the turn-taking holds when the caller interrupts mid-word — that is what the speech-to-speech index is trying to quantify, and it does not survive contact with a real phone line reliably enough to be taken on faith. And run the arithmetic on your own volume in audio-hours rather than in the units the two vendors invoice in, because on this particular pairing the expected price difference between "cheap Chinese TTS" and "Google flagship" was never as large as it looks, and after September 23 it is smaller still.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
