A generated hero card for 'Gemini 3.8 Live vs Qwen-Audio-3.1-TTS' on a white background with soft blue-and-cyan gradient accents. Two panels sit under a heading reading 'Two halves, refreshed eight days apart'. The left panel covers '15 Sept 2026 — Gemini 3.8 Live and Extended Thinking on the Live API'. The right panel covers '23 Sept 2026 — Qwen-Audio-3.1-TTS at Apsara, synthesis cut ~70%'. A footer line reads 'They meet in the middle of a pipeline, not at the same job.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Gemini 3.8 Live vs Qwen-Audio-3.1-TTS: Two Halves of a Voice Stack, Repriced Eight Days Apart

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Gemini 3.8 Live is Goog​le's speech-to-speech model, shipped on September 15, 2026 as gemini-3.8-live and gemini-3.8-live-extended-thinking on the Live API. Qwen-Audio-3.1-TTS is Aliba​ba's text-to-speech model, announced at the Apsara Conference in Hangzhou on September 23, 2026, eight days later, with a price cut of roughly 70% on speech synthesis attached to it. That eight-day gap is the reason this comparison is worth reading now rather than in July. Nothing about the two products' roles changed — one listens and talks, the other reads text aloud — but both halves of a production voice pipeline were rebuilt or repriced inside a single fortnight, and the arithmetic that used to settle this matchup quietly moved.

Two halves, refreshed eight days apart

Start with the thing that makes this pairing unusual: the two models do not compete. A voice product has a mouth and it has an ear, and these are one of each. Gemini 3.8 Live closes the loop — audio and video in, audio out, interruptions handled, tool calls running underneath the conversation. Qwen-Audio-3.1-TTS takes a string and a reference voice and returns a performance. It is a better mouth than it is anything else, and it is not trying to be an ear.

What makes the September timing interesting is that the two announcements are aimed at opposite ends of the same budget line. Goog​le's September 15 launch was a capability story with no price move attached; Aliba​ba's September 23 announcement was mostly a margin story, with new capability riding along. A team that built a hybrid pipeline in August has, in the space of eight days, been handed a reason to re-examine both legs — and only one of those legs got cheaper.

What the 3.1 release actually changed on the Qwen side

Aliba​ba did not ship one model on September 23. The 3.1 generation covers five endpoints — Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-ASR-Next, Qwen-Audio-3.1-TTS, Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-Realtime — and only one of them, the plain TTS tier, is a drop-in for anyone already calling the 3.0 TTS model.

On that tier, three things changed and one of them is a genuine migration cost. Delivery control moved from a fixed vocabulary of 86 inline tags to natural-language instruction: you describe the emotion, pace and style you want instead of inserting the right bracket at the right position. Cross-language timbre transfer arrived, letting one reference voice carry its identity across Mandarin, Cantonese, English and Japanese. And the model now tolerates noisy or echoed reference audio directly, without a separate denoise step. Language coverage is unchanged at a vendor-reported 16 languages plus 20 Chinese dialect regions, and — this is worth flagging because it gets smoothed over elsewhere — Aliba​ba's own material says the 3.1 tier adds seven languages, which does not reconcile with the 16 that published coverage of the July generation listed. Treat both counts as vendor-reported until you check the specific languages you need against Model Studio.

The genuinely new product in the release is Qwen-Audio-3.1-TTS-Next, which Aliba​ba calls an "Audiogen" model: one pass that produces voice, sound effects and background ambience together, for multi-speaker dialogue, podcast assembly and scene work. It lists at $0.848 per million input tokens and $1.696 per million output tokens on the Beijing node, capped at 3,000 characters of input, 240 seconds of generated audio for podcasts and 120 seconds otherwise, three requests per second, accepting up to three reference clips of 30 seconds each. It handles Chinese and English only. If your pipeline is a single narrator reading a script, that product is irrelevant to you; if it is two hosts and a music bed, it is the reason to look at this release at all.

The seam is the thing you are actually choosing

Because the two models do not do the same job, the decision is not a bake-off. It is a question about where you put the handoff, and how much you are willing to pay for the seam to be invisible.

There are two honest architectures. The first is one network: Gemini 3.8 Live handles the whole turn, hearing the caller, deciding what to say and saying it, and you accept the voice it gives you — which you cannot clone and cannot brand. The second is a cascade: recognition, then reasoning, then Qwen-Audio-3.1-TTS as the mouth, giving you a thousand voices including one recorded by your own talent, at the cost of latency across two or three vendor boundaries and the prosody that gets lost when a sentence is split across an API call.

Most teams end up with both, and that is not a failure of nerve. Use the realtime model where latency is felt — interruption, acknowledgement, the "mm-hm" that tells a caller you are still there — and the synthesis model where quality is heard: the long read-back of a policy, the confirmation number read at a deliberate pace, the branded voice that has to be identical on every single call. The hybrid is the correct answer for a large class of products. It is also where the cost model gets hard, because you are now metering two different units.

A screenshot of the Artificial Analysis Speech to Speech leaderboard in English, captured 25 September 2026, showing the AA Speech to Speech Index panel with index values of 82.6, 81.5, 81.3, 80.1, 76.0, 73.9 and 71.5, with the cost-per-hour-of-input-audio axis running to roughly $10.75. The bar labels are set vertically and are not legible at capture size. No Alibaba Qwen-Audio model appears on the board, because the board scores speech-to-speech models and Qwen-Audio-3.1-TTS is a synthesis model.

Priced per audio-hour, which is the only unit both fit

Goog​le bills the Live API by the minute; Aliba​ba bills the Qwen TTS tiers by the character and by the token. To compare them you have to pick a normaliser, and the only defensible one for voice is the hour of finished audio.

• Metered in — Gemini 3.8 Live per minute of audio, $0.005 in and $0.018 out vs Qwen-Audio-3.1-TTS per million characters, with the 3.0 Plus tier tracking near $27.60 and the Flash tier near $15 before the cut, and the TTS Flash tier listing at 1.5 yuan per million input tokens and 12 yuan per million output tokens after it
• Metered out — Gemini 3.8 Live two-way conversation runs about $1.38 for an hour of wall-clock talk at list, before any caching vs Qwen-Audio-3.1-TTS a one-way stream, with the 3.1-Next tier at $0.848 per million input tokens and $1.696 per million output on the Beijing node
• Hears the caller — Gemini 3.8 Live yes, native audio and video input vs Qwen-Audio-3.1-TTS no, text in and audio out
• Voices — Gemini 3.8 Live one voice, not cloneable, not brandable vs Qwen-Audio-3.1-TTS a thousand-plus with zero-shot cloning from noisy reference audio
• Languages — Gemini 3.8 Live a vendor-reported 97, switched mid-conversation vs Qwen-Audio-3.1-TTS a vendor-reported 16 plus 20 Chinese dialect regions, with timbre transfer across Mandarin, Cantonese, English and Japanese
• Delivery control — Gemini 3.8 Live prompt and conversation context vs Qwen-Audio-3.1-TTS natural-language instruction, replacing the 3.0 generation's 86 inline tags
• Independent score — Gemini 3.8 Live 82.6 on the Artificial Analysis Speech to Speech Index for the Extended Thinking variant and 76.0 for the standard one vs Qwen-Audio-3.1-TTS none; the nearest Qwen number is 1,259 Elo for the previous-generation 3.0 Plus tier on the Provider Voice Arena, which is a scoring of the predecessor and not of this model

The percentage cut is quoted against rates that vary by region and tier, so the headline 70% is not a promise about your traffic, and Aliba​ba Cloud also moved realtime transcription on the Singapore node from duration-based to token-based metering — which is a less predictable basis, since silence and multi-turn context both inflate token consumption. Read the new rate card against your own audio.

A screenshot of the Artificial Analysis Provider Voice Arena leaderboard in English, captured 25 September 2026, showing Cartesia Sonic 3.6 first at Elo 1,273 with a 17-point interval over 1,757 samples and $49.0 per million characters, Google Gemini 3.8 Flash TTS second at 1,260 on 1,999 samples at $16.5, Alibaba Qwen-Audio-3.0-TTS-Plus third at 1,259 on 1,447 samples at $27.6, and Google Gemini 3.8 Flash-Lite TTS sixth at 1,235 on 2,000 samples at $11.0. The board scores synthesis models and carries no row for Qwen-Audio-3.1-TTS or for any Gemini 3.8 Live endpoint.

What the 3.1 upgrade does not settle

Here is the part that is easy to get wrong in both directions. The 3.1 improvements do not make Qwen-Audio-3.1-TTS a candidate for the job Gemini 3.8 Live does — instruction-based control, timbre transfer and noisy-input tolerance all make it a better mouth, and none of them give it an ear. Equally, Gemini 3.8 Live's benchmark position does not tell you anything about whether your branded voice survives a sentence in Cantonese.

There is also a hole in the evidence that a price cut makes more tempting to ignore. Qwen-Audio-3.1-TTS has no independent score. The 1,259 Elo that appears next to the Qwen name on the Provider Voice Arena belongs to Qwen-Audio-3.0-TTS-Plus, the model being replaced, measured on 1,447 samples. The successor's own Arena row does not exist yet. If your sign-off process requires a third-party number — and for a brand voice it probably should — the newer model is the one without it, and the cheaper tier is the one you cannot yet defend. The single event that would settle this is Qwen-Audio-3.1-TTS appearing on a blind-listening board.

Where one key does and does not help

One consequence of the 3.1 release is easy to miss because it looks like a prompt-engineering detail rather than an infrastructure one. Replacing 86 hardcoded tags with free-form instruction turns your delivery directions into text artifacts. You now generate them, version them, and re-validate every one of them against the new model's interpretation — and the failure mode is drift, not an error, because two phrasings of "warm but not sentimental" will not land identically.

That is a text-inference problem wrapped around a speech pipeline, and it is where we are useful. OrcaRouter does not route Qwen-Audio-3.1-TTS or any Qwen-Audio model, and we do not route the Gemini 3.8 Live endpoints either — neither half of this stack is callable through us, and nothing here should be read as a claim that it is. What we do carry is the layer that writes and checks the direction text: qwen/qwen3.8-flash and google/gemini-3.8-flash among more than 200 models from a single key at provider list price with no markup, with a routing DSL that composes several models into one call and automatic failover when an upstream degrades. When you are about to re-validate ten thousand delivery prompts against a model that just changed how it reads them, generating the variants and scoring them for consistency is the part that should be one contract, not four.

What to measure before you choose

Three experiments answer more than any board. Take one reference voice and one paragraph of your own script through Qwen-Audio-3.1-TTS and listen for whether the instruction-based control gives you the same delivery twice — that is the migration risk, and it is cheaper to measure than to plan around. Take a real call recording with your own background noise through Gemini 3.8 Live and check whether the turn-taking holds when the caller interrupts mid-word — that is what the speech-to-speech index is trying to quantify, and it does not survive contact with a real phone line reliably enough to be taken on faith. And run the arithmetic on your own volume in audio-hours rather than in the units the two vendors invoice in, because on this particular pairing the expected price difference between "cheap Chinese TTS" and "Goog​le flagship" was never as large as it looks, and after September 23 it is smaller still.

A generated summary card for Gemini 3.8 Live vs Qwen-Audio-3.1-TTS on a white background with soft blue-and-cyan gradient accents. Five rows read 'Gemini 3.8 Live — speech-to-speech, Live API, 15 September 2026', 'Qwen-Audio-3.1-TTS — text-to-speech, Apsara, 23 September 2026', 'The gap: eight days, and only one of them got cheaper', 'Independent score: Gemini 82.6 on AA Speech to Speech; Qwen 3.1-TTS none', and 'Verdict: not substitutes — pick which half you are shopping for'. A footer line reads 'Vendor-reported figures where stated; independently measured where a board is named.' The OrcaRouter logo is composited in the bottom-right corner.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily