Title card reading 'Qwen3.8-LiveTranslate vs Gemini 3.5 Live Translate' with the subtitle 'One ships a number, the other ships a map'.
Guides & Insights

Qwen3.8-LiveTranslate vs Gemini 3.5 Live Translate: One Ships a Number, the Other Ships a Map

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two leading real-time speech translation models disagree about what they are willing to be measured on, and that disagreement is more useful than any spec comparison. Qwen3.8-LiveTranslate, released on September 19, 2026, leads with a single hard figure: average lagging of 2.3 seconds, down from 2.8 in the previous generation. Gemini 3.5 Live Translate, the vendor's speech-to-speech model announced in June 2026 and still carrying a preview model ID, leads with reach — 70-plus languages, automatic detection, mid-conversation switching, and integration into the vendor's Translate and Meet apps. One company published a latency number it can be held to. The other published a language count. If you are choosing between them, understanding why those are different kinds of promises matters more than which list is longer.

Two architectures that agree on the goal and nothing else

Both models exist to kill the same behaviour: the turn-based translator that waits for you to finish a sentence before it says anything. That pause is what makes machine interpretation feel like machine interpretation.

Qwen3.8-LiveTranslate gets there with an Interleave architecture on a Hybrid MoE backbone, split into a Thinker module that folds audio, video, source text and translation into one causal sequence, and a Talker module that re-synthesizes the translation in the original speaker's timbre. The claim is that collapsing recognition, translation and synthesis into a single sequence removes the latency and semantic loss that a three-stage pipeline pays at every module boundary.

Gemini 3.5 Live Translate takes the same end-to-end position from the other direction: it is built on Gemini 3 Pro as a native audio-to-audio model, streaming continuously through the Gemini Live API rather than executing a discrete ASR → MT → TTS cascade. In practice you hold a persistent bidirectional WebSocket, push 100 ms audio chunks up, and pull translated audio chunks down. Neither model does the old thing. Both now do the new thing. The differences are all in the second-order details.

The latency comparison is not the tie it looks like

Google describes Gemini 3.5 Live Translate as staying "a few seconds behind" the speaker. It has not published a LAAL figure, or any other named, reproducible latency metric, for the model. Qwen has published 2.3 seconds and defined what it means.

This asymmetry is routinely flattened into "both are around two to three seconds." That reading is not supported by anything either company has said. Google's phrase is compatible with 2 seconds and compatible with 4; the company has simply declined to be pinned. The comparison that can actually be made today is:

Published latency metric — Qwen3.8-LiveTranslate: LAAL 2.3 s, vendor-reported. Gemini 3.5 Live Translate: none published.

Independent latency measurement — neither, for either model. No third party has published a head-to-head.

The honest conclusion is that on latency, Qwen has made a falsifiable claim and Google has made a vague one. That is a point in Qwen's favour on transparency and exactly zero points in its favour on performance until somebody measures both. If latency is the deciding factor for your deployment, the current state of the evidence does not support a purchase decision in either direction.

Two-column comparison scoreboard. Qwen3.8-LiveTranslate column: latency LAAL 2.3s, languages 60 source and 29 spoken, speaker ID real-time diarization, input audio plus image, status generally available, watermark none stated. Gemini 3.5 Live Translate column: latency not published, languages 70-plus supported, speaker ID weak turn-taking, input audio only, status preview, watermark SynthID on all audio. Footer reads 'Qwen figures vendor-reported; Gemini per Google docs. No independent head-to-head.'

60 languages versus 70-plus is the wrong scoreboard

The language counts are quoted constantly and they mislead in both directions.

Qwen3.8-LiveTranslate recognizes 60 source languages and produces spoken output in 29 of them. Gemini 3.5 Live Translate supports 70-plus languages with automatic detection, and in Google Meet that expands to 2,000-plus language combinations where the previous ceiling was English-based translation across five languages.

Read the second number carefully before treating it as a threefold win. Google's 2,000-plus figure counts ordered pairs — every source language crossed with every target — which is a legitimate and genuinely useful measure of coverage but not comparable to a single-language count. And Google's list, while broader, is heavily weighted toward languages with large speaker populations: Afrikaans, Arabic, Bengali, Dutch, English, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Malay, Persian, Polish, Portuguese, Russian, Spanish, Swahili, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Zulu. Qwen's 29 spoken-output languages are narrower still, and neither vendor's list is a serious answer to the long tail — the regional dialects and low-resource languages where interpretation tools have historically failed hardest. Both companies say they are working on it. Neither has shipped it.

Where the two actually diverge: the room full of people

The one place the models make genuinely different promises is multi-speaker handling, and it is the difference most likely to decide a real deployment.

Qwen3.8-LiveTranslate ships real-time speaker diarization: each sentence is attributed to a speaker as it arrives, and voice replication is described as stable across turn changes. Qwen self-reports a diarization error rate of 9.7% on its own long multi-speaker test set, against 30.6% for Seed LiveInterpret 2.0 — a vendor comparison on a vendor-run evaluation, so treat it as a claim with a number attached rather than as a result. The model also emits source text and translation on the same timeline, which is what makes synchronized bilingual captions possible without a separate alignment pass.

Gemini 3.5 Live Translate takes the opposite emphasis. Its documented strength is prosody preservation — keeping the speaker's pitch, pacing, intonation and emotional register, so the translated voice carries the original's feel. Its documented weaknesses are precisely where diarization would help: inconsistent voice cloning that can drift after long pauses or land on the wrong gender, weak turn-taking with no clear end-of-sentence signal, trouble with strong accents and with closely related language pairs such as Spanish and Portuguese, and imperfect handling of background noise. Google also caps pure audio sessions at 15 minutes unless extended, and stamps every generated audio stream with a SynthID watermark that cannot currently be removed.

Multi-speaker attribution — Qwen: real-time diarization, claimed 9.7% DER. Gemini: documented weak turn-taking, no diarization claim.

Voice fidelity — Qwen: source-timbre reproduction via the Talker module. Gemini: prosody, pitch and pacing preservation; documented cloning drift.

Bilingual output — Qwen: source and translation emitted on the same timeline. Gemini: optional input and output transcription, configured per session.

Watermarking — Qwen: none stated. Gemini: SynthID on all generated audio, no removal path.

Session length — Qwen: not stated. Gemini: 15-minute cap on pure audio sessions.

Input types — Qwen: audio and image. Gemini: audio only, no text input.

Screenshot of Google's own Gemini Live API documentation for live translation, showing the speech-to-speech streaming setup, the supported-language list, and the documented session constraints for the preview model.

What each one costs to actually run

Qwen3.8-LiveTranslate is priced per million tokens over the WebSocket Realtime API. In the Singapore region: $7.50 audio input, $0.55 image input, $20.00 text output, $30.00 audio output. In Beijing: ¥40 / ¥3.3 / ¥100 / ¥160 per million — roughly $5.65 / $0.47 / $14.13 / $22.61. Its context is 53,248 tokens, and the rate limit is 10 requests per minute and 100,000 tokens per minute in both regions.

That 10 RPM ceiling is the most consequential number in the rate card. Gemini 3.5 Live Translate's commercial terms are not published the same way, which is itself a signal about maturity: Google is distributing it through the Live API and AI Studio in public preview, Google Meet in a private preview for selected Workspace customers, and the Translate app on Android and iOS. Preview pricing and preview rate limits are subject to change without notice, and Google has not committed to a GA date.

So the cost comparison right now is between a model with a published price and a hard concurrency ceiling, and a model with no published price and an unspecified ceiling. If you need to model unit economics this quarter, only one of those is budgetable.

Screenshot of Alibaba Cloud Model Studio documentation for the qwen3.8-livetranslate-flash-realtime model, showing the WebSocket Realtime API endpoint and the published per-million-token pricing for audio input, image input, text output and audio output.

What an independent test would have to show

Everything above is built from the two vendors' own announcements, because no one has run these models against each other. A result that would actually settle this matchup needs four things, and none of them exist yet:

• LAAL measured the same way on both models, on the same audio, in the same language pairs — Qwen's 2.3 s figure cannot be compared to anything Google has declined to state.

• Diarization error rate on a shared multi-speaker corpus, not on Qwen's own Omnilingua-MSpeaker set.

• Voice-cloning stability across long sessions, which is where Gemini's own documentation flags drift and Qwen makes its strongest claim.

• Behaviour at the concurrency limits — whether Qwen's 10 RPM holds up under real meeting load and whether Gemini's preview quotas survive contact with production.

Until that test exists, the defensible position is narrow: Qwen has published more and promised more specific things; Google has broader language coverage and a distribution footprint no API-only model can match.

Which one to pick

Pick on the shape of your problem, not on the scoreboard, because the scoreboard is not yet trustworthy.

If your workload is multi-party and attribution matters — meeting transcription, a panel, a courtroom, anything where knowing who said a translated sentence is part of the value — Qwen3.8-LiveTranslate is the only one of the two that claims the capability, and Gemini's documented turn-taking weakness points the same way. If your workload is broad-language, consumer-facing, and already inside Google's ecosystem, Gemini 3.5 Live Translate's 70-plus languages and its presence in the Translate app and Meet are advantages that no API-only competitor matches on distribution.

If you are building a product rather than buying a demo, note that both are narrow specialists. Neither runs an agent loop, neither summarizes, and Qwen3.8-LiveTranslate explicitly supports no function calling, no structured output, no context caching and no fine-tuning. The interpreter is the front of your pipeline, not the whole of it. OrcaRouter is not the place to call either translation model today — neither is on our catalogue, and we would rather say so than imply otherwise. What we do provide is the half of the stack that consumes the transcript: one key across 200-plus models at the provider's list price with zero markup passed through, so the summarizer, glossary enforcer or downstream agent reading that transcript can be swapped or failed over without touching a second vendor contract.

The bottom of it

Qwen3.8-LiveTranslate and Gemini 3.5 Live Translate are close enough on the fundamentals that the marketing has become the differentiator: one published a latency number and a rate card, the other published a language map and an ecosystem. Neither has submitted to an independent test. The version of this matchup worth reading is the one that gets written after somebody runs both on the same audio — and given how fast this category is moving, that may not be far off.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily