
Gemini 3.8 Live vs GLM-4-Voice: What 21 Months of Voice AI Actually Bought
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 362 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 183 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1285 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 119 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 224 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
GLM-4-Voice shipped on December 3, 2024. Gemini 3.8 Live was announced on September 15, 2026. That is 21 months, and it is the most important fact in this comparison — more important than any benchmark, because it determines what kind of question you are actually asking.
These two models are not rivals. Nobody deploying voice at scale in 2026 is choosing between them on quality. The real question is the one GLM-4-Voice poses and Gemini 3.8 Live cannot answer: what would it take to own this rather than rent it? That question has a genuine answer, and it has changed less in 21 months than you might expect.
What Google shipped, and what Zhipu shipped
Gemini 3.8 Live is a hosted, closed-weight live dialogue model. It handles near real-time visual input, automatically detects and transitions between 97 supported languages mid-conversation, and executes tools and API calls in the background while the conversation continues. It is available to developers through the Gemini API and Google AI Studio, in private preview in Gemini Enterprise, and in Search Live. Announced pricing is $0.005 per minute of audio input and $0.018 per minute of audio output through the Live API; Artificial Analysis's cost-per-hour-of-input-audio chart independently puts it at $1.50 per hour. Its sibling, Gemini 3.8 Live Extended Thinking, currently tops that same index at 82.6, while Gemini 3.8 Live itself sits at 76.0.
GLM-4-Voice is Zhipu AI's first end-to-end speech model, and it is a downloadable checkpoint. It is a three-part assembly: a speech tokenizer built on a Whisper-large-v3 encoder with a vector-quantised bottleneck at roughly 0.4B parameters, an autoregressive language model (GLM-4-Voice-9B, initialised from GLM-4-9B) at about 9B parameters, and a flow-matching decoder retrained from CosyVoice. It speaks Chinese and English, supports instruction-controlled emotion, tone, speech rate, and dialect — Cantonese, Chongqing Mandarin, Beijing speech among them — and tokenises speech at 12.5 Hz into a single codebook stream at roughly 175 bps, an order of magnitude lower than typical neural audio codecs.
The dimensions, side by side:
• Release — Gemini 3.8 Live: September 15, 2026. GLM-4-Voice: December 3, 2024.
• Weights — closed, hosted only vs downloadable, Apache-2.0 code with a custom weights licence.
• Languages — 97 with mid-conversation switching vs Chinese and English.
• Vision — near real-time visual input vs none.
• Tool calling — background tool and API execution mid-conversation vs no tool channel at all.
• Context — not published by Google vs 8K context, 4K max output.
• Self-hosting floor — not applicable vs 32 GB RAM plus a CUDA GPU officially; roughly 12 GB VRAM at int4.
• Watermarking — SynthID on all generated audio vs none described.

21 months bought capability, not just quality
The easy story is that a 2026 frontier model beats a 2024 open checkpoint. It does, and the margin is large — but the more useful observation is that the category changed shape rather than moving along one axis.
On Zhipu's own technical report, GLM-4-Voice scores 93.6% speech-to-text accuracy on Topic-StoryCloze, 82.9% speech-to-speech, a UTMOS naturalness of 4.45, and spoken QA of 5.40/10 general and 5.20/10 knowledge — all self-reported and unreproduced. Independent evaluation is scarcer and less flattering: on VoiceBench it records an overall 56.48, placing around 29th of 42 in one tabulation and 30th of 40 in another, with multi-turn comprehension described as weak in both languages. On the C³ multi-turn dialogue comprehension benchmark it lands at 11.09% in Chinese and 33.22% in English. VCB Bench found it led on emotional control with 93.0 on the Chinese SIF sub-metric, and strong text–speech alignment, but TELEVAL dialect handling came in at 4.57% without explicit instruction.
Those numbers describe a model that is good at vocal expression and poor at sustained comprehension. That is a 2024 speech model in one sentence.
Gemini 3.8 Live's 76.0 on Artificial Analysis's Speech to Speech Index is a composite over speech reasoning, agentic performance, arena preference, and task success rate. It is not that the newer model is better at the same task by a larger multiple. It is that the composite now includes agentic performance and task success at all — categories GLM-4-Voice cannot compete in because it has no tool channel. The 2024 model is being scored on a game it was never built to play.

The licence is the part people skip
If you are looking at GLM-4-Voice because it is open, read the terms before the weights finish downloading. The split is real and it is easy to misread:
• Code — Apache-2.0. Genuinely permissive.
• Weights — a custom licence that requires attribution and registration with Zhipu for commercial use.
"Open weights" and "Apache-2.0" are not the same claim, and a lot of secondary coverage of this model conflates them. If you intend to ship a commercial product on GLM-4-Voice, the registration requirement is a legal step, not a formality, and it should be on the critical path. One tracker goes further and describes the model as proprietary with conditional commercial use. Either way, the absence of a per-minute bill is not the absence of a commercial relationship.
The cost arithmetic, done honestly
The case for self-hosting is that a fixed monthly cost beats a per-minute one at volume. That case is real, and it is smaller than it looks once you count what you are operating.
GLM-4-Voice officially wants 32 GB of RAM and a CUDA GPU. Community int4 quantisations run in roughly 12 GB of VRAM, about 10–11 GB in practice, which is RTX 3060 territory, with a practical ceiling of around 30 seconds of speech per turn. A cloud A10 at roughly $0.50–$1.00 per hour equates to somewhere between $350 and $700 a month for a continuously running instance — before you have served a single user, and with no failover, no burst capacity, and no one to page when the decoder produces noise at 3am.
Against that, Gemini 3.8 Live at $1.50 per hour of input audio means $350 buys you roughly 233 hours of voice. That is not obviously worse, and it comes with capacity you did not provision. The crossover exists; it is higher than the headline comparison suggests, and it moves depending on whether your traffic is steady or bursty. Bursty traffic is the case where per-minute pricing wins decisively, because idle GPUs bill exactly the same as busy ones.
There is also a difference in what you get for the money. Community reports of quantised GLM-4-Voice deployments put continuous-dialogue latency at 38–120 ms, though those figures come from write-ups rather than from Zhipu. Gemini 3.8 Live's latency has not been published by Google at all. If low latency is your hard requirement, neither of these has a number you can plan against — one has third-party community measurements, the other has partner testimonials and no figure.

The middle option: own the logic, rent the capability
Framing this as self-host versus hosted misses the arrangement most teams actually end up in. GLM-4-Voice has no tool channel, so any product built on it needs a separate text model to do the looking-up — which means the self-hosted speech model sits in front of a hosted text model either way. The deployment decision and the capability decision are separable.
That split is where a router does real work. OrcaRouter carries 190 models behind a single key at provider list price with no markup, so the delegation target behind your speech front-end can be swapped on live traffic without rebuilding anything, and an upstream price cut reaches you the same day rather than at the next renewal. Automatic failover matters more than usual in a self-hosted setup, because when your own GPU instance is the thing that fell over, the backend model is still reachable and the session does not have to die with it. Two clarifications, since this comparison invites confusion: we do not host GLM-4-Voice, and the Gemini 3.8 Live endpoints are not on our router either — those come from their vendors. What is on the router is the text layer both of them hand work to.
The decision, and the lesson under it
Choose GLM-4-Voice if you need the model on your own hardware, if your traffic is genuinely steady at high volume, if you are building in Chinese or English only, and if you need to modify the model rather than call it. Budget the licence registration as a real task, and expect to solve multi-turn comprehension yourself — it is the model's documented weak point and no amount of fine-tuning around a 4K output ceiling will fully hide that.
Choose Gemini 3.8 Live if your requirements include vision, tool calling, more than two languages, or any traffic pattern you would describe as spiky. It is a preview model with unpublished context and latency figures, which is a real planning risk, but it is also the only one of the two that can look something up mid-sentence.
The general lesson is worth more than either verdict. Twenty-one months of voice AI produced a model that is not primarily a better voice model — it is a voice interface onto an agent. If your reason for wanting open weights was cost, the arithmetic is closer than the licence-free framing suggests. If your reason was control, that has not been commoditised at all, and GLM-4-Voice remains a legitimate answer to a question Gemini 3.8 Live does not address.
