
GPT-Live-1 vs Gemini 3.5 Live Translate: A Shipped Generalist Against a Preview Specialist
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-10$0.15 / $0.60 per 1M tokens
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
These two models get compared because both put real-time speech in an API, and the comparison falls apart the moment you read the model cards. GPT-Live-1 is a general-purpose full-duplex voice model that OpenAI moved into the developer API on September 10, 2026, three months after it shipped inside ChatGPT on July 8. Gemini 3.5 Live Translate is a single-purpose audio-to-audio translation model that Google DeepMind released on June 9, 2026, and its model ID still carries the word preview. One of these you can put in front of paying customers today under a published rate. The other one Google can change under you without a deprecation notice. That asymmetry, not the benchmark sheet, is what should decide the pick.
Two different machines wearing the same label
It is worth being blunt about how little these overlap. Gemini 3.5 Live Translate does translation. It takes streaming audio in one language and returns streaming audio in another, preserving the speaker's voice and tone, across 70-plus languages and more than 2,000 language pairs, with automatic language detection and bidirectional operation. It does not hold a conversation, call a function, or answer a question. It is a simultaneous-interpretation engine.
GPT-Live-1 is the opposite shape. It is a conversational model: it listens and speaks at once, decides many times a second whether to keep listening, pause, interrupt, or call a tool, and hands heavy work — web search, deeper reasoning, agentic tasks — to a separate reasoning model through the Responses API while the conversation continues. OpenAI's own published WebRTC example wires that delegation to GPT-5.6 Terra. Translation is one of the things it can do; it is not a feature Google's model has any equivalent of in the other direction.
So the practical question is narrower than the headline matchup suggests: if what you are building is interpretation, Google's model is purpose-built and OpenAI's is general-purpose. If what you are building is a voice agent that occasionally translates, GPT-Live-1 is the only one of the two that is even the right category of product.
What each one costs per minute of audio
This is where the specialist earns its keep, and the arithmetic is genuinely awkward for OpenAI.
• GPT-Live-1 — $0.05 per minute of voice, billed by the second rather than rounded up to the next whole minute. Reasoning and tool calls made by the delegated backend bill separately at that model's normal rates.
• Gemini 3.5 Live Translate — $3.50 per million audio input tokens and $21.00 per million audio output tokens, with billing computed at 25 tokens per second of audio. Combined, that lands at roughly $0.037 per minute of translated audio.
On pure translation minutes, Google is about a quarter cheaper. That is not a rounding difference at scale: 10,000 interpreted minutes a month costs about $500 on GPT-Live-1's voice layer against roughly $368 on Gemini 3.5 Live Translate. And the gap is real rather than an artifact of the metering change, because the two models bill on genuinely different units.
There is a catch in the other direction. The $0.05 headline for GPT-Live-1 is the price of the voice layer only. If your agent translates and also looks something up, or escalates a hard sentence to a reasoning model, you are paying for that delegation on top — and the delegated model is priced per token like any other frontier model. A translation-only workload never triggers that. A translation-plus-anything workload does, and the winner of the per-minute comparison stops being obvious.

The preview label is the risk nobody prices in
Gemini 3.5 Live Translate's model ID is gemini-3.5-live-translate-preview. It shipped to the Gemini Live API and Google AI Studio as a public preview, and simultaneously into the Google Translate app on Android and iOS and a private preview in Google Meet for select Workspace customers. Preview means what it has always meant at Google: no stability guarantee, no published deprecation window, no commitment that the rate you are paying survives the next release.
For a consumer app that is fine. For the workloads this model is actually aimed at — live interpretation in a call centre, a meeting product, a travel product — it is the single largest integration risk in the stack. You are building a production dependency on a model Google has explicitly not committed to.
GPT-Live-1 has the inverse problem, and it is smaller but not zero. It is generally available, at a published rate, with a documented endpoint, and it is not going away. But its API surface is narrow in a way that surprises teams who assume it drops into an existing OpenAI integration: there is exactly one model ID, one endpoint, and no path through Chat Completions, Responses, Realtime, Batch, Assistants, or fine-tuning. The only way in is the Live Sessions endpoint at v1/live/sessions over WebRTC, with event handling on a data channel. That is a new integration, not a parameter change.

Languages, and where the generalist is honestly weaker
Google's number is 70-plus languages with voice and tone preservation. OpenAI's own documentation is markedly less confident about its own coverage: the launch material notes that for certain languages the model may carry a non-native accent or show gaps in fluency, and it does not publish a language count that competes with Google's.
That is not a vendor being coy — it is what a generalist conversational model looks like next to a specialist trained on the problem. If your product's value proposition is "we interpret 40 languages well," the specialist is the honest choice and the generalist will embarrass you in the long tail. If your product is a voice agent for an English-speaking market with a translation feature bolted on, the generalist's language gaps will never come up.
Both models watermark generated audio with SynthID, which matters if you are in a jurisdiction or a customer contract that requires provenance on synthetic speech. That one is a wash.

Where the delegation architecture changes the build
The structural difference between these two is that GPT-Live-1 is really two models with a seam in the middle. The voice layer keeps the conversation alive; a separate reasoning model does the thinking. That seam is where your engineering effort goes, and it is also where the cost model gets complicated.
It is worth knowing that the delegated half is an ordinary text model you can route like any other. GPT-5.6 Terra, the backend in OpenAI's own example, is served through OrcaRouter at $2.00 per million input tokens and $12.00 per million output, with a 1M-token context window and both /v1/chat/completions and /v1/responses exposed. If you want to swap the reasoning brain behind a voice agent — for a cheaper model on simple calls, or a stronger one on escalations — that half of the stack has options, because it is a normal API call sitting alongside roughly 190 other models on one key.
The voice half does not. GPT-Live-1 is not on OrcaRouter, and neither is Gemini 3.5 Live Translate; both are available only from the vendor that made them. Where a router earns its place in an architecture like this is everything downstream of the audio: the text models doing the retrieval, the summarisation, the CRM write-back and the escalation, which is the half of the bill you can actually consolidate onto one key and one invoice.
What to actually pick
• Choose Gemini 3.5 Live Translate if translation is the product, the price per minute matters more than contract stability, and you can absorb a preview model changing underneath you. Budget roughly $0.037 a minute and keep an abstraction layer over the call so a GA model can replace it without a rewrite.
• Choose GPT-Live-1 if you need a conversational agent that happens to translate, if you need function calling and tool use inside the voice loop, or if a procurement team will ask what happens when the model is retired and you need an answer better than "nothing, probably."
• Do not choose either one because it won a benchmark. The benchmarks on both sides are not comparable — OpenAI's full-duplex numbers are its own, unreproduced by any third party, and Google's language claims are untested against a generalist on the same audio set.
One forward-looking note: the two surfaces are converging rather than colliding. Google's translation model already lives inside Gemini Live, and GPT-Live-1's voice layer is explicitly designed to have new frontier models dropped behind it. The reasonable expectation is that within a couple of release cycles the generalist's translation improves and the specialist's integration story becomes less of a gamble. If you are building today and the decision is close, that is an argument for the cheaper, more replaceable option, and for keeping the audio path isolated behind an interface either way.
