
GPT-Live-1 vs ElevenLabs: One Model That Listens, One Company That Sells Everything Else
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
The most useful number in this comparison is 90.5%. That is the task success rate Artificial Analysis measured for ElevenLabs Agents in its Speech Agent Arena — the highest in the top six of that board, achieved by a system that is not a single model at all but a cascade of speech recognition, a language model and a synthesiser stitched together by ElevenLabs' own turn-taking logic. GPT-Live-1, OpenAI's end-to-end full-duplex model, does not appear on that board, because it competes on a different one: the Speech to Speech Index, where it sits joint sixth of fourteen at 70.3%. Meanwhile ElevenLabs Eleven v3 remains the most widely deployed production voice in the industry, with more than 10,000 voices in its library and 70-plus languages.
The shorthand is that one side sells you a conversation and the other sells you a company. That shorthand is mostly right, and it hides the more interesting finding underneath: a well-engineered cascade still beats a native full-duplex model at finishing the task.
Two architectures, and the difference is where the seams are
GPT-Live-1 is a single model that takes audio and text in and produces audio and text out. It handles interruption natively — OpenAI's own testing, published with September's API release and unreproduced outside the company, reports 80.1% interactivity on Full Duplex Bench v1.5 against 45.4% for GPT-Realtime-2.1, and turn-taking latency of 0.798 seconds against 1.41. There are no seams, because there is nothing to seam together.
ElevenLabs Agents is the opposite bet, deliberately. ElevenLabs' documentation describes a pipeline: tuned speech recognition, a language model you choose or bring yourself, a low-latency synthesiser — typically something from the Flash line — and a proprietary turn-taking model on top. Barge-in is a per-agent configuration exposing turn eagerness and timeouts rather than a property of the model. In Artificial Analysis's benchmarked default configuration the stack reads Scribe v2 Realtime for speech recognition, a Gemini model for reasoning, and Eleven Flash v2 for synthesis.
That design has one enormous advantage and one structural cost. The advantage is that every stage is replaceable: a cheap model for simple turns, a frontier model for hard ones, a different synthesiser for a different locale. The cost is that latency accumulates at each seam, and the turn-taking decision has to be made from the outside.
Dimension by dimension
• Architecture — GPT-Live-1: one end-to-end model, audio in and audio out vs ElevenLabs Agents: cascaded speech recognition, language model and synthesis with a proprietary turn-taking layer
• Independent evidence — GPT-Live-1: 70.3% on the Artificial Analysis Speech to Speech Index, joint sixth of fourteen vs ElevenLabs Agents: Elo 937 on the Speech Agent Arena with a 90.5% task success rate across 676 samples, the best success rate in that board's top six
• Quality boards — GPT-Live-1: not ranked on the text-to-speech arenas, being a different kind of model vs ElevenLabs: Eleven v3 Conversational at 1,194 Elo and Eleven v3 at 1,166 on the Provider Voice board, priced at $50 and $100 per million characters respectively
• Price — GPT-Live-1: $0.05 per voice minute, billed per second, backend reasoning metered separately vs ElevenLabs: around $50 per million characters for Eleven v3 Conversational and about $100 for Eleven v3, with agents reported at roughly $0.08 per minute, rising past the concurrency ceiling, and language model and telephony costs billed separately
• Latency — GPT-Live-1: no model-level time-to-first-audio published; 0.798 seconds turn-taking latency in vendor testing vs ElevenLabs: about 75 milliseconds for Flash v2.5 and about 280 milliseconds for Eleven v3 Conversational, with vendor guidance of under 800 milliseconds to first audio at the agent level
• Voices and cloning — GPT-Live-1: twelve API presets, no cloning by design vs ElevenLabs: more than 10,000 voices in the library, instant cloning from a short sample, professional cloning from 30 to 180 minutes of audio with self-verification
• Languages — GPT-Live-1: major languages well served, uneven fluency beyond them in launch coverage vs ElevenLabs Eleven v3: 70-plus languages, against 32 for the Flash line and 90-plus for its speech recognition
• Breadth — GPT-Live-1: one endpoint, one model ID, concurrency tiers from 25 to 500 sessions and no free tier vs ElevenLabs: synthesis, speech recognition, dubbing, sound effects, music, a voice marketplace and an agents platform under one contract, with a free tier that carries no commercial licence


Where ElevenLabs is not just ahead, but uncontested
Three of the gaps in that list are not close, and any comparison that gives them a sentence is being unfair to the incumbent.
Cloning is the first. GPT-Live-1 ships twelve presets and, by design, no cloning at all — a deliberate safety posture rather than a missing feature. ElevenLabs offers instant cloning from a short reference sample and professional cloning trained on 30 to 180 minutes of audio, gated behind plan tiers and a self-verification step. If your product's voice is part of its identity, this is not a dimension on which GPT-Live-1 competes.
The voice library is the second. More than 10,000 voices, plus a licensed marketplace of iconic voices from estates and performers, is an asset no model release replicates. It is also the thing buyers most often underestimate: picking a voice is the last mile of a voice product, and having ten thousand to choose from compresses a branding exercise into an afternoon.
The third is breadth. Speech recognition, dubbing, sound effects, music, an agents platform — a team that needs four of those buys one contract instead of four. That is a strategy advantage, not a quality one, and it survives the other side winning a benchmark.
Both companies carry governance questions that belong in a procurement review rather than a footnote. ElevenLabs' professional cloning requires self-verification and its terms prohibit cloning another person's voice even with consent; the company also faces a set of class actions filed in May 2026 over training-data consent, in which the claims are unproven. OpenAI's position is simpler because it removed the feature that generates the risk.

The cascade tax, and where it is worth paying
A cascade's costs show up as latency and as orchestration. Vendor guidance for ElevenLabs puts an agent's time to first audio under 800 milliseconds at the median and under 1.5 seconds at the 95th percentile — figures that describe the whole pipeline, not the synthesiser's 75 milliseconds. Stacked on top are the language model's generation time and network transit. Some of that is recoverable by choosing your stages carefully; none of it is free.
The orchestration bill is softer but real. Every extra stage is another place a call can fail, another timeout to tune, another vendor to reconcile invoices with. That is the part a native end-to-end model removes entirely: there is one thing that can break, and it either answers or it does not.
Where the cascade pays for itself is the middle stage. The language model behind a voice agent is the component whose quality improves fastest and whose cost varies most, and being able to swap it without touching the voice layer is worth real money over a product's life. Roughly 190 models from eleven upstream providers sit behind a single OrcaRouter key at provider list price with zero markup, so a vendor price change reaches that layer the same day, and automatic failover keeps a backend timeout from ending a live call. Neither ElevenLabs' agents stack nor GPT-Live-1's Live Sessions endpoint is reachable that way — the voice layers belong to their vendors, and this is the layer underneath them.
Which one to pick
Choose GPT-Live-1 if the conversation mechanics are the product and you want one contract with one failure mode. Phone lines where callers talk over the agent, where a natural handoff matters more than a perfect voice, and where twelve good voices are enough. You accept a closed hosted model, no cloning, mid-table independent quality, and no free tier to prototype on.
Choose ElevenLabs if voice quality, cloning or breadth is the constraint. Narration, dubbing, multilingual products, anything where the voice is the brand, anything that needs speech recognition in the same contract. You accept a cascade to operate, prices roughly ten to twenty times GPT-Live-1's per unit of speech, and the fact that the interruption logic is yours to configure and yours to get wrong.
The 90.5% is the part worth carrying away. It says that a company that assembled three models and a turn-taking layer beat a company that trained one model to do all of it, on the metric closest to whether the call worked. Full duplex is a genuine architectural advance and it is not, on the current evidence, the thing that decides whether the caller gets what they wanted.
