Hero title card for 'GPT-Live-1 vs ElevenLabs', subtitled 'A single end-to-end model against a cascaded stack, measured where they actually meet', carrying three cards reading 'GPT-Live-1: $0.05 per voice minute, full duplex, no cloning', 'ElevenLabs Agents: Elo 937 on the Speech Agent Arena, 90.5% task success', 'ElevenLabs Eleven v3: 70+ languages, 10,000+ library voices', above a footer strip reading 'Arena figures per Artificial Analysis; pricing per vendor documentation, September 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

GPT-Live-1 vs ElevenLabs: One Model That Listens, One Company That Sells Everything Else

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most useful number in this comparison is 90.5%. That is the task success rate Artificial Analysis measured for ElevenLabs Agents in its Speech Agent Arena — the highest in the top six of that board, achieved by a system that is not a single model at all but a cascade of speech recognition, a language model and a synthesiser stitched together by ElevenLabs' own turn-taking logic. GPT-Live-1, Open​AI's end-to-end full-duplex model, does not appear on that board, because it competes on a different one: the Speech to Speech Index, where it sits joint sixth of fourteen at 70.3%. Meanwhile ElevenLabs Eleven v3 remains the most widely deployed production voice in the industry, with more than 10,000 voices in its library and 70-plus languages.

The shorthand is that one side sells you a conversation and the other sells you a company. That shorthand is mostly right, and it hides the more interesting finding underneath: a well-engineered cascade still beats a native full-duplex model at finishing the task.

Two architectures, and the difference is where the seams are

GPT-Live-1 is a single model that takes audio and text in and produces audio and text out. It handles interruption natively — Open​AI's own testing, published with September's API release and unreproduced outside the company, reports 80.1% interactivity on Full Duplex Bench v1.5 against 45.4% for GPT-Realtime-2.1, and turn-taking latency of 0.798 seconds against 1.41. There are no seams, because there is nothing to seam together.

ElevenLabs Agents is the opposite bet, deliberately. ElevenLabs' documentation describes a pipeline: tuned speech recognition, a language model you choose or bring yourself, a low-latency synthesiser — typically something from the Flash line — and a proprietary turn-taking model on top. Barge-in is a per-agent configuration exposing turn eagerness and timeouts rather than a property of the model. In Artificial Analysis's benchmarked default configuration the stack reads Scribe v2 Realtime for speech recognition, a Gem​ini model for reasoning, and Eleven Flash v2 for synthesis.

That design has one enormous advantage and one structural cost. The advantage is that every stage is replaceable: a cheap model for simple turns, a frontier model for hard ones, a different synthesiser for a different locale. The cost is that latency accumulates at each seam, and the turn-taking decision has to be made from the outside.

Dimension by dimension

• Architecture — GPT-Live-1: one end-to-end model, audio in and audio out vs ElevenLabs Agents: cascaded speech recognition, language model and synthesis with a proprietary turn-taking layer

• Independent evidence — GPT-Live-1: 70.3% on the Artificial Analysis Speech to Speech Index, joint sixth of fourteen vs ElevenLabs Agents: Elo 937 on the Speech Agent Arena with a 90.5% task success rate across 676 samples, the best success rate in that board's top six

• Quality boards — GPT-Live-1: not ranked on the text-to-speech arenas, being a different kind of model vs ElevenLabs: Eleven v3 Conversational at 1,194 Elo and Eleven v3 at 1,166 on the Provider Voice board, priced at $50 and $100 per million characters respectively

• Price — GPT-Live-1: $0.05 per voice minute, billed per second, backend reasoning metered separately vs ElevenLabs: around $50 per million characters for Eleven v3 Conversational and about $100 for Eleven v3, with agents reported at roughly $0.08 per minute, rising past the concurrency ceiling, and language model and telephony costs billed separately

• Latency — GPT-Live-1: no model-level time-to-first-audio published; 0.798 seconds turn-taking latency in vendor testing vs ElevenLabs: about 75 milliseconds for Flash v2.5 and about 280 milliseconds for Eleven v3 Conversational, with vendor guidance of under 800 milliseconds to first audio at the agent level

• Voices and cloning — GPT-Live-1: twelve API presets, no cloning by design vs ElevenLabs: more than 10,000 voices in the library, instant cloning from a short sample, professional cloning from 30 to 180 minutes of audio with self-verification

• Languages — GPT-Live-1: major languages well served, uneven fluency beyond them in launch coverage vs ElevenLabs Eleven v3: 70-plus languages, against 32 for the Flash line and 90-plus for its speech recognition

• Breadth — GPT-Live-1: one endpoint, one model ID, concurrency tiers from 25 to 500 sessions and no free tier vs ElevenLabs: synthesis, speech recognition, dubbing, sound effects, music, a voice marketplace and an agents platform under one contract, with a free tier that carries no commercial licence

A two-column comparison scoreboard titled 'GPT-Live-1 vs ElevenLabs — the scoreboard'. The left column, labelled GPT-Live-1, reads 'Architecture: one end-to-end full-duplex model', 'Price: $0.05 per voice minute plus backend', 'Independent score: 70.3% on the AA Speech to Speech Index, joint 6th of 14', 'Interactivity: 80.1% (vendor, unreproduced)', 'Voices: 12 presets, no cloning', 'Best at: interruption handling'. The right column, labelled ElevenLabs, reads 'Architecture: cascaded STT + LLM + TTS', 'Price: ~$50 to $100 per 1M characters; agents ~$0.08 per minute', 'Independent score: Elo 937 on the AA Speech Agent Arena, 90.5% task success', 'Latency: ~75 ms Flash v2.5, ~280 ms v3 Conversational (vendor)', 'Voices: 10,000+ library voices, cloning with verification', 'Best at: platform breadth and voice quality'. The footer reads 'ElevenLabs agent pricing is third-party reported; arena figures per Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.A screenshot of the Artificial Analysis Speech Agent Arena leaderboard, ranking cascaded voice-agent systems by Elo and task success rate. The top rows read Gemini 3.1 Flash Live Minimal at 1,046 Elo over 768 samples with a 74.6% task success rate, Gemini 3.1 Flash Live High at 1,014 with 71.8%, GPT-Realtime-1.5 at 1,000 with 85.1%, GPT Realtime (Aug '25) at 944 with 89.4%, and ElevenLabs Agents, labelled 'Default Cascaded System', at Elo 937 over 676 samples with a task success rate of 90.5% — the highest in the top six. The board continues with Amazon Bedrock Nova 2.0 Sonic at 917 and Grok Voice Think Fast 2.0 High at 908 with a 94.7% success rate.

Where ElevenLabs is not just ahead, but uncontested

Three of the gaps in that list are not close, and any comparison that gives them a sentence is being unfair to the incumbent.

Cloning is the first. GPT-Live-1 ships twelve presets and, by design, no cloning at all — a deliberate safety posture rather than a missing feature. ElevenLabs offers instant cloning from a short reference sample and professional cloning trained on 30 to 180 minutes of audio, gated behind plan tiers and a self-verification step. If your product's voice is part of its identity, this is not a dimension on which GPT-Live-1 competes.

The voice library is the second. More than 10,000 voices, plus a licensed marketplace of iconic voices from estates and performers, is an asset no model release replicates. It is also the thing buyers most often underestimate: picking a voice is the last mile of a voice product, and having ten thousand to choose from compresses a branding exercise into an afternoon.

The third is breadth. Speech recognition, dubbing, sound effects, music, an agents platform — a team that needs four of those buys one contract instead of four. That is a strategy advantage, not a quality one, and it survives the other side winning a benchmark.

Both companies carry governance questions that belong in a procurement review rather than a footnote. ElevenLabs' professional cloning requires self-verification and its terms prohibit cloning another person's voice even with consent; the company also faces a set of class actions filed in May 2026 over training-data consent, in which the claims are unproven. Open​AI's position is simpler because it removed the feature that generates the risk.

A screenshot of ElevenLabs' official models documentation page, describing Eleven v3 under a 'Flagship models' heading as 'Our most emotionally rich, expressive speech synthesis model', with the supporting lines 'Dramatic delivery and performance', '70+ languages supported', '5,000 character limit' and 'Support for natural multi-speaker dialogue'. The surrounding navigation lists ElevenLabs' wider product set, including text to speech, speech to text, music, text to dialogue, dubbing, sound effects, voices and voice agents.

The cascade tax, and where it is worth paying

A cascade's costs show up as latency and as orchestration. Vendor guidance for ElevenLabs puts an agent's time to first audio under 800 milliseconds at the median and under 1.5 seconds at the 95th percentile — figures that describe the whole pipeline, not the synthesiser's 75 milliseconds. Stacked on top are the language model's generation time and network transit. Some of that is recoverable by choosing your stages carefully; none of it is free.

The orchestration bill is softer but real. Every extra stage is another place a call can fail, another timeout to tune, another vendor to reconcile invoices with. That is the part a native end-to-end model removes entirely: there is one thing that can break, and it either answers or it does not.

Where the cascade pays for itself is the middle stage. The language model behind a voice agent is the component whose quality improves fastest and whose cost varies most, and being able to swap it without touching the voice layer is worth real money over a product's life. Roughly 190 models from eleven upstream providers sit behind a single OrcaRouter key at provider list price with zero markup, so a vendor price change reaches that layer the same day, and automatic failover keeps a backend timeout from ending a live call. Neither ElevenLabs' agents stack nor GPT-Live-1's Live Sessions endpoint is reachable that way — the voice layers belong to their vendors, and this is the layer underneath them.

Which one to pick

Choose GPT-Live-1 if the conversation mechanics are the product and you want one contract with one failure mode. Phone lines where callers talk over the agent, where a natural handoff matters more than a perfect voice, and where twelve good voices are enough. You accept a closed hosted model, no cloning, mid-table independent quality, and no free tier to prototype on.

Choose ElevenLabs if voice quality, cloning or breadth is the constraint. Narration, dubbing, multilingual products, anything where the voice is the brand, anything that needs speech recognition in the same contract. You accept a cascade to operate, prices roughly ten to twenty times GPT-Live-1's per unit of speech, and the fact that the interruption logic is yours to configure and yours to get wrong.

The 90.5% is the part worth carrying away. It says that a company that assembled three models and a turn-taking layer beat a company that trained one model to do all of it, on the metric closest to whether the call worked. Full duplex is a genuine architectural advance and it is not, on the current evidence, the thing that decides whether the caller gets what they wanted.