Hero title card for 'Yelp Put GPT-Live-1 Behind a Million Restaurant Calls', subtitled 'The first large-scale deployment of full-duplex voice, with no numbers attached', carrying three cards reading 'Yelp Host: 1,000,000+ restaurant calls since October 2025', 'GPT-Live-1: $0.05 per voice minute, backend reasoning billed separately', 'Disclosed impact: directional only — no call-handling or transfer figures', above a footer strip reading 'Yelp Host and Hatch figures are Yelp-reported; GPT-Live-1 pricing per OpenAI documentation.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Yelp Put GPT-Live-1 Behind a Million Restaurant Calls — and Won't Say What It Bought

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Yelp says it has handled more than one million restaurant calls through Yelp Host since October 2025, and that as of September 10, 2026 those calls run on GPT-Live-1, Open​AI's full-duplex voice model. The same announcement puts GPT-Live-1 inside Hatch, Yelp's AI communications platform for service businesses, which the company describes as managing tens of millions of leads across voice, text, email and web. That is the largest production deployment of a full-duplex speech model anyone has announced — and it arrived wrapped in the least quantified press release Yelp could have written. Yelp reports that call handling improved and that call transfers fell. It does not say by how much, on how many calls, or what any of it cost.

Both facts matter. The deployment is real evidence that a model that listens while it speaks survives contact with actual telephone traffic, which is more than any benchmark can establish. The silence is real evidence that the vendor's own numbers were not flattering enough to print — or that Yelp's disclosure obligations to investors are narrower than its marketing appetite.

What Yelp actually shipped

The architecture is the part worth understanding, because it is not a Yelp voice model. GPT-Live-1 is the front-end voice layer: it is what the caller talks to, and it is what decides when to keep listening, when to take the turn, and when to yield. Underneath sits Yelp's own business data and what the company calls purpose-built intelligence — the restaurant's hours, its reservation availability, its menu, its specials, its seating areas, and the reservation system's current state.

That split is the whole product. GPT-Live-1 does not know whether a table for four exists at 7:30 on a Friday. It knows how to have the conversation that finds out. The model's job is to make the exchange feel like a phone call rather than a menu tree; Yelp's job is to make the answer correct.

Yelp's stated behavioural claims are specific in kind and vague in degree. Callers can interrupt, change topics mid-sentence, or speak over background noise, and the model adapts. The system detects the caller's tone — excitement, frustration, impatience — and responds accordingly. Layering Yelp's language enhancements on top of GPT-Live-1 lets it detect and answer callers in nearly any language, which addresses a real restaurant problem: a missed call because nobody on shift speaks the caller's language is a lost booking, not a bad experience.

The two operational claims are the ones a restaurant operator would actually pay for. Yelp says callers now speak in full, natural sentences rather than short voice commands, and that post-call transcription accuracy improved, giving operators cleaner records. The first is a leading indicator that people stopped treating the agent as a machine to be talked around. The second is the one with compounding value, because a reservation line's output is not only a booking — it is a record, and a record nobody can read is a record nobody can act on.

A screenshot of the OpenAI developer documentation page for GPT-Live 1, headed '< Models' with a 'Default' badge and the summary 'Our premier model for natural, expressive voice conversations with smooth interruption handling'. A price panel reads '$0.05 Per minute' against audio and text for both input and output, with performance and speed marked 'Not specified', and a knowledge cutoff of Jul 31, 2025. The body text reads 'GPT-Live 1 is a full-duplex voice model for real-time conversations. It can listen and speak at the same time, and delegate reasoning and tool use to a backend agent.' A pricing section states 'Voice sessions cost $0.05 per minute, billed per second. Backend model and tool usage is billed separately.'

Akhil Kuduvalli Ramesh said the useful thing

Yelp's chief product officer gave the standard quote — every call more conversational and responsive, exceptional guest experiences, staff freed to serve guests instead of answering the phone — and then one sentence that is worth more than the rest of the release. Yelp Host, he said, captures "revenue opportunities they may have otherwise missed."

That is the honest framing for voice agents in hospitality, and it is a different claim from the usual labour-substitution pitch. A restaurant does not buy an AI host to fire a host. It buys one because the phone rings during service, nobody can answer it, and the caller books somewhere else. The one million calls Yelp Host has handled since October 2025 are the denominator for that argument — and Yelp pointedly did not give the numerator, which would be how many of those calls became reservations or orders.

Teri Yu, Open​AI's multimodal product lead, framed the same deployment the other way, describing customers using GPT-Live-1 for everything from reservation calls to scheduling repairs. Both readings are true. Yelp is selling the vertical; Open​AI is selling the horizontal.

The economics nobody put in the release

GPT-Live-1 costs $0.05 per minute of voice, billed per second, and the model was opened to developers on September 10, 2026 — the same day Yelp's announcement landed. The voice layer is the only thing that rate covers. Open​AI's documentation is explicit that backend calls made on the model's behalf, including the reasoning and tool use it delegates to another model, are billed at that model's normal rates.

Run that against Yelp's number. A restaurant reservation call is short — a couple of minutes, sometimes less. One million calls at two minutes each is roughly two million voice minutes, or about $100,000 of voice-layer spend across the full history of the product. That is a rounding error for Yelp, and it is the point: at $0.05 a minute the voice layer is not the expensive part of the system. The expensive parts are the reasoning behind each turn, the telephony, the integration into each restaurant's reservation book, and the humans who handle what the agent cannot.

It also means the per-minute rate is the wrong number to negotiate over. A voice agent that answers in one turn because it delegated correctly costs less than one that flounders for ninety seconds, even though both are billed by the second. Yelp's claim that transfers fell is, underneath the vagueness, a cost claim.

The reasoning behind the voice is a normal model call, and that layer is where a deployment like this is actually tunable. Around 190 models from eleven upstream providers sit behind a single OrcaRouter key at provider list price with zero markup, with automatic failover when a backend errors or times out — so the model doing Yelp-style reservation reasoning can be swapped or A/B-tested against live call traffic by changing one configuration value rather than rebuilding the integration. That is where the money and the quality both live; the voice layer's metered minutes are comparatively fixed.

A generated infographic card titled 'Yelp Host and Hatch on GPT-Live-1 — what was announced'. Two columns. The left column, labelled 'What Yelp shipped', reads 'GPT-Live-1 as the front-end voice layer', 'Yelp business data and reservation availability underneath', 'Live in Yelp Host and Hatch', 'Tone detection: excitement, frustration, impatience', 'Near-universal language detection'. The right column, labelled 'What Yelp disclosed', reads 'More than 1,000,000 calls since October 2025', 'Improved call handling — no figure given', 'Fewer call transfers — no figure given', 'Callers speaking in full sentences', 'Better post-call transcription accuracy'. A footer reads 'Yelp-reported deployment claims, September 10, 2026; no independent verification.' The OrcaRouter logo is composited in the bottom-right corner.

The data is the moat, not the model

Any competitor can buy GPT-Live-1 tomorrow. Toast, Square, Clover, ServiceTitan and Housecall Pro all sit close enough to the same customers, and Goog​le and Amazon's Alexa+ are working the same problem from the consumer side. Nothing about Yelp's position depends on exclusive access to the model, and Yelp's own framing — proprietary business data plus purpose-built intelligence, with GPT-Live-1 as the layer that talks — concedes that.

What Yelp owns is the reservation graph: which restaurants have tables, when, for how many, and the accumulated consumer intent behind a million calls. A model that can be interrupted is a prerequisite for collecting that data at scale, because a turn-based agent produces shorter calls, more hangups, and dirtier transcripts. That is why the deployment matters even with the numbers withheld. Full duplex is not a nicer user experience in this context; it is the data-collection mechanism.

A screenshot of the Artificial Analysis Speech to Speech Index, marked Updated, ranking fourteen native speech-to-speech models on a weighted average of speech reasoning, agentic performance, arena preference and task success. Bars run left to right: Grok Voice Think Fast 2.0 High at 79.0%, GPT-Realtime-2.1 High at 73.9%, GPT-Realtime-2 High at 73.6%, Grok Voice Think Fast 1.0 at 72.3%, Gemini 3.1 Flash Live High at 71.5%, GPT-Live-1 Mini and GPT-Live-1 level at 70.3%, then GPT-Realtime-2.1 Minimal at 68.5%, Qwen-Audio-Omni-3.0-Realtime-Plus at 66.8%, Qwen-Audio-Omni-3.0-Realtime-Flash at 64.2%, Gemini 3.1 Flash Live Minimal at 63.9%, GPT-Realtime-2 Minimal at 62.7%, GPT-Realtime Mini at 56.8% and Gemini 2.5 Flash at 52.8%. A companion panel headed 'Cost per Hour of Input Audio' charts a similar field.

What to watch

Three things would turn this from a credible deployment story into a measurable one. Yelp's next earnings call, if management puts a number on call deflection or reservation conversion. A second vertical announcing the same integration, which would show whether full-duplex voice is a general-purpose upgrade or a hospitality-specific one. And the first independent evaluation of GPT-Live-1 on a board that scores reasoning rather than conversational mechanics — the Artificial Analysis Speech to Speech Index currently puts it at 70.3%, level with GPT-Live-1 mini and joint sixth on a fourteen-model board, behind Gr​ok Voice Think Fast 2.0 High at 79.0% and GPT-Realtime-2.1 High at 73.9%. Yelp's own architecture is an implicit bet that the voice layer and the reasoning layer should be judged separately. Independent scoring, so far, agrees with the second half of that bet and not the first.

Until Yelp publishes what changed, the rest of us are reading a deployment announcement for its architecture rather than its results. That is still worth doing, because the architecture is the part other teams can copy this quarter.