
Yelp Put GPT-Live-1 Behind a Million Restaurant Calls — and Won't Say What It Bought
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Yelp says it has handled more than one million restaurant calls through Yelp Host since October 2025, and that as of September 10, 2026 those calls run on GPT-Live-1, OpenAI's full-duplex voice model. The same announcement puts GPT-Live-1 inside Hatch, Yelp's AI communications platform for service businesses, which the company describes as managing tens of millions of leads across voice, text, email and web. That is the largest production deployment of a full-duplex speech model anyone has announced — and it arrived wrapped in the least quantified press release Yelp could have written. Yelp reports that call handling improved and that call transfers fell. It does not say by how much, on how many calls, or what any of it cost.
Both facts matter. The deployment is real evidence that a model that listens while it speaks survives contact with actual telephone traffic, which is more than any benchmark can establish. The silence is real evidence that the vendor's own numbers were not flattering enough to print — or that Yelp's disclosure obligations to investors are narrower than its marketing appetite.
What Yelp actually shipped
The architecture is the part worth understanding, because it is not a Yelp voice model. GPT-Live-1 is the front-end voice layer: it is what the caller talks to, and it is what decides when to keep listening, when to take the turn, and when to yield. Underneath sits Yelp's own business data and what the company calls purpose-built intelligence — the restaurant's hours, its reservation availability, its menu, its specials, its seating areas, and the reservation system's current state.
That split is the whole product. GPT-Live-1 does not know whether a table for four exists at 7:30 on a Friday. It knows how to have the conversation that finds out. The model's job is to make the exchange feel like a phone call rather than a menu tree; Yelp's job is to make the answer correct.
Yelp's stated behavioural claims are specific in kind and vague in degree. Callers can interrupt, change topics mid-sentence, or speak over background noise, and the model adapts. The system detects the caller's tone — excitement, frustration, impatience — and responds accordingly. Layering Yelp's language enhancements on top of GPT-Live-1 lets it detect and answer callers in nearly any language, which addresses a real restaurant problem: a missed call because nobody on shift speaks the caller's language is a lost booking, not a bad experience.
The two operational claims are the ones a restaurant operator would actually pay for. Yelp says callers now speak in full, natural sentences rather than short voice commands, and that post-call transcription accuracy improved, giving operators cleaner records. The first is a leading indicator that people stopped treating the agent as a machine to be talked around. The second is the one with compounding value, because a reservation line's output is not only a booking — it is a record, and a record nobody can read is a record nobody can act on.

Akhil Kuduvalli Ramesh said the useful thing
Yelp's chief product officer gave the standard quote — every call more conversational and responsive, exceptional guest experiences, staff freed to serve guests instead of answering the phone — and then one sentence that is worth more than the rest of the release. Yelp Host, he said, captures "revenue opportunities they may have otherwise missed."
That is the honest framing for voice agents in hospitality, and it is a different claim from the usual labour-substitution pitch. A restaurant does not buy an AI host to fire a host. It buys one because the phone rings during service, nobody can answer it, and the caller books somewhere else. The one million calls Yelp Host has handled since October 2025 are the denominator for that argument — and Yelp pointedly did not give the numerator, which would be how many of those calls became reservations or orders.
Teri Yu, OpenAI's multimodal product lead, framed the same deployment the other way, describing customers using GPT-Live-1 for everything from reservation calls to scheduling repairs. Both readings are true. Yelp is selling the vertical; OpenAI is selling the horizontal.
The economics nobody put in the release
GPT-Live-1 costs $0.05 per minute of voice, billed per second, and the model was opened to developers on September 10, 2026 — the same day Yelp's announcement landed. The voice layer is the only thing that rate covers. OpenAI's documentation is explicit that backend calls made on the model's behalf, including the reasoning and tool use it delegates to another model, are billed at that model's normal rates.
Run that against Yelp's number. A restaurant reservation call is short — a couple of minutes, sometimes less. One million calls at two minutes each is roughly two million voice minutes, or about $100,000 of voice-layer spend across the full history of the product. That is a rounding error for Yelp, and it is the point: at $0.05 a minute the voice layer is not the expensive part of the system. The expensive parts are the reasoning behind each turn, the telephony, the integration into each restaurant's reservation book, and the humans who handle what the agent cannot.
It also means the per-minute rate is the wrong number to negotiate over. A voice agent that answers in one turn because it delegated correctly costs less than one that flounders for ninety seconds, even though both are billed by the second. Yelp's claim that transfers fell is, underneath the vagueness, a cost claim.
The reasoning behind the voice is a normal model call, and that layer is where a deployment like this is actually tunable. Around 190 models from eleven upstream providers sit behind a single OrcaRouter key at provider list price with zero markup, with automatic failover when a backend errors or times out — so the model doing Yelp-style reservation reasoning can be swapped or A/B-tested against live call traffic by changing one configuration value rather than rebuilding the integration. That is where the money and the quality both live; the voice layer's metered minutes are comparatively fixed.

The data is the moat, not the model
Any competitor can buy GPT-Live-1 tomorrow. Toast, Square, Clover, ServiceTitan and Housecall Pro all sit close enough to the same customers, and Google and Amazon's Alexa+ are working the same problem from the consumer side. Nothing about Yelp's position depends on exclusive access to the model, and Yelp's own framing — proprietary business data plus purpose-built intelligence, with GPT-Live-1 as the layer that talks — concedes that.
What Yelp owns is the reservation graph: which restaurants have tables, when, for how many, and the accumulated consumer intent behind a million calls. A model that can be interrupted is a prerequisite for collecting that data at scale, because a turn-based agent produces shorter calls, more hangups, and dirtier transcripts. That is why the deployment matters even with the numbers withheld. Full duplex is not a nicer user experience in this context; it is the data-collection mechanism.

What to watch
Three things would turn this from a credible deployment story into a measurable one. Yelp's next earnings call, if management puts a number on call deflection or reservation conversion. A second vertical announcing the same integration, which would show whether full-duplex voice is a general-purpose upgrade or a hospitality-specific one. And the first independent evaluation of GPT-Live-1 on a board that scores reasoning rather than conversational mechanics — the Artificial Analysis Speech to Speech Index currently puts it at 70.3%, level with GPT-Live-1 mini and joint sixth on a fourteen-model board, behind Grok Voice Think Fast 2.0 High at 79.0% and GPT-Realtime-2.1 High at 73.9%. Yelp's own architecture is an implicit bet that the voice layer and the reasoning layer should be judged separately. Independent scoring, so far, agrees with the second half of that bet and not the first.
Until Yelp publishes what changed, the rest of us are reading a deployment announcement for its architecture rather than its results. That is still worth doing, because the architecture is the part other teams can copy this quarter.
