
GPT-Live-1 vs Breeze TTS 2: The Open-Weights Voice Leader You Cannot Ship
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Breeze TTS 2 is the highest-rated open-weights text-to-speech model on Artificial Analysis's Provider Voices Speech Arena, holding 1,205 Elo and sitting seventh overall among more than a hundred rated systems — ahead of every commercially licensed engine that publishes weights. Its weights have been downloadable since August 25, 2026. It is also, under the licence those weights ship with, unusable in a product. GPT-Live-1 is the opposite trade in every dimension: closed, hosted, billed at $0.05 per minute of voice since it reached OpenAI's developer API on September 10, 2026, and the only one of the two that can hear a caller interrupt it.
Put those two sentences next to each other and the comparison resolves faster than most model matchups do. This is not a fight over which synthesises better speech. It is a fight over whether you are buying a voice or a conversation, and whether "open" means what you assumed it meant.
They are not the same kind of thing, and that is the point
Breeze TTS 2, from a roughly fifteen-person startup called BreezeBlue, is a 3-billion-parameter text-to-speech model. Text goes in, audio comes out. It does not hear anything. It has no opinion about when a turn should end, because it has no concept of a turn. Streamed over a WebSocket it will start producing audio before your text has finished generating, which is genuinely useful for keeping a voice agent's pacing tight — but the decision about when to speak was made by whatever wrote the text.
GPT-Live-1 is a full-duplex speech-to-speech model. Audio and text go in; audio and text come out; and the model decides many times a second whether to keep listening, pause, take the turn, or yield it. OpenAI's own Full Duplex Bench v1.5 numbers, published with the release and not reproduced outside the company, put its interactivity at 80.1% against 45.4% for GPT-Realtime-2.1, with turn-taking latency of 0.798 seconds against 1.41.
That asymmetry is not a knock on Breeze TTS 2. It is a warning about where each one belongs in a stack. If you are building a cascade — speech recognition feeding a language model feeding a synthesiser — Breeze TTS 2 is a candidate for the last third of it. If you are buying GPT-Live-1, you are buying the whole loop and handing over the turn-taking logic with it.
Dimension by dimension
• Interaction model — GPT-Live-1: full duplex, native barge-in, listens while speaking vs Breeze TTS 2: streaming text-to-speech, no audio input at all
• Weights — GPT-Live-1: closed, hosted only, one model ID and no dated snapshots vs Breeze TTS 2: 3B parameters on Hugging Face since August 25, 2026, PyTorch inference code Apache-2.0
• Licence — GPT-Live-1: commercial terms via OpenAI's API, nothing to review vs Breeze TTS 2: a research-and-non-commercial licence covering the weights, fine-tunes, LoRAs, merges, quantisations and self-hosted outputs
• Price unit — GPT-Live-1: $0.05 per voice minute, billed per second, reasoning backend metered separately vs Breeze TTS 2: $34 per million characters on the hosted endpoint, tiering down to $28 at the top published plan
• Quality evidence — GPT-Live-1: 70.3% on the Artificial Analysis Speech to Speech Index, joint sixth of fourteen scored models vs Breeze TTS 2: 1,205 Elo on the Provider Voices arena, seventh overall and first among open-weights models, per Artificial Analysis
• Languages — GPT-Live-1: launch coverage consistently flags uneven fluency outside the major languages vs Breeze TTS 2: 50 claimed by the vendor, with the model card tagged for English and Chinese
• Voice control — GPT-Live-1: twelve API voices, tone and pace steered by prompt, no cloning at all vs Breeze TTS 2: reference-audio cloning, reference-free voice design, voice direction, and inline vocal events
• Throughput — GPT-Live-1: metered by concurrent sessions, capped at 25 on tier 1 and 500 on tier 5, with no free tier vs Breeze TTS 2: roughly 45 characters per second on Artificial Analysis's measurement, which is slower than several commercial rivals and matters for long-form work

The licence is doing more work than the leaderboard
BreezeBlue's own model card is unusually blunt about this, which is to the company's credit. The inference code is Apache-2.0. The weights and any outputs you generate by self-hosting them are for research use, and commercial deployment requires either a paid hosted subscription or written authorisation. That restriction explicitly extends to derivative models, LoRAs, merges and quantisations — which closes the loophole people normally reach for.
So the "open-weights leader" framing needs a qualifier that headlines rarely carry. Breeze TTS 2 is open in the sense that you can download it, inspect it, benchmark it, and run it on your own hardware for evaluation. It is not open in the sense that lets a startup build a product on it without either paying BreezeBlue or buying legal review. Two of the three things people mean by open weights — verifiability and cost — you get. The third — freedom to ship — you do not.
That is a real difference from a model like GPT-Live-1, where the licence question is not "may I" but "how much." Both end in a bill. Only one of them ends in a lawyer's opinion first.

The two price units do not compare, so here is the arithmetic
$0.05 a minute and $34 per million characters look like they live in different universes because they do. To compare them you have to convert. Speech runs at roughly 150 words a minute, and English averages around five and a half characters per word once spaces are counted, so a minute of spoken output is roughly 800 to 900 characters. At Breeze TTS 2's $34 per million, that is about three cents of synthesis per minute — cheaper than GPT-Live-1's five, on the raw speech.
The comparison falls apart immediately afterwards, because GPT-Live-1's five cents includes the listening. It is also transcribing the caller, deciding when the turn ends, and managing the interruption — work that a cascade has to buy from somewhere else: a speech recognition model, a turn-detection model, and a language model to produce the text Breeze TTS 2 will read. Add those and the cascade's three cents of synthesis sits under a bill that is usually larger than the end-to-end model's.
The honest summary is that the cheaper voice is Breeze TTS 2 and the cheaper conversation is GPT-Live-1, and the difference between those two claims is everything a voice agent actually costs.

Where they would actually meet
A team building a phone agent today and unwilling to commit to one vendor's end-to-end stack would assemble exactly this cascade: Breeze TTS 2 or a comparable engine for the mouth, something else for the ears, and a routed language model for the thinking. The last of those three is the piece that changes most often — it is where quality improves fastest, where cost varies most, and where the model that wins this quarter is rarely the one that wins next.
That layer is a normal API call, and it is the one part of this stack worth keeping portable. Roughly 190 models from eleven upstream providers sit behind a single OrcaRouter key at provider list price with zero markup, so swapping the reasoning model behind a voice agent is a configuration change rather than an integration project, and automatic failover keeps a backend that times out from dropping the call. Neither GPT-Live-1's Live Sessions endpoint nor Breeze TTS 2's hosted API is reachable that way — the voice layers in this comparison sit outside a routing layer's reach, and it is worth being clear about that rather than implying otherwise.
Which one to pick
Choose GPT-Live-1 if the conversation is the product and you can accept a hosted, closed front end. Reservation lines, first-line support, anything where a caller talking over the agent is a normal event rather than an exception. You are buying the interruption handling, and you are buying it at a per-minute rate that stays predictable while the reasoning behind it is the part you tune.
Choose Breeze TTS 2 if you need a voice you can inspect, if you are building a product where the synthesis is one component among many, or if you need cloning and voice design that GPT-Live-1 does not offer at all. Go in knowing that the weights are for research, that the 50-language claim is a vendor claim against a card tagged for two languages, and that the latency figures are BreezeBlue's own measurements on a warmed fast path rather than an independent audit.
The instinct to treat this as a quality contest is the wrong instinct. Breeze TTS 2 is close to the top of the board it competes on and GPT-Live-1 competes on a different board entirely. The decision that actually costs money is the one underneath both of them: which model does the thinking, and whether you can change your mind about that in an afternoon.
