Title card for the GPT-Live-1 Speech to Speech Index article showing the three top scores: GPT-Live-1 with GPT-6 Astra at medium effort at 81.5, Grok Voice Think Fast 2.0 High at 81.3, and GPT-Live-1 with GPT-5.6 Sol at low effort at 80.1
Guides & Insights

GPT-Live-1 Tops the Speech to Speech Index at 81.5 — but the Ranking Belongs to Its Backend

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-Live-1, the full-duplex voice model that first shipped inside ChatGPT on July 8, 2026, is now first on the Artificial Analysis Speech to Speech Index at 81.5 — a row that exists only because the model was pointed at GPT-6 Astra to do its thinking. Run the same voice front end with GPT-5.6 Sol as the delegated backend at low reasoning effort and it scores 80.1, which is third. Second place, at 81.3, belongs to Grok Voice Think Fast 2.0 High.

Two honesties before the numbers. The model is not new: GPT-Live-1 reached the developer API on September 10, 2026, after two months as ChatGPT's default voice. What is new, measured on September 15, is that an outside evaluator has scored it — the one thing missing from every write-up of this model so far, including our own, where the only figures available were the ones OpenAI published about itself. And the 0.2-point margin over second place is smaller than the 1.4-point gap between GPT-Live-1's own two configurations, which is the most useful number on the board.

So the headline "GPT-Live-1 is the best voice model" is not quite what the index says. It says GPT-Live-1 with GPT-6 Astra at medium effort is, by a fifth of a point, and that the choice of backend moves the result more than any competing model does.

What the index actually measures

Artificial Analysis builds the Speech to Speech Index as an equal-weighted composite of four parts: Speech Reasoning, measured by Big Bench Audio; Agentic Performance, measured by τ-Voice; Arena Preference, a pairwise human-preference Elo; and Task Success Rate. Only models with all four components available get an index value — everything else shows a dash, which is why the table has far more rows than scores. The index itself launched on June 23, 2026 and has been revised twice since, most recently to swap Conversational Dynamics out for Task Success, so a score from one revision is not strictly comparable with a score from another.

Two further methodology details matter when you read the ranking. Arena Elo is frozen at whatever it was when a model first became publishable, so the preference component rewards arriving early rather than being best today. And the index scores a whole system rather than a single model: for GPT-Live-1, the number covers the voice front end plus whatever backend model it hands work to. That is a deliberate design choice, and it is the reason the same model appears twice with different scores.

The top of the field, as published:

• GPT-Live-1 (Astra, medium) — index 81.5, first

• Grok Voice Think Fast 2.0 High — index 81.3, second

• GPT-Live-1 (Sol, low) — index 80.1, third

• GPT-Realtime-2.1 High — index 73.9, fourth

• GPT-Realtime-2 High — index 73.6, fifth

• Grok Voice Think Fast 1.0 — index 72.3, sixth

• Gemini 3.1 Flash Live High — index 71.5, seventh

Artificial Analysis Speech to Speech leaderboard showing GPT-Live-1 with Astra at medium effort first at 81.5, Grok Voice Think Fast 2.0 High second at 81.3, and GPT-Live-1 with Sol at low effort third at 80.1, with the rest of the field from GPT-Realtime-2.1 High at 73.9 down to GPT-Realtime-2.1 Mini Minimal at 52.8

For scale, the model GPT-Live-1 replaces sits 7.6 points below it. That is a real generational gap, and it is not in dispute.

The 0.2-point win comes from exactly one component

Pull the composite apart and the two leaders stop looking like the same kind of model. On the components, per Artificial Analysis:

• Agentic performance (τ-Voice) — GPT-Live-1 67.9% vs Grok Voice Think Fast 2.0 High 56.5%

• Speech reasoning (Big Bench Audio) — 90% vs 97%

• Task success rate — 87.4% vs 94.6%

• Arena Elo — 1048 vs 1011

• Time to first audio — 1.34s vs 0.70s

• Cost per hour, backend included — $5.83 vs $4.80

Comparison scoreboard for GPT-Live-1 with Astra at medium effort against Grok Voice Think Fast 2.0 High, contrasting their Speech to Speech Index scores of 81.5 and 81.3, agentic performance of 67.9% and 56.5%, speech reasoning of 90% and 97%, task success of 87.4% and 94.6%, time to first audio of 1.34 and 0.70 seconds, and cost per hour of $5.83 and $4.80

GPT-Live-1 wins two of six lines and loses four, yet takes the index. The arithmetic is not a mystery: the τ-Voice gap is 11.4 points, far larger than any other separation in the table, and under equal weighting it drags the composite across the line. In plain terms, GPT-Live-1 is the better agent and Grok Voice Think Fast 2.0 High is the better listener and the faster one — and the index happens to weight the thing GPT-Live-1 is good at heavily enough to make it first.

If your product is a voice agent that books, resolves or transacts, the 11.4-point τ-Voice lead is the number that predicts your experience. If your product is a conversation that has to feel instant and never talk over the user, the 0.70-second time to first audio is the number that predicts your experience, and Grok wins it by roughly a factor of two. On regulated task execution there is a second warning: Grok's 94.6% task success rate is the highest in the visible table, and GPT-Live-1's 87.4% means roughly one task in eight does not complete. Both are improvements on the previous generation. Neither is a solved problem.

The #1 row is a routing configuration, not a model

Here is the part that deserves more attention than the leaderboard position. GPT-Live-1 is a full-duplex voice layer that handles listening, speaking, interruption and turn-taking itself, and delegates reasoning, tool calls and search to a separate backend model while the conversation continues. Artificial Analysis scored two of those pairings as separate entries. Astra at medium effort produced 81.5. Sol at low effort produced 80.1.

That 1.4-point spread is the story. The gap between first and second place is 0.2 points. The gap between GPT-Live-1's two scored configurations is seven times larger. Nothing about the voice model changed between those rows — only which model was doing the thinking, and at what effort level. A leaderboard that ranks systems is telling you, accurately, that the configuration is the product.

It also revises our own reporting. On September 11, 2026 we recorded GPT-Live-1 at 69.8% on this index, behind both GPT-Realtime-2.1 and Grok. That was the reading available at the time, and it supported a reasonable conclusion — that OpenAI had built the best conversational mechanics and not the best voice agent. The index now carries two delegated GPT-Live-1 configurations at 81.5 and 80.1. Artificial Analysis does not publish a changelog, and it revised the index's components in August, so we cannot say with certainty whether the earlier figure was an unscored or partially-scored configuration or a run under the older weighting; what is observable is that the delegated pairings now appear as their own complete rows, and that the answer changed when they did.

The practical consequence is that "which voice model" is the wrong question and "which voice model plus which backend, at which effort" is the right one. That is a routing question, and it is the one our own product exists to answer. OrcaRouter exposes GPT-Live-1's delegation backend, GPT-6 Astra, through the same OpenAI-compatible endpoint as 200+ other models, so swapping the thinking layer is a model-ID change rather than an integration. The routing DSL composes several models into a single call, and automatic failover means an unproven pairing is not a production bet — if the backend degrades, the request lands somewhere else instead of failing. To be precise about what we do not do: GPT-Live-1 itself is not in our catalogue and we do not serve it. The backend it delegates to is a different matter.

What an hour of this costs

GPT-Live-1's front end is billed at $0.05 per voice minute, charged by the second, which is $3.00 for an hour of open line. That is not the whole bill: delegated reasoning, tool calls and web search are metered separately at whatever the backend model charges. Artificial Analysis measured the all-in figure, which is why its cost column is the one to use:

• GPT-Live-1 (Sol, low) — $4.47 per hour

• Grok Voice Think Fast 2.0 High — $4.80 per hour

• GPT-Live-1 (Astra, medium) — $5.83 per hour

• GPT-Realtime-2.1 High — $10.75 per hour

The ranking winner is the most expensive of the three current options, and the configuration that won is 30% more per hour than the configuration that placed third. Read the other way, the split is instructive: $3.00 of every hour is the voice layer, and the remainder — $1.47 on Sol, $2.83 on Astra — is backend tokens. The thinking layer is somewhere between a third and a half of your bill, which is a large share for a component the leaderboard treats as a footnote.

OrcaRouter model page for GPT-6 Astra showing the model id openai/gpt-6-astra, a 1M-token context window, 128K max output, a p50 time to first token of 6.75 seconds, and pricing of $10.00 per million input tokens and $50.00 per million output tokens

Against the model it replaces, though, the whole family is roughly half price — the $0.05-per-minute front end is a genuine cut, not a repackaging. Because backend tokens bill at backend rates, the cost of a GPT-Live-1 deployment is mostly a function of which reasoning model you route to, and backends can be OpenAI's own or third-party models. On OrcaRouter those backends are passed through at provider list price with 0% markup, so a vendor price change on the thinking layer is live the same day rather than at the next billing cycle.

What the index does not settle

Three caveats, stated by the source rather than inferred. Both GPT-Live-1 entries and Grok Voice Think Fast 2.0 High are flagged as based on a single trial, which is thin for a 0.2-point decision — a re-run could reorder first and second without anything changing about either model. Arena Elo, as noted, was frozen at each model's first publishable measurement. And the index has no entry for how these systems behave on a long call: the components are short-horizon by construction, while the failure mode that actually kills voice agents in production is drift over a twenty-minute conversation.

Separately, and from the vendor rather than the index: GPT-Live-1's API shipped without image or video input, without structured outputs and without a free tier, and its rate limits are counted in concurrent sessions rather than requests per minute. Those are launch-day constraints that may already have moved; check them before you design around them.

What to do with this

If you are choosing an architecture this week rather than a model, the index rewards the delegated pattern on agentic work and the cascade pattern on latency. GPT-Live-1 at 81.5 is the strongest evidence yet that splitting conversation from reasoning is the right shape for task-completing voice agents. Grok Voice Think Fast 2.0 High at 81.3, with 0.70 seconds to first audio and a 94.6% task success rate, is the strongest evidence that the delegated shape is not the only one that works — and that it is not yet the fastest.

The decision that follows from the numbers is cheaper than it looks. Both configurations of GPT-Live-1, and the model that beat it on four of six components, sit within $1.36 an hour of each other. At that spread the right move is to stop reading leaderboards and run your own transcripts through both. The index tells you the shape of the trade-off; it cannot tell you which side of it your users are on.

Questions the number invites

Is GPT-Live-1 the best voice model now?

By the composite, yes — by 0.2 points and with a single trial behind it. By components it depends entirely on what you are building: it leads on agentic performance and arena preference, and trails on speech reasoning, task success rate and time to first audio. A 0.2-point composite lead resting on one 11.4-point component advantage is a defensible first place, not a decisive one.

Do you have to use GPT-6 Astra to get the 81.5?

For that exact row, yes — the 81.5 belongs to GPT-Live-1 configured with Astra at medium reasoning effort. Sol at low effort is a documented alternative at 80.1 and $1.36 per hour cheaper, and OpenAI supports bringing third-party models as the backend, so the leaderboard entries are two points on a curve rather than the only options. Which pairing is best for your workload is not something a public index can answer for you.

Why did the same model score 69.8 in our earlier coverage and 81.5 now?

The two figures cover different things. The earlier reading placed GPT-Live-1 on the index as a single entry; the current table carries its two delegated configurations as separate, fully-scored rows, with the Astra pairing at 81.5 and the Sol pairing at 80.1. Artificial Analysis revised the index's components in August and does not publish a changelog, so treat the delta as a change in what was measured rather than as evidence the model improved — and treat any single composite score for a delegating model as a statement about a configuration.