Hero-Titelkarte für Gemini 3.8 Live mit dem Kicker „OrcaRouter · Modell-Radar — Launch“ und dem Untertitel „Google teilt seine Sprachlinie in zwei auf — ein Modell für Skalierung, eines für Reasoning.“ Auf zwei Karten steht „Gemini 3.8 Live — Index 76,0, gebaut für Skalierung und Kosteneffizienz“ und „3.8 Live Extended Thinking — Index 82,6, mehrstufiges Reasoning“, mit Chips für 15. September 2026, 97 Sprachen im laufenden Gespräch, SynthID-Audio-Wasserzeichen und Live-API-Vorschau. Das OrcaRouter-Logo ist in der unteren rechten Ecke eingefügt.
Guides & Insights

Einführung von Gemini 3.8 Live und 3.8 Live Extended Thinking: Google teilt seine Sprachlinie in zwei

Autor

Elias Hawthorne

Veröffentlicht am

Neueste Modelle · 20Alle Modelle ansehen
Benchmarks: Artificial Analysis · täglich aktualisiert
Zurück zu allen Beiträgen

Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived on September 15, 2026 — one announcement, two models, and a genuine fork in the product line. The headline number is a 6.6-point gap on Artificial Analysis's Speech to Speech Index, 82.6 to 76.0. The number underneath it is 38. On the same board's τ-Voice agentic task completion measure, the reasoning variant resolves 68.6% of replica customer-service scenarios and the standard model resolves 30.1%. The 6.6 points are what the index says. The 38 points are what happens when you ask either model to actually get something done.

A day on from launch, the fork has a price attached. Both models carry published per-minute audio rates through the Live API, and Artificial Analysis's leaderboard puts cost-per-hour of input audio at $0.84 for the standard model against $3.50 for the reasoning variant — a little over four times as much per hour for those 6.6 index points. That trade is the whole decision, and reading it off the index alone gets it wrong.

Was dieses Release von der üblichen Auffrischung von Sprachmodellen unterscheidet, ist, dass die Aufteilung keine Größenklasse ist. Es ist eine Meinungsverschiedenheit darüber, was die Aufgabe eines Sprachagenten ist. Eines dieser Modelle ist eine konversationelle Schnittstelle. Das andere ist ein Reasoningsystem, das zufällig spricht.

Zwei Modelle, zwei Aufgaben.

Googles eigene Darstellung ist diesbezüglich ungewöhnlich klar. Gemini 3.8 Live wird beschrieben als „für Skalierung und Kosteneffizienz entwickelt, wobei konversationelle Intelligenz mit flüssigem Dialog und visueller Verankerung kombiniert wird." Gemini 3.8 Live Extended Thinking ist „für hochkomplexe Aufgaben entwickelt, mit erhöhter Intelligenz und mehrstufigem Reasoning."

Der praktische Unterschied zeigt sich eher in der Funktionsliste als im Datenblatt:

Gemini 3.8 Live — nahezu Echtzeit-Verarbeitung visueller Eingaben, automatische Umschaltung mitten im Gespräch zwischen 97 unterstützten Sprachen und Ausführung von Tools/APIs im Hintergrund, sodass das Modell eine Anfrage bestätigt und weiterspricht, während der Aufruf abgeschlossen wird.

Gemini 3.8 Live Extended Thinking — denkt und spricht gleichzeitig und beschreibt den Fortschritt mit Hinweisen wie „Lass mich das kurz prüfen…", während eine mehrstufige Aufgabe läuft.

Gemeinsam — alles generierte Audio trägt Googles SynthID-Wasserzeichen, beide sind durch eine Modellkarte abgedeckt, die unter dem Namen gemini-3-8-audio veröffentlicht wurde, und beide haben jetzt dokumentierte API-Modell-IDs: gemini-3.8-live und gemini-3.8-live-extended-thinking.

Das Extended-Thinking-Modell ist nicht einfach 3.8 Live mit einem längeren Denkbudget. Es ist dasjenige, das gleichzeitig ein Gespräch und einen Aufgabenplan aufrechterhalten kann, ohne dass das Gespräch ins Stocken gerät – und genau das ist der Punkt, an dem Voice-Agents im Produktivbetrieb scheitern. Ein Nutzer will keine Stille, während der Agent etwas nachschlägt. Er will, dass der Agent weiterredet.

A two-column scoreboard comparing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Index score 76.0 vs 82.6; built for scale and cost efficiency vs high-complexity reasoning; languages 97 mid-conversation on both; visual input near real-time on both; background tools yes on both; audio price $0.005 in / $0.018 out per minute vs not separately announced. Footer reads 'Index figures per Artificial Analysis, Sep 16 2026; pricing as announced.' The OrcaRouter logo is composited in the bottom-right corner.

The board says 6.6. The components say 38

The Speech to Speech Index is a weighted average of four underlying results: Speech Reasoning, measured by Artificial Analysis's own Big Bench Audio set; Agentic Performance, measured by running the τ-Voice customer-service benchmark; Arena Preference, taken from the Speech Agent Arena; and Task Success Rate. Only models with all four components receive an index score, which is why some rows on that board are blank.

Pulled apart, the two new models do not look like a 6.6-point pair at all:

Speech to Speech Index — Gemini 3.8 Live Extended Thinking (High) 82.6 vs Gemini 3.8 Live 76.0

Agentic performance (τ-Voice task completion) — 68.6% vs 30.1%

Speech reasoning (Big Bench Audio) — 98% vs 92%

Conversational dynamics — 91.9% vs 96.1%

Arena preference (Elo) — 990 vs 1083

Task success rate — 89.1% vs 93.2%

Time to first audio — 1.35s vs 1.18s

Cost per hour of input audio — $3.50 vs $0.84

Read it honestly and the standard model wins four of eight lines. It is faster to first audio, rated higher by the arena, more likely to complete a task, and rated better on conversational dynamics. It is also a third of the price. What it cannot do is finish the job when the job has steps: 30.1% against 68.6% is the largest single gap anywhere in this release, and it sits on the one axis that separates a voice agent from a voice interface.

Two caveats before that table gets treated as settled. The arena preference figure used in the index is frozen at the point a model becomes eligible for publication, so it does not track the live Elo on the Speech Agent Arena chart — read it as a snapshot, not a running score. And the τ-Voice margins at the top of the board are thin enough to be noise: the reasoning variant's 68.6% leads GPT-Live-1 at 67.9% (Astra backend, medium effort) by 0.7 points, with GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% behind it. The 38-point gap between the two Gemini models is not thin. The 0.7-point lead over GPT-Live-1 is.

Where it sits on the index

In the capture below, taken on September 16, 2026, the top of the board reads like this:

Gemini 3.8 Live Extended Thinking (High) — 82,6 (Erster Platz)

• GPT-Live-1 (Astra-Backend, mittlerer Aufwand) — 81.5

• Grok Voice Think Fast 2.0 High — 81.3

• GPT-Live-1 (Sol-Backend, geringer Aufwand) — 80.1

Gemini 3.8 Live — 76.0

• GPT-Realtime-2.1 High — 73.9

• Gemini 3.1 Flash Live High — 71.5

That is the case for the split in one screen. The Extended Thinking variant takes the top spot, running at the board's (High) reasoning-effort label — the same convention it uses for Grok Voice Think Fast 2.0 High and GPT-Realtime-2.1 High, and the reason the "(High)" suffix you may see attached to this model in coverage is a configuration label rather than a separate release. The standard variant lands fifth — below two GPT-Live-1 configurations and below Grok. Google's own blog describes the standard model as "highly cost-effective" and notes it "secured a second place in the Speech Agent Arena," which is a different board measuring a different thing. Both statements are true. Read together they say: 3.8 Live is a very good conversational model at a very good price, and it is not the frontier of voice intelligence.

Screenshot of the Artificial Analysis Speech to Speech leaderboard page, captured September 16, 2026. The AA-Speech to Speech Index bar chart shows Gemini 3.8 Live Extended Thinking at 82.6 in first place, GPT-Live-1 (Astra) at 81.5, Grok Voice Think Fast 2.0 at 81.3, GPT-Live-1 (Sol) at 80.1 and Gemini 3.8 Live at 76.0, alongside speed and cost-per-hour-of-input-audio panels.

Das Leck, das zuerst da war

Die Modelle wurden am Tag vor der Ankündigung auf einer Google-Cloud-Kontingentseite entdeckt, und wir berichteten damals über diese Sichtung als unbestätigten Slug ohne Modellkarte, ohne Preisangaben und ohne Bestätigung von Google. Dieser Beitrag lag beim Status richtig und beim Zeitplan falsch – die Bestätigung kam innerhalb von ungefähr 24 Stunden.

Die Lektion ist es wert, festgehalten zu werden, denn sie läuft dem üblichen Instinkt zuwider. Ein geleakter Slug einen Tag vor dem Launch ist ein Launch. Ein geleakter Slug ohne bestätigende Unterlagen über acht Wochen hinweg ist etwas ganz anderes. Das Signal, das sie voneinander unterschied, war nicht der Slug; es war das Vorhandensein einer Quota-Seite, die nur existiert, wenn Kapazität bereitgestellt wurde.

Screenshot of Google's official blog post titled 'Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking', dated Sep 15, 2026 and credited to Tom Ouyang and Malini Jaganathan of the Gemini Audio Team, with the summary describing the models as Google's most advanced live dialogue models with major upgrades in intelligence and parallel reasoning.

Was es pro Minute kostet

Zum Zeitpunkt der Ankündigung trennten Googles Materialien die beiden Modelle preislich nicht — für das Standardmodell wurde ein Preis angegeben, für die Reasoning-Variante nicht, was die obige Übersicht widerspiegelt. Diese Lücke hat sich inzwischen geschlossen. Beide Modelle haben inzwischen veröffentlichte Preise über die Live API: 0,005 $ pro Minute Audioeingabe und 0,018 $ pro Minute Audioausgabe, was Google am 16. September bestätigte, als die beiden eingeführt wurden. Dieselbe veröffentlichte Preisliste deckt auch die Token-Preise der Gemini 3 Live-Stufe ab — 0,75 $ pro 1 Mio. Text-Token bei der Eingabe und 4,50 $ bei der Ausgabe, 3,00 $ pro 1 Mio. Audio-Token bei der Eingabe und 12,00 $ bei der Ausgabe sowie 1,00 $ pro 1 Mio. für Bild- oder Videoeingabe — wobei diese Token-Zeilen für die Stufe und nicht pro Modell veröffentlicht werden, weshalb sie als Preisliste der Stufe zu lesen sind und nicht als Angebot für eine bestimmte Variante.

Artificial Analysis has also filled in the number that matters most for planning. Its leaderboard now carries a cost-per-hour-of-input-audio column — the cost to complete a fixed 40-question Big Bench Audio subset, normalised to an hourly rate — and both new models have a figure there:

Gemini 3.8 Live — $0.84/hour of input audio, the lowest paid rate on that board

Gemini 3.8 Live Extended Thinking (High) — 3,50 $/Stunde, und liegt damit weiterhin unter beiden Modellen, die es bei der Qualität übertrifft

• Grok Voice Think Fast 2.0 High — 4,80 $/Stunde

• GPT-Live-1 (Astra-Backend, mittlerer Aufwand) — 5,83 $/Stunde

• GPT-Realtime-2 (High) — 4,14 $/Stunde

• GPT-Realtime-2.1 High — 10,75 $/Stunde

One note on the captures above, since the costs panel in the leaderboard screenshot predates these two rows: neither $0.84 nor $3.50 appears in it, and its cheapest bar is $1.42. The figures in the list are read from the board's summary table as it stands on September 16, 2026, not from that image. The index values shown in the image are unaffected and match the table.

Read the two new numbers against the components above and the release stops being a two-model announcement and becomes a single decision with a published exchange rate. The reasoning variant buys 6.6 index points, six points of speech reasoning and 38.5 points of agentic task completion for a little over four times the hourly audio cost. Whether that is worth paying depends entirely on what your agent is doing — and the shape of the answer changed once the components were published. A voice agent whose job is to finish a multi-step task is buying the single largest improvement in this release, at $3.50 an hour while undercutting GPT-Live-1 Astra by about 40% and Grok Voice Think Fast 2.0 High by roughly a quarter, and scoring above both. A voice agent whose job is to converse — answer, hold a thread, take a message, be pleasant — is buying very little with extra reasoning depth, and at $0.84 an hour it is buying the cheapest competent voice model anyone currently publishes.

So sieht die Liste aus. Google hat sein Reasoning-Modell unter den zwei Modellen bepreist, die es schlägt, und sein Standardmodell unter allem, was sonst auf dem Tableau steht. Wenn Sie Sprachminuten in großem Volumen abwickeln, ist es diese Zahl, die Ihre Rechnung bestimmt, nicht der Index-Score — und die Tatsache, dass die beiden Varianten um den Faktor 4 auseinanderliegen, bedeutet, dass die Wahl, welche man aufruft, der mit Abstand größte Kosthebel im gesamten Stack ist.

„Private preview“ leistet in diesem Satz echte Arbeit.

Keines der beiden Modelle ist im vertraglichen Sinne allgemein verfügbar, aber die Entwickleroberfläche ist offen und jetzt bepreist, was ändert, worauf Sie planen können:

Gemini 3.8 Live — Entwickler erhalten es in der Gemini API, der Live API und Google AI Studio unter der dokumentierten ID gemini-3.8-live; Unternehmen erhalten eine private Vorschau in Gemini Enterprise, Gemini Enterprise for Customer Experience ist „demnächst verfügbar“; für alle ist es in Search Live verfügbar.

Gemini 3.8 Live Extended Thinking — dieselbe Entwickleroberfläche unter gemini-3.8-live-extended-thinking, plus einen breiteren Verbraucherpfad: Gemini Live, Docs Live für Google AI Pro- und Ultra-Abonnenten sowie Gmail Live und Keep Live für alle Google AI-Abonnenten.

Was noch fehlt, ist der Teil, den ein Unternehmen braucht: eine Zusage zur allgemeinen Verfügbarkeit, eine veröffentlichte Uptime-Zahl und einen Preis, den Google zugesagt hat, zu halten. Wenn Sie das brauchen, warten Sie. Wenn Sie prototypen, ist der Weg heute offen, mit einem echten Preisblatt dahinter, was eine deutlich bessere Position ist als eine Vorschau ohne veröffentlichten Preis. Die Integration jetzt zu bauen, ist vernünftig; eine umsatzkritische Telefonleitung auf ein Verfügbarkeitsziel zu stützen, zu dem Google sich nicht verpflichtet hat, ist es nicht, und die Lösung dafür wird unten beschrieben.

Wo die Angaben des Anbieters enden

Google published per-component figures of its own alongside the launch — 68.6% on τ-Voice, 35.1% on Sierra's τ³-Banking leaderboard, and 97.7% on Big Bench Audio, all credited to the Extended Thinking model. Those now divide into two groups, and the split matters.

Two of the three line up with measurements Artificial Analysis runs itself under its own harness. Its τ-Voice agentic component reads 68.6% for this model, against 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice Think Fast 2.0; its Big Bench Audio speech-reasoning component reads 98%, against Google's 97.7%. Artificial Analysis describes τ-Voice and Big Bench Audio as its own benchmarks, run across three trials where available. So the τ-Voice and reasoning numbers now sit on a public board you can go and read, not only in Google's announcement — treat them as corroborated rather than settled, since the exact figures Google quoted may well be that same run rather than a second one.

The Sierra number has no such backing and remains a vendor claim: 35.1% on Sierra's τ³-Banking leaderboard against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. It is also the number worth sitting with, because 35.1% means the leading voice model on this board fails roughly two of every three realistic banking task-completion attempts. Google also cites a ServiceNow EVA-Bench run performed on the Live API in Gemini Enterprise Agent Platform, and describes both models as pushing the Pareto frontier for complex workflows — neither of which has an independent reading attached.

There is a precedent worth holding onto here. xAI reported a Speech to Speech Index figure of 82.9 for Grok Voice Think Fast 2.0 in July. On the board captured above, that model sits at 81.3. Vendor-reported index scores do not always survive contact with a live leaderboard — usually because the index is revised, sometimes because the configuration tested was not the one that shipped. Verify the τ-Voice and Big Bench numbers on your own traffic, and read the agentic column as the one to test hardest.

Die Ebene, auf der diese Agenten tatsächlich laufen

Beide Modelle sitzen vor einem Backend. Hintergrund-Tool-Ausführung, mehrstufige Aufgabenplanung und das Delegationsmuster, das jeder ernsthafte Voice-Agent verwendet, laufen alle auf gewöhnliche Textmodell-Aufrufe hinaus – und das ist die Ebene mit der größten Kostenvarianz und dem geringsten Lock-in.

OrcaRouter bietet 190 Modelle über einen einzigen Schlüssel zum Listenpreis des Anbieters ohne Aufschlag, was bedeutet, dass eine Preissenkung eines vorgelagerten Labors bei uns am selben Tag wirksam wird und nicht erst bei der nächsten Vertragsverlängerung. Für einen Voice-Stack, bei dem das Frontend ein Vorschaumodell ist, sind die zwei nützlichen Eigenschaften, dass das Delegationsziel ausgetauscht werden kann, ohne die Voice-Integration anzufassen, und dass ein automatisches Failover einen Anruf am Leben hält, wenn ein Backend einen Fehler wirft oder ein Timeout auftritt. Um präzise zu sein, was wir hosten und was nicht: Die Gemini 3.8 Live-Endpunkte sind nicht auf unserem Router – diese stammen von Googles eigener API –, aber die Textmodelle, an die diese Agenten Arbeit abgeben, sind es sehr oft, darunter Gemini 3.8 Flash.

Was als Nächstes ansehen

Three things will settle the open questions in this release. The first is general availability and a rate Google has committed to hold — the prices are published now, but preview pricing has a habit of moving, and a rate card is not a contract. The second is whether the arena and task-success columns hold up, because the standard model's 1083 Elo and 93.2% task success are the strongest argument against paying 4× for the reasoning variant, and the index's arena figure is frozen at eligibility rather than live. The third is independent runs on the agentic gap itself: 68.6% against 30.1% is the largest claim in the release and the one most worth reproducing, because everything else about the two models is close.

Until then, the honest summary runs on two numbers rather than one. Gemini 3.8 Live Extended Thinking is the best-scoring voice model on the public board, leads the agentic component outright, and undercuts the two models directly behind it on hourly cost. Gemini 3.8 Live is the cheapest competent voice model on that board, is preferred by the arena and more reliable on shallow tasks, and sits 6.6 points back overall with a hard ceiling at the one thing agents get hired for. Google has shipped the same fork twice, four times apart on price, and made the choice unusually easy to price — provided you read the components and not just the index.

In diesem Artikel verglichen1

Aus diesem Artikel erkannt · Benchmarks: Artificial Analysis · täglich aktualisiert