Gemini 3.8 Live의 히어로 타이틀 카드. 킥커는 'OrcaRouter · 모델 레이더 — 출시', 부제는 'Google이 음성 라인을 둘로 나눈다 — 하나는 규모용, 하나는 추론용.' 두 카드에는 'Gemini 3.8 Live — Index 76.0, 규모와 비용 효율성을 위해 설계됨'과 '3.8 Live Extended Thinking — Index 82.6, 다단계 추론'이라고 적혀 있으며, 칩으로는 2026년 9월 15일, 대화 중 97개 언어, SynthID 오디오 워터마크, Live API 프리뷰가 있다. OrcaRouter 로고는 오른쪽 아래 모서리에 합성되어 있다.
Guides & Insights

Gemini 3.8 Live 및 3.8 Live Extended Thinking 소개: Google, 음성 라인을 둘로 나누다

작성자

Elias Hawthorne

게시일

최신 모델 · 20모든 모델 보기
벤치마크: Artificial Analysis · 매일 업데이트
모든 게시물로 돌아가기

Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived on September 15, 2026 — one announcement, two models, and a genuine fork in the product line. The headline number is a 6.6-point gap on Artificial Analysis's Speech to Speech Index, 82.6 to 76.0. The number underneath it is 38. On the same board's τ-Voice agentic task completion measure, the reasoning variant resolves 68.6% of replica customer-service scenarios and the standard model resolves 30.1%. The 6.6 points are what the index says. The 38 points are what happens when you ask either model to actually get something done.

A day on from launch, the fork has a price attached. Both models carry published per-minute audio rates through the Live API, and Artificial Analysis's leaderboard puts cost-per-hour of input audio at $0.84 for the standard model against $3.50 for the reasoning variant — a little over four times as much per hour for those 6.6 index points. That trade is the whole decision, and reading it off the index alone gets it wrong.

이번 릴리스가 일반적인 음성 모델 업데이트와 다른 점은, 그 구분이 크기 등급이 아니라는 것입니다. 그것은 음성 에이전트의 역할이 무엇인지에 대한 관점 차이입니다. 이 모델들 중 하나는 대화형 인터페이스입니다. 다른 하나는 말을 하기도 하는 추론 시스템입니다.

두 모델, 두 작업

이 점에 관한 구글 자신의 설명은 이례적으로 명확하다. Gemini 3.8 Live는 "규모와 비용 효율성을 위해 설계되었으며, 대화형 지능과 유연한 대화, 시각적 그라운딩을 결합한" 제품이라고 설명된다. Gemini 3.8 Live Extended Thinking은 "향상된 지능과 다단계 추론을 갖춘, 높은 복잡도의 작업을 위해 설계되었다"고 설명된다.

실질적인 차이는 사양서가 아니라 기능 목록에서 드러납니다:

Gemini 3.8 Live — 거의 실시간에 가까운 시각 입력 처리, 지원되는 97개 언어 전반에 걸친 대화 중 자동 전환, 그리고 백그라운드 도구/API 실행으로, 모델이 요청을 접수한 뒤 호출이 완료될 때까지 계속 대화를 이어갑니다.

Gemini 3.8 Live Extended Thinking — 추론과 발화를 동시에 하며, 다단계 작업이 실행되는 동안 “잠시 확인해 볼게요…”와 같은 신호로 진행 상황을 내레이션합니다.

공통 — 생성된 모든 오디오에는 Google의 SynthID 워터마크가 포함되어 있으며, 둘 다 다음 이름으로 공개된 모델 카드의 적용을 받습니다: gemini-3-8-audio. 또한 둘 다 이제 문서화된 API 모델 ID를 갖추고 있습니다: gemini-3.8-livegemini-3.8-live-extended-thinking.

Extended Thinking 모델은 단순히 더 긴 사고 예산을 가진 3.8 Live가 아니다. 그것은 대화가 멈추지 않으면서 대화와 작업 계획을 동시에 유지할 수 있는 모델이다 — 이는 음성 에이전트가 프로덕션에서 실패하게 만드는 바로 그 요인이다. 사용자는 에이전트가 무언가를 찾아보는 동안 침묵을 원하지 않는다. 에이전트가 계속 말하기를 원한다.

A two-column scoreboard comparing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Index score 76.0 vs 82.6; built for scale and cost efficiency vs high-complexity reasoning; languages 97 mid-conversation on both; visual input near real-time on both; background tools yes on both; audio price $0.005 in / $0.018 out per minute vs not separately announced. Footer reads 'Index figures per Artificial Analysis, Sep 16 2026; pricing as announced.' The OrcaRouter logo is composited in the bottom-right corner.

The board says 6.6. The components say 38

The Speech to Speech Index is a weighted average of four underlying results: Speech Reasoning, measured by Artificial Analysis's own Big Bench Audio set; Agentic Performance, measured by running the τ-Voice customer-service benchmark; Arena Preference, taken from the Speech Agent Arena; and Task Success Rate. Only models with all four components receive an index score, which is why some rows on that board are blank.

Pulled apart, the two new models do not look like a 6.6-point pair at all:

Speech to Speech Index — Gemini 3.8 Live Extended Thinking (High) 82.6 vs Gemini 3.8 Live 76.0

Agentic performance (τ-Voice task completion) — 68.6% vs 30.1%

Speech reasoning (Big Bench Audio) — 98% vs 92%

Conversational dynamics — 91.9% vs 96.1%

Arena preference (Elo) — 990 vs 1083

Task success rate — 89.1% vs 93.2%

Time to first audio — 1.35s vs 1.18s

Cost per hour of input audio — $3.50 vs $0.84

Read it honestly and the standard model wins four of eight lines. It is faster to first audio, rated higher by the arena, more likely to complete a task, and rated better on conversational dynamics. It is also a third of the price. What it cannot do is finish the job when the job has steps: 30.1% against 68.6% is the largest single gap anywhere in this release, and it sits on the one axis that separates a voice agent from a voice interface.

Two caveats before that table gets treated as settled. The arena preference figure used in the index is frozen at the point a model becomes eligible for publication, so it does not track the live Elo on the Speech Agent Arena chart — read it as a snapshot, not a running score. And the τ-Voice margins at the top of the board are thin enough to be noise: the reasoning variant's 68.6% leads GPT-Live-1 at 67.9% (Astra backend, medium effort) by 0.7 points, with GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% behind it. The 38-point gap between the two Gemini models is not thin. The 0.7-point lead over GPT-Live-1 is.

Where it sits on the index

In the capture below, taken on September 16, 2026, the top of the board reads like this:

Gemini 3.8 Live Extended Thinking (High) — 82.6 (1위)

• GPT-Live-1 (Astra 백엔드, 중간 노력) — 81.5

• Grok Voice Think Fast 2.0 High — 81.3

• GPT-Live-1 (Sol 백엔드, 낮은 노력) — 80.1

Gemini 3.8 Live — 76.0

• GPT-Realtime-2.1 High — 73.9

• Gemini 3.1 Flash Live High — 71.5

That is the case for the split in one screen. The Extended Thinking variant takes the top spot, running at the board's (High) reasoning-effort label — the same convention it uses for Grok Voice Think Fast 2.0 High and GPT-Realtime-2.1 High, and the reason the "(High)" suffix you may see attached to this model in coverage is a configuration label rather than a separate release. The standard variant lands fifth — below two GPT-Live-1 configurations and below Grok. Google's own blog describes the standard model as "highly cost-effective" and notes it "secured a second place in the Speech Agent Arena," which is a different board measuring a different thing. Both statements are true. Read together they say: 3.8 Live is a very good conversational model at a very good price, and it is not the frontier of voice intelligence.

Screenshot of the Artificial Analysis Speech to Speech leaderboard page, captured September 16, 2026. The AA-Speech to Speech Index bar chart shows Gemini 3.8 Live Extended Thinking at 82.6 in first place, GPT-Live-1 (Astra) at 81.5, Grok Voice Think Fast 2.0 at 81.3, GPT-Live-1 (Sol) at 80.1 and Gemini 3.8 Live at 76.0, alongside speed and cost-per-hour-of-input-audio panels.

가장 먼저 도착한 유출

해당 모델들은 발표 하루 전 Google Cloud 할당량 페이지에서 발견되었으며, 당시 우리는 이를 모델 카드도, 가격 정보도, Google의 확인도 없는 미확인 슬러그로 보도했습니다. 그 글은 상태에 대해서는 맞았지만 시점에 대해서는 틀렸습니다 — 확인은 약 24시간 이내에 이루어졌습니다.

이 교훈은 기록해 둘 가치가 있다. 그것은 일반적인 본능에 어긋나기 때문이다. 출시 하루 전에 유출된 슬러그는 출시나 다름없다. 8주 동안 이를 뒷받침하는 서류가 전혀 없는 유출된 슬러그는 완전히 다른 이야기다. 이 둘을 구분한 신호는 슬러그가 아니라, 용량이 프로비저닝된 경우에만 존재하는 할당량 페이지의 존재였다.

Screenshot of Google's official blog post titled 'Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking', dated Sep 15, 2026 and credited to Tom Ouyang and Malini Jaganathan of the Gemini Audio Team, with the summary describing the models as Google's most advanced live dialogue models with major upgrades in intelligence and parallel reasoning.

분당 비용은 얼마인가요

발표 당시 구글의 자료는 두 모델의 가격을 구분하지 않았습니다 — 표준 모델에는 요금이 제시되었고 추론 변형 모델에는 제시되지 않았으며, 위의 스코어보드에 기록된 것이 바로 그것입니다. 그 공백은 이후 메워졌습니다. 이제 두 모델 모두 Live API를 통해 공개된 가격을 적용받습니다: 오디오 입력 분당 $0.005, 오디오 출력 분당 $0.018이며, 구글은 두 모델이 출시된 9월 16일에 이를 확인했습니다. 동일한 공개 요금표는 Gemini 3 Live 티어의 토큰 가격도 다룹니다 — 텍스트 토큰 100만 개당 입력 $0.75, 출력 $4.50, 오디오 토큰 100만 개당 입력 $3.00, 출력 $12.00, 이미지 또는 동영상 입력 100만 개당 $1.00 — 다만 이 토큰 항목들은 모델별이 아니라 티어 기준으로 공개된 것이므로, 특정 변형에 대한 견적이 아니라 해당 티어의 요금표로 읽어야 합니다.

Artificial Analysis has also filled in the number that matters most for planning. Its leaderboard now carries a cost-per-hour-of-input-audio column — the cost to complete a fixed 40-question Big Bench Audio subset, normalised to an hourly rate — and both new models have a figure there:

Gemini 3.8 Live — $0.84/hour of input audio, the lowest paid rate on that board

Gemini 3.8 Live Extended Thinking (High) — $3.50/hour, 품질에서 앞서는 두 모델보다도 여전히 저렴합니다

• Grok Voice Think Fast 2.0 High — $4.80/시간

• GPT-Live-1 (Astra 백엔드, 중간 수준 노력) — 시간당 $5.83

• GPT-Realtime-2 (높음) — 시간당 $4.14

• GPT-Realtime-2.1 High — 시간당 $10.75

One note on the captures above, since the costs panel in the leaderboard screenshot predates these two rows: neither $0.84 nor $3.50 appears in it, and its cheapest bar is $1.42. The figures in the list are read from the board's summary table as it stands on September 16, 2026, not from that image. The index values shown in the image are unaffected and match the table.

Read the two new numbers against the components above and the release stops being a two-model announcement and becomes a single decision with a published exchange rate. The reasoning variant buys 6.6 index points, six points of speech reasoning and 38.5 points of agentic task completion for a little over four times the hourly audio cost. Whether that is worth paying depends entirely on what your agent is doing — and the shape of the answer changed once the components were published. A voice agent whose job is to finish a multi-step task is buying the single largest improvement in this release, at $3.50 an hour while undercutting GPT-Live-1 Astra by about 40% and Grok Voice Think Fast 2.0 High by roughly a quarter, and scoring above both. A voice agent whose job is to converse — answer, hold a thread, take a message, be pleasant — is buying very little with extra reasoning depth, and at $0.84 an hour it is buying the cheapest competent voice model anyone currently publishes.

그것이 목록의 형태다. 구글이 자사의 추론 모델을 그 모델이 앞서는 두 모델보다 낮은 가격에 책정했고, 자사의 표준 모델은 이 보드의 모든 모델보다 낮은 가격에 책정했다. 음성 분을 대량으로 운영하고 있다면, 청구서를 움직이는 숫자는 지수 점수가 아니라 바로 이것이다. 그리고 두 변형이 4배 차이로 벌어져 있다는 사실은 어느 쪽을 호출할지 고르는 것이 스택에서 단일 최대 비용 레버라는 뜻이다.

"Private preview"가 그 문장에서 실제로 제 몫을 톡톡히 하고 있다

두 모델 중 어느 것도 계약적 의미에서 일반 제공되는 것은 아니지만, 개발자 인터페이스는 열려 있고 이제 가격이 책정되어 있으므로, 무엇을 계획할 수 있는지가 달라집니다:

Gemini 3.8 Live — 개발자는 문서화된 ID를 통해 Gemini API, Live API 및 Google AI Studio에서 사용할 수 있습니다: gemini-3.8-live. 기업은 Gemini Enterprise에서 비공개 프리뷰를 이용할 수 있으며, Gemini Enterprise for Customer Experience는 "출시 예정"입니다. 모든 사용자는 Search Live에서 사용할 수 있습니다.

Gemini 3.8 Live Extended Thinking — 동일한 개발자 서피스를 gemini-3.8-live-extended-thinking에서 제공하며, 더 넓은 소비자 경로도 추가됩니다: Gemini Live, Google AI Pro 및 Ultra 구독자를 위한 Docs Live, 그리고 모든 Google AI 구독자를 위한 Gmail Live 및 Keep Live.

아직 빠져 있는 것은 기업이 필요로 하는 부분입니다: 정식 출시(GA) 약속, 공개된 가동률 수치, 그리고 Google이 유지하겠다고 약속한 요금입니다. 이런 것들이 필요하다면 기다리세요. 프로토타이핑 중이라면, 실제 요금표가 뒷받침되는 경로가 오늘 열려 있으며, 이는 공개된 가격이 없는 프리뷰보다 실질적으로 더 나은 상황입니다. 지금 통합을 구축하는 것은 합리적입니다. 하지만 수익에 결정적인 전화선을 Google이 약속하지 않은 가용성 목표에 올리는 것은 그렇지 않으며, 그에 대한 해결 방안은 아래에 설명되어 있습니다.

벤더의 주장이 끝나는 곳

Google published per-component figures of its own alongside the launch — 68.6% on τ-Voice, 35.1% on Sierra's τ³-Banking leaderboard, and 97.7% on Big Bench Audio, all credited to the Extended Thinking model. Those now divide into two groups, and the split matters.

Two of the three line up with measurements Artificial Analysis runs itself under its own harness. Its τ-Voice agentic component reads 68.6% for this model, against 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice Think Fast 2.0; its Big Bench Audio speech-reasoning component reads 98%, against Google's 97.7%. Artificial Analysis describes τ-Voice and Big Bench Audio as its own benchmarks, run across three trials where available. So the τ-Voice and reasoning numbers now sit on a public board you can go and read, not only in Google's announcement — treat them as corroborated rather than settled, since the exact figures Google quoted may well be that same run rather than a second one.

The Sierra number has no such backing and remains a vendor claim: 35.1% on Sierra's τ³-Banking leaderboard against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. It is also the number worth sitting with, because 35.1% means the leading voice model on this board fails roughly two of every three realistic banking task-completion attempts. Google also cites a ServiceNow EVA-Bench run performed on the Live API in Gemini Enterprise Agent Platform, and describes both models as pushing the Pareto frontier for complex workflows — neither of which has an independent reading attached.

There is a precedent worth holding onto here. xAI reported a Speech to Speech Index figure of 82.9 for Grok Voice Think Fast 2.0 in July. On the board captured above, that model sits at 81.3. Vendor-reported index scores do not always survive contact with a live leaderboard — usually because the index is revised, sometimes because the configuration tested was not the one that shipped. Verify the τ-Voice and Big Bench numbers on your own traffic, and read the agentic column as the one to test hardest.

이 에이전트들이 실제로 실행되는 계층

두 모델 모두 백엔드 앞에 위치합니다. 백그라운드 도구 실행, 다단계 작업 계획, 그리고 모든 본격적인 음성 에이전트가 사용하는 위임 패턴은 모두 일반적인 텍스트 모델 호출로 귀결됩니다. 그리고 바로 그 계층이 비용 변동이 가장 크고 종속성이 가장 적은 계층입니다.

OrcaRouter는 단일 키로 190개 모델을 제공업체 정가에 마크업 없이 제공하며, 이는 상위 연구소의 가격 인하가 다음 계약 갱신 시점이 아니라 같은 날 우리 쪽에 반영된다는 뜻입니다. 프런트엔드가 프리뷰 모델인 음성 스택에서 유용한 두 가지 특성은, 음성 통합을 건드리지 않고도 위임 대상을 교체할 수 있다는 점, 그리고 백엔드에서 오류가 발생하거나 타임아웃이 나더라도 자동 페일오버가 통화를 계속 유지해 준다는 점입니다. 우리가 호스팅하는 것과 하지 않는 것을 정확히 밝히자면, Gemini 3.8 Live 엔드포인트는 우리 라우터에 없습니다. 그것은 Google 자체 API에서 제공되는 것이지만, 이런 에이전트가 작업을 넘기는 텍스트 모델은 대개 우리 쪽에 있으며, Gemini 3.8 Flash도 그중 하나입니다.

다음에 볼 콘텐츠

Three things will settle the open questions in this release. The first is general availability and a rate Google has committed to hold — the prices are published now, but preview pricing has a habit of moving, and a rate card is not a contract. The second is whether the arena and task-success columns hold up, because the standard model's 1083 Elo and 93.2% task success are the strongest argument against paying 4× for the reasoning variant, and the index's arena figure is frozen at eligibility rather than live. The third is independent runs on the agentic gap itself: 68.6% against 30.1% is the largest claim in the release and the one most worth reproducing, because everything else about the two models is close.

Until then, the honest summary runs on two numbers rather than one. Gemini 3.8 Live Extended Thinking is the best-scoring voice model on the public board, leads the agentic component outright, and undercuts the two models directly behind it on hourly cost. Gemini 3.8 Live is the cheapest competent voice model on that board, is preferred by the arena and more reliable on shallow tasks, and sits 6.6 points back overall with a hard ceiling at the one thing agents get hired for. Google has shipped the same fork twice, four times apart on price, and made the choice unusually easy to price — provided you read the components and not just the index.

이 글에서 비교한 모델1

이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트