Thẻ tiêu đề hero cho Gemini 3.8 Live với dòng kicker 'OrcaRouter · radar mô hình — ra mắt' và phụ đề 'Google tách dòng giọng nói của mình thành hai — một model để mở rộng quy mô, một để suy luận.' Hai thẻ ghi 'Gemini 3.8 Live — Index 76.0, được thiết kế để mở rộng quy mô và tối ưu chi phí' và '3.8 Live Extended Thinking — Index 82.6, suy luận nhiều bước', với các chip cho ngày 15 tháng 9 năm 2026, 97 ngôn ngữ giữa cuộc trò chuyện, dấu thủy vân âm thanh SynthID và bản xem trước Live API. Logo OrcaRouter được ghép ở góc dưới cùng bên phải.
Guides & Insights

Giới thiệu Gemini 3.8 Live và 3.8 Live Extended Thinking: Google tách dòng sản phẩm giọng nói của mình thành hai

Tác giả

Elias Hawthorne

Ngày đăng

Mô hình mới nhất · 20Xem tất cả mô hình
Benchmark: Artificial Analysis · cập nhật hằng ngày
Quay lại tất cả bài viết

Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived on September 15, 2026 — one announcement, two models, and a genuine fork in the product line. The headline number is a 6.6-point gap on Artificial Analysis's Speech to Speech Index, 82.6 to 76.0. The number underneath it is 38. On the same board's τ-Voice agentic task completion measure, the reasoning variant resolves 68.6% of replica customer-service scenarios and the standard model resolves 30.1%. The 6.6 points are what the index says. The 38 points are what happens when you ask either model to actually get something done.

A day on from launch, the fork has a price attached. Both models carry published per-minute audio rates through the Live API, and Artificial Analysis's leaderboard puts cost-per-hour of input audio at $0.84 for the standard model against $3.50 for the reasoning variant — a little over four times as much per hour for those 6.6 index points. That trade is the whole decision, and reading it off the index alone gets it wrong.

Điều khiến bản phát hành này khác với đợt làm mới mô hình giọng nói thông thường là việc phân chia không phải là một bậc kích thước. Đó là sự bất đồng về việc công việc của một tác nhân thoại là gì. Một trong những mô hình này là giao diện đàm thoại. Mô hình còn lại là một hệ thống suy luận mà tình cờ biết nói.

Hai mẫu, hai công việc

Cách Google tự định hình về vấn đề này rõ ràng một cách lạ thường. Gemini 3.8 Live được mô tả là "được xây dựng để mở rộng quy mô và tối ưu chi phí, kết hợp trí tuệ đối thoại với hội thoại lưu loát và nền tảng thị giác." Gemini 3.8 Live Extended Thinking thì "được xây dựng cho các tác vụ có độ phức tạp cao, với trí tuệ được gia tăng và khả năng suy luận nhiều bước."

Sự khác biệt thực tế thể hiện rõ trong danh sách tính năng hơn là trên bảng thông số kỹ thuật:

Gemini 3.8 Live — xử lý đầu vào hình ảnh gần như theo thời gian thực, tự động chuyển đổi ngay giữa cuộc trò chuyện trên 97 ngôn ngữ được hỗ trợ, và thực thi công cụ/API chạy nền để mô hình xác nhận yêu cầu và vẫn tiếp tục trò chuyện trong khi cuộc gọi hoàn tất.

Gemini 3.8 Live Extended Thinking — vừa suy luận vừa nói cùng lúc, tường thuật tiến trình bằng những câu như "Để tôi kiểm tra nhé…" trong khi một tác vụ gồm nhiều bước đang chạy.

Dùng chung — mọi âm thanh được tạo đều mang dấu thủy vân SynthID của Google, cả hai đều được bao gồm trong một thẻ mô hình được công bố dưới tên gemini-3-8-audio, và cả hai hiện đều có ID mô hình API được ghi trong tài liệu: gemini-3.8-livegemini-3.8-live-extended-thinking.

Mô hình Extended Thinking không chỉ đơn thuần là 3.8 Live với ngân sách suy nghĩ dài hơn. Đây là mô hình có thể duy trì đồng thời cả một cuộc hội thoại lẫn một kế hoạch tác vụ mà cuộc hội thoại không bị đình trệ — đó chính xác là điều khiến các tác nhân thoại thất bại trong vận hành thực tế. Người dùng không muốn sự im lặng trong khi tác nhân tra cứu thông tin. Họ muốn tác nhân tiếp tục nói.

A two-column scoreboard comparing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Index score 76.0 vs 82.6; built for scale and cost efficiency vs high-complexity reasoning; languages 97 mid-conversation on both; visual input near real-time on both; background tools yes on both; audio price $0.005 in / $0.018 out per minute vs not separately announced. Footer reads 'Index figures per Artificial Analysis, Sep 16 2026; pricing as announced.' The OrcaRouter logo is composited in the bottom-right corner.

The board says 6.6. The components say 38

The Speech to Speech Index is a weighted average of four underlying results: Speech Reasoning, measured by Artificial Analysis's own Big Bench Audio set; Agentic Performance, measured by running the τ-Voice customer-service benchmark; Arena Preference, taken from the Speech Agent Arena; and Task Success Rate. Only models with all four components receive an index score, which is why some rows on that board are blank.

Pulled apart, the two new models do not look like a 6.6-point pair at all:

Speech to Speech Index — Gemini 3.8 Live Extended Thinking (High) 82.6 vs Gemini 3.8 Live 76.0

Agentic performance (τ-Voice task completion) — 68.6% vs 30.1%

Speech reasoning (Big Bench Audio) — 98% vs 92%

Conversational dynamics — 91.9% vs 96.1%

Arena preference (Elo) — 990 vs 1083

Task success rate — 89.1% vs 93.2%

Time to first audio — 1.35s vs 1.18s

Cost per hour of input audio — $3.50 vs $0.84

Read it honestly and the standard model wins four of eight lines. It is faster to first audio, rated higher by the arena, more likely to complete a task, and rated better on conversational dynamics. It is also a third of the price. What it cannot do is finish the job when the job has steps: 30.1% against 68.6% is the largest single gap anywhere in this release, and it sits on the one axis that separates a voice agent from a voice interface.

Two caveats before that table gets treated as settled. The arena preference figure used in the index is frozen at the point a model becomes eligible for publication, so it does not track the live Elo on the Speech Agent Arena chart — read it as a snapshot, not a running score. And the τ-Voice margins at the top of the board are thin enough to be noise: the reasoning variant's 68.6% leads GPT-Live-1 at 67.9% (Astra backend, medium effort) by 0.7 points, with GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% behind it. The 38-point gap between the two Gemini models is not thin. The 0.7-point lead over GPT-Live-1 is.

Where it sits on the index

In the capture below, taken on September 16, 2026, the top of the board reads like this:

Gemini 3.8 Live Extended Thinking (High) — 82.6 (vị trí đầu tiên)

• GPT-Live-1 (backend Astra, mức nỗ lực trung bình) — 81.5

• Grok Voice Think Fast 2.0 High — 81.3

• GPT-Live-1 (backend Sol, nỗ lực thấp) — 80.1

Gemini 3.8 Live — 76.0

• GPT-Realtime-2.1 High — 73.9

• Gemini 3.1 Flash Live High — 71.5

That is the case for the split in one screen. The Extended Thinking variant takes the top spot, running at the board's (High) reasoning-effort label — the same convention it uses for Grok Voice Think Fast 2.0 High and GPT-Realtime-2.1 High, and the reason the "(High)" suffix you may see attached to this model in coverage is a configuration label rather than a separate release. The standard variant lands fifth — below two GPT-Live-1 configurations and below Grok. Google's own blog describes the standard model as "highly cost-effective" and notes it "secured a second place in the Speech Agent Arena," which is a different board measuring a different thing. Both statements are true. Read together they say: 3.8 Live is a very good conversational model at a very good price, and it is not the frontier of voice intelligence.

Screenshot of the Artificial Analysis Speech to Speech leaderboard page, captured September 16, 2026. The AA-Speech to Speech Index bar chart shows Gemini 3.8 Live Extended Thinking at 82.6 in first place, GPT-Live-1 (Astra) at 81.5, Grok Voice Think Fast 2.0 at 81.3, GPT-Live-1 (Sol) at 80.1 and Gemini 3.8 Live at 76.0, alongside speed and cost-per-hour-of-input-audio panels.

Vụ rò rỉ đã đến trước tiên

Các mô hình đã được phát hiện trên một trang hạn mức Google Cloud vào ngày trước thông báo, và vào thời điểm đó, chúng tôi đã đưa tin về lần xuất hiện này như một slug chưa được xác nhận, không có thẻ mô hình, không có giá và không có xác nhận từ Google. Bài đăng đó đúng về trạng thái nhưng sai về mốc thời gian — xác nhận đã đến trong khoảng 24 giờ.

Bài học này đáng được ghi lại, vì nó đi ngược lại bản năng thường thấy. Một slug bị rò rỉ một ngày trước khi ra mắt là một đợt ra mắt. Một slug bị rò rỉ mà không có giấy tờ xác thực nào trong suốt tám tuần là một câu chuyện hoàn toàn khác. Tín hiệu phân biệt hai trường hợp không phải là slug; mà là sự hiện diện của một trang hạn ngạch, thứ chỉ tồn tại khi dung lượng đã được cấp phát.

Screenshot of Google's official blog post titled 'Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking', dated Sep 15, 2026 and credited to Tom Ouyang and Malini Jaganathan of the Gemini Audio Team, with the summary describing the models as Google's most advanced live dialogue models with major upgrades in intelligence and parallel reasoning.

Chi phí mỗi phút là bao nhiêu

Khi công bố, tài liệu của Google không tách biệt hai mô hình về giá — mô hình tiêu chuẩn được đưa ra một mức giá còn biến thể suy luận thì không, đó là những gì bảng điểm ở trên ghi lại. Khoảng trống đó kể từ đó đã được lấp đầy. Cả hai mô hình hiện đều có mức giá được công bố thông qua Live API: 0,005 USD mỗi phút âm thanh đầu vào và 0,018 USD mỗi phút âm thanh đầu ra, điều mà Google đã xác nhận vào ngày 16 tháng 9 khi cả hai được triển khai. Cùng bảng giá được công bố đó cũng bao gồm mức giá theo token của hạng Gemini 3 Live — 0,75 USD mỗi 1 triệu token văn bản đầu vào và 4,50 USD đầu ra, 3,00 USD mỗi 1 triệu token âm thanh đầu vào và 12,00 USD đầu ra, và 1,00 USD mỗi 1 triệu cho đầu vào hình ảnh hoặc video — mặc dù các dòng token đó được công bố cho hạng chứ không phải theo từng mô hình, nên hãy đọc chúng như bảng giá của hạng chứ không phải như báo giá cho một biến thể cụ thể.

Artificial Analysis has also filled in the number that matters most for planning. Its leaderboard now carries a cost-per-hour-of-input-audio column — the cost to complete a fixed 40-question Big Bench Audio subset, normalised to an hourly rate — and both new models have a figure there:

Gemini 3.8 Live — $0.84/hour of input audio, the lowest paid rate on that board

Gemini 3.8 Live Extended Thinking (High) — 3,50 USD/giờ, vẫn thấp hơn cả hai mô hình mà nó vượt trội về chất lượng

• Grok Voice Think Fast 2.0 High — $4,80/giờ

• GPT-Live-1 (backend Astra, nỗ lực trung bình) — $5.83/giờ

• GPT-Realtime-2 (High) — 4,14 đô la/giờ

• GPT-Realtime-2.1 High — $10.75/giờ

One note on the captures above, since the costs panel in the leaderboard screenshot predates these two rows: neither $0.84 nor $3.50 appears in it, and its cheapest bar is $1.42. The figures in the list are read from the board's summary table as it stands on September 16, 2026, not from that image. The index values shown in the image are unaffected and match the table.

Read the two new numbers against the components above and the release stops being a two-model announcement and becomes a single decision with a published exchange rate. The reasoning variant buys 6.6 index points, six points of speech reasoning and 38.5 points of agentic task completion for a little over four times the hourly audio cost. Whether that is worth paying depends entirely on what your agent is doing — and the shape of the answer changed once the components were published. A voice agent whose job is to finish a multi-step task is buying the single largest improvement in this release, at $3.50 an hour while undercutting GPT-Live-1 Astra by about 40% and Grok Voice Think Fast 2.0 High by roughly a quarter, and scoring above both. A voice agent whose job is to converse — answer, hold a thread, take a message, be pleasant — is buying very little with extra reasoning depth, and at $0.84 an hour it is buying the cheapest competent voice model anyone currently publishes.

Đó là hình dạng của danh sách. Google đã định giá mô hình suy luận của mình thấp hơn hai mô hình mà nó đánh bại, và định giá mô hình tiêu chuẩn của mình thấp hơn mọi thứ trên bảng. Nếu bạn đang chạy số phút thoại với khối lượng lớn, đây là con số làm thay đổi hóa đơn của bạn, chứ không phải điểm chỉ số — và việc hai biến thể nằm cách nhau 4× nghĩa là chọn cái nào để gọi chính là đòn bẩy chi phí lớn nhất trong toàn bộ stack.

"Private preview" đang thực sự đóng vai trò quan trọng trong câu đó

Cả hai mô hình đều không được cung cấp rộng rãi theo nghĩa hợp đồng, nhưng bề mặt dành cho nhà phát triển đã mở và hiện đã được định giá, điều này thay đổi những gì bạn có thể dựa vào để lập kế hoạch:

Gemini 3.8 Live — các nhà phát triển có thể dùng nó trong Gemini API, Live API và Google AI Studio với ID được ghi trong tài liệu là gemini-3.8-live; các doanh nghiệp được dùng bản xem trước riêng tư trong Gemini Enterprise, còn Gemini Enterprise for Customer Experience thì "sắp ra mắt"; tất cả mọi người đều có thể dùng nó trong Search Live.

Gemini 3.8 Live Extended Thinking — cùng một bề mặt dành cho nhà phát triển dưới gemini-3.8-live-extended-thinking, cùng với một hướng tiếp cận rộng hơn dành cho người tiêu dùng: Gemini Live, Docs Live cho người đăng ký Google AI Pro và Ultra, và Gmail Live cùng Keep Live cho tất cả người đăng ký Google AI.

Điều còn thiếu là phần mà một doanh nghiệp cần: cam kết về tính khả dụng chung, một con số thời gian hoạt động được công bố, và một mức giá mà Google đã hứa sẽ giữ. Nếu bạn cần những thứ đó, hãy chờ. Nếu bạn đang tạo nguyên mẫu, con đường đã mở ngay hôm nay với một bảng giá thực sự phía sau, một vị thế tốt hơn đáng kể so với bản xem trước không có giá được công bố. Xây dựng tích hợp ngay bây giờ là hợp lý; đặt một đường dây điện thoại trọng yếu về doanh thu lên một mục tiêu khả dụng mà Google chưa cam kết thì không, và cách khắc phục điều đó được mô tả bên dưới.

Nơi những tuyên bố của nhà cung cấp kết thúc

Google published per-component figures of its own alongside the launch — 68.6% on τ-Voice, 35.1% on Sierra's τ³-Banking leaderboard, and 97.7% on Big Bench Audio, all credited to the Extended Thinking model. Those now divide into two groups, and the split matters.

Two of the three line up with measurements Artificial Analysis runs itself under its own harness. Its τ-Voice agentic component reads 68.6% for this model, against 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice Think Fast 2.0; its Big Bench Audio speech-reasoning component reads 98%, against Google's 97.7%. Artificial Analysis describes τ-Voice and Big Bench Audio as its own benchmarks, run across three trials where available. So the τ-Voice and reasoning numbers now sit on a public board you can go and read, not only in Google's announcement — treat them as corroborated rather than settled, since the exact figures Google quoted may well be that same run rather than a second one.

The Sierra number has no such backing and remains a vendor claim: 35.1% on Sierra's τ³-Banking leaderboard against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. It is also the number worth sitting with, because 35.1% means the leading voice model on this board fails roughly two of every three realistic banking task-completion attempts. Google also cites a ServiceNow EVA-Bench run performed on the Live API in Gemini Enterprise Agent Platform, and describes both models as pushing the Pareto frontier for complex workflows — neither of which has an independent reading attached.

There is a precedent worth holding onto here. xAI reported a Speech to Speech Index figure of 82.9 for Grok Voice Think Fast 2.0 in July. On the board captured above, that model sits at 81.3. Vendor-reported index scores do not always survive contact with a live leaderboard — usually because the index is revised, sometimes because the configuration tested was not the one that shipped. Verify the τ-Voice and Big Bench numbers on your own traffic, and read the agentic column as the one to test hardest.

Lớp mà các tác nhân này thực sự chạy trên

Cả hai mô hình đều nằm phía trước một backend. Thực thi công cụ chạy nền, lập kế hoạch tác vụ nhiều bước và mô thức ủy quyền mà mọi voice agent nghiêm túc đều sử dụng, tất cả đều quy về các lệnh gọi mô hình văn bản thông thường — và đó chính là tầng có mức biến động chi phí lớn nhất cùng mức khóa chặt thấp nhất.

OrcaRouter cung cấp 190 mô hình từ một khóa duy nhất với giá niêm yết của nhà cung cấp và không có phụ phí, nghĩa là khi một phòng lab thượng nguồn giảm giá thì bên chúng tôi áp dụng ngay trong cùng ngày, thay vì phải chờ đến kỳ gia hạn hợp đồng tiếp theo. Với một hệ thống thoại mà phần giao diện phía trước là một mô hình xem trước, có hai đặc tính hữu ích: mục tiêu ủy quyền có thể được hoán đổi mà không cần đụng đến phần tích hợp thoại, và cơ chế chuyển đổi dự phòng tự động giữ cho cuộc gọi không bị ngắt khi một backend gặp lỗi hoặc hết thời gian chờ. Nói rõ về những gì chúng tôi có và không vận hành: các endpoint Gemini 3.8 Live không nằm trên router của chúng tôi — chúng đến từ API riêng của Google — nhưng các mô hình văn bản mà những agent này chuyển giao công việc cho thì rất thường có mặt trên đó, trong đó có Gemini 3.8 Flash.

Xem gì tiếp theo

Three things will settle the open questions in this release. The first is general availability and a rate Google has committed to hold — the prices are published now, but preview pricing has a habit of moving, and a rate card is not a contract. The second is whether the arena and task-success columns hold up, because the standard model's 1083 Elo and 93.2% task success are the strongest argument against paying 4× for the reasoning variant, and the index's arena figure is frozen at eligibility rather than live. The third is independent runs on the agentic gap itself: 68.6% against 30.1% is the largest claim in the release and the one most worth reproducing, because everything else about the two models is close.

Until then, the honest summary runs on two numbers rather than one. Gemini 3.8 Live Extended Thinking is the best-scoring voice model on the public board, leads the agentic component outright, and undercuts the two models directly behind it on hourly cost. Gemini 3.8 Live is the cheapest competent voice model on that board, is preferred by the arena and more reliable on shallow tasks, and sits 6.6 points back overall with a hard ceiling at the one thing agents get hired for. Google has shipped the same fork twice, four times apart on price, and made the choice unusually easy to price — provided you read the components and not just the index.

So sánh trong bài viết này1

Phát hiện từ bài viết này · Benchmark: Artificial Analysis · cập nhật hằng ngày