
隆重推出 Gemini 3.8 Live 與 3.8 Live Extended Thinking:Google 將其語音產品線一分為二
- deepseek新DeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- openai新OpenAI: GPT-6 Astra2026-09-0453智能77程式
- google新Google: Gemini 3.8 Flash2026-09-0241智能76程式
- qwen新Qwen: Qwen3.8 Max (0902)2026-09-0240智能72程式
- anthropicAnthropic: Claude Fable 5.12026-09-0153智能82程式
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百萬 tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72程式
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 每百萬 tokens
- z-aiZ.ai: GLM 5.32026-08-1845智能75程式
- obsidianQwen3.8 27B2026-08-1534智能68程式
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236智能69程式
- grokSpaceXAI: Grok 4.62026-08-1244智能77程式
- metaMeta: Muse Spark 1.22026-08-0540智能72程式
- qwenQwen: Qwen3.8 Max2026-08-0340智能72程式
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135智能69程式
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 每百萬 tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451智能78程式
- googleGoogle: Gemini 3.6 Flash2026-07-2134智能69程式
Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived on September 15, 2026 — one announcement, two models, and a genuine fork in the product line. The headline number is a 6.6-point gap on Artificial Analysis's Speech to Speech Index, 82.6 to 76.0. The number underneath it is 38. On the same board's τ-Voice agentic task completion measure, the reasoning variant resolves 68.6% of replica customer-service scenarios and the standard model resolves 30.1%. The 6.6 points are what the index says. The 38 points are what happens when you ask either model to actually get something done.
A day on from launch, the fork has a price attached. Both models carry published per-minute audio rates through the Live API, and Artificial Analysis's leaderboard puts cost-per-hour of input audio at $0.84 for the standard model against $3.50 for the reasoning variant — a little over four times as much per hour for those 6.6 index points. That trade is the whole decision, and reading it off the index alone gets it wrong.
這次發布之所以不同於常見的語音模型更新,在於這種區分並非尺寸分級。它關乎對語音代理職責為何的分歧。這些模型之中,一個是對話介面。另一個則是碰巧會說話的推理系統。
兩種型號,兩種任務
Google 自己對此的表述異常清楚。Gemini 3.8 Live被描述為「專為規模化與成本效益而打造,結合對話式智慧、流暢對話與視覺接地。」Gemini 3.8 Live Extended Thinking則被形容為「專為高複雜度任務而打造,具備更強的智慧與多步驟推理能力。」
實際差異會顯示在功能清單上,而不是規格表上:
• Gemini 3.8 Live — 近乎即時的視覺輸入處理、在 97 種支援語言間自動於對話中途切換,以及背景工具/API 執行,讓模型確認請求後能繼續交談,同時呼叫在背景完成。
• Gemini 3.8 Live Extended Thinking — 同時推理並開口說明,在多步驟任務執行期間以「讓我確認一下……」等提示語即時描述進度。
• 共同 — 所有生成的音訊都帶有 Google 的 SynthID 浮水印,兩者皆涵蓋於以 gemini-3-8-audio 為名發布的模型卡中,且兩者現在都有已記載的 API 模型 ID:gemini-3.8-live 與 gemini-3.8-live-extended-thinking。
Extended Thinking 模型並不只是擁有更長思考預算的 3.8 Live。它是能夠同時維持對話與任務計畫,而不會讓對話停滯的模型——這正是讓語音代理在正式環境中失敗的原因。使用者不希望在代理查資料時出現沉默。他們希望代理繼續說話。

The board says 6.6. The components say 38
The Speech to Speech Index is a weighted average of four underlying results: Speech Reasoning, measured by Artificial Analysis's own Big Bench Audio set; Agentic Performance, measured by running the τ-Voice customer-service benchmark; Arena Preference, taken from the Speech Agent Arena; and Task Success Rate. Only models with all four components receive an index score, which is why some rows on that board are blank.
Pulled apart, the two new models do not look like a 6.6-point pair at all:
• Speech to Speech Index — Gemini 3.8 Live Extended Thinking (High) 82.6 vs Gemini 3.8 Live 76.0
• Agentic performance (τ-Voice task completion) — 68.6% vs 30.1%
• Speech reasoning (Big Bench Audio) — 98% vs 92%
• Conversational dynamics — 91.9% vs 96.1%
• Arena preference (Elo) — 990 vs 1083
• Task success rate — 89.1% vs 93.2%
• Time to first audio — 1.35s vs 1.18s
• Cost per hour of input audio — $3.50 vs $0.84
Read it honestly and the standard model wins four of eight lines. It is faster to first audio, rated higher by the arena, more likely to complete a task, and rated better on conversational dynamics. It is also a third of the price. What it cannot do is finish the job when the job has steps: 30.1% against 68.6% is the largest single gap anywhere in this release, and it sits on the one axis that separates a voice agent from a voice interface.
Two caveats before that table gets treated as settled. The arena preference figure used in the index is frozen at the point a model becomes eligible for publication, so it does not track the live Elo on the Speech Agent Arena chart — read it as a snapshot, not a running score. And the τ-Voice margins at the top of the board are thin enough to be noise: the reasoning variant's 68.6% leads GPT-Live-1 at 67.9% (Astra backend, medium effort) by 0.7 points, with GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% behind it. The 38-point gap between the two Gemini models is not thin. The 0.7-point lead over GPT-Live-1 is.
Where it sits on the index
In the capture below, taken on September 16, 2026, the top of the board reads like this:
• Gemini 3.8 即時延伸思考(高)—— 82.6(第一名)
• GPT-Live-1(Astra 後端,中等投入)— 81.5
• Grok Voice Think Fast 2.0 High — 81.3
• GPT-Live-1(Sol 後端,低投入)— 80.1
• Gemini 3.8 Live — 76.0
• GPT-Realtime-2.1 High — 73.9
• Gemini 3.1 Flash Live High — 71.5
That is the case for the split in one screen. The Extended Thinking variant takes the top spot, running at the board's (High) reasoning-effort label — the same convention it uses for Grok Voice Think Fast 2.0 High and GPT-Realtime-2.1 High, and the reason the "(High)" suffix you may see attached to this model in coverage is a configuration label rather than a separate release. The standard variant lands fifth — below two GPT-Live-1 configurations and below Grok. Google's own blog describes the standard model as "highly cost-effective" and notes it "secured a second place in the Speech Agent Arena," which is a different board measuring a different thing. Both statements are true. Read together they say: 3.8 Live is a very good conversational model at a very good price, and it is not the frontier of voice intelligence.

搶先一步的洩密
這些模型在公告前一天被發現在 Google Cloud 的配額頁面上,而我們當時報導了這次發現,將其視為一個未經證實的代號,沒有模型卡、沒有定價,也沒有 Google 的確認。那篇文章對狀態的判斷是對的,但對時間線的判斷是錯的——確認在大約 24 小時內就到來了。
這項教訓值得記錄下來,因為它違背了一般直覺。發表前一天外洩的網址代稱,就是發表。一個外洩的網址代稱,若八週內都沒有可佐證的文件,就完全是另一回事。區分兩者的訊號並不是那個網址代稱,而是配額頁面的存在;這種頁面只有在容量已完成佈建時才會存在。

每分鐘的費用
在發布宣布時,Google 的資料並未在價格上區分這兩款模型——標準模型有給出費率,推理版本則沒有,這正是上方計分板所記錄的情況。這個落差後來已經補上。這兩款模型現在都透過 Live API 公布定價:音訊輸入每分鐘 $0.005,音訊輸出每分鐘 $0.018,Google 在 9 月 16 日兩者推出時確認了這一點。同一份公布的價目表也涵蓋 Gemini 3 Live 級別的 token 定價——文字 token 每 1M 輸入 $0.75、輸出 $4.50,音訊 token 每 1M 輸入 $3.00、輸出 $12.00,以及圖片或影片輸入每 1M $1.00——不過那些 token 列是針對該級別而非個別模型公布,因此請把它們解讀為該級別的價目表,而不是特定變體的報價。
Artificial Analysis has also filled in the number that matters most for planning. Its leaderboard now carries a cost-per-hour-of-input-audio column — the cost to complete a fixed 40-question Big Bench Audio subset, normalised to an hourly rate — and both new models have a figure there:
• Gemini 3.8 Live — $0.84/hour of input audio, the lowest paid rate on that board
• Gemini 3.8 Live Extended Thinking(高)— 每小時 $3.50,仍低於它憑品質勝過的兩個模型
• Grok Voice Think Fast 2.0 High — 每小時 $4.80
• GPT-Live-1(Astra 後端,中等投入)— 每小時 $5.83
• GPT-Realtime-2 (High) — 每小時 $4.14
• GPT-Realtime-2.1 High — 每小時 $10.75
One note on the captures above, since the costs panel in the leaderboard screenshot predates these two rows: neither $0.84 nor $3.50 appears in it, and its cheapest bar is $1.42. The figures in the list are read from the board's summary table as it stands on September 16, 2026, not from that image. The index values shown in the image are unaffected and match the table.
Read the two new numbers against the components above and the release stops being a two-model announcement and becomes a single decision with a published exchange rate. The reasoning variant buys 6.6 index points, six points of speech reasoning and 38.5 points of agentic task completion for a little over four times the hourly audio cost. Whether that is worth paying depends entirely on what your agent is doing — and the shape of the answer changed once the components were published. A voice agent whose job is to finish a multi-step task is buying the single largest improvement in this release, at $3.50 an hour while undercutting GPT-Live-1 Astra by about 40% and Grok Voice Think Fast 2.0 High by roughly a quarter, and scoring above both. A voice agent whose job is to converse — answer, hold a thread, take a message, be pleasant — is buying very little with extra reasoning depth, and at $0.84 an hour it is buying the cheapest competent voice model anyone currently publishes.
這就是這份清單的樣貌。Google 將其推理模型的定價訂得低於它擊敗的兩個模型,也將其標準模型的定價訂得低於排行榜上的所有模型。如果你要大量處理語音分鐘數,真正會影響你帳單的是這個數字,而不是指數分數——而這兩種變體相差 4 倍的事實,意味著該選哪一個來呼叫,是整個堆疊中最大的單一成本槓桿。
「Private preview」在那句話裡起了實際作用。
就合約意義而言,這兩個模型都尚未正式全面提供,但開發者介面已開放,且現在已有定價,這改變了你可以據以規劃的內容:
• Gemini 3.8 Live — 開發人員可透過 Gemini API、Live API 和 Google AI Studio,以文件記載的 ID gemini-3.8-live 取得;企業可在 Gemini Enterprise 中獲得私人預覽,而 Gemini Enterprise for Customer Experience「即將推出」;所有人則可在 Search Live 中使用。
• Gemini 3.8 Live Extended Thinking——在 gemini-3.8-live-extended-thinking 下提供相同的開發者介面,並加上更廣泛的消費者使用途徑:Gemini Live、專為 Google AI Pro 與 Ultra 訂閱者推出的 Docs Live,以及為所有 Google AI 訂閱者提供的 Gmail Live 與 Keep Live。
目前仍然欠缺的,是企業所需的那一塊:正式推出(GA)的承諾、公布的正常運作時間數據,以及 Google 承諾維持的費率。如果你需要這些,就再等等。如果你是在做原型開發,這條路今天已經走得通,而且背後有真實的價目表支撐,這比一個沒有公布價格的預覽版處於實質上更有利的位置。現在就建置整合是合理的;但把攸關營收的電話線路,押在 Google 尚未承諾的可用性目標上就不合理了,而針對這一點的解決辦法說明如下。
供應商宣稱的終點
Google published per-component figures of its own alongside the launch — 68.6% on τ-Voice, 35.1% on Sierra's τ³-Banking leaderboard, and 97.7% on Big Bench Audio, all credited to the Extended Thinking model. Those now divide into two groups, and the split matters.
Two of the three line up with measurements Artificial Analysis runs itself under its own harness. Its τ-Voice agentic component reads 68.6% for this model, against 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice Think Fast 2.0; its Big Bench Audio speech-reasoning component reads 98%, against Google's 97.7%. Artificial Analysis describes τ-Voice and Big Bench Audio as its own benchmarks, run across three trials where available. So the τ-Voice and reasoning numbers now sit on a public board you can go and read, not only in Google's announcement — treat them as corroborated rather than settled, since the exact figures Google quoted may well be that same run rather than a second one.
The Sierra number has no such backing and remains a vendor claim: 35.1% on Sierra's τ³-Banking leaderboard against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. It is also the number worth sitting with, because 35.1% means the leading voice model on this board fails roughly two of every three realistic banking task-completion attempts. Google also cites a ServiceNow EVA-Bench run performed on the Live API in Gemini Enterprise Agent Platform, and describes both models as pushing the Pareto frontier for complex workflows — neither of which has an independent reading attached.
There is a precedent worth holding onto here. xAI reported a Speech to Speech Index figure of 82.9 for Grok Voice Think Fast 2.0 in July. On the board captured above, that model sits at 81.3. Vendor-reported index scores do not always survive contact with a live leaderboard — usually because the index is revised, sometimes because the configuration tested was not the one that shipped. Verify the τ-Voice and Big Bench numbers on your own traffic, and read the agentic column as the one to test hardest.
這些代理實際運行的層
這兩個模型都位在後端之前。背景工具執行、多步驟任務規劃,以及每個認真的語音代理都會採用的委派模式,全都歸結為一般的文字模型呼叫——而這正是成本變異最大、鎖定程度最低的層級。
OrcaRouter以單一金鑰提供 190 個模型,價格為供應商定價且無加成,這意味著上游實驗室的降價當天就會在我們這邊生效,而不是等到下一次合約續約。對於前端是預覽模型的語音堆疊而言,兩項有用的特性是:委派目標可以更換而無須觸碰語音整合,以及自動容錯移轉能在後端出錯或逾時時讓通話保持不中斷。確切說明我們有託管與沒有託管什麼:Gemini 3.8 Live 端點不在我們的路由器上——那些來自 Google 自家的 API——但這些代理交辦出去的文字模型往往在我們這裡,Gemini 3.8 Flash 便是其中之一。
接下來要看什麼
Three things will settle the open questions in this release. The first is general availability and a rate Google has committed to hold — the prices are published now, but preview pricing has a habit of moving, and a rate card is not a contract. The second is whether the arena and task-success columns hold up, because the standard model's 1083 Elo and 93.2% task success are the strongest argument against paying 4× for the reasoning variant, and the index's arena figure is frozen at eligibility rather than live. The third is independent runs on the agentic gap itself: 68.6% against 30.1% is the largest claim in the release and the one most worth reproducing, because everything else about the two models is close.
Until then, the honest summary runs on two numbers rather than one. Gemini 3.8 Live Extended Thinking is the best-scoring voice model on the public board, leads the agentic component outright, and undercuts the two models directly behind it on hourly cost. Gemini 3.8 Live is the cheapest competent voice model on that board, is preferred by the arena and more reliable on shallow tasks, and sits 6.6 points back overall with a hard ceiling at the one thing agents get hired for. Google has shipped the same fork twice, four times apart on price, and made the choice unusually easy to price — provided you read the components and not just the index.
本文中的比較1
根據本文內容識別 · 基準測試:Artificial Analysis · 每日更新
