
เปิดตัว Gemini 3.8 Live และ 3.8 Live Extended Thinking: Google แยกสายผลิตภัณฑ์เสียงออกเป็นสอง
- deepseekใหม่DeepSeek: DeepSeek V4.1 Flash2026-09-1040ความฉลาด
- openaiใหม่OpenAI: GPT-6 Astra2026-09-0453ความฉลาด77การเขียนโค้ด
- googleใหม่Google: Gemini 3.8 Flash2026-09-0241ความฉลาด76การเขียนโค้ด
- qwenใหม่Qwen: Qwen3.8 Max (0902)2026-09-0240ความฉลาด72การเขียนโค้ด
- anthropicAnthropic: Claude Fable 5.12026-09-0153ความฉลาด82การเขียนโค้ด
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 ต่อ 1 ล้านโทเค็น
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642ความฉลาด72การเขียนโค้ด
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 ต่อ 1 ล้านโทเค็น
- z-aiZ.ai: GLM 5.32026-08-1845ความฉลาด75การเขียนโค้ด
- obsidianQwen3.8 27B2026-08-1534ความฉลาด68การเขียนโค้ด
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236ความฉลาด69การเขียนโค้ด
- grokSpaceXAI: Grok 4.62026-08-1244ความฉลาด77การเขียนโค้ด
- metaMeta: Muse Spark 1.22026-08-0540ความฉลาด72การเขียนโค้ด
- qwenQwen: Qwen3.8 Max2026-08-0340ความฉลาด72การเขียนโค้ด
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135ความฉลาด69การเขียนโค้ด
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 ต่อ 1 ล้านโทเค็น
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451ความฉลาด78การเขียนโค้ด
- googleGoogle: Gemini 3.6 Flash2026-07-2134ความฉลาด69การเขียนโค้ด
Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived on September 15, 2026 — one announcement, two models, and a genuine fork in the product line. The headline number is a 6.6-point gap on Artificial Analysis's Speech to Speech Index, 82.6 to 76.0. The number underneath it is 38. On the same board's τ-Voice agentic task completion measure, the reasoning variant resolves 68.6% of replica customer-service scenarios and the standard model resolves 30.1%. The 6.6 points are what the index says. The 38 points are what happens when you ask either model to actually get something done.
A day on from launch, the fork has a price attached. Both models carry published per-minute audio rates through the Live API, and Artificial Analysis's leaderboard puts cost-per-hour of input audio at $0.84 for the standard model against $3.50 for the reasoning variant — a little over four times as much per hour for those 6.6 index points. That trade is the whole decision, and reading it off the index alone gets it wrong.
สิ่งที่ทำให้การเปิดตัวครั้งนี้แตกต่างจากการรีเฟรชโมเดลเสียงตามปกติ คือการแยกออกเป็นสองแบบไม่ได้เกิดจากระดับขนาด แต่เป็นความเห็นที่ไม่ลงรอยกันเกี่ยวกับหน้าที่ของเอเจนต์เสียง หนึ่งในโมเดลเหล่านี้เป็นอินเทอร์เฟซสำหรับการสนทนา ส่วนอีกโมเดลเป็นระบบการให้เหตุผลที่บังเอิญพูดได้
สองโมเดล สองงาน
การวางกรอบของ Google เองนั้นชัดเจนเป็นพิเศษในเรื่องนี้ Gemini 3.8 Liveถูกอธิบายว่า "สร้างขึ้นเพื่อรองรับการขยายขนาดและความคุ้มค่าด้านต้นทุน ผสมผสานความฉลาดในการสนทนากับบทสนทนาที่ลื่นไหลและการยึดโยงกับภาพ" Gemini 3.8 Live Extended Thinkingถูกออกแบบมา "สำหรับงานที่มีความซับซ้อนสูง พร้อมความฉลาดที่เพิ่มขึ้นและการให้เหตุผลแบบหลายขั้นตอน"
ความแตกต่างในทางปฏิบัติจะปรากฏในรายการความสามารถ มากกว่าที่จะอยู่ในเอกสารข้อมูลจำเพาะ:
• Gemini 3.8 Live — การประมวลผลอินพุตภาพแบบเกือบเรียลไทม์ การสลับภาษาโดยอัตโนมัติกลางบทสนทนาทั่ว 97 ภาษาที่รองรับ และการทำงานของเครื่องมือ/API เบื้องหลัง เพื่อให้โมเดลรับทราบคำขอและพูดต่อไปในขณะที่การเรียกนั้นดำเนินไปจนเสร็จ
• Gemini 3.8 Live Extended Thinking — คิดและพูดไปพร้อมกัน เล่าความคืบหน้าด้วยสัญญาณอย่าง "ขอฉันตรวจสอบก่อน…" ขณะที่งานแบบหลายขั้นตอนกำลังดำเนินอยู่
• ใช้ร่วมกัน — เสียงที่สร้างขึ้นทั้งหมดมีลายน้ำ SynthID ของ Google ทั้งสองแบบครอบคลุมอยู่ใน model card ที่เผยแพร่ภายใต้ชื่อ gemini-3-8-audio และตอนนี้ทั้งสองแบบมี ID โมเดล API ที่จัดทำเอกสารไว้อย่างเป็นทางการแล้ว ได้แก่ gemini-3.8-live และ gemini-3.8-live-extended-thinking.
โมเดล Extended Thinking ไม่ใช่แค่ 3.8 Live ที่มีงบการคิดยาวขึ้นเท่านั้น แต่เป็นโมเดลที่สามารถดำเนินบทสนทนาและแผนงานไปพร้อมกันได้โดยที่บทสนทนาไม่หยุดชะงัก — ซึ่งนั่นคือสิ่งที่ทำให้เอเจนต์เสียงล้มเหลวในการใช้งานจริง ผู้ใช้ไม่ต้องการความเงียบในขณะที่เอเจนต์กำลังค้นหาข้อมูล พวกเขาต้องการให้เอเจนต์พูดต่อไป

The board says 6.6. The components say 38
The Speech to Speech Index is a weighted average of four underlying results: Speech Reasoning, measured by Artificial Analysis's own Big Bench Audio set; Agentic Performance, measured by running the τ-Voice customer-service benchmark; Arena Preference, taken from the Speech Agent Arena; and Task Success Rate. Only models with all four components receive an index score, which is why some rows on that board are blank.
Pulled apart, the two new models do not look like a 6.6-point pair at all:
• Speech to Speech Index — Gemini 3.8 Live Extended Thinking (High) 82.6 vs Gemini 3.8 Live 76.0
• Agentic performance (τ-Voice task completion) — 68.6% vs 30.1%
• Speech reasoning (Big Bench Audio) — 98% vs 92%
• Conversational dynamics — 91.9% vs 96.1%
• Arena preference (Elo) — 990 vs 1083
• Task success rate — 89.1% vs 93.2%
• Time to first audio — 1.35s vs 1.18s
• Cost per hour of input audio — $3.50 vs $0.84
Read it honestly and the standard model wins four of eight lines. It is faster to first audio, rated higher by the arena, more likely to complete a task, and rated better on conversational dynamics. It is also a third of the price. What it cannot do is finish the job when the job has steps: 30.1% against 68.6% is the largest single gap anywhere in this release, and it sits on the one axis that separates a voice agent from a voice interface.
Two caveats before that table gets treated as settled. The arena preference figure used in the index is frozen at the point a model becomes eligible for publication, so it does not track the live Elo on the Speech Agent Arena chart — read it as a snapshot, not a running score. And the τ-Voice margins at the top of the board are thin enough to be noise: the reasoning variant's 68.6% leads GPT-Live-1 at 67.9% (Astra backend, medium effort) by 0.7 points, with GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% behind it. The 38-point gap between the two Gemini models is not thin. The 0.7-point lead over GPT-Live-1 is.
Where it sits on the index
In the capture below, taken on September 16, 2026, the top of the board reads like this:
• Gemini 3.8 Live Extended Thinking (High) — 82.6 (อันดับที่หนึ่ง)
• GPT-Live-1 (แบ็กเอนด์ Astra, ความพยายามระดับปานกลาง) — 81.5
• Grok Voice Think Fast 2.0 High — 81.3
• GPT-Live-1 (แบ็กเอนด์ Sol, ความพยายามต่ำ) — 80.1
• Gemini 3.8 Live — 76.0
• GPT-Realtime-2.1 High — 73.9
• Gemini 3.1 Flash Live High — 71.5
That is the case for the split in one screen. The Extended Thinking variant takes the top spot, running at the board's (High) reasoning-effort label — the same convention it uses for Grok Voice Think Fast 2.0 High and GPT-Realtime-2.1 High, and the reason the "(High)" suffix you may see attached to this model in coverage is a configuration label rather than a separate release. The standard variant lands fifth — below two GPT-Live-1 configurations and below Grok. Google's own blog describes the standard model as "highly cost-effective" and notes it "secured a second place in the Speech Agent Arena," which is a different board measuring a different thing. Both statements are true. Read together they say: 3.8 Live is a very good conversational model at a very good price, and it is not the frontier of voice intelligence.

การรั่วไหลที่ไปถึงก่อน
โมเดลเหล่านี้ถูกพบเห็นบนหน้าโควตาของ Google Cloud หนึ่งวันก่อนการประกาศ และเราได้รายงานการพบเห็นนั้นในตอนนั้นในฐานะ slug ที่ยังไม่ได้รับการยืนยัน โดยไม่มีโมเดลการ์ด ไม่มีข้อมูลราคา และไม่มีการยืนยันจาก Google โพสต์นั้นถูกต้องเกี่ยวกับสถานะ แต่ผิดเกี่ยวกับไทม์ไลน์ — การยืนยันมาถึงภายในประมาณ 24 ชั่วโมง
บทเรียนนี้ควรค่าแก่การบันทึก เพราะมันขัดกับสัญชาตญาณตามปกติ สลักที่รั่วไหลหนึ่งวันก่อนเปิดตัวถือเป็นการเปิดตัว สลักที่รั่วไหลโดยไม่มีเอกสารยืนยันใด ๆ เป็นเวลาแปดสัปดาห์กลับเป็นอีกเรื่องหนึ่งโดยสิ้นเชิง สัญญาณที่แยกสองกรณีนั้นไม่ใช่ตัวสลัก แต่คือการมีอยู่ของหน้าโควตา ซึ่งจะมีขึ้นก็ต่อเมื่อมีการจัดสรรความจุแล้วเท่านั้น

ค่าใช้จ่ายต่อนาที
ตอนที่ประกาศ เอกสารของ Google ไม่ได้แยกสองโมเดลในเรื่องราคา — โมเดลมาตรฐานนั้นมีการกำหนดอัตราไว้ ส่วนเวอร์ชัน reasoning ไม่มี ซึ่งเป็นสิ่งที่ตารางคะแนนด้านบนบันทึกไว้ ช่องว่างนั้นได้รับการปิดไปแล้วนับตั้งแต่นั้น ปัจจุบันทั้งสองโมเดลมีราคาที่เผยแพร่อย่างเป็นทางการผ่าน Live API แล้ว: $0.005 ต่อนาทีของเสียงอินพุต และ $0.018 ต่อนาทีของเสียงเอาต์พุต ซึ่ง Google ยืนยันเมื่อวันที่ 16 กันยายน ขณะที่ทั้งสองเริ่มเปิดตัว ตารางอัตราที่เผยแพร่ชุดเดียวกันนี้ยังครอบคลุมราคาต่อโทเคนของระดับ Gemini 3 Live ด้วย — $0.75 ต่อ 1 ล้านโทเคนข้อความขาเข้า และ $4.50 ขาออก, $3.00 ต่อ 1 ล้านโทเคนเสียงขาเข้า และ $12.00 ขาออก และ $1.00 ต่อ 1 ล้านสำหรับอินพุตภาพหรือวิดีโอ — แม้ว่าแถวโทเคนเหล่านั้นจะเผยแพร่ในระดับ tier มากกว่าจะแยกตามโมเดล จึงควรอ่านเป็นอัตราของ tier นั้น ไม่ใช่เป็นใบเสนอราคาสำหรับเวอร์ชันใดเวอร์ชันหนึ่งโดยเฉพาะ
Artificial Analysis has also filled in the number that matters most for planning. Its leaderboard now carries a cost-per-hour-of-input-audio column — the cost to complete a fixed 40-question Big Bench Audio subset, normalised to an hourly rate — and both new models have a figure there:
• Gemini 3.8 Live — $0.84/hour of input audio, the lowest paid rate on that board
• Gemini 3.8 Live Extended Thinking (High) — 3.50 ดอลลาร์ต่อชั่วโมง ยังต่ำกว่ารุ่นทั้งสองที่มันเอาชนะได้ในด้านคุณภาพ
• Grok Voice Think Fast 2.0 High — $4.80/ชั่วโมง
• GPT-Live-1 (แบ็กเอนด์ Astra, ความพยายามปานกลาง) — $5.83/ชั่วโมง
• GPT-Realtime-2 (High) — $4.14/ชั่วโมง
• GPT-Realtime-2.1 High — $10.75/ชั่วโมง
One note on the captures above, since the costs panel in the leaderboard screenshot predates these two rows: neither $0.84 nor $3.50 appears in it, and its cheapest bar is $1.42. The figures in the list are read from the board's summary table as it stands on September 16, 2026, not from that image. The index values shown in the image are unaffected and match the table.
Read the two new numbers against the components above and the release stops being a two-model announcement and becomes a single decision with a published exchange rate. The reasoning variant buys 6.6 index points, six points of speech reasoning and 38.5 points of agentic task completion for a little over four times the hourly audio cost. Whether that is worth paying depends entirely on what your agent is doing — and the shape of the answer changed once the components were published. A voice agent whose job is to finish a multi-step task is buying the single largest improvement in this release, at $3.50 an hour while undercutting GPT-Live-1 Astra by about 40% and Grok Voice Think Fast 2.0 High by roughly a quarter, and scoring above both. A voice agent whose job is to converse — answer, hold a thread, take a message, be pleasant — is buying very little with extra reasoning depth, and at $0.84 an hour it is buying the cheapest competent voice model anyone currently publishes.
นั่นคือลักษณะของรายการนี้ Google ตั้งราคาโมเดลการให้เหตุผลของตนไว้ต่ำกว่าโมเดลสองตัวที่มันเอาชนะได้ และตั้งราคาโมเดลมาตรฐานของตนไว้ต่ำกว่าทุกอย่างบนบอร์ดนี้ ถ้าคุณใช้งานนาทีเสียงในปริมาณมาก นี่คือตัวเลขที่ขยับบิลของคุณ ไม่ใช่คะแนนดัชนี — และการที่สองตัวแปรนี้มีราคาต่างกัน 4 เท่า หมายความว่าการเลือกว่าจะเรียกตัวไหนคือคันโยกด้านต้นทุนที่ใหญ่ที่สุดในสแต็ก
"Private preview" กำลังทำหน้าที่จริง ๆ ในประโยคนั้น
ไม่มีโมเดลใดพร้อมใช้งานทั่วไปในความหมายเชิงสัญญา แต่ช่องทางสำหรับนักพัฒนานั้นเปิดแล้วและขณะนี้มีการกำหนดราคาแล้ว ซึ่งเปลี่ยนสิ่งที่คุณสามารถวางแผนได้:
• Gemini 3.8 Live — นักพัฒนาสามารถเข้าถึงได้ใน Gemini API, Live API และ Google AI Studio ภายใต้ ID ที่ระบุไว้ในเอกสาร gemini-3.8-live; องค์กรต่างๆ ได้รับการพรีวิวแบบส่วนตัวใน Gemini Enterprise โดย Gemini Enterprise for Customer Experience จะ "เร็ว ๆ นี้"; ทุกคนสามารถเข้าถึงได้ใน Search Live
• Gemini 3.8 Live Extended Thinking — ชุดเครื่องมือสำหรับนักพัฒนาเดียวกันภายใต้ชื่อ gemini-3.8-live-extended-thinking พร้อมเส้นทางสำหรับผู้ใช้ทั่วไปที่กว้างขึ้น: Gemini Live, Docs Live สำหรับสมาชิก Google AI Pro และ Ultra และ Gmail Live กับ Keep Live สำหรับสมาชิก Google AI ทุกคน
สิ่งที่ยังขาดไปคือส่วนที่องค์กรธุรกิจต้องการ ได้แก่ คำมั่นความพร้อมใช้งานทั่วไป ตัวเลขความพร้อมใช้งานที่เผยแพร่ และอัตราค่าบริการที่ Google สัญญาว่าจะคงไว้ หากคุณต้องการสิ่งเหล่านั้น ให้รอ หากคุณกำลังทำต้นแบบ เส้นทางเปิดอยู่แล้ววันนี้โดยมี rate card จริงรองรับ ซึ่งเป็นสถานะที่ดีกว่าอย่างมีนัยสำคัญเมื่อเทียบกับพรีวิวที่ไม่มีราคาเผยแพร่ การสร้างการผสานรวมตอนนี้สมเหตุสมผล แต่การนำสายโทรศัพท์ที่สำคัญต่อรายได้ไปวางบนเป้าหมายความพร้อมใช้งานที่ Google ยังไม่ได้ให้คำมั่นนั้นไม่สมเหตุสมผล และวิธีแก้สำหรับเรื่องนั้นอธิบายไว้ด้านล่าง
ซึ่งเป็นจุดสิ้นสุดของข้อกล่าวอ้างจากผู้ขาย
Google published per-component figures of its own alongside the launch — 68.6% on τ-Voice, 35.1% on Sierra's τ³-Banking leaderboard, and 97.7% on Big Bench Audio, all credited to the Extended Thinking model. Those now divide into two groups, and the split matters.
Two of the three line up with measurements Artificial Analysis runs itself under its own harness. Its τ-Voice agentic component reads 68.6% for this model, against 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice Think Fast 2.0; its Big Bench Audio speech-reasoning component reads 98%, against Google's 97.7%. Artificial Analysis describes τ-Voice and Big Bench Audio as its own benchmarks, run across three trials where available. So the τ-Voice and reasoning numbers now sit on a public board you can go and read, not only in Google's announcement — treat them as corroborated rather than settled, since the exact figures Google quoted may well be that same run rather than a second one.
The Sierra number has no such backing and remains a vendor claim: 35.1% on Sierra's τ³-Banking leaderboard against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. It is also the number worth sitting with, because 35.1% means the leading voice model on this board fails roughly two of every three realistic banking task-completion attempts. Google also cites a ServiceNow EVA-Bench run performed on the Live API in Gemini Enterprise Agent Platform, and describes both models as pushing the Pareto frontier for complex workflows — neither of which has an independent reading attached.
There is a precedent worth holding onto here. xAI reported a Speech to Speech Index figure of 82.9 for Grok Voice Think Fast 2.0 in July. On the board captured above, that model sits at 81.3. Vendor-reported index scores do not always survive contact with a live leaderboard — usually because the index is revised, sometimes because the configuration tested was not the one that shipped. Verify the τ-Voice and Big Bench numbers on your own traffic, and read the agentic column as the one to test hardest.
เลเยอร์ที่เอเจนต์เหล่านี้ทำงานอยู่จริง ๆ
โมเดลทั้งสองตั้งอยู่ด้านหน้าของแบ็กเอนด์ การทำงานของเครื่องมือเบื้องหลัง การวางแผนงานหลายขั้นตอน และรูปแบบการมอบหมายงานที่วอยซ์เอเจนต์จริงจังทุกตัวใช้ ล้วนแปลงเป็นการเรียกใช้โมเดลข้อความธรรมดา — และนั่นคือเลเยอร์ที่มีความแปรปรวนของต้นทุนมากที่สุดและมีการผูกติดกับผู้ให้บริการน้อยที่สุด
OrcaRouter ให้บริการ 190 โมเดลจากคีย์เดียว ในราคาตามรายการของผู้ให้บริการโดยไม่บวกเพิ่ม ซึ่งหมายความว่าเมื่อแล็บต้นทางลดราคา ฝั่งเราจะมีผลทันทีในวันเดียวกัน แทนที่จะรอถึงการต่อสัญญาครั้งถัดไป สำหรับสแตกเสียงที่ส่วนหน้าสุดเป็นโมเดลพรีวิว คุณสมบัติที่มีประโยชน์สองข้อคือ เป้าหมายที่ส่งต่องานไปให้สามารถสลับได้โดยไม่ต้องแตะการผสานรวมระบบเสียง และการสลับสำรองอัตโนมัติช่วยให้สายยังอยู่เมื่อแบ็กเอนด์เกิดข้อผิดพลาดหรือหมดเวลา เพื่อให้ชัดเจนว่าเรโฮสต์อะไรและไม่โฮสต์อะไร: ปลายทาง Gemini 3.8 Live ไม่ได้อยู่บนเราเตอร์ของเรา — สิ่งเหล่านั้นมาจาก API ของ Google เอง — แต่โมเดลข้อความที่เอเจนต์เหล่านี้ส่งงานต่อไปให้มักอยู่บนเราเตอร์ของเรา โดยมี Gemini 3.8 Flash รวมอยู่ด้วย.
สิ่งที่ควรดูต่อไป
Three things will settle the open questions in this release. The first is general availability and a rate Google has committed to hold — the prices are published now, but preview pricing has a habit of moving, and a rate card is not a contract. The second is whether the arena and task-success columns hold up, because the standard model's 1083 Elo and 93.2% task success are the strongest argument against paying 4× for the reasoning variant, and the index's arena figure is frozen at eligibility rather than live. The third is independent runs on the agentic gap itself: 68.6% against 30.1% is the largest claim in the release and the one most worth reproducing, because everything else about the two models is close.
Until then, the honest summary runs on two numbers rather than one. Gemini 3.8 Live Extended Thinking is the best-scoring voice model on the public board, leads the agentic component outright, and undercuts the two models directly behind it on hourly cost. Gemini 3.8 Live is the cheapest competent voice model on that board, is preferred by the arena and more reliable on shallow tasks, and sits 6.6 points back overall with a hard ceiling at the one thing agents get hired for. Google has shipped the same fork twice, four times apart on price, and made the choice unusually easy to price — provided you read the components and not just the index.
การเปรียบเทียบในบทความนี้1
ตรวจพบจากบทความนี้ · เบนช์มาร์ก: Artificial Analysis · อัปเดตทุกวัน
