
MAI-Transcribe-2-Streaming vs Gemini 3.5 Transcribe Live: Same $9, Different Trade
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 221 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 106 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
MAI-Transcribe-2-Streaming and Gemini 3.5 Transcribe Live cost the same money and are built on opposite theories of what a streaming transcriber is for. Both land at roughly $9.00 per 1,000 minutes of streamed audio. Microsoft's model, released October 1, 2026, is a precision instrument: a claimed 2.5% streaming word error rate at 0.13 seconds to final transcript, first for accuracy on both final and partial transcripts by Microsoft's own account, measured on Artificial Analysis's streaming methodology but not yet plotted as a row on that board. The other model, in public preview since August 26, 2026 and served through the Live API, is a coverage instrument: 85+ languages with mid-stream switching, a free tier, and 4.00% streaming word error rate at 0.40 seconds, the figure Artificial Analysis measures for it. Neither is the obvious pick, and the reason is that they are not really competing for the same workload.
Here is the head-to-head as the numbers actually read — vendor-reported figures labeled, tracker figures attributed, and the parts where one model wins on a dimension the other does not contest at all.
The one-line version
• Accuracy and latency — MAI-Transcribe-2-Streaming, on the published numbers. Microsoft claims 2.5% streaming WER at 0.13 seconds to final transcript, first for accuracy on both final and partial transcripts, and 2.5% on the first partial at 0.12 seconds. Against that, Artificial Analysis has Gemini 3.5 Transcribe Live at 4.00% at 0.40 seconds. Microsoft's claimed edge is 1.5 points and roughly three times faster to settle. The caveat is that the tracker has not yet plotted MAI-Transcribe-2-Streaming itself — the board's current leader is Grok Voice Transcribe 2.0 at 2.73% and 0.49 seconds, with Muse Voice Transcribe at 3.06% and 0.16 seconds — so the 2.5% is a vendor figure read against a field the tracker does measure. It is not a close call on the axis both models publish, but only one side of it is independently published.
• Language breadth — Gemini 3.5 Transcribe Live. 85+ languages with automatic detection and the ability to switch language mid-stream without restarting the session, which is Google's headline claim for the Live API. MAI-Transcribe-2-Streaming covers 60 languages with automatic continuous language detection. Microsoft's floor is higher; Google's ceiling is wider, and the 25-language gap is mostly in languages where a single-language streaming model is not usable at all.
• Session shape — Gemini 3.5 Transcribe Live, and this one cuts both ways. The Live API is a WebSocket session capped at 10 minutes, which is fine for a voice agent turn and awkward for an hour-long meeting; Microsoft publishes no equivalent session ceiling for its streaming model, which matters if your unit of work is a call, not a conversational turn.
• Cost of entry — Gemini 3.5 Transcribe Live. There is a free tier, and Google's blended rate works out to about $0.009 per minute on the token-based pricing. Microsoft's $0.54 per audio-hour is an introductory rate Microsoft says runs through the end of the year, with no standard price disclosed — the same structure as its September batch launch.

Why the accuracy gap is bigger than it looks
A 1.5-point word error rate gap sounds like a rounding difference until you price it. At 4.00% streaming WER, Gemini 3.5 Transcribe Live gets roughly 40 words wrong per 1,000; at Microsoft's claimed 2.5%, MAI-Transcribe-2-Streaming gets 25. On the volumes a real-time pipeline produces — a support organization running 5,000 hours of calls a month is 300,000 minutes, more than 45 million words at conversational rates — that spread is millions of incorrect words annually, concentrated in the entities that downstream systems depend on: account names, ticket numbers, dosages, addresses.
The latency difference is the harder one to sell. 0.13 seconds versus 0.40 seconds after end of speech is perceptible in a live caption stream — under about 0.2 seconds reads as simultaneous, and a third of a second reads as the transcript chasing the speaker — but it is not perceptible in an agent-assist panel that refreshes on a human timescale. If your consumer is a person reading along, Microsoft's latency lead is the product. If your consumer is a model that gets the final transcript when the turn ends, it is a number.
Both figures carry the same caveat, and it is worth stating once: these are single-run numbers from a tracking board 37 models deep, and the gap between first and third on that board is under a point. A re-run can shuffle the top of the streaming leaderboard in a way that a two-point gap on the batch board cannot. Nor is Microsoft's 2.5% on the board at all yet — it is a vendor claim measured on the tracker's methodology, which is a different thing from a row the tracker publishes.
The limits are where the decision actually lives
Feature limits separate these two faster than benchmarks do, and they run in opposite directions.
Gemini 3.5 Transcribe Live carries three documented constraints: a 10-minute session cap on the Live API, no speaker diarization, and no word-level timestamps. The first is architectural — streaming over a WebSocket session has a ceiling by design — and it means long-form live transcription has to be stitched across sessions, with the seam handling left to you. The second two are absolute: if your product needs to attribute speech to a speaker, or to place a word at a timestamp for search or subtitle sync, Gemini 3.5 Transcribe Live does not do it at any price.
MAI-Transcribe-2-Streaming's constraints are less documented, which is itself information. Microsoft makes no diarization claim for the streaming model and no word-timestamp claim either; the batch MAI-Transcribe-2 bundles diarization and returns word-level timestamps, but that is the batch product, and assuming the streaming SKU inherits them is exactly the kind of inference launch material exists to prevent. What Microsoft does claim is 60 languages with continuous language detection and Pareto-frontier placement on accuracy versus latency — both vendor statements that the tracker's data is consistent with.
So the honest feature read is: neither streaming model publishes diarization or word timestamps. Gemini 3.5 Transcribe Live has a published session cap and a free tier; MAI-Transcribe-2-Streaming has a published language count and no published cap. If your requirements include diarization, neither of these is the answer and you are looking at batch transcription with a streaming layer bolted in front.
What the price parity is really telling you
Two vendors arriving at the same $9 per 1,000 minutes from opposite directions is the most interesting fact in this comparison. Google got there through token-based pricing on a Live API session: audio in at $3.50 per million tokens, text out at $21 per million, and a rough consumption of 175 text tokens per minute, which blends to about $0.009 a minute. Microsoft got there by listing a flat $0.54 per audio-hour for a model in a family whose batch sibling runs at $0.10. One is a metered real-time session; the other is a flat per-minute rate with an introductory discount.
Those structures behave differently as you scale. A flat rate is predictable and easy to budget; a token-metered session moves with how much the model talks back and how long the session idles open. For a high-volume pipeline, the flat rate is the friendlier invoice; for a prototype or a low-volume feature, the free tier and per-token billing are the friendlier entry.

What neither structure changes is the layer above the transcript. Whichever model produces it, a live transcript is input to summarization, classification, redaction, and routing, and those steps have their own cost and quality curves — usually run across several models rather than one. That is the layer OrcaRouter covers: one API across 200+ models, provider list prices passed through at 0% markup, automatic failover, and a routing DSL that can send a transcript down different models per workload. Neither MAI-Transcribe-2-Streaming nor Gemini 3.5 Transcribe Live is on that roster — they are served through Microsoft's and Google's own speech stacks respectively — and the point is not to route them, it is to keep everything downstream of them portable.

Which one to pick
Choose MAI-Transcribe-2-Streaming when accuracy and settle time are the product: live captions a human reads, real-time compliance or quality scoring where a misheard entity has a cost, or long sessions that Gemini's 10-minute ceiling would force you to stitch. Choose Gemini 3.5 Transcribe Live when coverage is the product: a global deployment across languages you cannot enumerate in advance, mid-stream language switching, or a prototype where the free tier and per-token billing remove the commitment. Teams that need both are not stuck — the two are different APIs with different session models, and the honest architecture is to route by workload rather than standardize on one.
The claim worth taking from this comparison is not that Microsoft won. It is that streaming transcription now has two credible frontier options priced identically, which means the decision has moved from "what is available" to "what does this specific workload need." That is a better problem to have.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
