
MAI-Transcribe-2-Streaming Is Live: Microsoft Takes the Streaming Transcript Crown at 2.5% WER
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 221 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 106 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
MAI-Transcribe-2-Streaming is Microsoft AI's answer to the one question its September release left open, and it answered it in 28 days. Released on October 1, 2026, MAI-Transcribe-2-Streaming is a real-time speech-to-text model that transcribes as the words arrive rather than after the file finishes. Microsoft says it ranks first for accuracy on both final and partial transcripts on Artificial Analysis's AA-WER Streaming evaluation, at 2.5% word error rate with the final transcript ready 0.13 seconds after end of speech. That is a vendor claim on a tracker's methodology rather than a row the tracker publishes — as of drafting, the board's 37 models do not include it, and the model it would displace is Grok Voice Transcribe 2.0 (Streaming) at 2.73% and 0.49 seconds. The model it is built on, MAI-Transcribe-2, shipped on September 3 as a batch model at 2.0% WER with no real-time SKU; the streaming variant closes that gap without giving up much of the accuracy that made the batch version notable. What it does give up is the price: streaming is 5.4× the batch rate at launch.
The framing Microsoft wants is a clean sweep of the streaming board. The framing the numbers support is narrower and more interesting — a model claiming to lead on accuracy and latency at once, on a frontier where the board's most accurate entry is four times slower to settle than its fastest ones and none of those fast ones is under 3% word error, priced at parity with the incumbent rather than below it. Here is the launch as it happened, with vendor claims labeled.
What Microsoft shipped on October 1
MAI-Transcribe-2-Streaming is served through Microsoft's speech stack in the Azure AI Speech service, with documentation under the model's own name and the same API-only, no-weights distribution model as the batch release. It transcribes 60 languages with automatic continuous language detection, so a session that switches languages mid-stream does not need to be declared up front. Microsoft positions it on the accuracy-versus-latency Pareto frontier of Artificial Analysis's streaming chart — the vendor's own reading of the tracker's data, and the reading the tracker's own chart supports.
The streaming model was not the whole release. Microsoft shipped two voice models alongside it: MAI-Voice-2.1, covering 23 languages across 26 locales, and MAI-Voice-2.1-Flash, which handles up to 45 seconds of audio at roughly 150 ms of latency. Read as a set, the October 1 drop is a full real-time conversational stack — hear, understand, speak — rather than a single model launch.

Reading the streaming leaderboard
AA-WER Streaming measures two things per model: how wrong the finished transcript is, and how long after the speaker stops it takes to finish. Microsoft's claim is that MAI-Transcribe-2-Streaming leads on both: 2.5% word error rate on the final transcript at 0.13 seconds after end of speech, and the same 2.5% on the first partial transcript at 0.12 seconds, which Microsoft says is first for accuracy on both. Read that as a vendor figure — as of drafting, the tracker's own streaming board has not yet plotted the model, so there is no independent MAI-Transcribe-2-Streaming row to read. What the board does show is the field it is claiming to beat, and the field is closer than a "crown" framing suggests:
• Grok Voice Transcribe 2.0 — 2.73% final-transcript WER at 0.49 seconds after end of speech, per Artificial Analysis, and the current leader on the board. That is 0.23 points behind Microsoft's claimed 2.5%, and roughly four times slower to settle — the difference between a caption that keeps up and one that lags.
• Muse Voice Transcribe — 3.06% at 0.16 seconds, per Artificial Analysis. The closest competitor on latency: about 30 milliseconds behind, and 0.5 points worse on accuracy than Microsoft's claim.
• ElevenLabs Scribe v2 Realtime — 3.59% at 0.14 seconds, per Artificial Analysis. Effectively tied on speed, more than a point behind on accuracy.
• Gemini 3.5 Transcribe Live — 4.00% streaming WER at 0.40 seconds, the figure Artificial Analysis published and Google cites for its own Live API model. That is a 1.5-point gap, and it is the comparison this release is aimed at.
The first-partial column is where the claim is least like the rest of the board. On the same tracker, Grok Voice Transcribe 2.0's first partials sit at 3.36% and Gemini 3.5 Transcribe Live's at 5.77% — partials are supposed to be worse than the final transcript, often by a point or more, because the model is committing before the sentence ends. A first partial that matches the final number is the genuinely unusual part of Microsoft's launch, and it is the part the tracker's own data cannot yet confirm. Cite it as a claim until a row exists.
The honest caveat is that 0.13 seconds and 0.16 seconds are the same product experience to a human reading captions, and 2.5% against the 2.73% currently leading the board is about 2 errors per 1,000 words. A lead that size is real but small, and whether it is a lead anyone can feel depends on the workload. Where it does bite is the tail: a contact-center stream running eight hours a day at 2.73% versus 2.5% is tens of thousands of extra wrong words a year, and word errors in a transcript are not evenly distributed — they cluster on exactly the proper nouns and numbers that make a transcript useful.

What $9.00 per 1,000 minutes buys
Streaming costs more than batch, and Microsoft priced it at the top of the real-time tier rather than under it. MAI-Transcribe-2-Streaming lists at $0.54 per audio-hour, which works out to $9.00 per 1,000 minutes of streamed audio, described as an introductory rate good through the end of the year. For comparison, the batch MAI-Transcribe-2 is $0.10 per audio-hour — $1.67 per 1,000 minutes — so real-time transcription costs roughly 5.4× the file-based rate for the same underlying family and 0.5 points more word error. That is not a markup Microsoft is hiding; it is the honest price of a persistent connection and a partial transcript that has to be right before the sentence is over.
Against the rest of the streaming tier the picture is mixed. MAI-Transcribe-2-Streaming ties Gemini 3.5 Transcribe Live at roughly $9 per 1,000 minutes — more accurate, same money. It sits above ElevenLabs Scribe v2 Realtime and Deepgram Flux at about $6.50 per 1,000 minutes, above Cartesia Ink-2 at about $4, and roughly three times Muse Voice Transcribe at about $3. So the launch is an accuracy-and-latency play at parity pricing, not a price war; teams already running a cheaper streaming model will be trading dollars for points of WER.
That trade is worth naming precisely, because it is the same trade the batch launch inverted. In September Microsoft made accuracy cheap. In October it made real-time accuracy cost the same as everyone else's. If your pipeline only needs transcripts eventually, nothing about October 1 changes your bill. If it needs them while the call is still happening, the calculus just moved to quality.

Building on a transcript is the other half of that decision, and it is where a multi-model layer earns its keep. A real-time transcript is rarely the end of the pipeline — it gets summarized, scored, searched, routed, and escalated, and those steps want different models at different costs. OrcaRouter exists for that layer: one API across 200+ models, provider list prices passed through at 0% markup, automatic failover, and a routing DSL that can send a live transcript to one model for compliance screening and another for meeting notes without a second integration. MAI-Transcribe-2-Streaming itself is not on that roster — it is served through Microsoft's own speech stack — but the streaming transcript it produces is exactly the kind of high-volume, latency-sensitive input that argues for keeping the layer above it model-agnostic.
What is not proven yet
Three things are genuinely open. First, diarization: the batch model bundles speaker attribution with no published diarization error rate, and Microsoft's streaming material does not make a diarization claim at all — for a real-time model feeding live agent-assist or captions, who-said-what is often the feature that matters more than raw WER. Second, the multilingual breakdown: 60 languages with automatic continuous language detection is a wide claim, and a streaming model's per-language latency can vary far more than its batch counterpart's, because partial-transcript quality depends on how quickly the model commits to a language. Microsoft has published no per-language streaming table. Third, streaming WER is measured on clean benchmark audio; the failure mode that hurts real deployments is a noisy far-field stream, and no independent far-field streaming comparison existed on launch day.
Where it leaves your stack
The interesting thing about this launch is not that Microsoft has the most accurate streaming transcript. It is the cadence: batch in September, streaming in October, in the same family, on the same docs surface, with a pricing structure that now spans $1.67 to $9.00 per 1,000 minutes for the same underlying capability. That gives teams a real choice for the first time — pay for real-time only where real-time matters, and batch everything else — but it also means the transcription layer is now something you tune per workload rather than standardize once.
Read the headline claim literally and it holds as a vendor statement: Microsoft says MAI-Transcribe-2-Streaming is first on both final-transcript and first-partial accuracy, and that it sits on the Pareto frontier. The asterisks are that the tracker's board has not yet plotted the model, the lead it claims over the current leader is small enough to disappear in a re-run, the price advantage is zero, and the diarization story has not been told. Those are the things to watch, not the crown.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
