Article hero card reading 'MAI-Transcribe-2 Is Live' with the badge 'NEWS · MICROSOFT AI' and the subtitle 'Microsoft's $0.10-an-hour bid to own speech-to-text', with a stat strip showing AA-WER 2.0% (#2 per Artificial Analysis), speed ~411x real time (#2) and a $0.10-per-hour launch price, over a white-to-blue gradient with the OrcaRouter logo bottom right.
Guides & Insights

MAI-Transcribe-2 Is Live: Microsoft's $0.10-an-Hour Bid to Own Speech-to-Text

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MAI-Transcribe-2 is the model Microsoft AI is betting turns speech-to-text into a commodity rather than a premium feature, and the numbers it shipped with are engineered to make that argument for it. Released on September 3, 2026 into public preview on Microsoft Foundry, the MAI Playground, and Microsoft's own API surfaces, MAI-Transcribe-2 is the third Microsoft speech model in five months — after MAI-Transcribe-1 in April at $0.36 per audio-hour and MAI-Transcribe-1.5 in June — and the first one that leads every managed rival it names on the two metrics teams actually shop on. On Artificial Analysis's non-streaming word-error-rate leaderboard it sits at 2.0% AA-WER, in second place, and at roughly 411× real time, also second, which together put it on the accuracy-speed Pareto frontier: on that board, no model is both more accurate and faster. The launch price undercuts its own predecessor by about 72%.

The "fastest, most accurate and cheapest speech recognition model in the world" is Microsoft's own headline for the release, and it deserves to be read with the source material's asterisks attached. This is the launch story as it actually happened, with the vendor claims labeled and the parts that still need proof separated out.

What Microsoft actually shipped

MAI-Transcribe-2 is a pre-recorded, batch-oriented transcription model aimed squarely at long-form audio — meetings, calls, media archives, contact-center recordings. Microsoft positions it for exactly the workloads where throughput and per-minute cost dominate. It transcribes in 60 languages with automatic language identification, includes speaker diarization that attributes words to the right person, returns word-level timestamps, and supports keyword biasing for names, abbreviations, and domain terminology. Two configurable transcription styles cover the usual split: "verbatim," which keeps fillers and false starts for compliance-grade transcripts, and "clean," which strips disfluencies for captions and notes. Microsoft also calls out code-switching in mixed-language speech such as Hinglish and Spanglish, and robustness in noisy real-world audio.

Screenshot of the top of the Microsoft AI announcement for MAI-Transcribe-2 (dated September 3, 2026), showing the headline 'MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world' and the opening paragraph naming new features — diarization, configurable transcription styles, word-level timestamps — and comparisons against Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large and Scribe V2.

The availability story matters as much as the feature list. Microsoft is not gating this behind its enterprise speech stack: the model is demoable today from the MAI Playground and served through Microsoft Foundry in public preview, with the Foundry documentation living under the Azure AI Speech service. That is a deliberate distribution choice — put the frontier model where a developer can hit it with a key, not where an enterprise sales cycle has to approve it first. Microsoft has not published weights, and nothing in the launch material suggests it will.

Reading the headline literally

"Fastest, most accurate and cheapest in the world" is three claims of different strengths, and the launch post itself quietly downgrades two of them. The honest version of each:

• Accuracy — On Artificial Analysis's leaderboard, MAI-Transcribe-2 ranks second at 2.0% AA-WER, not first; Microsoft's own body text says second. The "most accurate" phrasing only survives if you read it as "among managed batch APIs that track at this level," and even then it is a claim about a four-day-old model, not a settled crown. Microsoft separately reports it first on the FLEURS multilingual benchmark at a 5.2% average WER across its 60 languages — that figure is from Microsoft's launch post, not yet independently re-derived per language.

• Speed — Per Artificial Analysis, MAI-Transcribe-2 runs at roughly 411× real time, second on that board. Curiously, Microsoft's own announcement never prints the absolute number; it leads instead with friendlier relative multipliers — "up to 10× faster processing than leading competitors, especially for long-form audio," and specifically 10× faster than OpenAI's GPT-Transcribe, 7× faster than ElevenLabs' Scribe v2, and 5× faster than Gemini 3.5 Transcribe. Those ratios are Microsoft's framing of the same Artificial Analysis data, and the base rates matter: "5× faster than Gemini 3.5 Transcribe" sounds like a modest gap until you translate it into minutes of compute per hour of audio.

• Cheapest — True only inside two fences Microsoft states clearly. The $0.10 rate is a "limited-time offer until the end of the year," and no post-promotion standard price is disclosed anywhere in the launch material. And it is cheapest among high-accuracy managed APIs; a self-hosted open model like Whisper Large v3 Turbo still transcribes for effectively the cost of electricity. The commodity argument is real — it is just narrower than the headline.

Screenshot of the Artificial Analysis Speech-to-Text leaderboard summary table showing Word Error Rate and Median Speed Factor columns for models including MAI-Transcribe-2 and MAI-Transcribe-1.5 under Microsoft AI, alongside rows for ElevenLabs Scribe v2 and Gemini 3.5 Transcribe.

What $1.67 per 1,000 minutes buys

Priced as $0.10 per audio-hour, MAI-Transcribe-2 comes to about $1.67 per 1,000 minutes of audio. The comparison that makes the number land is against the other managed transcription APIs that compete on accuracy. GPT-Transcribe lists at $4.50 per 1,000 minutes; Gemini 3.5 Transcribe's pre-recorded tier works out to roughly $5 per 1,000 minutes on Google's token-based pricing; ElevenLabs' Scribe v2 lists higher still. On 1,000 hours of audio a month — a serious but not enormous contact-center or media workload — the spread between MAI-Transcribe-2 at the early-bird rate and GPT-Transcribe at list is on the order of $170 a month, every month, before any accuracy difference. That is the price gap that makes transcription stop being a line item teams optimize and start being a line item they ignore.

Two caveats belong in the same breath. First, a promotional price is a pricing experiment until it is not: teams that build a product on the $0.10 rate should know the standard rate is undisclosed, and the honest planning number for a 2027 budget is somewhere between the launch offer and whatever Microsoft names when the offer expires. Second, "cheapest at this accuracy tier" says nothing about total cost of ownership — the model is API-only, so the comparison that matters for high-volume teams is $1.67 per 1,000 minutes against the fully loaded cost of running their own open-weights stack, not just against other APIs.

Compressed to a scoreboard, the launch position looks like this — every figure below carries the sourcing label attached to it in the body above.

A single-column scoreboard titled 'MAI-Transcribe-2 — the scoreboard' listing AA-WER 2.0% (#2 per Artificial Analysis), speed ~411x real time (#2), price $1.67 per 1,000 minutes early-bird through 2026, 60 languages with automatic language ID, bundled diarization with unpublished error rate, and no streaming SKU announced, with the OrcaRouter logo bottom right.

What is not proven yet

For all the specificity of the launch claims, three things are genuinely open. The first is streaming: MAI-Transcribe-2 is a batch model, Microsoft's speed story is about file throughput, and no real-time SKU has been announced or demonstrated — the "411×" is not a live-transcription latency number. The second is diarization quality: speaker attribution is bundled, but Microsoft publishes no diarization error rate, and that is the feature most likely to look worse in independent testing than the WER does. The third is the multilingual breakdown: a single 60-language FLEURS average can hide wide per-language variance, and Microsoft has not published the per-language table. Each is a reason the story is not finished on day four.

Where it leaves your transcription stack

None of this requires choosing MAI-Transcribe-2 today, which is exactly why the pricing deserves attention from teams that are not in the market for a new vendor. Speech-to-text is rarely the end of a pipeline — transcripts get summarized, actioned, searched, and routed, and that downstream layer is where a multi-model setup earns its keep. A platform like OrcaRouter exists for that half of the stack: one API across 200+ models, provider list prices passed through at 0% markup, automatic failover, and a routing DSL that lets a transcript feed a different model per workload — meeting minutes to one summarizer, compliance review to another — without a second integration. MAI-Transcribe-2 itself is not on that roster today; it lives on Microsoft's own API and the third-party platforms Microsoft lists. But the commodity price it just set is the best argument yet for building the part above the transcript to be model-agnostic.

The launch price is the tell. Microsoft is not trying to win speech-to-text on features alone — it is trying to make high accuracy so cheap that the category stops being a line item. MAI-Transcribe-2 deserves the attention it is getting; just read "fastest, most accurate and cheapest in the world" with the three asterisks attached: second on AA-WER, promotional pricing through year-end, and a streaming story that has not been told yet.