Article hero card reading 'MAI-Transcribe-2 vs GPT Transcribe' with the badge 'MODEL COMPARISON' and the subtitle 'Microsoft undercuts OpenAI's transcript on every number', with chips comparing 2.0% vs 3.31% AA-WER, $1.67 vs $4.50 per 1,000 minutes and ~411x vs ~34x real time, over a white-to-blue gradient with the OrcaRouter logo bottom right.
Guides & Insights

MAI-Transcribe-2 vs GPT Transcribe: Microsoft Undercuts OpenAI's Transcript on Every Number

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Three numbers separate MAI-Transcribe-2 from GPT Transcribe, and every one of them points in the same direction. On Artificial Analysis's non-streaming word-error-rate leaderboard, MAI-Transcribe-2, released by Microsoft on September 3, 2026, measures 2.0% AA-WER against 3.31% for GPT Transcribe, OpenAI's pre-recorded transcription API that launched five weeks earlier on July 28, 2026. On the same board's speed factor, MAI-Transcribe-2 runs at roughly 411× real time against GPT Transcribe's roughly 34×. And on price, MAI-Transcribe-2 lists at $0.10 per audio-hour — about $1.67 per 1,000 minutes as a limited-time launch offer — against GPT Transcribe's $4.50 per 1,000 minutes. That last gap is a 63% undercut, and it lands on a product OpenAI repriced 25% downward just weeks ago. This is what a commodity market looks like the moment a second frontier lab decides to compete on price.

The gap, one number at a time

Accuracy first, because it is the number neither vendor can hide behind a marketing claim. The 2.0% and 3.31% come from the same independent harness, Artificial Analysis, run against both models — which makes the 1.3-point spread the cleanest fact in this whole comparison: roughly thirteen fewer errors per thousand words at the same WER line. GPT Transcribe's 3.31% was itself an improvement of about 0.7 points over its predecessor, GPT-4o Transcribe, and OpenAI's own multilingual runs showed it roughly halving Whisper-era error rates. Microsoft has now taken the same board's second spot outright, and on multilingual FLEURS it reports an average 5.2% WER across 60 languages — that figure is from Microsoft's launch post, not yet independently re-derived per language.

Speed next. GPT Transcribe's roughly 34× real time is typical for an asynchronous transcription endpoint: a one-hour file in about two minutes of compute. MAI-Transcribe-2's roughly 411× real time transcribes the same hour in under nine seconds of compute. On the Artificial Analysis numbers the true gap is closer to 12×; Microsoft markets it as "10× faster than OpenAI's GPT-Transcribe," which is the vendor rounding down its own advantage rather than up — an unusual tell. The honest caveat is that MAI-Transcribe-2's figure is four days old and batch-file throughput, not live latency, while GPT Transcribe's has had five weeks of production traffic behind it. But nothing about the architecture of the two products suggests the speed gap will close; they are built for the same batch job with very different efficiency.

Price last, because it is the number that changes decisions. GPT Transcribe lists at $0.0045 per minute — $4.50 per 1,000 minutes, about $0.27 per audio-hour — after OpenAI's July cut from GPT-4o Transcribe's $0.006. MAI-Transcribe-2 lists at $0.10 per audio-hour until the end of the year, with no standard price disclosed for after the promotion. On 1,000 hours of audio a month, the spread is about $170: roughly $100 at MAI-Transcribe-2's early-bird rate against roughly $270 at GPT Transcribe's list price. OpenAI's 25% cut in July looked aggressive; Microsoft's 63% undercut of the post-cut price is a different category of move, and it is the whole reason this comparison exists.

A two-column scoreboard titled 'MAI-Transcribe-2 vs GPT Transcribe — the scoreboard': left column MAI-Transcribe-2 lists 2.0% AA-WER, ~411x real time, $1.67 early-bird per 1,000 minutes, 60 languages, bundled diarization and verbatim/clean styles; right column GPT Transcribe lists 3.31% AA-WER, ~34x real time, $4.50 per 1,000 minutes, languages not published, no diarization and caller-side post-processing; footer notes AA-WER and speed per Artificial Analysis and vendor list prices.

Why OpenAI's number was the credible one

There is a reason GPT Transcribe's 3.31% AA-WER carried more weight than most launch numbers: it had been independently audited, and the audit survived contact with hard audio. The same Artificial Analysis run that produced the headline figure also surfaced GPT Transcribe's weak spot — a 6.10% WER on the Earnings22 earnings-call set, the hardest public benchmark in the category — which told buyers exactly where not to point it. Five weeks of production traffic followed, and OpenAI's price cut showed it was competing on cost rather than defensively protecting margin. GPT Transcribe also inherits the practical constraint of its API: it is a batch-only product, files up to 25 MB per request, with no streaming tier at the $4.50 price, no diarization, no vocabulary biasing, and languages OpenAI has not published. It transcribes; everything above the words is the caller's problem.

That clarity is what made GPT Transcribe easy to trust, and it is exactly the clarity a four-day-old model has not earned yet. Microsoft has published no diarization error rate for MAI-Transcribe-2, no per-language accuracy table behind its 60-language FLEURS average, and no streaming SKU. None of those omissions means the model is weak — bundled diarization and a 60-language footprint are real product surface area that GPT Transcribe does not even offer — but each is a place where an independent rerun could revise the story. The accuracy number is AA-measured and real; the product around it is unproven in exactly the ways that only production traffic reveals.

The feature sets are not the same shape

• Diarization — MAI-Transcribe-2 bundles speaker attribution with no published error rate; GPT Transcribe does not offer it at all, so the transcript arrives as one undifferentiated stream.

• Control — MAI-Transcribe-2 offers keyword biasing for names and jargon plus verbatim and clean transcription styles; GPT Transcribe offers neither, returning raw words and punctuation-free text for the caller to post-process.

• Timestamps — both return time-aligned output; MAI-Transcribe-2's are word-level out of the box.

• Languages — MAI-Transcribe-2 publishes 60 with automatic identification; GPT Transcribe does not publish its language list, which is itself a planning problem for multilingual workloads.

• Streaming — GPT Transcribe's $4.50 tier is batch-only, and its streaming sibling, GPT Live Transcribe, is a separate product at $17 per 1,000 minutes; MAI-Transcribe-2 has no streaming product at all, announced or demonstrated.

• Shape — both are managed APIs for pre-recorded files, which is why this matchup is fair; the difference is that Microsoft put more of the useful post-processing into the base product at one-third the price.

Screenshot of the Artificial Analysis Speech-to-Text leaderboard summary table showing Word Error Rate and Median Speed Factor columns for models including MAI-Transcribe-2 and MAI-Transcribe-1.5 under Microsoft AI, alongside rows for ElevenLabs Scribe v2 and Gemini 3.5 Transcribe.

Who should move, and when

If you transcribe high-volume batch audio and the early-bird rate survives contact with your own files, MAI-Transcribe-2 is the rational default right now: better measured WER, roughly twelve times the throughput, and a 63% lower per-1,000-minute price on a product that bundles diarization and style control OpenAI charges nothing extra for — because it does not offer them. The move costs little: a test against a slice of your own hardest audio settles most of the uncertainty, and the $1.67 rate makes the experiment nearly free through year-end.

If your workload is production-critical, regulated, or already wired around OpenAI, the five-week head start and the published Earnings22 weakness are worth more than the price gap. GPT Transcribe is the model whose failure modes you can name; MAI-Transcribe-2's are still being discovered. The prudent pattern is the one a routing layer exists to support: keep the proven model as the default and send a measured slice of traffic to the challenger, comparing transcripts on your own data before promoting anything. That trial-and-promote discipline is the whole point of OrcaRouter's automatic failover and routing DSL — new models earn production traffic in percentages, not by decree, and the same one-key setup that reaches 200+ models makes the rest of the stack above the transcript provider-agnostic while the STT contest plays out.

The launch post behind MAI-Transcribe-2's numbers frames everything as a comparison to named rivals rather than as absolute figures — which is a useful discipline check for anyone reading it: take the relative claims, find the absolute numbers on the independent board, and label the rest vendor framing.

Screenshot of the top of the Microsoft AI announcement for MAI-Transcribe-2 (dated September 3, 2026), showing the headline 'MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world' and the opening paragraph naming new features — diarization, configurable transcription styles, word-level timestamps — and comparisons against Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large and Scribe V2.

The $1.67 rate has an expiration date and the 2.0% has a four-day reputation, and OpenAI will answer both eventually. But the shape of this market just changed: for the first time, the cheapest high-accuracy managed transcription API is also the most accurate and the fastest, and it is not OpenAI's. Even if MAI-Transcribe-2's standard price lands well above the launch offer, the accuracy-per-dollar bar in managed transcription has moved — and it moved on Microsoft's terms.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily