
Voxtral Mini 4B Realtime Arabic vs MAI-Transcribe-2-Streaming: Two October Arrivals, Two Evidence Problems
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 118 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 366 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Seven days separate these two real-time speech models, and they arrived in opposite ways. MAI-Transcribe-2-Streaming is Microsoft AI's first streaming transcription model, announced on October 1, 2026 with a launch post, an Azure preview listing, a price of $0.54 per audio-hour introductory through the end of the year, and a claim to be first on Artificial Analysis's streaming board for both final and partial transcript accuracy at 2.5% word error rate with the final text ready 0.13 seconds after end of speech. Voxtral Mini 4B Realtime Arabic is Mistral AI's Arabic-dialect streaming model, published to Hugging Face on October 8, 2026 with no announcement at all — no blog post, no price, no service, just Apache 2.0 weights and a model card reporting 8.82% average Character Error Rate across seven Arabic benchmarks at a 480-millisecond delay. The Microsoft model is a finished product with an unverifiable headline. The Mistral model is a verifiable artifact with no product around it. Both are worth understanding precisely because both leave the central question — how well does this handle Arabic — answered only by the vendor.
The comparison is worth making anyway, because the two models frame the same engineering decision from opposite ends. Microsoft is selling breadth and a managed service: 60 languages with automatic continuous language detection, running inside Azure AI Speech, priced per hour, in preview. Mistral is giving away depth: sixteen named Arabic varieties, small enough to run on one accelerator, with a normalisation script so you can audit the scoring yourself. If you are standing up an Arabic transcription pipeline this month, you are choosing between those two positions, and the choice turns on evidence you do not have yet in either case.
What each vendor actually said, and what the boards say
The evidence gap is the most important feature of this matchup, so it goes first.
• MAI-Transcribe-2-Streaming — 2.5% final-transcript word error rate at 0.13 seconds after end of speech, and the same 2.5% on first partials at 0.12 seconds, both presented as first place for accuracy on Artificial Analysis's AA-WER Streaming methodology. Microsoft's own framing. As of the drafting of this piece, the tracker's streaming board has not plotted a row for the model, so the claim sits on the tracker's methodology without a tracker-published number behind it.
• Voxtral Mini 4B Realtime Arabic — 8.82% average Character Error Rate across seven Arabic benchmarks at a 480-millisecond configured delay. Mistral's own card, with the evaluation code supplied. No third party has reproduced it, and the model is two days old.
So the two headline numbers in this article are, respectively, a vendor claim on someone else's ruler and a vendor claim with its own ruler attached. Neither is an independent measurement. What differs is that Mistral's number arrives with the normalisation script you need to re-derive it, and Microsoft's arrives with a board that has not yet put it on the chart. If you are making a procurement decision, that distinction matters more than the 2.5-versus-8.82 spread, which is meaningless anyway — word error rate and character error rate do not convert, and Arabic orthography widens their gap considerably.
The one comparison inside Mistral's card that does convert is the one against its own offline sibling: Voxtral Transcribe Arabic at 7.91% CER on the same seven benchmarks with the same normalisation, so streaming costs this model 0.91 percentage points. Microsoft published an equivalent figure for its own family when it shipped the batch MAI-Transcribe-2 on September 3, 2026 at 2.0% word error rate — the streaming variant gives up half a point against the batch model while adding real-time delivery. Read those two vendor-internal deltas side by side and you have the only apples-to-apples statement in this article: on their own benchmarks, each lab's streaming SKU trails its batch sibling, by about one point in one case and half a point in the other, and both are measuring different languages.

Language coverage: 60 against 16, and the same unnamed gap on both sides
• MAI-Transcribe-2-Streaming — 60 languages with automatic continuous language detection, so a session that switches languages mid-stream does not need the language declared up front. Microsoft's launch material gives the count and not the list; the Azure language-support documentation for the MAI-Transcribe family is where the coverage is itemised, and it documents Arabic among the supported languages for the family.
• Voxtral Mini 4B Realtime Arabic — sixteen Arabic varieties, named individually: Modern Standard Arabic, Moroccan, Libyan, Tunisian, Algerian and Hassaniya Arabic, Gulf, Najdi, Omani and Sanaani Arabic, Egyptian and Sudanese Arabic, Mesopotamian and both South and North Levantine Arabic, and Chadian Arabic. Outside that set the model is documented as out of scope, with a pointer to the general Voxtral Mini 4B Realtime.
Neither number tells you what you need. Microsoft's 60 is a breadth count that includes Arabic without describing Arabic performance; the launch post states the language total once and moves on. Mistral's 16 is an enumeration of training targets with a single aggregate error rate across seven benchmark sets, which means the dialect-level breakdown — the thing that actually determines whether your Moroccan call audio transcribes — is not published either. Two models, two vendors, and in both cases the honest answer to "how good is it at Gulf Arabic specifically" is: not stated.
What the enumeration does give you is a prior. A model with Chadian Arabic and Hassaniya Arabic as named targets was tuned against a North African distribution; Mistral co-developed it with Morocco's Ministère de la Transition Numérique et de la Réforme Administrative and built two of the seven benchmarks with that ministry — ISMA and Darija in the Wild — specifically to measure the hard cases. A general 60-language model improves on Modern Standard Arabic first, because that is where the volume of training signal sits. If your audio is Darija or Gulf Arabic, those are different problems, and only one of these two vendors has put a number on either.
Latency and the shape of the delay
• MAI-Transcribe-2-Streaming — 0.13 seconds from end of speech to final transcript, and first partials at 0.12 seconds, which is fast enough that the partial and the final are effectively the same latency. Microsoft's claim, measured on the tracker's methodology.
• Voxtral Mini 4B Realtime Arabic — 480 milliseconds of configured delay at the published accuracy point, with the delay adjustable. That is roughly three and a half times the Microsoft figure, and it is the parameter the operator chooses rather than a fixed property.
Three-tenths of a second is a real difference for a voice agent that needs to respond mid-utterance and close to invisible for a human reading captions. It is also the difference between a dial and a number. Mistral's delay parameter means you can run tighter and accept more errors or looser and take the accuracy back — the base model this checkpoint was fine-tuned from advertises a range from 240 milliseconds up to 2.4 seconds — while Microsoft's operating point is what Azure ships. That makes Microsoft's figure easier to reason about and Mistral's easier to tune, and it means any comparison of the two accuracy numbers is really a comparison of two chosen points on one curve and one fixed point on another.
The feature surface differs just as sharply, and it is worth checking against your pipeline requirements before the accuracy discussion gets interesting.
• MAI-Transcribe-2-Streaming — delivered through Azure AI Speech with the family's documented feature set: diarization, word-level timestamps, keyword and phrase biasing, verbatim versus clean output styles, and automatic language identification. The streaming SKU's own documentation for diarization and timestamps is thinner than the batch model's, so confirm those per-endpoint rather than assuming parity.

• Voxtral Mini 4B Realtime Arabic — the card documents transcription with a causal audio encoder, vLLM serving over a WebSocket at /v1/realtime, native Transformers support from version 5.2.0, and a normalisation module. It does not document diarization or word-level timestamps for this checkpoint.
Delivery model: a preview service against downloadable weights
• MAI-Transcribe-2-Streaming — Azure preview, which means no SLA and a vendor recommendation against production dependence. Priced at $0.54 per audio-hour as an introductory rate running through the end of 2026, with the standard rate not published; the batch MAI-Transcribe-2 launched at $0.10 per audio-hour on the same limited-time basis. Streaming costs 5.4 times the batch rate.
• Voxtral Mini 4B Realtime Arabic — Apache 2.0, approximately 4.4 billion parameters in BF16, downloadable, fine-tunable, and servable on a single accelerator. No price, no SLA, and no support commitment, because there is no service behind it.
Those two sentences contain the decision. A preview service with an introductory rate and an undisclosed standard price is a commitment that expires; the 5.4× streaming premium and the through-2026 window are both Microsoft's to change, and the batch-to-streaming ratio is the number to watch when the standard rate lands. Apache 2.0 weights with no announcement are the opposite: an artifact nobody can take back, tied to a vendor relationship nobody has described. Whether the checkpoint gets maintained is unknowable, but the file you have today keeps working regardless.
The industry-specific constraint is the one that usually decides these questions in practice. An Arabic call-centre deployment inside a regulated institution may not be able to send audio to a preview cloud endpoint at all; a team that has already standardised on Azure AI Speech may not want a second on-premises stack for one language family. Both constraints are legitimate, and neither has anything to do with the 2.5% and 8.82% figures that headline the two launches.
Above the transcription layer, the same point applies to both. OrcaRouter routes neither of these speech models — this is a comparison of two systems we do not serve — but the summariser, the classifier or the agent that consumes the transcript is routable: more than 200 models on one API key at provider list price with 0% markup, so a vendor price change reaches our side the same day, and automatic failover if a provider degrades. Where that helps specifically here is the trial phase: pointing both transcription candidates at one downstream pipeline, behind one key, is how you run the head-to-head without standing up two integrations.
The test that would settle this
Neither vendor has published Arabic-specific performance, and neither has been independently evaluated on dialectal audio. The comparison therefore reduces to three facts and one recommendation.
The facts: Microsoft's streaming model is faster at 0.13 seconds and claims better aggregate accuracy on a board that has not yet plotted it; Mistral's model publishes a dialect list, an Arabic error rate, and the normalisation code to re-derive it, at 480 milliseconds; and the two accuracy figures cannot be compared because one is word error rate and the other is character error rate on a different corpus.
The recommendation: whoever is making this decision should run both on their own audio, because that is the only measurement either vendor has left open. Build the evaluation set from real recordings — the dialect mix you actually receive, the codec you actually use, the domain vocabulary you actually need — and score both with Mistral's normalisation against a reference you trust. A two-hour test on your own recordings will tell you more than both launch posts combined, and it is the only way to find out whether the 0.13-second advantage is worth a preview dependency and a 5.4× streaming premium, or whether a downloadable 4.4B model at 480 milliseconds is already good enough for the job.

OrcaRouter serves neither of these speech models, and this piece does not suggest otherwise - what it does cover is the language layer that reads the transcript: more than 200 models behind one key at provider list price with 0% markup, so a vendor price change on the downstream model reaches you the day it is announced.
Automatic failover lets both transcription candidates feed one downstream pipeline instead of two integrations, which is also the cheapest way to run the head-to-head.
