
Voxtral Mini 4B Realtime Arabic vs Gemini 3.5 Transcribe Live: Dialects vs Breadth
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 118 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 366 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two streaming speech-to-text models, two completely different bets about what the hard part of transcription is. Voxtral Mini 4B Realtime Arabic is a 4.4-billion-parameter fine-tune that Mistral AI pushed to a public repository on October 8, 2026 without a launch post, a blog entry or a line in the company news index, trained specifically on Modern Standard Arabic and fifteen named dialects, released under Apache 2.0 with downloadable BF16 weights. Gemini 3.5 Transcribe Live is Gogle's streaming endpoint of Gemini 3.5 Transcribe, announced in August 2026 with the full apparatus of a platform release, covering 85-plus locales, and closed — a preview API with no weights and a ten-minute session cap. One model went deep on a single language family and said so in its own limitations section; the other went wide and left the accent-level guarantees unstated. Which one you want depends on an axis that neither vendor's marketing puts first: what your audio actually sounds like.
This is not a symmetric comparison, and it is worth saying that up front rather than burying it. The Mistral model is 48 hours old at the time of writing, has no vendor announcement, no independent evaluation, and no published price. The Google model has been in public preview since late August and is measured on a third-party tracker. Treat the Mistral side as a set of claims from a model card and the Google side as a set of measurements plus a set of claims, and the comparison stays honest.
What each model is, before the numbers
• Voxtral Mini 4B Realtime Arabic — a streaming ASR model fine-tuned from Voxtral-Mini-4B-Realtime-2602, roughly 4.4B parameters in BF16, taking 16 kHz audio through a causal audio encoder and a language decoder. It emits text as audio arrives, with a configurable transcription delay, and it is Apache 2.0 with weights on Hugging Face. Mistral co-developed it with Morocco's Ministère de la Transition Numérique et de la Réforme Administrative under a partnership signed in January 2026, and describes the model as a sovereign AI asset.
• Gemini 3.5 Transcribe Live — the streaming half of a two-model release. The Live API endpoint carries continuous microphone audio over a WebSocket, returns partial transcripts while the speaker is still talking, and settles on a final transcript shortly after each utterance ends. The batch sibling, gemini-3.5-transcribe, serves pre-recorded files up to an hour. Both are in public preview, both are API-only, and neither has published weights.
The first structural difference is not accuracy. It is that one of these is a file you can download tonight and the other is a service you have to be accepted into using.

Language coverage: 16 against 85-plus, and the gap is not the story
Counting languages makes the Mistral model look narrow and the Google model look enormous. The counts measure different things, and reading them side by side without that caveat is the most common mistake this comparison invites.
• Voxtral Mini 4B Realtime Arabic — 16 Arabic varieties, named individually on the model card by dialect family: Modern Standard Arabic, plus Moroccan, Libyan, Tunisian, Algerian and Hassaniya Arabic from the Maghreb; Gulf, Najdi, Omani and Sanaani from the Gulf; Egyptian and Sudanese from the Nile Valley; Mesopotamian and both South and North Levantine; and Chadian Arabic. The card's own framing is that the model targets Arabic transcription, and that code-switching — speakers alternating languages mid-conversation — is the capability it was built to get right rather than a side effect.
• Gemini 3.5 Transcribe Live — automatic language detection across 85-plus locales, with intra-sentence and cross-sentence code-switching. Google's documentation lists Arabic within that set; the locale table names Arabic (Egypt) as the Arabic entry, and the wider Arabic locale coverage is not itemised in the model's own documentation.
Read that honestly and the shape of the trade appears. Google's 85-plus is a breadth claim: if your audio is in a language other than Arabic, the Google endpoint covers it and the Mistral model will refuse it — the card says outright that for other languages you should use the original Voxtral Mini 4B Realtime, not this checkpoint. Mistral's 16 is a depth claim. It is not claiming to transcribe Arabic; it is claiming to transcribe Moroccan Darija, Gulf Arabic, Egyptian Arabic and Chadian Arabic as distinct targets, which is a materially harder problem than transcribing Modern Standard Arabic and hoping the dialect falls out of it. An English-weighted, breadth-first benchmark will never show you that difference, because the dialect sets are not on it.
For a Moroccan call centre or a Gulf customer-support line, a model that handles 16 dialects and nothing else can beat a model that handles 85 locales at an unstated per-dialect accuracy. For a European company transcribing meetings in five languages, the equation reverses instantly. Neither vendor is wrong; they picked different customers.

The numbers you can compare, and the ones you cannot
Here is where the asymmetry bites. The two models report different units, on different corpora, with different evidence behind them, and no third party has run them against each other.
• Voxtral Mini 4B Realtime Arabic — 8.82% average Character Error Rate across seven Arabic benchmarks, at a configured transcription delay of 480 milliseconds. Per Mistral's own card, and not independently reproduced.
• The model it is chasing — Voxtral Transcribe Arabic at 7.91% CER on the same seven benchmarks by the same measurement, so the realtime model lands 0.91 percentage points behind an offline Arabic model from the same vendor. That gap is Mistral's own comparison, which makes it unusually informative: the vendor is telling you what streaming costs in accuracy on this language family.
• Gemini 3.5 Transcribe Live — 4.0% streaming word error rate, measured by Artificial Analysis and cited by Google, with a final transcript arriving 0.40 seconds after the end of speech. Also measured: 2.6% on the same tracker's non-streaming board, and 5.50% streaming on FLEURS per Google's own reporting.
Do not put 8.82% CER next to 4.0% WER and conclude anything. Character error rate and word error rate are different denominators — Arabic orthography makes the gap between them larger than it is for Latin-script languages — and the corpora are not the same, the normalisation is not the same, and one number comes with a reproducible script while the other comes from a tracker's fixed mix that is weighted toward English agent-talk.
The more useful comparison is the one inside Mistral's own card: 8.82% streaming against 7.91% offline, on identical benchmarks with identical normalisation. That is a clean measurement of what low-latency buys and costs on Arabic specifically, and it is worth more than any cross-vendor percentage. What nobody has published is a like-for-like Arabic run of the Google and Mistral models on the same audio. Until someone does, "which is more accurate in Arabic" is an open question, and the only defensible answer is: try both on your recordings.
One detail in Mistral's card deserves a mention because it cuts against the vendor's own interest. It notes that Gemini's safety filters block transcription of some FLEURS audio samples. That is a real, checkable observation about a competitor's pipeline, and it is also exactly the kind of thing that never appears on a leaderboard — a model that refuses to transcribe a sample scores as a failure or as absent, and neither reading is fair. If your audio includes material a content filter might catch, that is a deployment risk no WER figure will surface.
Latency: a dial against a fixed number
Streaming models are usually compared on a single latency figure. These two do not have one in common.
• Voxtral Mini 4B Realtime Arabic — configurable transcription delay. The 8.82% CER figure is quoted at 480 milliseconds, and the delay is adjustable, which means the accuracy number is one point on a curve the operator chooses rather than a property of the model. Its base model advertises a range from 240 milliseconds up to 2.4 seconds.
• Gemini 3.5 Transcribe Live — 0.40 seconds from end of speech to a final transcript, measured by Artificial Analysis, with partials arriving continuously during speech. Google separately claims a 70% improvement in time-to-final-transcription over Chirp 3, which is a vendor figure and unreproduced.
Both land in the same neighbourhood — roughly half a second — but the engineering meaning differs. The Mistral model hands the latency-accuracy trade to you as a parameter, so you can run at 240 ms when sub-second captions matter more than the last few points of accuracy and push the delay out when they do not. The Google endpoint presents a fixed operating point with Google's own content-filter behaviour attached. For a voice agent, a dial is worth more than a slightly better fixed number, because the right setting depends on your audio and your users and neither vendor can pick it for you.
The real divide: open weights against a preview endpoint
The licence and distribution difference is larger than any benchmark gap in this comparison, and it decides most real deployments before accuracy is discussed.
• Voxtral Mini 4B Realtime Arabic — Apache 2.0, BF16 weights on Hugging Face, servable with vLLM over a WebSocket at /v1/realtime, and natively supported in Transformers from version 5.2.0. The card ships a Python example and a normalisation module so you can reproduce the Arabic CER computation yourself. Nothing about the pipeline requires Mistral's servers, and the model is small enough that the operational question is which GPU, not whether you can afford one.
• Gemini 3.5 Transcribe Live — API only, public preview, no published price commitment beyond list rates, and a ten-minute session cap that means anything longer has to be chunked, reconnected and stitched by your own code. Three constraints reported alongside the launch matter for Arabic deployments specifically: the streaming endpoint returns no speaker diarization and no word-level timestamps, both of which exist on the batch endpoint but not this one. If your Arabic transcript needs to know who said what, or needs to be time-aligned to audio, the Live endpoint cannot produce it regardless of language.
Those two constraints are the strongest argument for the small open model in a real Arabic deployment, and they have nothing to do with accuracy. A call-centre pipeline wants speaker attribution, a compliance pipeline wants timestamps, and an offline Arabic model with a normalisation script you control gives you both without arguing with an endpoint contract. Mistral's model card also lists that as a limitation rather than a feature — it targets transcription, it is a fine-tune of a general realtime model, and it is honest that a broader model exists upstream if you need one.
What each one costs, and what is unknown
• Gemini 3.5 Transcribe Live — roughly $0.009 per minute of audio blended, derived from list prices of $3.50 per million input tokens and $21 per million output tokens, with audio assessed at 25 tokens per second and the transcript at about 175 tokens per minute. The batch endpoint is about $0.005 per minute. Both figures are estimates from Google's assumed tokenisation, so they move if the accounting changes. There is a free tier today.
• Voxtral Mini 4B Realtime Arabic — no published price, because there is no published service. Cost is whatever your hardware costs. A 4.4-billion-parameter BF16 model fits on a single modern accelerator, so the marginal cost per audio-hour is electricity and utilisation, and the fixed cost is the machine. That is cheaper than any per-minute API at volume and more expensive than any per-minute API at zero volume, which is the ordinary shape of the open-weights trade.
The unknown that matters most on the Mistral side is not price. It is whether this repository is the beginning of a product line or a one-off deliverable against the Morocco partnership. A model that ships quietly with no announcement, no pricing and no service is a model with no support commitment behind it. Apache 2.0 means the licence cannot be withdrawn for the weights you already have, but it says nothing about whether the checkpoint gets maintained, and the card's own base-model pointer suggests the general realtime model is still the supported line.
For the layer above the transcript, a routing layer earns its keep in a way that has nothing to do with transcription. Neither of these models is a speech-to-text service we host — OrcaRouter routes neither endpoint — but the summariser, the intent classifier or the agent that consumes the stream is a model you can route: one API key across more than 200 models at provider list price with 0% markup, so a vendor price cut lands on our side the same day, plus automatic failover if a provider degrades. If you are already stitching two speech vendors together for dialect coverage, putting the downstream language models behind one key rather than two contracts is the cheap half of the integration.
Which one to build on
If your audio is Arabic — one dialect, several, or code-switched between Arabic and something else — the Mistral model is the more serious candidate, and the reason is structural rather than numerical. It names the dialects it was built for, it is small enough to run on your own hardware, it returns raw text you can post-process with your own normalisation, and it does not sit behind a ten-minute session cap or a content filter you cannot inspect. The cost is that it does one language family and refuses the rest, and that nobody has published an independent evaluation of it yet.
If your audio spans languages, or if you need a service with a support path and a published rate card today, Gemini 3.5 Transcribe Live is callable now, measured on a third-party tracker, and covers the languages the Mistral checkpoint will reject. The costs are the session cap, the missing diarization and timestamps on the streaming endpoint, and the fact that the Arabic-specific accuracy is a claim you will have to test yourself — the locale list names Egypt, and dialect performance is not itemised.
The thing to watch is not a benchmark. It is whether Mistral announces this model at all, and if so, with what price and what language roadmap. A quiet repository with an Apache 2.0 licence is a useful artifact and a weak commitment. If an Arabic dialect roadmap follows — Levantine, Gulf, Egyptian as separate checkpoints — the small-model route becomes a genuinely complete answer for the region, and the breadth argument gets much harder to make.

Neither of these is a speech model we host, but the layer above the transcript is routable - more than 200 models behind one API key at provider list price with 0% markup, so a vendor price cut on the reasoning model downstream of your transcript is live the same day.
Automatic failover means a stitched speech pipeline has one fewer failure domain, because the summariser or intent classifier consuming the transcript can be rerouted instead of going down with a provider.
