
Voxtral-Mini-4B-Realtime-Arabic: Mistral Shipped an Arabic Dialect ASR Model and Said Nothing
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 118 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 366 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Fifty-nine seconds is the entire public announcement for Voxtral-Mini-4B-Realtime-Arabic. At 11:36:06 UTC on 8 October 2026 a repository named mistralai/Voxtral-Mini-4B-Realtime-Arabic appeared on Hugging Face. At 11:37:05 — fifty-nine seconds later — a second repository, mistralai/LIDstral-Arabic, appeared beside it. Both carry the same modification timestamp, both sat at zero downloads and three likes when this was written on 10 October, and neither has a launch post, a pricing page, a blog entry or a line in Mistral's news index behind it. Voxtral-Mini-4B-Realtime-Arabic is a streaming speech-to-text model for Arabic — Modern Standard Arabic plus fifteen named dialects — fine-tuned from Voxtral-Mini-4B-Realtime-2602 and published under Apache 2.0 with downloadable BF16 weights. That much the repository proves. Whether Mistral intends to support it, price it, or say anything about it at all is not yet knowable, and this piece keeps those two categories apart on purpose.
The short version: a known-good realtime ASR architecture was pointed at one language family and its dialects, the weights and both serving paths are public and runnable today, and the company behind it has not spoken. What follows is a read of what actually exists — the repository timestamps, the config files, the model card's own numbers — plus a clear line around what none of that tells you.
What is on the hub, and how it got there
The primary artifact is a single Hugging Face repository, mistralai/Voxtral-Mini-4B-Realtime-Arabic, carrying a model card, safetensors weights in BF16, and the processor and config files needed to load it. Its creation timestamp is 2026-10-08T11:36:06Z. Its last modification is 2026-10-09T17:11:35Z — the repository was still being touched the day after it appeared.
The timing of the sibling is the detail that turns a single quiet upload into something worth reading. mistralai/LIDstral-Arabic was created at 2026-10-08T11:37:05Z, under a minute later, with an identical modification timestamp and identical engagement counts, and a pipeline tag of text-classification rather than speech recognition. Two repositories, two different tasks, sixty seconds apart. That is not how a one-off experiment gets published; it is how a batch gets pushed.
The wider gap is what makes the pairing odd. Before 8 October, the most recent Mistral model repository of any kind was Shieldstral-1.0-3B on 2026-07-16 — roughly twelve weeks of an empty hub, broken by two uploads inside a minute. An Arabic dialect ASR checkpoint and an Arabic language-identification classifier arriving together, unnamed and unexplained, is the signature of a release that is coordinated internally and unannounced externally. It is not the signature of a leak, and it is worth being precise about why: a leak is weights without a licence, without config, without a serving path. This has all three.

The model card is written to ship. It documents usage, benchmarks, limitations and licence in the usual format, and it names the co-developer: Morocco's Ministère de la Transition Numérique et de la Réforme Administrative, under the strategic partnership signed between Mistral and Morocco in January 2026, with the model described as a sovereign AI asset. That framing is consistent with a deliverable that exists because a contract requires it to exist, which is one plausible reason weights can be public while a launch post is absent. It is a plausible reason. It is not a confirmed one, and the card does not say it.
The dialect list is the actual product decision
The card carries sixteen language tags. Only one of them is a language in the ordinary sense.
• Modern Standard Arabic — the ar tag, the variety every Arabic ASR system already handles.
• Maghreb — Moroccan (ary), Libyan (ayl), Tunisian (aeb), Algerian (arq) and Hassaniya, the Mauritania variety (mey).
• Gulf — Gulf Arabic across Bahrain, Kuwait, Qatar and the UAE (afb), Najdi (ars), Omani (acx) and Sanaani, the Yemeni variety (ayn).
• Nile Valley — Egyptian (arz) and Sudanese (apd).
• Levantine and Iraq — Mesopotamian (acm), South Levantine for Jordan and Palestine (ajp) and North Levantine for Lebanon and Syria (apc).
• Other — Chadian Arabic (shu).
Read that as a specification and the point is not "it does Arabic." Plenty of models do Arabic. The point is that Moroccan Darija, Gulf Arabic, Egyptian Arabic and Chadian Arabic are listed as separate training targets rather than as accents one Modern Standard Arabic model is expected to absorb. That is a materially harder problem, and it is the problem that actually breaks Arabic speech systems in production — a Moroccan call centre and a Gulf support line are not the same audio, and a model tuned on fusḥā tends to fail on both. The card also states that code-switching, where a speaker alternates languages mid-conversation, was a training focus rather than a side effect.
The limitation is stated just as plainly, and it is worth quoting the substance rather than softening it: the model targets Arabic transcription, and for other languages the card directs you to the original Voxtral-Mini-4B-Realtime-2602. This is a specialist checkpoint, not a general one. There is no ambiguity in it and no hint that a multilingual version is coming.
The architecture, read from the config rather than the prose
The prose on a model card is marketing-shaped; the config files are mechanical, and they are where the operational constraints live. All of the following is read directly from the repository's config.json, params.json and processor_config.json.
• Two stacks, not one — a causal audio encoder of 32 layers at 1280 hidden width and 32 attention heads, feeding a language decoder of 26 layers at 3072 hidden width with 32 query heads against 8 key-value heads. The card's "approximately 4.4 billion parameters" is the sum; the encoder is the smaller half of it.
• Audio front end — 16 kHz input, 128 mel bins, a 400-sample FFT with a 160-sample hop, giving an encoder frame rate of 12.5 per second, or one frame every 80 milliseconds. The architecture is registered as voxtral_realtime with a VoxtralRealtimeForConditionalGeneration head.
• The delay is a token count, not a milliseconds field — default_num_delay_tokens is 6. Six encoder frames at 80 milliseconds each is 480 milliseconds, which is exactly the delay the card quotes its accuracy figure at. The latency figure in the marketing is therefore a config value you can change, and the card's headline number is one point on a curve rather than a property of the model.
• Encoder attention is windowed to 750 frames — about sixty seconds of audio — while the decoder runs a sliding window of 8,192 tokens against a maximum position embedding of 131,072. The encoder window is the first number to check against your own workload: a streaming session longer than a minute is operating on a rolling window, not on unlimited context, and how the model behaves when a long recording scrolls past that boundary is not something the card addresses.
• Quantisation is not shipped — the published dtype is bfloat16 throughout. Anyone who wants an int8 or fp8 deployment builds it themselves.
The 8.82% figure, and who is standing behind it
Every accuracy number attached to this model comes from the model card. Nothing here has been reproduced by a third party, and the numbers below should be read as the vendor's own claims, not as measurements.
• Average character error rate — 8.82% across seven Arabic benchmarks, at the default 480-millisecond transcription delay, temperature 0.0.
• Against its own offline sibling — Voxtral Transcribe Arabic sits at 7.91% on the same seven benchmarks, so the realtime model lands 0.91 percentage points behind a non-streaming Arabic model from the same vendor. That comparison is Mistral's own, which is what makes it useful: it is a clean measurement of what sub-second latency costs on this language family, on identical data with identical scoring.
• New benchmarks in the mix — the seven include ISMA and Darija in the Wild, which the card says were co-developed with Morocco's MTNRA. A vendor introducing its own evaluation sets alongside a model is normal, and it also means part of the benchmark suite has no external history to compare against.
• Normalisation is the load-bearing caveat — the CER figures are computed with a normalisation scheme also co-developed with MTNRA, covering dialect-specific spelling, clitics, loanwords and spoken numbers, and the card states outright that scores may differ under other normalisation methods. The code ships in the repository as normalization.py. Read that as an invitation to check, and also as a warning: two teams can run this model on the same audio and publish different CER numbers without either being wrong.
• One footnote cuts against a competitor — the card notes that Gemini's safety filters block transcription of some FLEURS audio samples. That is a specific, checkable observation about a rival pipeline, and it is the kind of behaviour a leaderboard flattens: a sample a model refuses to transcribe reads as a failure or as missing data, and neither reading is a fair score.

Two things are missing from the evidence base and both matter. There is no independent evaluation, so no figure here has survived contact with a third party. And there is no price, because there is no service — nothing is being sold, so there is nothing to price. A model shipping with claimed numbers and no commercial offer is a model whose numbers nobody outside the lab has had a reason to test.
Running it today
This is the part that separates a real checkpoint from a placeholder, and the repository is unusually complete for something unannounced. Two supported paths ship on the card.
• vLLM — install vLLM with the audio extras from mistral-common, then serve the checkpoint by name, and stream audio over a WebSocket to the /v1/realtime endpoint. The card points at the vLLM realtime speech-to-text example as the thing to adapt, which means the streaming contract is a documented, maintained interface rather than something glued together for the demo.
• Transformers — native support from version 5.2.0, loading through AutoProcessor and VoxtralRealtimeForConditionalGeneration in bfloat16, with a full Python example on the card. The sample audio it uses is an Arabic bank call, assets/bank-take-2.wav, which is at least an honest choice of demo clip for this model.
Both paths are self-hosted. There is no hosted endpoint for this checkpoint anywhere we can verify, no vendor API, and no published rate. Support in the wider ecosystem is genuinely thin by design: native Transformers support at 5.2.0 is the newest thing here, and the vLLM path depends on the audio extras. An operator who wants this in production supplies the GPU, the serving process and the WebSocket client, and inherits a checkpoint with no support commitment attached to it.
![A screenshot of the vLLM documentation page for realtime speech-to-text (captured October 10, 2026), showing the heading 'Realtime Speech-to-Text with Voxtral Realtime', the install commands 'pip install --upgrade vllm mistral-common[audio]' and 'vllm serve mistralai/Voxtral-Mini-4B-Realtime-Arabic', the instruction to stream audio over WebSocket to the /v1/realtime endpoint, and a following section on the OpenAI Realtime Client.](https://cms.orcarouter.ai/api/media/file/4-1756.png)
That gap is where a routing layer is genuinely useful, and it is worth being exact about the boundary. OrcaRouter does not serve Voxtral-Mini-4B-Realtime-Arabic — the checkpoint is not on our catalogue, and we will not imply otherwise. What a router does reach in this deployment is everything downstream of the transcript: the summariser, the intent classifier, the agent that consumes the stream. Those are models you can call through one API key across 200-plus models at provider list price with 0% markup, with automatic failover when a provider degrades and a routing DSL for composing several models into one call. If you are standing up a self-hosted Arabic ASR service, the speech half of that stack is yours to run; the language half does not need a second contract, a second SDK or a second billing relationship to go with it.
What would turn this into a story
This is a real artifact and an incomplete release, and the incompleteness is the interesting part. Apache 2.0 cannot be withdrawn from the weights already downloaded, so the licence risk is settled. What is not settled is everything the licence does not cover.
Three things would change the picture, and none of them exist yet. An announcement: a blog post, a pricing tier or a line in Mistral's news index would explain whether this is a product or a contract deliverable, and it would almost certainly carry the detail the card omits. An independent run: the 8.82% figure and the 0.91-point gap to Voxtral Transcribe Arabic are the claims most worth reproducing, and the shipped normalisation code makes that unusually easy for anyone who wants to try. And a second checkpoint: a Levantine, Gulf or Egyptian model arriving separately would turn a single dialect-specialist into a family, which is the difference between a promising research artifact and something a regional deployment can actually standardise on.
Until at least one of those lands, the accurate summary of Voxtral-Mini-4B-Realtime-Arabic is this: a complete, runnable, Apache-2.0 Arabic dialect ASR checkpoint, published with a full card and two supported serving paths, arriving with no announcement from the company that made it, no independent evaluation, no price and no support commitment — alongside an Arabic language-identification model that appeared fifty-nine seconds later and suggests more of this is coming rather than less.
