Generated hero title card for a comparison of Voxtral Mini 4B Realtime Arabic and Grok Voice Transcribe 2.0, subtitled 'a label leaderboard leader meets a 16-dialect Arabic specialist', badged 'Mistral repository, 8 October 2026', with three stat chips reading 8.82% Arabic character error rate at 480ms against 2.7% word error rate at 0.49s, 16 named dialects against 38+ undocumented languages, and Apache 2.0 weights against $0.20 per audio-hour, and the OrcaRouter logo composited in a white strip at the bottom right.
Guides & Insights

Voxtral Mini 4B Realtime Arabic vs Grok Voice Transcribe 2.0: A Leaderboard Leader Meets a Dialect Specialist

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The honest starting point for this comparison is that Voxtral Mini 4B Realtime Arabic and G​rok Voice Transcribe 2.0 are not competing for the same job, and the fastest way to see that is to look at what each vendor was willing to publish. Mistral AI put Voxtral Mini 4B Realtime Arabic on Hugging Face on October 8, 2026 with no announcement — Apache 2.0 weights, a model card naming sixteen Arabic varieties, and a single accuracy figure of 8.82% average Character Error Rate across seven Arabic benchmarks at a 480-millisecond delay. x​AI's G​rok Voice Transcribe 2.0 shipped on September 18, 2026 with a launch post, a price, and a third-party leaderboard result: 2.7% word error rate on final transcripts with the text settled 0.49 seconds after the speaker stops, first on Artificial Analysis's streaming speech-to-text board when it landed. One model tells you exactly which dialects it handles and refuses to guess outside them. The other tells you it covers 38-plus languages and does not itemise a single one.

That asymmetry is the whole comparison. If you are buying transcription capacity in bulk for mixed-language audio, the xAI model is the more complete product by a wide margin. If your audio is Arabic — and specifically if it is dialectal Arabic, not broadcast Modern Standard Arabic — the Mistral checkpoint is aimed at a problem that the leaderboard the xAI model tops does not measure, and the fact that it is second on that board rather than first would be a mistake to read as a verdict.

The price arithmetic, because it is unusually clean here

xAI publishes a rate. Mistral publishes a download. That difference drives everything downstream, so start with the number that exists.

• Grok Voice Transcribe 2.0 — $0.10 per audio-hour over REST for recorded files, $0.20 per audio-hour over the streaming WebSocket. That is roughly $1.67 per thousand minutes for batch and $3.33 per thousand minutes streaming, unchanged from Grok Voice Transcribe 1.0 at the same price points.

• Voxtral Mini 4B Realtime Arabic — no published price, because there is no published service. The weights are Apache 2.0 and roughly 4.4 billion parameters in BF16, which means the cost is the GPU you attach to it and the utilisation you get out of it. At sustained volume that is cheap; at zero volume it is the most expensive option in this article, because you are paying for idle hardware.

The break-even is not subtle. A $0.20-per-hour streaming rate is about a third of a cent per minute. Any accelerator you would run a 4.4B model on costs more than that per hour unless you are pushing sustained traffic through it, which means the open-weights route wins on cost only after you have enough Arabic volume to keep a machine busy — and loses on every other axis while you are getting there, because the API has no capex, no capacity planning and no ops burden.

Screenshot of the Hugging Face model card for mistralai/Voxtral-Mini-4B-Realtime-Arabic, captured 10 October 2026, showing the model name, the mistralai organisation with 199k followers, the Automatic Speech Recognition pipeline tag, three likes, and the repository file listing that includes model.safetensors and normalization.py.

The one place the arithmetic flips early is compliance. The Mistral checkpoint processes audio on hardware you control, so nothing leaves your network, and the card ships a normalisation module so the Arabic scoring pipeline is yours to audit. Grok Voice Transcribe 2.0 runs in xAI's regions and, per its documentation, processes audio in real time without retaining it or using it for training, backed by SOC 2 Type II and HIPAA eligibility. Both stories are defensible; only one of them lets a regulator inspect the code path.

What $0.20 an hour actually buys, and the Mistral model does not ship

The feature gap here is larger than the accuracy gap and gets discussed far less. Grok Voice Transcribe 2.0 is a product with an interface; Voxtral Mini 4B Realtime Arabic is a checkpoint that emits text.

• Speaker diarization — included on Grok Voice Transcribe 2.0, plus multichannel transcription of up to eight independent channels. The Mistral card lists no diarization capability; the model produces a transcript, not an attributed one.

• Word-level timestamps — Grok carries them with confidence scores attached, so you can threshold low-confidence spans rather than trusting a whole transcript equally. Mistral's card does not claim timestamps for this checkpoint.

• Key-term biasing — up to 100 domain terms per request on the Grok endpoint, which is how you make a model get product names, account codes and place names right without fine-tuning. The Mistral route to the same outcome is fine-tuning on your own data, which Apache 2.0 permits and the hosted endpoint does not.

• Inverse text normalisation — Grok formats numbers, dates, currencies, phone numbers and email addresses into written form across 25 languages. The Mistral card includes a normalisation script, but its stated purpose is scoring Arabic benchmarks, not formatting output for display.

• Interface surface — REST for files up to 500 MB, WebSocket streaming at wss://api.x.ai/v1/stt with interim results roughly every 500 ms, a documented 10 requests per second and 100 concurrent streaming sessions, and a single region for now. On the Mistral side, vLLM serving with a WebSocket at /v1/realtime, or native Transformers from version 5.2.0, with no rate limit other than your hardware.

If you need diarization and timestamps, the comparison is over before accuracy enters it. That is not a knock on the Mistral model — a 4B Arabic specialist trained to produce accurate dialectal text is a different engineering target from a general transcription service — but it does mean the two only become interchangeable if your pipeline treats the transcript as raw material for your own post-processing.

The language list problem

Both vendors describe language support in a way that leaves the same question unanswered, from opposite directions.

• Grok Voice Transcribe 2.0 — 38-plus languages, automatic language detection with mid-recording language switching in a single pass, across 12 audio formats from 8 kHz to 48 kHz. The documentation gives the count and not the list. Arabic is not named individually anywhere in the material, which means Arabic support is implied by the count and not confirmed by the vendor.

• Voxtral Mini 4B Realtime Arabic — sixteen Arabic varieties named explicitly: Modern Standard Arabic; Moroccan, Libyan, Tunisian, Algerian and Hassaniya from the Maghreb; Gulf, Najdi, Omani and Sanaani from the Gulf; Egyptian and Sudanese from the Nile Valley; Mesopotamian and both South and North Levantine; and Chadian Arabic. Everything outside that set is out of scope, and the card says to use the general realtime model instead.

So the question "which of these handles Arabic better" has a strange structure. The Mistral model is a documented Arabic specialist with a published Arabic error rate and no independent verification. The xAI model is a documented leaderboard leader whose vendor has not confirmed that Arabic is in scope at all. A 38-language count that includes Arabic and a sixteen-dialect list that only includes Arabic are both answering the question "does it do Arabic" — but only one of them is answering "does it do Darija."

Screenshot of the xAI documentation page for the grok-voice-transcribe-2.0 model, captured 10 October 2026, showing the Speech to Text entry in the Voice documentation navigation, the model page slug grok-voice-transcribe-2.0, its audio-to-text modality, and the us-east-1 region listing with rate-limit sections.

The leaderboard does not help here. Artificial Analysis's speech-to-text evaluation is a fixed mix weighted toward English agent-talk, which is the right design for ranking general transcription and the wrong one for surfacing a dialectal Arabic win. A model could be genuinely better on Gulf Arabic and Moroccan Arabic and show up as merely good on that board. This is not a criticism of the board; it is a reason not to use it as the only input for a regional deployment.

Accuracy, in units that do not convert

• Grok Voice Transcribe 2.0 — 2.7% final-transcript word error rate with 0.49 seconds to final, and 3.4% on first partials, both measured on Artificial Analysis's AA-WER v2 index, which blends a private agent-talk set with VoxPopuli and Earnings22. The first-partial figure is the notable one: it fell from 18.3% in Grok Voice Transcribe 1.0, which is the difference between a partial transcript you can route on and one you can only display.

• Voxtral Mini 4B Realtime Arabic — 8.82% average Character Error Rate across seven Arabic benchmarks at 480 milliseconds of delay, on the vendor's own evaluation with a normalisation script supplied alongside it. Mistral reports that this comes within 0.91 percentage points of Voxtral Transcribe Arabic at 7.91% CER, which is the offline Arabic model from the same lab — a comparison worth more than a cross-vendor one because the benchmarks, normalisation and measurement are identical on both sides.

Putting 2.7% next to 8.82% is not a comparison; it is a category error. Word error rate counts wrong words against a reference, character error rate counts wrong characters, and Arabic orthography — where the same word tolerates multiple legitimate spellings and clitics attach freely — widens the gap between the two metrics well beyond what Latin-script languages show. The corpora are different, the normalisation is different, and only one figure comes from a third party. What the two numbers can legitimately tell you is that the xAI model is a measured, well-ranked general transcriber and the Mistral model is a claimed, Arabic-specific one. They cannot tell you which transcribes your audio better.

The comparison that does convert is the one inside Mistral's own card: 8.82% streaming against 7.91% offline, same seven benchmarks, same normalisation, same lab. Streaming costs about a point of accuracy on Arabic relative to a batch model. Whether that is a good trade is your call, and it is now an informed one.

Where the small model wins, and it is not on the scoreboard

Three things about the Mistral checkpoint are worth more than the numbers suggest, and none of them appears on any leaderboard.

The first is the dialect taxonomy itself. Naming sixteen varieties as distinct training targets, including Hassaniya and Chadian Arabic, is a statement about where the model's errors were measured. A general multilingual model that includes Arabic as one of 38 languages improves on Modern Standard Arabic first, because that is where the bulk of the training signal and the benchmark weight sits. A model built for the Morocco partnership alongside a North African ministry is optimising for a different distribution, and it was co-developed with two benchmarks — ISMA and Darija in the Wild — built specifically to measure that.

The second is the configurable delay. The 8.82% figure is quoted at 480 milliseconds, and the delay is adjustable rather than fixed, which turns the accuracy number into a point on a curve the operator picks. For a live captioning job you can shorten it and accept more errors; for a call-recording pipeline you can lengthen it and take the accuracy back. Grok Voice Transcribe 2.0 presents a fixed operating point, and the vendor's own upgrade path shows why that has consequences — moving from 1.0 to 2.0 cost about a third of a second in both first-partial and final latency, and the fix available to you is pinning the old model, not adjusting a dial.

The third is that the Apache 2.0 grant is unconditional for the weights you have. Whatever happens to the checkpoint upstream — whether it gets announced, priced, deprecated or never mentioned again — the artifact remains downloadable, fine-tunable and shippable inside your own product. A hosted endpoint's rate card and its model retirement schedule are both the vendor's to change, and xAI has already signalled its intent to make Grok Voice Transcribe 2.0 the default and deprecate 1.0, which will move every integration that never passed a model parameter onto the new latency profile at once.

On the routing layer above the transcript, one practical note: neither model is speech-to-text we host, and this article is not claiming otherwise. What OrcaRouter covers is the language-model layer that consumes the stream — the summariser, the classifier, the agent — behind one API key across more than 200 models at provider list price with 0% markup, so a vendor price change lands here the same day. Teams running two speech vendors for language coverage are already carrying two contracts and two auth paths; collapsing the downstream half onto one key is the cheap win.

What to do with this

Build on Grok Voice Transcribe 2.0 if your audio is multilingual, if you need diarization, timestamps or key-term biasing out of the box, or if you want a service with a published rate and a support path today. It is the more complete product, it is measured by a third party, and the streaming rate of $0.20 per audio-hour is low enough that the open-weights cost argument does not bite until you are running serious volume.

Build on Voxtral Mini 4B Realtime Arabic if your audio is Arabic and specifically dialectal, if you need the model inside your own network for compliance reasons, or if you intend to fine-tune — because only one of these allows it. Accept three things first: no diarization, no timestamps, and no independent evaluation published by anyone. Two days after the weights appeared, that last one is not a red flag; it is simply where the evidence stands.

What would change this comparison fastest is an Arabic-specific evaluation on real dialectal audio, run outside both vendors, with the same normalisation applied to both systems. Until that exists, the decision rests on a vendor's own dialect list against a vendor's own unstated one, and the only reliable test is your recordings.

Generated two-column comparison scoreboard for Voxtral Mini 4B Realtime Arabic and Grok Voice Transcribe 2.0 across six shared rows: price none published and self-host only against $0.10 per audio-hour over REST and $0.20 streaming; accuracy 8.82% Arabic character error rate vendor-reported against 2.7% final word error rate per Artificial Analysis; delay configurable with 480ms quoted against 0.49s to final and 0.49s to first partial; languages 16 named Arabic varieties against 38+ with none itemised by the vendor; diarization and timestamps not documented against included with confidence scores; and distribution Apache 2.0 weights on Hugging Face against API only. A footer reads that Mistral figures are vendor-reported and unverified while Grok figures are per Artificial Analysis, and that character error rate and word error rate are not comparable. The OrcaRouter logo sits in a white strip at the bottom right.

Neither endpoint is one we serve — this is a comparison of two systems outside our catalogue — but the models that consume the transcript are routable: more than 200 models on a single API key at provider list price with 0% markup, so a rate change on the summariser or the classifier lands the same day it is announced.

Automatic failover is what keeps a two-vendor speech setup from carrying two single points of failure into the language layer that feeds off it.