Generated hero title card for a comparison of Voxtral Mini 4B Realtime Arabic and Voxtral 4B TTS, subtitled 'two ends of the same Arabic voice loop, two different licences', badged 'speech in versus speech out', with three stat chips reading 8.82% Arabic character error rate at 480ms against 70ms time to first audio, 16 Arabic dialects against 9 languages including Arabic, and Apache 2.0 weights against CC BY-NC 4.0 weights, and the OrcaRouter logo composited in a white strip at the bottom right.
Engineering & Research

Voxtral Mini 4B Realtime Arabic vs Voxtral 4B TTS: Two Ends of the Same Conversation

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

These two models share a name, a parameter count and a licence file that behaves differently, and that is where the resemblance stops. Voxtral Mini 4B Realtime Arabic, pushed to Hugging Face on October 8, 2026 with no accompanying announcement, takes audio in and produces Arabic text — a streaming fine-tune of Voxtral-Mini-4B-Realtime-2602 built for Modern Standard Arabic and fifteen named dialects, Apache 2.0, 8.82% average Character Error Rate at a 480-millisecond delay. Voxtral 4B T​TS, announced by Mistral AI in March 2026, does the reverse: text in, speech out, twenty preset voices plus adaptation from a short reference clip, nine languages including Arabic, and a licence — CC BY-NC 4.0 — that forbids commercial use of the weights themselves. They are not alternatives and they are not variants. They are the two halves of a voice loop, and the reason to look at them together is that the decisions you make about one of them constrain the other more than most teams expect.

Most comparison pages will tell you these models are unrelated because one is ASR and one is TTS, then stop. That is true and not useful. What is actually worth knowing is where they meet: a voice agent needs both, the two licences point in opposite directions, the two hardware footprints are similar while the throughput characteristics are not, and the Arabic coverage claims are of completely different shapes. Everything below is about those four meeting points.

What each model is, in the terms that matter for a pipeline

• Voxtral Mini 4B Realtime Arabic — streaming automatic speech recognition. 16 kHz audio in, text out, roughly 4.4 billion parameters in BF16, served through vLLM with a WebSocket at /v1/realtime or natively in Transformers from version 5.2.0. A causal audio encoder and a language decoder, a configurable transcription delay quoted at 480 milliseconds for the published accuracy figure, and a normalisation module shipped alongside so the Arabic character error rate can be re-derived. Apache 2.0.

• Voxtral 4B TTS — streaming and batch text-to-speech. Text in, 24 kHz audio out in WAV, PCM, FLAC, MP3, AAC or Opus, roughly 4 billion parameters, with a 20-voice preset roster and voice adaptation from a reference clip as short as three seconds. Mistral's own benchmark on a single H200 running vLLM-Omni 0.18.0 with a 500-character input and a ten-second reference reports 70 milliseconds of latency and a real-time factor of 0.103 at concurrency 1, rising to 331 milliseconds and 0.237 at concurrency 16, and 552 milliseconds at concurrency 32. Consulted as an API at $0.016 per thousand characters.

The first observation from those two paragraphs is that the pair is small. Both models are in the 4-billion-parameter range and both fit on a single accelerator in BF16, which means an Arabic voice agent built entirely on open weights is a two-GPU deployment at most, and plausibly a one-GPU deployment with quantisation on both sides. That is the practical reason to look at them as a set rather than as two separate product decisions.

Screenshot of the Hugging Face model card for mistralai/Voxtral-Mini-4B-Realtime-Arabic, captured 10 October 2026, showing the model name, the mistralai organisation with 199k followers, the Automatic Speech Recognition pipeline tag, three likes, and the repository file listing that includes model.safetensors and normalization.py.

The licence asymmetry is the finding, not a footnote

Here is the thing that surprises people who assume models from the same lab with the same naming convention carry the same terms.

• Voxtral Mini 4B Realtime Arabic — Apache 2.0. Commercial use, modification, redistribution and fine-tuning are all permitted, subject to the standard restriction against infringing third-party rights.

• Voxtral 4B TTS — CC BY-NC 4.0, inherited from the reference-voice datasets behind it. Non-commercial. Mistral's card is explicit that the model inherits the licence of its voice references, drawn from sources including EARS, CML-TTS, IndicVoices-R and an Arabic natural-audio set. The weights are downloadable and the licence is not a commercial one.

This is a hard constraint on the exact architecture everybody reaches for first. If you build an Arabic voice agent whose speech-to-text leg is Voxtral Mini 4B Realtime Arabic and whose text-to-speech leg is Voxtral 4B TTS, you have a pipeline where half the components are freely commercial and half are not, and the restrictive half governs the product. The ASR model being Apache 2.0 does not rescue the TTS leg. Running the pair locally does not change the licence, and neither does not-shipping the weights — the constraint attaches to the use, not the distribution.

Screenshot of the Hugging Face model card for mistralai/Voxtral-4B-TTS-2603, captured 10 October 2026, showing the model name, the mistralai organisation, the Text-to-Speech pipeline tag, 922 likes, a language badge, and the CC BY-NC 4.0 licence inherited from the reference-voice datasets.

The commercial route Mistral offers for the speech side is the hosted API at $0.016 per thousand characters, which is a service agreement rather than a licence grant. For a product team the practical consequence is that the freely commercialisable half of the loop is the half that listens, and the half that speaks is either an API dependency or a licensing conversation. That inverts what most teams assume when they set out to build an open voice stack, and it is worth discovering before you have written the integration rather than after.

The Arabic claims do not line up, and only one of them is a coverage promise

Both models list Arabic. The listings mean different things.

• Voxtral Mini 4B Realtime Arabic — Arabic is the entire scope. Sixteen varieties enumerated as training targets: Modern Standard Arabic, Moroccan, Libyan, Tunisian, Algerian and Hassaniya Arabic, Gulf, Najdi, Omani and Sanaani Arabic, Egyptian and Sudanese, Mesopotamian and both South and North Levantine, and Chadian. Anything outside that set is documented as out of scope. One aggregate accuracy figure, 8.82% CER, averaged across seven Arabic benchmarks.

• Voxtral 4B TTS — Arabic is one of nine supported languages, alongside English, French, Spanish, German, Italian, Portuguese, Dutch and Hindi. The card describes "diverse dialects" within that support, without enumerating them, and there is no per-language quality figure published for Arabic or for any of the other eight.

Read together, the pair gives you a detailed statement about how well your system will hear Arabic and almost no statement at all about how well it will speak it. That asymmetry has a real consequence for a Darija or Gulf Arabic voice agent: the input side has a published benchmark and a dialect list you can evaluate against, and the output side has a language badge and a voice roster. Nobody has published a dialectal Arabic intelligibility measurement for Voxtral 4B TTS, and the voice adaptation feature does not resolve it — adapting a voice changes who the speech sounds like, not which dialect it produces.

There is a second, subtler point about the ASR model's own accuracy figure that carries across to the TTS side. Mistral reports 8.82% CER streaming against 7.91% for the offline Voxtral Transcribe Arabic on the same seven benchmarks with the same normalisation, so streaming costs about a point. The equivalent trade on the TTS side — what streaming inference costs in output quality versus batch — is not quantified in the material at all. The latency numbers exist; the quality delta does not.

Throughput: one model is built to be pushed, the other is not

Mistral published throughput data for exactly one of these models, and the numbers explain why.

• Voxtral 4B TTS — measured on one H200 with vLLM-Omni, a 500-character input and a ten-second audio reference, throughput runs from about 119 characters per second of GPU time at concurrency 1 to roughly 879 at concurrency 16 and about 1,431 at concurrency 32, with latency rising from 70 milliseconds to 331 and then 552. That is a throughput curve you can plan capacity against, and it is the shape you would expect from a model designed as a production voice service.

• Voxtral Mini 4B Realtime Arabic — the card reports no throughput figure, no concurrency behaviour and no hardware benchmark. What it does state is the architecture: a causal encoder and a sliding-window attention scheme on both the encoder and the decoder, described on the base model as supporting effectively unbounded streaming. The accuracy point is quoted at 480 milliseconds of delay, which implies the model keeps up with real-time audio on appropriate hardware, and that is the extent of the published performance envelope.

The gap is not a criticism of either model so much as a statement about maturity. The TTS model has a published SLO table because it was released as a product with a blog post, a demo, an API and a price. The Arabic ASR model has no published performance envelope because it arrived as a repository with a card. If you are sizing an Arabic transcription fleet today, the throughput number you need does not exist yet, and the only way to get it is to run the vLLM serving path on your own hardware with your own audio.

That is also where a routing layer changes the shape of the deployment rather than just the billing. OrcaRouter routes neither of these models — we host neither Mistral speech endpoint, and this article is not claiming otherwise — but the language-model layer between them is exactly the layer that benefits from being routed: one API key across more than 200 models at provider list price with 0% markup, so a vendor price change lands on our side the same day it is announced, and automatic failover so a single upstream provider's bad afternoon is a reroute rather than an outage in your voice loop. A voice agent assembled from two self-hosted Mistral models plus one hosted reasoning model has three failure domains; the routing layer is what makes the middle one boring.

Cost, side by side, with the caveat that one side has no price

• Voxtral 4B TTS — $0.016 per thousand characters through Mistral's API, or the cost of the GPU if you self-host under terms that permit your use case. For a rough sense of scale, a thousand characters is a couple of hundred words, so a minute of spoken output lands in the low fractions of a cent at list price, before any concurrency advantage from running it yourself.

• Voxtral Mini 4B Realtime Arabic — no published price, no hosted service, no rate card. Cost is hardware and utilisation. A 4.4B BF16 model on one modern accelerator is a fixed monthly cost that you amortise across as much audio as the machine can process, which is cheap at sustained volume and pure waste at occasional volume.

The comparison that probably matters more is against the alternatives you would reach for instead. On the listening side, the Arabic checkpoint is competing with per-hour streaming APIs whose rates are published and low — Microsoft's MAI-Transcribe-2-Streaming at $0.54 per audio-hour introductory, xAI's Grok Voice Transcribe 2.0 at $0.20 per audio-hour streaming — which means an Arabic transcription service can be bought for a few tenths of a cent per minute and the open-weights case has to be made on dialect accuracy, data residency or fine-tuning rather than on unit cost. On the speaking side, a self-hosted TTS model at zero marginal cost is a real advantage over per-character APIs at volume, which is why the CC BY-NC licence is the constraining factor rather than the price.

One structural note for anyone budgeting a price-based switch: a router that passes provider list price through at 0% markup is the surface where a vendor rate change shows up the same day it is announced rather than at the next repricing cycle. That matters more on the listening side than the speaking side, because transcription is priced per hour of audio and that is a number vendors move.

How to think about picking between them

You are not picking between them. The question is which half of the loop you can build on open weights, and the licence answers it.

If your requirement is a self-contained Arabic voice agent where audio never leaves your infrastructure, the listening leg is available to you unconditionally: Apache 2.0, downloadable, fine-tunable, with a dialect list you can test against and a published Arabic error rate. The speaking leg is where the constraint lands, and you have two honest options — move the speech synthesis to a hosted service under a commercial agreement, or remove Voxtral 4B TTS from the architecture and find a permissively licensed synthesis model for Arabic.

If your requirement is a production transcription service and Arabic is the language, the decision is between a downloadable 4.4B checkpoint with a published dialect taxonomy and one aggregate error rate, and a hosted streaming API with a published rate, a published latency and a third-party board. The open model wins on control, residency and fine-tuning; the hosted options win on evidence, support and operational simplicity. Two days after the weights appeared, nobody outside Mistral has measured them, and that is a reason to run your own benchmark rather than a reason to dismiss the checkpoint.

What would most change this picture is an Arabic synthesis release under terms that permit commercial use — the same pattern the listening side already follows. Until that exists, the loop stays half-open, and the practical work is deciding which half you are willing to rent.

Generated two-column comparison scoreboard for Voxtral Mini 4B Realtime Arabic and Voxtral 4B TTS across six shared rows: task speech to text against text to speech with 24 kHz audio output; Arabic claim 16 named dialects at 8.82% character error rate against 1 of 9 languages with no quality figure; licence Apache 2.0 permitting commercial use against CC BY-NC 4.0 non-commercial weights; service price none published against $0.016 per 1,000 characters; published throughput none against 119 to 1,431 characters per second per GPU; and hardware roughly 4.4B in BF16 on a single accelerator against roughly 4B in BF16 on a single accelerator. A footer reads that all figures are Mistral's own and unreproduced, and that the two models are complementary rather than competing. The OrcaRouter logo sits in a white strip at the bottom right.

We host neither of these Mistral speech models and this article does not claim otherwise; what the voice loop needs above them is the language layer, and that is routable - one API key across more than 200 models at provider list price with 0% markup, so a vendor rate change lands on our side the same day it is announced.

Automatic failover keeps the middle of a three-component voice agent - one upstream reasoning model between two self-hosted Mistral models - from being the part that takes the product down.