
Eleven v4 vs VoxCPM2: A Ranked Model Against a Checkpoint You Cannot Rent
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The most useful fact in this matchup is a row that does not exist. ElevenLabs Eleven v4 holds rank 1 on Artificial Analysis' Provider Voice Arena at an Elo of 1,319 across 1,674 appearances, at $80.00 per million characters. VoxCPM2, the OpenBMB checkpoint released in April 2026, does not appear anywhere on either speech board — not ranked, not unranked, not as a preview entry. That is not a verdict on quality. It means that whichever number you were hoping to compare against 1,319, nobody has produced it, and the comparison has to be made on something other than preference. What is left is ownership, fine-tuning, and a licensing question that a hosted endpoint does not let you ask.
What each side actually hands you
These are different classes of product, and the difference is not cosmetic.
• Access — Eleven v4 is a hosted endpoint reached through ElevenLabs' own API, billed per character. VoxCPM2 is a checkpoint on Hugging Face under openbmb/VoxCPM2; you download it and run it, and there is no first-party endpoint to call.
• Licence — VoxCPM2 ships Apache 2.0, explicitly free for commercial use. Eleven v4 is a commercial API whose terms are ElevenLabs'. That difference matters if your legal review asks whether you may fine-tune, redistribute or ship weights; with the checkpoint the answer is written down and static.
• Size and hardware — VoxCPM2 is 2B parameters on a MiniCPM-4 backbone and the card quotes roughly 8 GB of VRAM at bfloat16, with a maximum sequence length of 8,192 tokens. ElevenLabs publishes no parameter count for Eleven v4, and a hosted model's hardware footprint is not a thing you manage.
• Output chain — VoxCPM2 accepts 16 kHz reference audio and emits 48 kHz through AudioVAE V2's built-in super-resolution, with no external upsampler in the path. Hosted TTS hands you a stream at whatever rate the vendor chose.
• Independent score — Eleven v4 has one, and it is first place. VoxCPM2 has none. Absence is not a low score; it is an unscored model, and the two should never be spoken of in the same sentence as if the second were a number.

The number that replaces the missing Elo
If you cannot compare preference, compare what you can compute, and for a self-hosted model that is throughput. VoxCPM2's card quotes a real-time factor of about 0.30 in standard PyTorch on an RTX 4090, falling to about 0.13 under Nano-VLLM. Real-time factor is seconds of compute per second of audio, so those two figures mean one GPU-hour yields roughly 3.3 hours of finished speech in the first case and about 7.7 hours in the second.

Two consequences follow immediately. The serving stack matters more than the model: the same checkpoint on the same card more than doubles its output when you swap the inference engine, so any cost model that uses the 0.30 figure is overstating your self-hosted cost by a factor of about 2.3. And utilisation is the whole game: a card that sits idle does not beat a per-character API at any price, because the API's bill scales with what you say while the card's bill scales with when you bought it.
Which is why the honest version of this decision is not "which sounds better". It is "how many hours of audio do you produce a month, and is the hardware busy". Below the volume where a card stays loaded, the API wins and it is not close. Above it, on hardware you already own, the checkpoint's marginal cost per thousandth hour is the same as its cost per first hour — and that is a structural advantage no rate card can match.
The capability a hosted endpoint will not sell you
This is where the comparison stops being a mirror image, because there is one thing VoxCPM2 does that Eleven v4 fundamentally cannot: you can retrain it. The model card documents both full SFT and LoRA fine-tuning on as little as five to ten minutes of audio, with a supplied training script and configuration for each.
That changes the shape of a voice programme. A hosted API gives you a fixed model plus whatever cloning the vendor exposes — ten seconds of reference audio in Eleven v4's case for instant clones, with a professional tier above it. A LoRA-tunable checkpoint gives you a model that has never seen anyone else's data and whose behaviour you can pin by version. For a brand voice, a regulated domain, or a language the vendor's cast handles poorly, that is a different category of solution, not a cheaper one.
The trade is explicit and worth stating: with VoxCPM2 you accept that no third party has ever scored the model, that quality guarantees are ones you build and measure yourself, and that the 8,192-token sequence limit and the card's own warning — that voice design and style control vary between runs and that one to three generations are recommended — are now your QA problem.
What Eleven v4 gives you that the checkpoint cannot
Going the other way, the hosted model's advantages are the ones that show up on a release calendar rather than a spec sheet.
• A measured ranking — first on a board with 1,674 samples behind the estimate, and an interval of roughly ±19 points around it. Whatever that board measures, it is independent of both vendors and it is the only third-party quality signal in this comparison.
• Performance direction in the text — inline audio tags such as [laughs], [whispers] and [said angrily in French accent], plus natural-language delivery prompts and improved International Phonetic Alphabet support. Getting a mood change out of a self-hosted model means a re-generation or a fine-tune; here it is a token in the script.
• Coverage you do not have to verify — 90-plus languages against VoxCPM2's 30, with the caveat that the checkpoint's list is explicit and named (including Chinese dialects) while the vendor's "90-plus" is a count without a list.
• No operations surface — no GPU, no serving stack, no version drift, and a promotional rate of $0.022 per 1,000 characters until October 12 against a list of $0.08 after.

Neither of these is ours, and that is the point of this section
OrcaRouter hosts neither model. There is no ElevenLabs endpoint and no VoxCPM2 endpoint in our catalogue, and the honest routing answer is that Eleven v4 is reached through ElevenLabs' own API while VoxCPM2 is something you deploy on your own hardware. We are not going to imply otherwise to make a comparison tidier.
Where a single key does help is the decision before the decision. The reason self-hosting conversations stall is that the alternative is never evaluated: an API is one call and a checkpoint is a serving stack, so the API wins by default rather than by arithmetic. OrcaRouter's contribution is that the hosted half of that comparison is cheap to test — 200-plus models behind one API key with provider list prices passed through at 0% markup, so the hosted option is evaluated at its real rate rather than an estimated one, with automatic failover if you put a candidate behind a live path before you have finished deciding.
How to decide, in one pass
Pick Eleven v4 if the voice has to be good on the first take and you want it done this week. You get the top of an independent board, a documented direction syntax, the widest documented language count in this comparison, and zero infrastructure. You pay $0.08 per 1,000 characters from October 13 onwards, and you accept that you cannot retrain it, version-pin it, or keep the audio on your own metal.
Pick VoxCPM2 if the model has to be yours. Apache 2.0, 2B parameters on about 8 GB of VRAM, 48 kHz output with no upsampler in the chain, and a fine-tuning path that runs on five to ten minutes of audio — those are capabilities a hosted endpoint does not offer at any price, and the absence of an arena row is the cost of them. Run the throughput numbers on your own card before you commit, because the 0.30 and 0.13 real-time factors are the model card's own measurements and your serving stack will not reproduce them exactly.
And if the honest answer is that you do not know your monthly audio volume, that is the number to go and find. It decides this comparison. Neither board will tell you.
