A generated hero card for 'Gemini 3.8 TTS vs VoxCPM2' on a white background with soft blue-and-cyan gradient accents. The left card 'Gemini 3.8 Flash TTS' lists an API endpoint, Elo 1,260 at rank 2, and $16.50 per 1M characters normalised; the right card 'VoxCPM2' lists Apache 2.0, 2B parameters, roughly 8 GB VRAM to run, a real-time factor of 0.30 in plain PyTorch and no hosted endpoint. A footer line reads 'Elo per Artificial Analysis Provider Voice Arena, Sept 2026; VoxCPM2 figures per its model card.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Gemini 3.8 TTS vs VoxCPM2: Renting a Voice or Running One, in Real Numbers

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is no leaderboard row that puts Gemini 3.8 Flash TTS and VoxCPM2 side by side, and that absence is the most useful fact in this comparison. Gemini 3.8 Flash TTS is a hosted endpoint with an Elo of 1,260 at rank 2 on the Artificial Analysis Provider Voice Arena and a published rate card in tokens. VoxCPM2 is an OpenBMB checkpoint — 2 billion parameters, Apache 2.0, released April 2026 — that no inference provider hosts. You cannot rent it. You run it.

So the question is not which model sounds better. It is whether your workload is better served by paying per character for someone else's GPUs or by owning the GPUs and paying for them whether or not they are working. That is an arithmetic question, and unlike most model comparisons it has an answer you can compute before you commit.

A generated comparison scoreboard for Gemini 3.8 Flash TTS and VoxCPM2, showing access (a hosted API against self-hosted only), size (undisclosed against 2B parameters), licence (vendor terms against Apache 2.0), price ($16.50 per 1M characters against a GPU-hour cost), throughput (not applicable against a real-time factor of 0.13 to 0.30) and independent Elo (1,260 against none).

One of these you rent, the other you operate

Start with what each one actually hands you.

• Access — Gemini 3.8 Flash TTS is API-only, called through the Gemini API and the Google surfaces named in its September 23, 2026 announcement. VoxCPM2 has no first-party endpoint in production use and no inference provider serving it; the checkpoint is the product.

• Licence — VoxCPM2 ships under Apache 2.0, which permits commercial use without a separate agreement. That is a genuinely permissive licence, and it is not the default for open speech weights, several of which ship under research-only terms that route commercial use back through a hosted API.

• Size — VoxCPM2 is 2B parameters and fits in roughly 8 GB of VRAM at reduced precision for development, with serving peak around 22 GB on a 24 GB card. Google has published no parameter count for either Gemini 3.8 TTS model.

• Voice creation — Gemini 3.8 Flash TTS does voice design from a written description, voice replication from a 30-second sample, and two-speaker dialogue in one call. VoxCPM2 is a single-voice text-to-speech model; a conversation is two runs stitched.

• Independent score — Flash TTS has one. VoxCPM2 does not appear among the ranked rows on the Provider Voice Arena captured on September 24, 2026. Absence from a board is not a verdict on quality — it usually means nobody has entered the model — but it does mean there is no third-party preference number to weigh.

A screenshot of the Hugging Face model card for OpenBMB/VoxCPM2, showing the Apache 2.0 licence, the 2 billion parameter checkpoint size, the 30-language coverage, 48 kHz output and the 2 million hours of training audio recorded on the card.

That last point cuts both ways and is worth stating plainly rather than glossing: if you need an independent quality signal before you commit, only one of these two supplies one.

The number that decides it: real-time factor

Self-hosting a speech model converts a per-character bill into a per-GPU-hour bill, and the exchange rate between those two units is the model's real-time factor — how many seconds of compute it takes to produce one second of audio.

VoxCPM2's published figures are 0.30 in plain PyTorch and 0.13 under Nano-VLLM. Read those as throughput rather than latency: at 0.13, one GPU-hour of wall clock produces something on the order of 7.7 hours of finished audio. At 0.30, the same GPU-hour produces closer to 3.3 hours. Those are the model card's own numbers, measured by OpenBMB, and they are the single most important input into any build-versus-buy decision here.

Three things follow from them.

• The serving stack matters more than the model. Moving from plain PyTorch to Nano-VLLM more than doubles throughput on identical hardware. Any cost model built on the 0.30 figure overstates the self-hosted cost by a factor of about 2.3.

• Utilisation is the whole game. A rented GPU that is idle 90% of the time is not cheaper than an API at any price. The break-even only appears when you are keeping the card busy.

• Batching changes the arithmetic again. Real-time factor is measured per stream; a server handling concurrent requests amortises the fixed cost across them, which is exactly the regime where self-hosting starts to win.

What an hour of audio costs, both ways

Put the two billing models on the same axis. One hour of finished audio is 3,600 seconds of output.

On the rented side, Google bills in tokens rather than characters: $0.50 per million text tokens in and $9.00 per million audio tokens out through December 31, 2026, then $18.00; Flash-Lite TTS is $0.50 and $6.00, then $12.00. Batch and Flex run $0.25 and $4.50 for Flash TTS against $0.25 and $3.00 for Flash-Lite TTS, and Priority runs $0.90 and $16.00. Artificial Analysis normalises the standard rates to $16.50 and $11.00 per million characters, which is the board's conversion rather than a Google quote, and it is the figure to use when comparing against a model that bills in characters.

On the owned side there is no per-character rate at all. The cost is the GPU-hour, multiplied by however many hours you need at the throughput above, plus the serving stack, plus the people who keep it running. VoxCPM2's advantage is that the marginal cost of the thousandth hour is the same as the first — and its disadvantage is that the marginal cost of the first hour is a GPU.

Which is why the decision is unusually clean:

• Below a certain volume, the API wins and it is not close. You are paying single-digit dollars per hour of audio against a capital cost, and the promotional Google rate makes that gap wider until January 1, 2027.

• Above that volume, on hardware you already own and can keep busy, the Apache 2.0 checkpoint wins decisively — and unlike a research-licensed checkpoint, nothing about the licence stops you shipping it.

• In between, the honest answer is that you are buying optionality rather than savings. Self-hosting buys the ability to pin a version, keep audio on your own infrastructure, and stop caring about a vendor's rate change.

Where each one is the right answer

Pick Gemini 3.8 Flash TTS when you want the better-documented product. It has an independent score, a published price, two-speaker dialogue in a single call, voice design and replication, and no operational surface beyond an API key. Its sibling Flash-Lite TTS gives you a second price point inside the same interface at a normalised $11.00 per million characters. What you accept in exchange is a rate that doubles on January 1, 2027 and no ability to run the model anywhere.

Pick VoxCPM2 when the operational work is the point. Apache 2.0 means commercial use is permitted outright, the 2B size means a single accelerator is enough, and the 0.13 real-time factor under Nano-VLLM means a modest card produces several hours of audio per hour of wall clock. What you accept is that nobody else is going to run it for you: no hosted endpoint, no arena row, no vendor rate card, and every quality guarantee you get is one you build and measure yourself.

A screenshot of the Artificial Analysis Provider Voice Arena leaderboard captured September 24, 2026, showing Google Gemini 3.8 Flash TTS at rank 2 with an Elo of 1,260 across 1,999 samples and 8 arena voices, with Gemini 3.8 Flash-Lite TTS at rank 6 and no VoxCPM2 row on the board.

Neither model is hosted by OrcaRouter, and that is worth saying directly rather than implying otherwise. The place to call Gemini 3.8 Flash TTS is Google's own API and the surfaces Google named; VoxCPM2 is a checkpoint you deploy. What OrcaRouter covers is the layer above that decision — 200-plus models behind one key at provider list price with no markup, so a vendor price change is live on our side the same day rather than at the next billing cycle, automatic failover so a model under evaluation never has to be a production dependency, and a routing DSL that composes several models into one call when a single voice does not carry the whole job.

What to watch next

Two things would change this comparison materially, and neither has happened yet.

The first is someone hosting VoxCPM2. Its absence from the inference market is the main reason the two models are being compared on economics rather than on quality — if a provider picks it up, the rent-versus-own arithmetic stops being the deciding factor and the missing arena row becomes the thing you want filled in.

The second is the January rate change on Google's side. At the promotional rate the API is cheap enough that most workloads should not be self-hosting anything; at the post-promotional rate the gap narrows to the point where a team with existing GPU capacity and steady volume should run the numbers again rather than assume the API still wins.

Until one of those moves, the deciding input is not a leaderboard. It is how many hours of audio a month you produce, and whether the card is busy.