A generated hero card titled 'HeyGen Voice Debuts at #1 on the Controlled Voice Arena' with the line '1,202 Elo — ahead of Qwen-Audio-3.1-TTS-Plus and Eleven v4 Turbo' and the note 'Same eight cloned reference voices — Artificial Analysis, 9 Oct 2026'. The OrcaRouter logo appears in the bottom-right corner.
Guides & Insights

HeyGen Voice Debuts at #1 on the Controlled Voice TTS Arena — Beating Qwen and ElevenLabs on Cloned Voices

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On 9 October 2026, Artificial Analysis added HeyGen Voice to its Controlled Voice arena and the model went straight to the top of it. HeyGen Voice — the first text-to-speech model from HeyGen to be evaluated on the board — scored 1,202 Elo, ahead of Qwen-Audio-3.1-TTS-Plus at 1,186 and ElevenLabs' Eleven v4 Turbo at 1,167. Cartesia Sonic 3.6, which leads the site's open Provider Voices board, sits fifth here at 1,143. That is a genuine result on a leaderboard built to remove the single biggest confound in voice evaluation, and it is also a result with a very specific meaning that is easy to overread: HeyGen Voice is the best voice in this arena on eight cloned reference voices, not yet a measured champion of the board with ninety-six models on it.

What the Controlled Voice board measures, and why it matters

Artificial Analysis runs two speech boards and they answer different questions. The Provider Voices arena lets each vendor supply its own native voices — the polished demo voices a company picks to represent itself — and ranks models on how good those sound. Its leaderboard is broad and mature: ninety-six entries, with Eleven v4 Turbo on top at 1,328 Elo and Cartesia Sonic 3.6 fourth at 1,280.

The Controlled Voice arena does something narrower and, for anyone shipping a product, more useful. Every model on it generates speech from the same eight cloned voices — four US, four UK — so the comparison is not "whose marketing voice is nicer" but "how does each engine handle the same voice, the same text and the same recording conditions." It is the closest thing the industry has to a like-for-like test of the engine rather than the demo reel, and it is the board that matters if you plan to clone a voice and hand it to a TTS model.

HeyGen Voice now leads it. The margin is 16 Elo over Qwen-Audio-3.1-TTS-Plus and 35 over Eleven v4 Turbo, with a reported 95% confidence interval of roughly 1,185 to 1,219 on the HeyGen figure. Sixteen Elo is a real gap but not a chasm — on a board this dense, it is the difference between first and second, not between usable and unusable.

A generated single-column scoreboard titled 'HeyGen Voice — the scoreboard' with rows reading 'Controlled Voice Elo: 1,202', '95% confidence interval: roughly 1,185 to 1,219', 'Rank in the controlled field: #1 of 42 ranked entries', 'Provider Voices Elo: not listed, insufficient votes', "Price on the arena's basis: $30.00 per 1M characters, derived from credit pricing" and 'Elo anchor: the board baseline is 1,000 points', with a footer reading 'Elo per Artificial Analysis, 9 October 2026. Price shown is the arena's conversion of credit pricing.' The OrcaRouter logo appears in the bottom-right corner.

Where HeyGen Voice is not yet ranked

The honest edge of this result is that HeyGen Voice does not appear on the Provider Voices board at all, and its absence is not a signal about quality. Artificial Analysis notes that models can be omitted for not yet having enough votes. A model days old, with a single differently-scored arena behind it, has simply not accumulated the blind-listener volume the bigger board requires.

So the correct reading on 10 October 2026 is: HeyGen Voice is the top-ranked voice in the controlled, cloned-voice comparison, and an unknown quantity against the wider field. Anyone quoting "the best TTS model in the world" off this number is quoting a smaller board than the sentence implies.

A screenshot of Artificial Analysis's Provider Voice Arena leaderboard, captured 10 October 2026, showing the ranked model list with Eleven v4 Turbo first at 1,328 Elo, Eleven v4 second at 1,320, Qwen-Audio-3.1-TTS-Plus third at 1,296 and Sonic 3.6 fourth at 1,280, alongside samples, release dates and API pricing columns and 8-voice arena badges.

The two numbers behind the Elo

• Controlled Voice Elo — HeyGen Voice: 1,202 (95% CI roughly 1,185–1,219). Qwen-Audio-3.1-TTS-Plus: 1,186. Eleven v4 Turbo: 1,167.

• Provider Voices Elo — HeyGen Voice: not listed (insufficient votes). Eleven v4 Turbo: 1,328 at #1. Sonic 3.6: 1,280 at #4.

• Rank in the controlled field — HeyGen Voice: #1 of 42 ranked entries (#42 is OpenVoice v2 at 812).

• Price on the arena's basis — HeyGen Voice: $30.00 per 1M characters, the arena's own conversion of HeyGen's credit pricing. Qwen-Audio-3.1-TTS-Plus: $19.30 per 1M characters. Eleven v4 Turbo: $40.00 per 1M characters.

• Elo anchor — OpenAudio S1 is the fixed reference point at 1,000, so HeyGen Voice is 202 points above the anchor.

• Who else is in the top five — Eleven v4 at 1,160 and Sonic 3.6 at 1,143 complete the top five behind the leaders.

A screenshot of HeyGen's developer documentation sidebar, captured 10 October 2026, showing the Voice Management group with the entries Voices, Browse Voices, Design a Voice, Voice Clone and Text to Speech, above the Realtime Video and Avatar Management groups.

What HeyGen actually shipped

The engine behind the entry is HeyGen's in-house TTS, exposed in the company's v3 API. A GET /v3/voices call takes an engine filter whose values include orca — HeyGen's own voice engine — alongside starfish and a pass-through to ElevenLabs voices. Speech rendering runs through POST /v3/voices/speech; per-request tuning exposes speed from 0.5 to 1.5 and pitch from −50 to +50 semitones, with a BCP-47 locale hint.

The catalogue is listed as 300+ pre-built voices across dozens of languages. Custom voices come in three flavours: text-to-voice design that generates up to three ranked options from a description capped at 1,000 characters, an instant clone from a single recording, and a professional clone trained on twenty minutes or more of audio. That last one is the part the Controlled Voice leaderboard is effectively stress-testing — those eight reference voices are clones, and clone quality is where TTS engines separate.

Engines are also named in the docs as selectable per voice, which means HeyGen does not make you take its model: you can route to a third-party voice engine through its API where you prefer one. That is a vendor hedging its own model, and it is worth knowing when you read a "#1 voice" claim about HeyGen.

What this changes for anyone choosing a voice

Three practical consequences follow from a controlled-voice result landing, and none of them is "switch everything".

First, if your use case is a cloned brand voice — one speaker, many lines — this is now the board to read, and HeyGen Voice is the model to test first. A leaderboard that holds the voice constant is measuring the thing you will actually be doing.

Second, if your use case is a wide cast of native-sounding characters in languages your clone does not cover, the Provider Voices board and its ninety-six entries are still the relevant reference, and HeyGen Voice has no rank there yet. Nothing in a cloned-voice result tells you how the engine handles a hundred distinct voices.

Third, price now has a number, but not one the vendor published. The arena lists HeyGen Voice at $30.00 per 1M characters, an Artificial Analysis conversion of a credit system that HeyGen's own page never translates into a voice rate. That puts the new leader under Eleven v4 Turbo's $40.00 and above the Qwen tier's $19.30 — a mid-field price attached to a first-place score, and one that is derived rather than quoted.

The routing question

Voice selection has a property most model choices do not: the failure is audible to the end user within a syllable. That makes the migration cheap to test and expensive to get wrong in production. Running a new engine behind the same interface as the incumbent — one API surface, one key, provider list price passed through rather than marked up, with automatic failover to a second voice provider when a request fails — is how you get a real answer about your own audio instead of an Elo chart's answer about eight reference voices. OrcaRouter's routing layer exists for exactly that shape of decision: put the challenger on a fraction of traffic, keep the incumbent warm, and let failover cover the window where the new model is unproven. The leaderboard tells you where to look. It cannot tell you how your script sounds.

What to watch next

Two numbers will settle this story. The first is whether HeyGen Voice accumulates enough votes to enter the Provider Voices board, and where it lands against Eleven v4 Turbo's 1,328 — a model that leads that board yet trails in the controlled arena, which is precisely the gap the two boards exist to expose. The second is whether HeyGen publishes a voice rate of its own, so that the $30.00 the arena has derived can be checked against a vendor number rather than inferred from credits. Until both exist, a #1 debut on the controlled board is a strong claim about cloned-voice quality and an open question about everything else.