
HeyGen Voice vs Cartesia Sonic 3.6: Cloned-Voice Champion Against the Sub-90ms Incumbent
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 112 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 51 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 61 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 404 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
HeyGen Voice and Cartesia Sonic 3.6 hold two different number-one claims on Artificial Analysis, and the entire comparison lives in the gap between them. HeyGen Voice tops the Controlled Voice arena at 1,202 Elo — the board where every model speaks the same eight cloned voices, four US and four UK. Cartesia Sonic 3.6 is fourth on the Provider Voices arena at 1,280 Elo, where each vendor supplies its own native voices, and Cartesia's own site carries a badge claiming a #1 ranking across both the Speech Arena and the Speech to Text leaderboard, a vendor-displayed claim the live Artificial Analysis table currently does not support at the top position. One engine wins the test where the voice is held constant, the other leads the test where the production budget is real.
Start with what each model is actually optimised for
Sonic 3.6 is built to be consumed in a stream. Cartesia's own product page leads with "first audio in under 90ms" — vendor-reported, unreproduced — and the same page describes a real-time TTS API with first-class streaming. It covers 44 languages from a single model, interprets emotional subtext by default, accepts non-verbal cues such as laughter written into the transcript, clones a voice from roughly ten seconds of audio, localises existing audio into other languages, and takes custom pronunciation dictionaries for domain terms and proper nouns. Every one of those features is aimed at the same customer: something that talks back, fast, in a language the caller understands.
HeyGen Voice is built to be consumed as a performance. It is HeyGen's in-house engine, exposed through the company's v3 API, with 300+ pre-built voices across dozens of languages and three routes to a custom voice — description-driven voice design returning up to three ranked options from a prompt capped at 1,000 characters, an instant clone from one recording, and a professional clone trained on twenty minutes or more. Per-request tuning covers speed from 0.5 to 1.5 and pitch across ±50 semitones. Nothing in HeyGen's published material promises a latency figure at all.
That asymmetry is the whole article. One model publishes a millisecond number and no per-character price; the other publishes a price and no latency number.
The scoreboard

• Controlled Voice Elo (same 8 cloned voices) — HeyGen Voice: 1,202, #1 of 42. Sonic 3.6: 1,143, #5.
• Provider Voices Elo (native voices) — HeyGen Voice: not ranked, insufficient votes. Sonic 3.6: 1,280, #4 of 96.
• Published price — Sonic 3.6: $49.00 per 1M characters, vendor rate card. HeyGen Voice: $30.00 per 1M characters on the arena's own conversion of HeyGen's credit pricing; HeyGen publishes no voice rate itself.
• Latency — Sonic 3.6: under 90ms to first audio (vendor-reported, unreproduced). HeyGen Voice: no figure published by the vendor or measured by Artificial Analysis.
• Languages — Sonic 3.6: 44 from a single model. HeyGen Voice: 300+ voices across "dozens of languages", not enumerated as a count.
• Custom voice path — Sonic 3.6: ~10-second clone. HeyGen Voice: voice design from a text description, instant clone, or a professional clone trained on 20+ minutes.

The Elo gap, read correctly
A 59-point lead in the controlled arena inverts to a 137-point deficit on the provider board, and the inversion is not a contradiction. The controlled board removes the vendor's ability to pick a flattering voice and asks how the engine handles a voice it did not choose. HeyGen Voice wins that. The provider board asks how good a vendor's own best voices sound, which is closer to what a customer hears in a sales demo — and Sonic 3.6 does well there, fourth of ninety-six and 38 points behind the top-ranked Eleven v4 Turbo.
Neither number is the general-purpose truth. If you are cloning a specific person and need the engine to render that person convincingly, the controlled result is the better predictor, and HeyGen Voice currently leads it. If you are picking a voice from a library for an agent that has to sound distinctive at scale, the provider result is the better predictor, and Sonic 3.6 has a rank where HeyGen Voice has none.
Worth noting for anyone holding older figures: both of these boards move as votes accumulate. Earlier in 2026 Sonic 3.6 was reported leading the controlled-voice board outright; the current table has it fifth there at 1,143 with HeyGen Voice, Qwen-Audio-3.1-TTS-Plus, and two Eleven v4 tiers above it. A leaderboard citation is a timestamp, not a permanent property of the model.
Where the pricing comparison breaks
Sonic 3.6's $49.00 per million characters is the arena's own benchmark price, drawn from the vendor's published API rate at default settings. That makes a five-figure monthly audio bill calculable before you write a line of integration code.
HeyGen's number is softer, but it exists. The arena's price chart lists HeyGen Voice at $30.00 per million characters — nineteen dollars under Sonic 3.6 — and that figure is Artificial Analysis's conversion of HeyGen's credit pricing, not a rate HeyGen publishes. HeyGen's own page sells credits (Creator at $29/month for 600 credits, Pro at $49 for 1,000, Business at $149 for 1,500) and states plainly that consumption varies by model, duration, and generation complexity, without ever translating those credits into characters. So the model that wins the controlled-voice board is priced well below the one that ranks fifth, and the gap is a derived number you will want to confirm against your own text rather than a rate card you can audit.
That is not an argument against HeyGen Voice. It is an argument for testing it on the traffic you can afford to measure, because the only way to learn its real cost in your pipeline is to run your own text through it.
Choosing between them
Pick Sonic 3.6 when the voice has to answer. A sub-90ms first-audio claim, a streaming-native API, 44 languages from one model, and — critically — a published rate card make it the lower-risk production choice for conversational agents, IVR, and anything where a pause is a bug. Its controlled-voice rank of #5 is respectable, not a weakness, and its fourth place on the broader board is the strongest independent signal either model carries.
Test HeyGen Voice first when the voice has to perform. Narration, brand-read content, character work, and any project where you are cloning one speaker and need that speaker's clone to be the best-sounding thing on the wire — that is the job the controlled arena was constructed to measure, and it is the job HeyGen Voice just won. The professional-clone path trained on twenty-plus minutes is a materially different product from a ten-second instant clone, and it is the one that lines up with the eight cloned reference voices in the arena.
Pick neither on the strength of a single Elo number. The two boards disagree about these models, which means the boards are telling you something real: quality depends on the voice and the workload, not on the engine alone.

Running both without committing
The practical path through a disagreement like this is to stop treating it as a procurement decision. Put the two engines behind one interface, split traffic, and judge on your own audio — a handful of lines of narration and a live agent loop will tell you more than fifty Elo points will. OrcaRouter routes both through a single API key at provider list price with no markup, so the vendor's own rate card is what you pay and a Sonic 3.6 price move shows up the same day rather than at renewal. Automatic failover matters more here than it does for most model swaps, because a voice provider that stalls mid-sentence is audible to the caller, and the routing DSL lets you pin one engine for latency-sensitive turns and the other for pre-rendered narration without maintaining two integrations.
What would settle it
Two things. A published per-character or per-credit voice rate from HeyGen would make the cost half of this comparison as concrete as the quality half, and HeyGen Voice accumulating enough votes to enter the Provider Voices arena would tell us whether its controlled-voice lead survives contact with native voices — the test Sonic 3.6 has already passed at fourth of ninety-six. Until then, Sonic 3.6 is the incumbent with a price and a latency figure, and HeyGen Voice is the challenger with the better cloned-voice evidence and no cost to put next to it.
