A generated hero card titled 'HeyGen Voice vs Sesame Preview' contrasting a documented speech API carrying a measured Elo against a consumer demo with no public endpoint, marked with the Artificial Analysis date of 9 October 2026. The OrcaRouter logo appears in the bottom-right corner.
Guides & Insights

HeyGen Voice vs Sesame Preview: The Voice You Can Buy Against the Voice You Can Only Talk To

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

This is not a model comparison, and pretending otherwise would be the only real error available here. HeyGen Voice is an API engine — ranked first on Artificial Analysis's Controlled Voice arena at 1,202 Elo, exposed through HeyGen's v3 endpoints, sold by the credit, clone-able from a single recording. Sesame Preview is a free consumer voice assistant you reach by opening a web page and speaking to one of four named personas, with no public endpoint, no model identifier, and no price at which it can be rented. One of them you can ship. The other is the reason you know what a good conversational voice sounds like, and never the thing you deploy.

What each one actually is

Sesame Preview is a demonstration of conversational speech. You reach it through the company's own site, you talk to Maya, Miles, Simone or Charlie, and the experience is deliberately optimised for friendliness and expressivity — the company says as much when describing why the companions behave the way they do. It is English-only today, with the team noting that multilingual ability shows up incidentally from data contamination and "does not perform well yet," and that expanding past twenty languages is a future plan. The company's own framing is a work in progress: "We're not there yet."

Underneath the demo is the Conversational Speech Model, a multimodal text-and-speech architecture built from two autoregressive Llama-style transformers. It ships in three sizes — Tiny with a 1B backbone and 100M decoder, Small at 3B and 250M, Medium at 8B and 300M — each trained at a 2,048-token sequence length, roughly two minutes of audio, for five epochs over about one million hours of mostly English speech. Key research components are open-sourced under Apache 2.0, and there is a GitHub repository for them. The demo — the thing people actually praise — is not that. The demo has no public API.

HeyGen Voice is the inverse arrangement. There is no free consumer app to try; there is a documented API with a voice catalogue, a speech endpoint, clone paths and a bill. What you cannot do is open it in a browser and have a conversation with it on a whim.

Why the two get compared anyway

They get compared because both are held up as evidence that synthetic speech has crossed something. Sesame Preview reset expectations for what conversational audio could sound like — the reactions when it landed were about how much it sounded like a person rather than a text-to-speech engine. HeyGen Voice's 1,202 on the controlled board is the same claim in numeric form: first place among forty-two entries when every model speaks the same eight cloned reference voices.

The difference is what you can do with the claim afterwards. A board position is reproducible, testable, and attached to an endpoint. A demo impression is none of those things — and, importantly, Sesame Preview does not appear on the controlled board at all, so there is no Elo to compare against the 1,202. The comparison people want to make is unavailable, not because Sesame lost it but because the model was never entered.

The scoreboard

A generated two-column scoreboard titled 'HeyGen Voice vs Sesame Preview — the scoreboard'. The HeyGen Voice column reads 'Controlled Voice Elo: 1,202, #1 of 42', 'Access: documented v3 API with a speech endpoint', 'Price: $30.00 per 1M characters on the arena's conversion of credit pricing', 'Voice selection: 300+ pre-built voices, design and cloning', 'Languages: dozens of languages, count unpublished' and 'Open weights: none'. The Sesame Preview column reads 'Controlled Voice Elo: not listed', 'Access: consumer web and iOS app, no public endpoint', 'Price: free, not purchasable at any price', 'Voice selection: four fixed personas', 'Languages: English only, multilingual underperforming' and 'Open weights: CSM under Apache 2.0 in 1B, 3B and 8B'. A footer notes the arena price is a conversion of credit pricing rather than a vendor-quoted rate.

• Controlled Voice Elo — HeyGen Voice: 1,202, #1 of 42. Sesame Preview: not listed.

• Access — HeyGen Voice: documented v3 API, POST /v3/voices/speech. Sesame Preview: consumer web app and iOS app; no public endpoint.

• Price — HeyGen Voice: $30.00 per 1M characters on the arena's own conversion of credit pricing; HeyGen publishes no voice rate. Sesame Preview: free, and not purchasable at any price.

• Voice selection — HeyGen Voice: 300+ pre-built voices, text-described voice design, instant and professional cloning. Sesame Preview: four fixed personas.

• Languages — HeyGen Voice: "dozens of languages", count unpublished. Sesame Preview: English only, with multilingual support underperforming today by the company's own account.

• Open weights — HeyGen Voice: none. Sesame Preview: CSM released under Apache 2.0 in 1B, 3B and 8B sizes.

A screenshot of Artificial Analysis's Controlled Voice arena leaderboard, captured 10 October 2026, showing the language selector set to 'English (US & UK)' and the ranked field with HeyGen Voice first at 1,202 Elo and Qwen-Audio-3.1-TTS-Plus second at 1,186; no Sesame entry appears anywhere on the board.

The Apache 2.0 detail is the one that matters

Buried under the "no API" verdict is the fact that Sesame open-sourced the research model under Apache 2.0 — a permissive licence, in three sizes, with a 1B variant small enough to run without a data centre. That is a genuinely different proposition from the demo, and it is the only route by which anything resembling Sesame Preview's voice ends up in a shipped product.

Two caveats keep it honest. The released CSM components are research artefacts: the company notes that the models handle speech content but not conversation structure, and points toward fully duplex models as future work — so the released weights are not the interactive experience. And an 8B backbone running locally is a different cost structure from $19.30 per million characters of hosted inference, which is what the cheapest top-tier rival on the arena charges. Self-hosting has no per-character bill and a very real per-GPU-hour one.

For a builder, that reframes the question. If what impressed you about Sesame Preview was the conversation, the open weights do not give you it. If what impressed you was the voice quality of the speech model itself, they might — and that is testable in a way a demo is not.

What shipping on HeyGen Voice looks like

The practical differences are the ordinary ones. HeyGen's API takes a voice selection, text, and optional speed between 0.5 and 1.5 and pitch across ±50 semitones, with a BCP-47 locale hint. Custom voices come three ways: a description-driven design route that returns up to three ranked options from a prompt capped at 1,000 characters, an instant clone from one recording, and a professional clone trained on twenty minutes or more. The engine filter on the voices endpoint also accepts non-HeyGen engines, so the platform is not a walled garden around its own model.

None of that exists on the Sesame side, and it is not a shortcoming of Sesame Preview — it is what the thing is. A demo optimised to show what is possible should not carry an SLA.

The bill, though, is the softer half. HeyGen sells credits — 600 for $29/month, 1,000 for $49, 1,500 for $149 — and states that consumption varies by model, duration and complexity, without publishing a voice rate. The $30.00 per million characters the arena lists is Artificial Analysis's conversion of those credits, not a quote from HeyGen. So the model you can ship is priced, but on someone else's arithmetic — which is an argument for a measured pilot rather than a criticism of the engine.

A screenshot of HeyGen's developer documentation sidebar, captured 10 October 2026, showing the Voice Management group with the entries Voices, Browse Voices, Design a Voice, Voice Clone and Text to Speech.

What actually to do with a demo you admire

The useful move is to separate the two things Sesame Preview demonstrates. The conversational behaviour — turn-taking, memory, the sense of talking to someone — is the product of a system that has not been released and has no endpoint. The speech quality is the part with Apache 2.0 weights behind it, and the part you can evaluate on your own audio without asking anyone's permission.

For everything in between, the migration path is the mundane one: ship on an engine that has an API and a rank, and design so that adopting Sesame's model if it ever ships commercially is a configuration change rather than a rewrite. That means an abstraction over the voice provider from day one — one interface, one key, engine selected by config. OrcaRouter's routing layer is built for precisely this: many providers behind a single endpoint at provider list price with no markup, automatic failover when a vendor stalls, and a routing DSL that lets you swap the engine behind a voice without touching application code. The point is not that it hosts a Sesame model — it does not — but that the day a voice you admire becomes purchasable, the cost of trying it should be a line in a config file.

The honest close

Sesame Preview is not a competitor to HeyGen Voice and it should not be evaluated as one. It is a free, English-only, four-persona demonstration that happens to have reset expectations for conversational audio and happens to have released permissively-licensed research weights in three sizes. HeyGen Voice is the top-ranked engine on the only speech leaderboard that holds the voice constant, sold by an API you can call today, at a price you cannot yet calculate.

If you need a voice you can ship this quarter, the decision is not between these two — it is between HeyGen Voice and the other priced engines on the arena. If you are waiting on Sesame, wait on the weights, not the app.