
Realtime-Venus vs Sesame Preview: The Thing You Can Get Is Not the Thing That Is Good
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Realtime-Venus and Sesame Preview are built by labs with completely different ideas about what a voice model is for, and they have stumbled into the identical trap from opposite directions. Sesame Preview is a free consumer app — iOS since May 2026, Android since early August — with four conversational agents, live web search during calls, and a voice that reviewers have called best-in-class. The open weights Sesame published are a different, smaller model that the company itself does not present as the voice you heard in the demo. Realtime-Venus, the 9B full-duplex system from Ant Group's Venus Team and Tsinghua University, went the other way: the downloadable artefacts are the real thing, and there is no product at all. In both cases the gap between what is available and what is impressive is the most useful thing to understand before you plan anything around either one.
What each one actually is

Sesame Preview is a product. Sesame is a San Francisco lab founded in 2023 by Brendan Iribe, Ankit Kumar and Ryan Brown, backed by a16z, Sequoia, Spark and Matrix, and its consumer voice assistant ships four personalities — Maya, Miles, Simone and Charlie — each with a distinct temperament and voice. The agent can search the web mid-conversation, take notes and reminders, and send a summary after a call. Memory is per-agent and syncs across web and mobile when you are signed in, with an incognito mode that reads past memories without writing new ones. It is free, it runs no ads, and the company says it does not sell data.
Realtime-Venus is a research release. Two 9B checkpoints under Apache-2.0 on Hugging Face — Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction — plus a technical report submitted to arXiv on 12 September 2026, a project page, and a companion repository holding the Realtime-Venus-Harness component. There is no app, no API, no price, and no announcement from the vendor. The repository is not deployed by any inference provider, and downloads are not tracked.
The trap, in both directions
Sesame's version is the classic open-washing shape, and Sesame is more honest about it than most. The CSM checkpoint the lab released under Apache-2.0 is a base speech generation model — a Llama backbone feeding a Mimi-codec decoder — and the documentation is explicit that it is not the app's voice and not an end-to-end speech-to-speech system. It cannot generate text, and it is trained on primarily English data with limited capacity outside that. What powers Sesame Preview is a larger production model that has not been released. So if you went looking for the voice from the demo, you found the research lineage of it and not the thing itself. When one of the four agents answers your call, you are hearing something you cannot download.
Realtime-Venus's version is the mirror image, and arguably the stranger one. Everything published is genuine: real weights, real code, real report, a permissive licence. What is missing is any way to use it that does not involve owning a GPU. Two 9B checkpoints in BF16 are roughly eighteen gigabytes of weights apiece before activation memory, and both are designed to run a continuous one-second streaming loop with a live audio and video input. That is a server-class deployment. A consumer app that ships the same capability to a phone is doing something this release does not attempt, and the gap between a downloadable model and a runnable one is exactly where most of the people who want it will get stuck.
Measurability is the sharper difference
Neither model has an independent voice-quality score, but the reasons differ.
• Sesame Preview — no arena entry, no published evaluation of the production voice. Its reputation rests on reviewers and on the original demo's effect on listeners, which is real evidence of a kind and completely unquantified.
• CSM-1B — the open checkpoint is a base generation model; it is not benchmarked against conversational systems because it is not one.
• Realtime-Venus — one self-reported voice figure, a VoiceBench AlpacaEval score of 4.81 that its report describes as matching the best comparison number. No third party has evaluated it.
• What neither gives you — a number produced by someone with no stake in the result. If voice quality is your buying criterion, the entire comparison is unscoreable, and the honest move is to listen to both and decide for yourself rather than defer to a table.
Turn-taking, where both have something to prove
Sesame's reputation was built on exactly this. The demo that made the lab famous was a voice with breath, hesitation and interruption timing convincing enough that listeners assumed a person was on the line. That is a claim about conversational dynamics, and it is the thing the product is sold on.
Realtime-Venus makes the same claim with numbers attached. On Full-Duplex-Bench v1.5 its audio checkpoint reports responding to 75% of user interruptions, with continuation rates of 97% under backchannels, 88% under other-directed speech and 86% under background speech — stated to exceed Gemini 3.1 Live and GPT-4o on all three continuation metrics. Those figures come from its own report and nobody has reproduced them.
So the comparison is a demo with no number against a number with no demo, which is a genuinely awkward position to be in and a fair description of how this whole category reports on itself right now. The one thing that can be said with confidence is that neither claim has survived an outside party, and 97% continuation under backchannels is high enough that it deserves one before it gets quoted as settled.

What you can actually build
Sesame Preview gives a builder nothing. There is no API, no model identifier, no rate card and no stated plan for one. Calls cap at thirty minutes when logged in and five otherwise, English is the only officially supported language, and the agent will research and remember but not execute tasks — it will not book the flight. The free tier is the product. That is a perfectly good consumer offer and a dead end for integration.
Realtime-Venus gives a builder everything except a place to run it. The weights are permissively licensed, the code loads with standard Transformers calls, full-duplex streaming is entered with a single method switch, and the documented loop is one second of streaming prefill followed by streaming generation, repeated. Long-video memory is enabled with a memory-minutes parameter and is described as training-free, which is unusual for a capability that normally requires fine-tuning. Two caveats are documented rather than left to be discovered: a frame cap that silently truncates long video unless a constant is raised before the utility import, and a CJK font requirement for non-Latin subtitles rendered into duplex video.
What none of that changes is that the harness — the component that executes the asynchronous delegations which are the most distinctive part of the architecture — lives in a separate GitHub repository, is far less documented than the model, and is the part whose production readiness is hardest to judge. In a real deployment the delegation leg is ordinary text inference sitting behind a conversational front end. OrcaRouter carries neither Realtime-Venus nor Sesame Preview; both voice layers are outside what we serve, and we say that plainly rather than implying a hosting relationship that does not exist. Nearly 200 text models behind one key at provider list price with no markup is the half of that architecture we do cover, with automatic failover so a delegated task survives an upstream outage, a routing DSL that composes several models into one call, and model fusion where one model's judgment is not enough. It is the purchasable part of a design whose interesting part is not for sale.

The honest recommendation
If you are a consumer who wants the best-sounding conversational voice available today and does not need to integrate it, Sesame Preview is free, it is on both mobile platforms, and its reputation is deserved enough that reviewers keep saying so. Use it and enjoy it.
If you are a researcher or a well-resourced engineering team that needs an open audio-visual full-duplex stack you can inspect and fine-tune, Realtime-Venus is one of the very few real options, and the fact that its benchmarks are self-reported does not change that the weights are genuinely open and the architecture is genuinely documented.
If you are a product team looking for a voice layer to ship this quarter, neither of these is your answer, and the useful takeaway from the pairing is why. One lab built something excellent and kept the good part closed; the other published everything and left you to supply the infrastructure. The category has proved it can make a voice worth listening to and has not yet decided whether that voice should be purchasable in any form other than an app.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
