
Gemini 3.8 TTS vs Sesame Preview: One of These Is a Download, the Other Is an App
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Search for a comparison between Gemini 3.8 Flash TTS and Sesame Preview and you will find plenty of writing that treats them as two products in the same category. They are not, and the difference is not a technicality. Gemini 3.8 Flash TTS is an API endpoint: it has a model identifier, a published rate card, a leaderboard row at rank 2 with an Elo of 1,260 across 1,999 samples on the Artificial Analysis Provider Voice Arena, and a documented request format. Sesame Preview is a free consumer voice-assistant app with named agents, multi-turn conversations and a 30-minute call cap. It has no API, no model ID, no billing plan and no score on that board.
The one Sesame artifact you can actually build on is a different thing again: CSM-1B, an Apache 2.0 checkpoint on Hugging Face that generates audio codes from text using a Llama backbone and a Mimi audio decoder. It is not the voice people are talking about when they describe how human the app sounds. So the comparison resolves into a question that is answerable: do you want to rent a speech model, or do you want to download one?

Two things called Sesame, and only one of them is buildable
The naming is the source of most of the confusion, so separate it before going further.
• Sesame Preview — the app. Free to use, agents named Maya, Miles, Simone and Charlie, multi-turn conversation with context that persists across turns, web search and reminders as capabilities, English only, a 30-minute cap on a call. There is no developer surface: nothing to authenticate against, nothing to meter, no endpoint to call.
• CSM-1B — the checkpoint. Published on Hugging Face under sesame/csm-1b with an Apache 2.0 licence, a Llama backbone and a smaller decoder that produces Mimi audio codes. Roughly 2 billion parameters, English, F32 weights, and it runs through Hugging Face Transformers. The model card records the 1B variant shipping in March 2025 and native Transformers support arriving in May 2025.
Those two share a company and a research lineage, and that is roughly where the relationship ends. The app is a product; the checkpoint is an artifact from the research that preceded it. Reviewers who praise the app's conversational memory are describing a system with retrieval and state management around a speech model, not a property of the 2B checkpoint you can download.
What is actually in the checkpoint
CSM-1B is a text-to-speech model in the architectural sense that matters for anyone planning to run it: it takes text and audio context as input and emits RVQ audio codes, which the decoder turns into waveform. The practical facts from the model card:
• Licence — Apache 2.0. This is the single most consequential line on the page. Unlike research-only weight licences, Apache 2.0 permits commercial use without a separate agreement, which makes CSM-1B one of the few genuinely usable open speech checkpoints.
• Size — about 2B parameters at F32. Large enough to need real hardware for anything concurrent, small enough to run on a single accelerator for development.
• Language — English, per the card. Not a multilingual model.
• Traction — 141,461 downloads in the last month at the time of writing, with roughly 245,000 likes on the repository. That is a widely-used checkpoint by open speech-model standards, and it is the strongest argument that the artifact has a real user base independent of the app's press.

What the card does not give you is a quality claim you can compare against anything. There is no arena row, no vendor benchmark reproduced by a third party, and no published evaluation of how the checkpoint's output relates to the app's. Anyone who tells you the downloadable model sounds like the app is inferring it from the name.
Why there is no head-to-head, and why that is not a research gap
There is no Gemini 3.8 TTS versus Sesame Preview comparison on the leaderboards because there cannot be one. The arena ranks models by having listeners compare clips produced by a hosted endpoint. One side of this pairing does not have an endpoint, so it cannot be entered.
That is worth being precise about, because "no head-to-head exists" is often written up as if the data were missing. It is not missing. It is undefined. The two things being compared do not share a measurable surface:
• Gemini 3.8 Flash TTS is scored on its native hosted voices. Sesame Preview has no native voices to score in that sense — it has agents.
• Gemini 3.8 Flash TTS publishes a rate card in tokens. Sesame Preview is free and publishes nothing, because there is nothing to bill.
• Gemini 3.8 Flash TTS takes a script and returns audio. Sesame Preview takes a conversation and returns a conversation, with whatever retrieval and memory the app layers on top.
The closest thing to a legitimate comparison is between Gemini 3.8 Flash TTS and CSM-1B, and even that is uneven: one has an independent Elo and a vendor rate card, the other has neither. What you can compare is the shape of the commitment — a metered API against a downloadable checkpoint — and that comparison does not need a leaderboard.
The only side of this with numbers on it
Everything quantitative in this pairing sits on Google's side, which is itself the point.
On the Artificial Analysis Provider Voice Arena, captured September 24, 2026, Gemini 3.8 Flash TTS is at rank 2 with an Elo of 1,260 across 1,999 samples and 8 arena voices, released September 23, 2026. Its sibling Gemini 3.8 Flash-Lite TTS is at rank 6 with 1,235. Google bills Flash TTS at $0.50 per million text tokens in and $9.00 per million audio tokens out through December 31, 2026, then $18.00, with Flash-Lite TTS at $0.50 and $6.00, then $12.00; Artificial Analysis normalises those to $16.50 and $11.00 per million characters so they can share a board with models that bill in other units. That normalisation is the board's arithmetic, not a Google quote.
Against that, CSM-1B has no price because it has no meter. It has a download count, a licence and a parameter count. If your decision framework needs an Elo, this comparison will not supply one for the Sesame side, and the honest conclusion is that the two options are being evaluated on different criteria rather than that one of them is unproven.

If what you wanted was the app experience
A lot of people searching this comparison are not choosing a model at all. They used the Sesame app, liked how the conversation felt, and want that in their product. That is a different problem, and the download will not solve it.
What makes the app feel the way it does is mostly not the speech model. Persistent context across turns, the decision to search the web mid-conversation, the timing of when an agent speaks — those are orchestration, and none of them live in a 2B checkpoint. Downloading CSM-1B gives you a voice. Building the app's behaviour on top of it is your engineering, and the 30-minute call cap and the absence of an API mean you cannot rent the finished version either.
If what you wanted was a speech model you can call, the honest answer is that Gemini 3.8 Flash TTS is the option with a documented interface, a published price, an independent score, and two-speaker dialogue in a single request. If what you wanted was weights you can keep, CSM-1B under Apache 2.0 is the one of the two you can actually hold — and the trade is that you supply the hardware, the serving stack and every quality guarantee yourself.
OrcaRouter hosts neither of these. Google's model is called through Google's own API and the surfaces named in its September 23 announcement; CSM-1B is a checkpoint you run. What OrcaRouter is for is the layer above that choice: 200-plus models behind a single key at provider list price with no markup, so a vendor rate change is live the same day, automatic failover so an experimental route is not a production dependency, and a routing DSL that composes models into one call when a single voice is not the whole job.
The short version of this comparison is that it is not really a comparison. One side has a rate card and an Elo; the other has a licence and a download count. Pick the one whose column you can actually fill in.
