
Inworld Realtime TTS-2 vs Cartesia Sonic 3.6: The Voice That Listens vs the King of the Arenas
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
There are two competing theories of why a synthetic voice feels human, and right now the two best commercial examples of each are Inworld Realtime TTS-2 and Cartesia Sonic 3.6. Cartesia's theory is that humanity lives in the audio: make the timbre, the prosody and the timing indistinguishable from a person's, and listeners will rank you first — which is what Sonic 3.6 did in August, jumping to the top of Artificial Analysis's Provider Voice arena on Elo 1,282. Inworld's theory is that humanity lives in the listening: TTS-2 conditions its delivery on the actual audio of the conversation that came before it, so it does not just sound like a person, it responds like one. This week the two theories collided on the scoreboard: the same re-scoring that put Inworld Realtime TTS-2 at #2 on the provider board, on Elo 1,252, also put it at #1 on the Controlled Voice arena — the board Cartesia had led — at 1,123 against Sonic 3.6's 1,119. One model holds the provider crown. The other just took the controlled one.
This is not a matchup where one side is obviously wrong. The two models were built for overlapping but not identical jobs, and the price difference — Sonic 3.6 at $49.0 per million characters against $25 on demand for Inworld Realtime TTS-2 — is real enough that the decision has a budget component as well as a quality one.
Two theories of a human-sounding voice
Cartesia shipped Sonic 3.6 quietly in August 2026, a point release that improved on Sonic 3.5 without changing the price or the architecture story. The improvement showed up where Cartesia needed it: the model jumped to Elo 1,282 on Artificial Analysis's Provider Voice arena — where every model uses its own native voices — and it held the top of the Controlled Voice arena through August. On the controlled board every model is cloned onto the same eight reference speakers, so that crown is a verdict on the synthesis engine itself, independent of voice talent. After Inworld's general-availability re-score this week, Cartesia still leads the provider board, but Sonic 3.6 now sits #2 on the controlled board at Elo 1,119, four points behind Inworld Realtime TTS-2 on 1,123.
Inworld Realtime TTS-2 comes from a different tradition. Its predecessor line, TTS 1.5, spent early 2026 near the top of the same arenas, and the research preview of TTS-2 measured Elo 1,209 on the provider board in June. The GA model has climbed since: on the current boards it sits at 1,252 on Provider Voice (#2, behind Sonic 3.6's 1,282) and at 1,123 on Controlled Voice (#1). The flagship's real differentiator, though, is not its Elo; it is that the model takes the prior turns of a conversation as audio input and lets that shape how it delivers the next line. Cartesia calibrates emotion from the transcript. Inworld listens to how the last thing was said.
The scoreboard, honestly labeled
Elo figures are from Artificial Analysis's speech arenas, read from the live boards in September 2026. Latency figures are vendor-reported unless noted, and no independent latency audit has been published for either model.
• Independent quality — Cartesia Sonic 3.6: Elo 1,282 (#1, Provider Voice) and 1,119 (#2, Controlled Voice) vs Inworld Realtime TTS-2: Elo 1,252 (#2, Provider Voice) and 1,123 (#1, Controlled Voice)
• Price per 1M characters — $49.0 (Artificial Analysis-tracked) vs $25 on demand, $20.8 blended as tracked by Artificial Analysis, down to $12.50 at volume (vendor)
• Time to first audio — sub-90 ms (vendor-stated) vs sub-200 ms median, sub-100 ms at p99 in Inworld's launch tests (vendor-stated); Inworld's low-latency answer to Sonic is the separate TTS-2 Flash tier at ~25 ms time-to-first-byte
• Languages — 42 localized vs 100+ languages for delivery steering, with Inworld's one-voice-identity claim extending to 200+ locales (vendor-stated)
• Instant voice cloning — ~10 seconds of audio vs 5–15 seconds, plus a professional-cloning beta that takes ~10 minutes of audio for higher-fidelity clones
• Delivery control — inline transcript tags like [laughter], plus volume and speed modulation vs plain-English steering such as "[calm, reassuring]" or [whisper], request-level instructions, and three stability modes from Stable to Expressive

Why "listens to the conversation" is a different axis
The scoreboard above measures how a voice sounds in isolation. The dimension where the two models genuinely diverge does not show up on any leaderboard, because it only exists across multiple turns.
Inworld Realtime TTS-2 is built for the case where a person just said something to it. The model ingests the audio of prior turns — not just a transcript of them — reasons about tone, pacing and emotional state, and selects a response state before it speaks. If a caller is frustrated, the delivery shifts; if they are relaxed, it shifts back. That is why Inworld markets the model for agents, companions and games rather than for narration: the feature is meaningless in a one-shot text-to-audio call, and meaningful in a conversation.
Cartesia Sonic 3.6 works differently. It is a fast, high-quality streaming engine, and Cartesia's own claim is automatic emotional calibration from the transcript — it reads the words and modulates accordingly. What it does not do is listen to the audio of the preceding turns. Within a single utterance the two can sound comparable; across a conversation, Inworld is modeling something Sonic does not attempt. If your product is a one-turn voiceover, that modeling is overhead. If it is a live agent, it is the product.

Speed, price, and the Flash caveat
On pure latency, Cartesia has held the bragging rights: sub-90 ms time-to-first-audio is vendor-stated, but it is the figure the company has shipped since the Sonic 3 generation, and no one has published a contradicting audit. Inworld's flagship answers in a similar conversational class at sub-200 ms, which is fine for live agents. The real Inworld latency play is the Flash tier that shipped alongside the GA model — a separate model at ~25 ms time-to-first-byte, aimed at high-volume workloads where Sonic-class cost and latency matter more than expressiveness. The caveat is that Flash is a stripped-down member of the family: instant cloning and inline non-verbals, but no professional cloning and no full natural-language steering.
Price cuts the other way. Sonic 3.6 costs $49.0 per million characters — roughly double Inworld's $25 on-demand rate, and nearly four times the $12.50 top tier. The quality gap between them, on the boards where both appear, does not justify that multiple for a high-volume buyer. For a real-time voice agent generating thousands of characters per conversation, the per-character gap is the difference between a healthy unit economy and a thin one.
Which to pick
• Pick Cartesia Sonic 3.6 if the provider-board crown is what you want — the highest-scored voice using its own native voices, Elo 1,282, for short self-contained utterances like assistant answers, game lines and status readouts. At $49 per million characters you pay for the crown, and it remains the model to beat on that board.
• Pick Inworld Realtime TTS-2 if the product is a multi-turn conversation where delivery should respond to the person on the other end — support calls, a companion, an interview, a tutoring session. It now leads the controlled board, where every model speaks the same eight reference voices, and it does so at roughly half the price while listening to the audio of the conversation.
• Build the switch into routing if you cannot know in advance which kind of voice a call needs. Neither model is on OrcaRouter as of this writing — Sonic is reached through Cartesia's own API and several third-party platforms, Inworld through its own API — but a routing layer is exactly the structure that lets a conversation start on one voice and fail over to the other, with each provider's list price passed through at 0% markup. The day one of these boards flips again — and this week they did — the choice becomes a config change instead of a rewrite.
The short version
This is a genuine fork, and the honest answer is that the fork is about turn-taking, not audio quality. If you need the best-scored standalone voice using its own native voices, Cartesia Sonic 3.6 holds the provider crown on Elo 1,282 and charges $49 per million characters for it. If you need a voice that responds to how it was spoken to, Inworld Realtime TTS-2 went GA this week at $25 on demand, leads the controlled board where the two meet on equal voices, and carries a conversational-awareness feature no arena fully measures. The scoreboards are now split between the two models — which is the most honest possible picture of a market where "best" depends on the test.

