
Claude Opus 5.5 and Gemini 3.8 Flash TTS Produced a Narrated News Video: The Brief, the Two Rate Cards, and What the Voice Costs
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 111 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 982 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 43 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two models, two vendors, one finished video, and a prompt short enough to paste into a chat box. On October 2, 2026, the Chinese-language AI writer 宝玉 (X/@dotey) published two renders of the same gossip roundup — one in English, one in Chinese — and said both were built end to end by Claude Opus 5.5, which found its own material and wrote its own narration, with the voice supplied by Gemini 3.8 Flash TTS using an API key the author provided. Neither model is new: Anthropic shipped Claude Opus 5.5 on September 22, 2026, and Google's release notes date the Gemini 3.8 text-to-speech models to the same day. What is days old is the packaging — the producer brief itself, and a run of repositories created between September 30 and October 3 that turn that brief into a Claude Code skill you can install. This is a piece about the workflow, not about a launch.
What the brief actually specifies
The template is one sentence with three slots. Rendered into English from the original Chinese, it reads: make me a gossip video about {topic} (supplementary material: {links/screenshots, optional}; de-identified version: remove {person's name}, optional). Everything interesting about the format is hidden in those three slots, because a brief that short is only delegatable if the recipient can do three things without being told: find its own source material, decide what the story is, and hand back a finished artifact rather than a plan.
The topic is a free-text string, so the model is doing editorial selection — what counts as the story, which beats to include, what order they go in. That is a different job from summarisation, and it is the part of the pipeline that the operator has least control over. The optional supplementary material is the escape hatch: paste a link or a screenshot and the model works from a fixed source instead of searching, which is the difference between a reproducible render and a different video every time you run it. The third slot, 去敏, is the one worth dwelling on and we will come back to it.
Read the brief as a specification and it tells you what it assumes about the harness. It assumes the model can reach the outside world — the discovery step is delegated, not described — and what the model can reach is a property of the tools the operator mounts, not of Claude Opus 5.5 itself. It assumes an image or rendering stack, because the delivered artifact is a video and Claude Opus 5.5 does not emit video. And it assumes a voice.
The division of labour is the story
That last assumption is the structural fact of this workflow. Our own catalogue entry for anthropic/claude-opus-5.5 lists its input modalities as text, image and file, and its output modality as text — one modality out. The model cannot synthesise speech, and it cannot hear a track once it exists. Every narrated video built this way therefore has at least two model dependencies with two rate cards, two vendors and two credentials, and the second one has to be supplied by a human. Opus 5.5 can write the script and wire the render; it cannot open a Google account or provision a key. The author of the October 2 post says exactly that: the Gemini key was his.
The constraint has a second edge that is easy to miss until you ship. Because Claude Opus 5.5 takes no audio in, the model that wrote the narration cannot check the narration. Mispronounced names, odd emphasis, a sentence the synthesiser renders flat — none of that reaches the model that authored it. Quality control on the voice is a human listening to the output, every time, and that is the step that keeps this from being a fully unattended pipeline no matter how good the script is. Where the model's own output is a program, the review is a listen, not a diff.

What the voice costs — and it is not the expensive part
Google's pricing page publishes the conversion that makes this arithmetic clean: audio tokens correspond to 25 tokens per second of output. The two renders in the October 2 post run 351 seconds and 332 seconds, read from the tweet's own media metadata — roughly five minutes fifty-one and five minutes thirty-two. At 25 tokens a second those are 8,775 and 8,300 audio tokens, and at the current Flash TTS audio rate of $9.00 per million tokens, that is about eight cents of narration per video and roughly fifteen cents for both language versions together.
Put that next to the reasoning side of the same job. Claude Opus 5.5's list rate is $4 per million input tokens and $20 per million output, with thinking tokens billed as output, and cache reads at $0.20 per million. A six-minute video's narration is a rounding error against the tokens the model spends deciding what the story is, reading sources and writing the script that the voice reads. The practical conclusion is not "TTS is cheap" — it is that in this pipeline the voice is the one line item you do not have to optimise, and the effort belongs on the part that costs dollars.
Two rate-card details will change that arithmetic, so plan against them now:
• The promotional window closes — Flash TTS standard is $0.50 per million text tokens in and $9.00 per million audio tokens out through December 31, 2026, then $1.00 and $18.00. The same video's narration costs about sixteen cents in January, exactly double.
• There is a cheaper tier with a real trade-off — Gemini 3.8 Flash-Lite TTS bills $0.50 and $6.00 on the same schedule, rising to $1.00 and $12.00, covering 101 languages against Flash TTS's 130. For a single-narrator explainer in English, Lite is two-thirds the audio rate and the language ceiling is irrelevant; for the bilingual render in the October 2 post it is not obviously irrelevant, which is the decision the tier split is asking you to make.
Google's Batch and Flex tier is $0.25 in and $4.50 out, and Priority is $0.90 and $16.20 — worth knowing if your renders are queued rather than interactive, because nothing about a gossip roundup requires the answer in real time.

This week the brief became a skill
A prompt in a tweet is a demonstration; a repository is a thing you can install. The repositories that appeared around this format are the part that is genuinely inside the last seven days, and they are worth reading separately rather than as a set, because each one makes a different bet about how much of the pipeline should be fixed in code.
• hoangduong92/idea-to-film-skill — created September 30, 2026, MIT licensed, and the closest match to the October 2 post: a Claude Code skill that takes an idea to a narrated short-film MP4 with code-drawn animation, music, sound effects, a Gemini TTS voice and subtitles. The voice is named in the description, which is the honest way to ship this: the second vendor is a declared dependency, not a hidden one.
• samuelstroschein/opus-video-generator — created October 2, MIT licensed, and the most opinionated about credentials: it makes launch videos in your browser on your own Claude subscription, run through npx. That is a different answer to the key problem above — keep the human's existing entitlement in the loop instead of minting a new API key for an agent.
• makevoid/motion-graphics-music-video-skill-noimage — created October 2, MIT licensed, and the one with a discipline worth copying: it states in its own description that it uses no image model and no video model, producing motion graphics from a music track and a prompt. When the only generative step is code, the output is reproducible and every asset licence in the finished piece is yours to reason about.
• RohanRatwani/explainer-video-opus-5.5 — created October 2, MIT licensed, targeting the 16:9 and Shorts explainer rather than the gossip roundup, and it names the timing problem explicitly: every animation timed to the narration. That ordering matters, because it means the audio is rendered before the visuals are placed.
• zhuyansen/awesome-opus-5.5-video — created September 27, no licence declared, and by far the most visible of the group at 67 stars, with the others in single digits or at zero. It is an index rather than a tool, in English and Chinese, and its existence is a useful signal about the shape of the interest: what people want first is a catalogue of what others claim to have made, and the tooling follows.
None of these is a stable dependency. They are days old, they are single-author, and their licence terms are the only thing standing between you and a piece of footage you cannot use. But the direction is the point: the prompt in the tweet is the prototype, and the skill in the repository is what a team would actually adopt.
Two language versions, one story
The October 2 post is deliberately bilingual — an English render and a Chinese one of the same topic. That is unusual for a demo and it exposes something real about this format, which is that gossip does not travel by translation. The beats that carry a story in one market (who said what to whom, which denial was issued, what the timeline looks like) are not the beats that carry it in another, so the second version is a re-write with a different editorial judgment, not a translation pass. What does transfer is the skeleton: the same story structure, the same rendering stack, the same voice.
The cost consequence is small in the right direction. On the arithmetic above, adding the second language costs about eight cents of additional audio against a story the model has already researched once. When the expensive part of a pipeline is the reasoning and the cheap part is the output, producing a second version in another language is close to free — which is an argument for treating localisation as a first-class output of the workflow rather than a separate project.
De-identification is an edit, not a delete
The third slot in the brief is the one that should slow you down. 去敏 — de-sensitise, a de-identified version that removes a named person — is offered as an optional parameter, as if deleting a name were a filter you switch on. It is not. A gossip roundup about an identifiable event, with the name removed, is usually still about that person: the reader supplies the name from context, and everything from the headline to the thumbnail can carry the identifier. Stripping the string changes what the video says it is about without changing who it is about, and if your reason for stripping the name is a legal or reputational one, the string is not the part that was creating the exposure.
Two things follow for anyone shipping this format. First, the sourcing duty does not transfer away from you because the presenter is synthetic: a generated roundup that reads from other people's reporting inherits the obligation to say where the reporting came from, and the brief in the October 2 post does not contain a sourcing slot at all — that is a gap the operator has to fill by hand. Second, the artifact is legitimately synthetic and a listener is not told so unless you tell them. A label is one line of text and it is the difference between a demo and a piece of content that misleads a viewer about what they are watching. If you want the parameters to carry that weight, put the disclosure in the brief and treat it as non-optional, the way the voice's vendor should be named rather than hidden.

Running it on one key, and the one key it does not yet cover
If you are building this pipeline, the integration cost is not the model calls — it is the second credential. Claude Opus 5.5 is routable today as anthropic/claude-opus-5.5 at Anthropic's own list price of $4 and $20 per million tokens, passed through at 0% markup, so a vendor price move reaches your bill the same day rather than at the next contract review. Routing it alongside the rest of your stack through one OpenAI-compatible endpoint is what removes the second SDK, and automatic failover is what lets you try a newer or less proven model on an internal render without putting the production path behind it.
The honest boundary is the voice. Google's Gemini 3.8 Flash TTS and Flash-Lite TTS endpoints are not on our catalogue today — the public model API returns a not-found for google/gemini-3.8-flash-tts — so the workflow in the October 2 post genuinely needs the author's own Google key, and that is a real constraint rather than a temporary one we can wave at. The TTS endpoint we do carry is the older google/gemini-3.1-flash-tts-preview, at $1.00 per million text tokens in and $20.00 per million audio tokens out: usable in the same pipeline shape, priced for the older generation, and not the model this workflow was built on. Pick your path knowingly — one key for the reasoner and a second for the voice, or one key and the older TTS.
What to watch
The thing that would make this format boring, in the good sense, is an audio modality on the producing model. Until Claude Opus 5.5 — or whatever replaces it — can hear the track it commissioned and adjust it, the voice stays a manually reviewed step, and these pipelines stay human-in-the-loop by construction rather than by policy. Watch the skill repositories over the next fortnight rather than the demos: the ones that survive will be the ones that declare their vendors, carry a licence, and let you re-run the same brief and get the same video. The prompt in the October 2 post will look like the easy part, because it was.
For a pipeline that already holds its own material and script, run your own number rather than borrowing one from X: at 25 tokens a second, every minute of finished narration is 1,500 audio tokens and about 1.35 cents at today's Flash TTS rate — a figure you can check against a rate card before you commit a production slot to a synthetic host.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
