
Gemini Omni 1.1 Flash Goes Production-Ready: 40-Second Scenes, Frame Control, and a 4K Finish
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
Hand a video model the last second of a clip and ask it to continue the shot, and you are gambling on its memory. That is exactly what Gemini Omni 1.1 Flash — the production update Gogle shipped on August 27, 2026 — was built to eliminate: instead of looking back one second of footage when extending a scene, it now reads up to ten. Along the way Gogle pulled the model's most-requested controls out of the experimental column: first and last frame conditioning, a 360p drafting tier that costs about a third as much as the standard 720p output, and final renders up to 4K. The price of the workhorse tier did not move — a 720p clip still costs $0.10 per second — so the story of this release is control and cost, not a sticker shock.
Gemini Omni 1.1 Flash is the successor to the Gemini Omni Flash preview that reached developers on June 30, 2026, and it inherits everything that made the original unusual: multimodal input that accepts text, images, audio, and video in a single call, native synchronized audio on every clip, and conversational multi-turn editing where you refine a scene by talking to the model instead of re-prompting from scratch. What 1.1 adds is the missing tooling — length, framing, resolution, and iteration cost are now inputs a builder can actually set.
The lookback problem, solved
The single most consequential change is context. Earlier versions of the model referenced only the final second of a clip when continuing it, which is why extensions drifted in character, lighting, and camera motion. Gemini Omni 1.1 Flash analyzes up to 10 seconds of prior video, and extensions run in 10-second increments up to a cumulative total of 40 seconds. The practical difference is between "a set of adjacent clips" and "a sequence that reads as one scene." For a 30-second ad or a looping brand asset, that is the difference between assembly and direction.
• Scene extension. Continue an existing clip in 10-second chunks to a 40-second ceiling, with up to 10 seconds of lookback context to keep the continuation consistent.
• First/last frame control. Specify the starting and ending frame of a shot and the model generates continuous video between the two keyframes — the primitive behind camera orbits, zoom transitions, and seamless loops.
• Video references. Up to three seconds of reference footage can be included in the multimodal input to hold character, costume, and motion style across shots.
• Native audio. Generated clips still carry synchronized dialogue, sound effects, and ambient sound, so an output is a finished short rather than a silent asset waiting for a sound pass.
Veo's controls, Omni's pipeline
The announcement leans on the phrase "your favorite creative controls from Veo," and it is an accurate description of where these features came from. Veo 3.1, Gogle's separate one-shot high-fidelity video generator, already had keyframe-style conditioning and upscaling in its toolset; Gemini Omni 1.1 Flash brings that control vocabulary into the multimodal Omni pipeline — where input can be images, audio, and reference video, and where editing is conversational rather than a fresh generation every time. The two models are not rivals inside Gogle's lineup; the Gemini API documentation positions Gemini Omni Flash as the recommended default for video generation, with Veo as the higher-fidelity specialist.
That matters for a practical reason: the controls now sit in the model you are most likely to be calling for everyday video work, not in a separate premium tier you have to learn a second API to reach. Frame control, drafting, and upscaling share one endpoint and one pricing model.
A draft-to-final workflow
The 360p drafting tier is the quietest feature in the announcement and possibly the most valuable. Previews at 360p render up to 60% faster than standard 720p and cost roughly a third as much, which changes the economics of iteration outright. On the per-second rates people are quoting, a 10-second 360p draft runs about $0.30 against $1.00 at 720p — three rounds of exploration for the price of one committed render.
The intended flow is obvious: storyboard and iterate in 360p, then render the one shot that works at 720p, 1080p, or 4K. For a team producing social cuts or ad variants, the draft tier converts a cost-constrained workflow into an exploration-constrained one. Speed and cost figures here are vendor-reported — Gogle's own claims, unreproduced so far — but the arithmetic on top of them is simple.

The pricing picture
The $0.10-per-second figure that keeps getting quoted is Gogle's published 720p standard tier, unchanged from the preview. Input is $1.50 per million tokens regardless of whether those tokens are text, images, audio, or video, and text output is $9.00 per million tokens — the same structure as the original. The 360p tier works out to about $0.03 per second, which is what Gogle's "roughly one-third the cost" framing implies and what reseller listings now show.
What Gogle has not published is a per-second rate for the new 1080p and 4K tiers. Reseller listings quote $0.15 per second for 1080p and $0.30 per second for 4K, but those are marketplace estimates, not Gogle's numbers — treat any high-resolution production budget as provisional until the official rate appears on the Gemini API pricing page. For a sense of scale, a 40-second fully extended scene at 720p is up to $4.00, a 10-second 4K finish is around $3.00 on the reseller-quoted rate, and the same 10 seconds drafted at 360p is about $0.30.
Per-second pricing is the headline, but the number you actually pay is set by the platform you buy through — Gogle sells at list price, and resellers and API gateways add their own margins on top. That is exactly why the 0% markup policy matters on the infrastructure side: OrcaRouter passes provider list price straight through, so when a model's rate changes, the price on our side changes the same day, with no margin layer between you and the vendor's published figure. We do not route Gemini Omni 1.1 Flash yet — Gogle has not opened the Omni video API to third-party routing, and we do not claim to host a model we do not — but the pricing discipline is the point, and when that API tier opens, the model will appear at list price.

Where you can call it
Gemini Omni 1.1 Flash is live now under the model ID gemini-omni-1.1-flash through the Gemini API in Gogle AI Studio, billed as production-ready rather than a preview. It is also available through the Gemini Enterprise Agent Platform for production workloads — the same interface Adobe, Figma, GMI Cloud, and Runway are already integrating with, per the announcement. On the consumer side, Gogle Flow carries the full feature set for all AI Plus, Pro, and Ultra subscribers globally, and scene extension is live in the Gemini app. This is not a paper launch: API, agent platform, and consumer surfaces all moved on the same day.
If you tried the June preview and bounced off its hard 10-second ceiling or its inability to hold a scene together, this release answers those exact complaints — 40 seconds of extension, working video references, and a proper resolution ladder. The API shape is also worth knowing: generation and editing run through the Interactions API, which is what makes the conversational, multi-turn editing model possible, and a response_format setting can pin the output resolution — 360p for drafts, full res for finals.

What is still not proven
For all the new control, the trade-offs that defined the original carry over. Output is still short-form video with synchronized audio, not a Veo 3.1 replacement — the two models target different jobs. Character consistency across scene changes is improved by the 10-second lookback and reference-video input, but it remains a generative weakness worth testing on your own footage rather than assuming. Every generated clip carries a SynthID watermark, which matters for tools that need provenance guarantees.
And the honest caveat: there is no independent benchmark for Gemini Omni 1.1 Flash on any public video leaderboard as of this writing. The 60% speedup for 360p drafts, the quality at 1080p and 4K, the reliability of the first/last frame conditioning — all vendor-reported, none reproduced by a third party. Anyone quoting you a quality score for this model today is citing a number nobody outside Gogle's labs has produced.
Who should move now
Three groups get value from this release immediately. Teams already building on the Omni API should switch model IDs and re-test their extension and frame-control paths — the upgrade is a request-field change, not a re-architecture. Anyone who needs short-form video at volume — ad variants, social cuts, looping brand assets — should price the 360p draft-to-final workflow, because that is where the cost math is genuinely new. And teams that were waiting on the preview should try the model now, because the two blockers that made the preview unusable for production — ten-second clips and no scene continuity — are precisely what this update removes.
The group that should hold off is equally well defined: if you need a third-party benchmark or an official 4K per-second rate before you commit budget, nobody has them yet, and that is worth saying plainly rather than papering over. The release is real, the controls are real, and the pricing gap is the only part still moving. Watch the Gemini API pricing page for the 1080p and 4K rates, and watch for the first independent evaluations — the moment someone else runs this model, the vendor-reported caveat on every quality claim in this article gets a check mark or a red flag.
