
Gemini 3.7 Flash Is Live: Half the Price, 1M Context, and What's Actually Confirmed
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxNEWMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2186 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
Gemini 3.7 Flash is live. Google shipped the next model in its Flash tier on August 13, 2026 — just three weeks after Gemini 3.6 Flash — and announced it at $0.75 per million input tokens and $3.75 per million output, exactly half of Gemini 3.6 Flash's $1.50/$7.50. The leak signals this blog tracked a week ago turned out to be accurate on every point that mattered: the model was real, the "3.7" name survived to the official model ID, and the rumored half-price was the announced price.
Gemini 3.7 Flash is a multimodal model Google's model card describes as its latest and most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution. Google's launch framing leads with gains in coding, knowledge work, and web development — vendor claims, examined below — while the specs are already on the record: the same 1 million-token context window the Flash tier has standardized on, a 64K-token output, and availability through the Gemini API (including Google AI Studio), Google Antigravity, Gemini Enterprise, and Gemini Spark for AI Pro and Ultra subscribers. What the model is, what it costs, and where it runs are now a matter of record; the open question is what it actually delivers.
One piece of discipline from the leak coverage carries straight into a launch piece: Google's performance claims are vendor-reported until an independent measurement confirms them. The first independent number arrived within a day — Artificial Analysis scores Gemini 3.7 Flash at 56 on its Intelligence Index at high reasoning, four points above the 52 Gemini 3.6 Flash scored at launch and, per AA, on the Pareto frontier of intelligence versus speed. Google's own benchmark figures for coding, knowledge work, and web development remain the vendor's word. Treat the price, the context window, and the availability as facts. Treat the performance deltas as promises under test.
From leak to launch: the signals that checked out
Because this model was first tracked here as an unverified leak, it is worth recording how the leak aged — it is the difference between a rumor and a pattern the Flash tier has now demonstrated repeatedly.
• The August 7 X post from @chetaslua said groundwork for Gemini 3.7 Flash was "in place on their side," with no date. Six days later the model shipped. The "sooner rather than later" read was right.
• The August 7-8 references to gemini-3.7-flash inside Google's own Python GenAI SDK on GitHub pointed at a real model ID — the same ID now present in the Gemini API.
• The August 13 report that Gemini 3.7 Flash was staged on Gemini Enterprise in a hidden state with a possible same-day release was accurate about the date, and the rumored $0.75/$3.75 price was accurate about the price.
None of that makes leak coverage redundant — these same signals could have resolved to "delayed" or "renamed." But a tier that ships this fast, this quietly, and this cheaply is exactly why Google keeps doing its most interesting work in Flash.
What's confirmed versus what's still a claim
Now that the model is live, the ledger from the leak post flips: most of what was "reported" has moved into the confirmed column, and the remaining uncertainty is about performance, not existence.
• Confirmed — Google has published a Gemini 3.7 Flash model card; the model is available through the Gemini API, Google AI Studio, Google Antigravity, Gemini Enterprise, and Gemini Spark; the announced price is $0.75/$3.75 per million tokens through December 31, 2026; the context window is 1M tokens with a 64K-token output.
• Independently measured — Artificial Analysis put Gemini 3.7 Flash on its Intelligence Index within a day of launch: 56 at high reasoning (53 at medium, 51 at low), four points above the 52 Gemini 3.6 Flash scored at launch, with the gains AA attributes largely to agentic evaluations.
• Vendor-reported — Google's own eval figures for the headline use cases: 43.6% on FrontierCode 1.1 (up from 34.4%), 65.3% on DeepSWE v1.1, an Elo of 1588 on the web-development arena where Google says it ranks #1, 34.0% on GDP.pdf, and 30.4% on AutomationBench. These are Google's numbers until an outside party reproduces them.
• Not yet known — whether Google extends the introductory rate past December 31, 2026 (it doubles to $1.50/$7.50 on January 1, 2027), whether the gains change the calculus that has kept Gemini 3.5 Pro unreleased, and how the model holds up on workloads the vendor did not advertise.
For context on those "gains over Gemini 3.6 Flash" claims, this is the independent baseline Google is being measured against:

Gemini 3.6 Flash's independently measured Artificial Analysis Intelligence Index of 52 (ranked #26 of 185 models at launch) and its $1.50/1M input price are the numbers Gemini 3.7 Flash has to beat — and the 3.7 price point already undercuts the 3.6 input price by half before any performance comparison begins.
Why the price is the headline
The half-price is the most consequential part of this launch, and the arithmetic is simple. On a 1M-token input / 1M-token output workload, Gemini 3.6 Flash costs $1.50 plus $7.50, or $9.00. Gemini 3.7 Flash at the announced rate costs $0.75 plus $3.75, or $4.50 — a 50% cut on the same volume, and if Google's claimed efficiency improvements hold, the effective cost per completed task drops further still.
Two caveats worth writing down. First, the $0.75/$3.75 rate is introductory and time-boxed: Google has confirmed it runs through December 31, 2026, and doubles to $1.50/$7.50 on January 1, 2027. Second, some platforms are currently running promotional pricing below the announced rate, so check the live price at the point of call for what you will actually pay — the announced rate is the reliable floor, but it is a floor with a date on it.
The Gemini 3.5 Pro backdrop still matters
Google's flagship Gemini 3.5 Pro still has not shipped as of August 14, 2026 — months after I/O, with Bloomberg reporting in mid-July that it was behind schedule over coding capability and later reports describing internal deadline slippage. The Flash tier keeps delivering while the Pro tier stalls, which is why a half-price, faster Flash is a bigger event than a typical point release: it is the proof that Google's workhorse line is where the company is actually competing.
Don't confuse it with Qwen 3.7 Flash
Both "3.7 Flash" models now exist as shipped models, and they could not be more different.
• Gemini 3.7 Flash — Google's new Flash-tier model at $0.75/$3.75 per million tokens, 1M context, multimodal, announced August 13, 2026.
• Qwen 3.7 Flash — Alibaba's vision-language model that shipped quietly in late July 2026 at roughly $0.03/$0.13 per million tokens with a 1M context window, aimed at multimodal agents and computer-use tasks.
The price gap alone — roughly 25x on input tokens — separates them, and the vendor does too. When someone says "I'm running 3.7 Flash," the first question is still which one.
Should you move from Gemini 3.6 Flash?
The honest answer splits in two. For cost, the announced half-price is a real lever: if you pay meaningful volume on Gemini 3.6 Flash, the same workload on Gemini 3.7 Flash costs half as much on tokens, before any claimed efficiency gains. That is a switch worth modeling on your own call patterns. For trust, the picture improved within a day of launch — the independent Artificial Analysis Index is already in at 56 versus 3.6 Flash's 52 — but the specific coding, knowledge-work, and web-development deltas Google is selling are still vendor-reported, and Gemini 3.6 Flash remains the stable, documented default. The standard play is to keep a proven model as the default and route a slice of traffic to the new one, comparing quality on your own tasks.
Neither move requires a project if you are on a routing layer. Gemini 3.6 Flash and Gemini 3.5 Flash are both available today through a single OpenAI-compatible endpoint on OrcaRouter at provider list price with no markup — so a Google price cut anywhere in the tier shows up on our side the same day, and a later switch from one Flash model to another is a model-name change rather than an integration. Automatic failover is the pattern for evaluating a just-shipped model however it is served: point a slice of real traffic at it, let the router fall back to a proven model on errors, rate limits, or obvious quality regressions, and keep the old model as the safety net until the new one earns the traffic. The routing DSL and model fusion let you compose the new model with others in a single call, so testing it against your own workload is a configuration change.
What to watch next
• The next independent check — Artificial Analysis has already published its Intelligence Index (56 at high reasoning), so the open question has moved from "is the gain real" to whether Google's coding and web-development benchmark claims survive external reproduction and hold up on your own workloads.
• The price — the $0.75/$3.75 rate is now a dated fact: it holds through December 31, 2026, then doubles. Watch whether Google extends it, and whether third-party promotional pricing normalizes to it.
• Gemini 3.5 Pro — if Google's flagship finally ships, the "Flash is where Google delivers" framing of the past month may shift.
FAQ
Is Gemini 3.7 Flash released?
Yes — on August 13, 2026. Google has published a Gemini 3.7 Flash model card describing it as its latest and most capable Flash model, the gemini-3.7-flash ID is served through the Gemini API, and Google lists it across Google AI Studio, Google Antigravity, Gemini Enterprise, and Gemini Spark. This is the model whose preparation was reported here as a leak on August 7.
How much does Gemini 3.7 Flash cost?
$0.75 per million input tokens and $3.75 per million output at the announced introductory price — half of Gemini 3.6 Flash's $1.50/$7.50, per Google's announcement. The rate is confirmed through December 31, 2026, and doubles to $1.50/$7.50 on January 1, 2027. Some platforms currently run promotional pricing below the announced rate, so the live rate at the point of call is what you will pay.
Should I upgrade from Gemini 3.6 Flash to Gemini 3.7 Flash?
The price is half and the first independent score is in — Artificial Analysis Index 56 at high reasoning, up from 3.6 Flash's 52 — but Google's headline coding, knowledge-work, and web-development deltas are still vendor-reported. Standard practice: keep Gemini 3.6 Flash as the default, evaluate Gemini 3.7 Flash on your own workload with failover, and switch when it earns the traffic.
Gemini 3.7 Flash is the Flash tier doing exactly what the Flash tier does — shipping fast, shipping cheaper, and this time arriving with an independent score that beats its predecessor. The leak was right about the name, the date, and the price; the performance claims are still under test. If cost-per-task is your lever, the half-price makes Gemini 3.7 Flash worth benchmarking today. If stability is your priority, Gemini 3.6 Flash remains the proven choice — and moving later is a model-name change, not a project.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
