
Gemini 3.7 Flash vs Gemini 3.1 Pro: Did Google's Flash Tier Make Its Own Flagship Obsolete?
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
There is an uncomfortable question sitting inside Google's 2026 model lineup, and it is aimed at every team currently paying for Gemini 3.1 Pro: why is the cheap Flash model outscoring the flagship on the independent index? Gemini 3.7 Flash, released August 13, 2026 at $0.75 per million input tokens and $3.75 per million output, scores 56 on the Artificial Analysis Intelligence Index at high reasoning. Gemini 3.1 Pro, the preview Google shipped on February 19, 2026 at $2.00 / $12.00 — and $4.00 / $18.00 above 200K context — measures in the upper-40s on the same index today. The flagship that topped the index at launch has been overtaken by its own workhorse tier, and this article is about what that means for the two models, and for the teams that have to choose between them.
The premise: Google's flagship tier is stuck in preview
The context matters more than the numbers. Gemini 3.1 Pro has been in preview since February 2026 — seven months as of this writing — and its successor, Gemini 3.5 Pro, was announced at Google I/O in May as "coming next month," then missed June, July, and August. As of early September the flagship tier has no released successor, no date, and reports of possible cancellation. Meanwhile the Flash tier shipped three generations in the same window: Gemini 3.6 Flash on July 21, Gemini 3.7 Flash on August 13, and the agentic video understanding capability across both on September 1. The comparison is not really "Pro vs Flash." It is "the tier Google is shipping" against "the tier Google has stopped shipping."
The spec contrast
• Price — Gemini 3.7 Flash $0.75 / $3.75 per 1M tokens (promo through Dec 31, 2026, then $1.50 / $7.50) vs Gemini 3.1 Pro $2.00 / $12.00 under 200K context, rising to $4.00 / $18.00 above 200K. Three to four times cheaper at list, before the reasoning-tier controls.
• Context / max output — 1M / 64K vs 1M / 64K. Identical ceilings.
• Inputs — both text, image, audio, video, PDF. The multimodal surface is the same; the difference is what the model does with video (below).
• Reasoning control — Gemini 3.7 Flash low / medium / high thinking. Gemini 3.1 Pro low / medium / high with dynamic thinking on by default.
• Independent score — Gemini 3.7 Flash 56 on the Artificial Analysis Intelligence Index vs Gemini 3.1 Pro in the upper-40s on the current index version (it scored 57 at launch under the previous version).
• Output speed — Gemini 3.7 Flash roughly 285 tokens/sec at high reasoning vs Gemini 3.1 Pro roughly 112–125 tokens/sec — about two and a half times faster.
• Status — Gemini 3.7 Flash stable, GA. Gemini 3.1 Pro preview, with preview's short deprecation window.
The index, and the benchmark mess underneath it
The headline is the index: Artificial Analysis, the independent cross-vendor evaluator, now ranks Gemini 3.7 Flash at 56 and Gemini 3.1 Pro in the upper-40s — the workhorse ahead of the flagship, a reversal of the positions at 3.1 Pro's launch, when its 57 topped the index. That is the single most important fact in this matchup, and it is independently sourced rather than vendor-reported.
Below the index, the benchmark ledger is a mess of non-comparable harnesses, and reading it naively produces the wrong answer. Gemini 3.1 Pro's published record — 80.6% on SWE-bench Verified, 77.1% on ARC-AGI-2, 94.3% on GPQA Diamond, a 2887 Elo on LiveCodeBench Pro — was reported by Google at its February launch, uses different benchmark versions than the Flash numbers, and was never uniformly reproduced. Gemini 3.7 Flash's 43.6% on FrontierCode 1.1, 65.3% on DeepSWE v1.1, and 1588 Elo on the web-development arena are similarly vendor-reported on the newer harnesses. Where the two overlap on an independent aggregate — the AA Index — the Flash wins. Anyone telling you "Pro beats Flash on SWE-bench" is comparing two versions of two different tests and not saying so.


The video capability gap
The capability that makes the comparison feel lopsided is video. Google's new agentic video understanding mode — announced September 1, 2026 — applies to Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite. Gemini 3.1 Pro is not on the list. The Flash model reads video the way the new mode defines reading video: the model decides what to watch, at what speed, and through frames, audio, or transcript, loading only the relevant segments, with vendor-reported savings of up to 66% on cost and 88% on tokens. Gemini 3.1 Pro technically accepts video input, but it uses the older static sampling path and gets none of the new tooling. For the flagship's users, the newest video capability in the Gemini family is shipping on the cheaper model first, and the Pro tier has no announced date for receiving it.
Speed, and the preview-risk problem
Speed compounds the price gap the same way it does in every Gemini 3.7 Flash matchup: roughly three times the output rate on a model that already costs a third of the price. But the second difference is the one teams forget to price: status. Gemini 3.1 Pro is a preview, and previews come with a short deprecation notice window when the successor lands. A production workload on a preview model is a standing risk that the model ID changes, the pricing changes, or the capability moves with little warning. Gemini 3.7 Flash is GA — stable, documented, and not scheduled to be replaced on a whim, even if Google's near-monthly Flash cadence means a 3.8 Flash could follow quickly. The 3.1 Pro preview's deprecation risk is not hypothetical; it is the current state of a tier whose successor has been delayed for three months.
Where Gemini 3.1 Pro still wins
The honest case for 3.1 Pro has three pillars, and none of them is raw capability. First, the custom-tools endpoint: Gemini 3.1 Pro is served with a dedicated bash/custom-tools API surface that Flash does not match, and teams that have built agent tooling against it have a working system today. Second, tuned workloads: a pipeline already tuned to 3.1 Pro's dynamic-thinking defaults, cache patterns, and eval scores is a moving target, and switching costs real engineering time even when the destination is faster and cheaper. Third, the specific tasks where Pro's longer production record still has no Flash equivalent — the launch-generation benchmarks like GPQA and ARC-AGI-2 that Flash has never been measured on. If your workload is one of those, "Flash outranks Pro on the index" does not settle your question; only a test on your own eval does.
The verdict: what the Flash-tier takeover means
The direction is unambiguous. For a new workload today, Gemini 3.7 Flash is the better default Gemini: cheaper, faster, GA-stable, higher on the independent index, and it gets the new video capability first. Gemini 3.1 Pro retains a real but narrow constituency — custom-tools users, tuned pipelines, and tasks where its older benchmark record still applies. The broader story is that Google's flagship tier is stalled while its workhorse tier is where the company is actually competing — and the teams still paying flagship prices for 3.1 Pro are paying for a status that the company itself has stopped defending.

For teams that want to migrate without a project, the move between Gemini models is a model-name change rather than an integration. Gemini 3.1 Pro is served on OrcaRouter at its $2.00 / $12.00 list price passed through at 0% markup, alongside Gemini 3.6 Flash and the rest of the routable Gemini tier behind one key — so a workload can send a slice of traffic to a Flash model, keep 3.1 Pro as the fallback for the tasks where it still wins, and let automatic failover absorb the risk that either Google model changes. Gemini 3.7 Flash itself is available through Google's own API and several third-party platforms. The routing DSL and model fusion are exactly how you would run a like-for-like comparison between the two Gemini models on your own eval before switching. The choice the lineup is pushing you toward is clear: default to the Flash tier, and keep Pro only for the narrow cases where it demonstrably earns its premium.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
