
Gemini 3.7 Flash vs DeepSeek V4 Flash: Two "Flash" Models, Opposite Answers to the Same Question
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
The two most confusingly named models of 2026 are both called Flash, and they could not answer the same question more differently. Gemini 3.7 Flash, released August 13, 2026, is Google's closed workhorse: roughly 340 output tokens per second, an Intelligence Index of 56, and a promotional price of $0.75 per million input and $3.75 per million output. DeepSeek V4 Flash is the open-weights half of DeepSeek's V4 family, shipping in April and updated at the end of July: MIT-licensed, downloadable, and priced at $0.14 per million input and $0.28 per million output on DeepSeek's own API — about a fifth of Google's promotional input price and a thirteenth of the output price. Both are "flash" models, built to be fast and cheap for high-volume work. The word is where the similarity ends.
The question they answer differently is the defining one of this generation: do you rent intelligence or own it? DeepSeek V4 Flash is a 284-billion-parameter mixture-of-experts model with roughly 13B active per token, a 1M-token context window and a 384K output ceiling, released under an MIT license — the weights are a ~160GB download, and the July 31 update (V4-Flash-0731) added materially stronger agentic capabilities. Gemini 3.7 Flash is a proprietary model built on Gemini 3.6 Flash with a 1M context, a 64K output cap, multimodal input, and no path to self-hosting. One of these models can be air-gapped, fine-tuned and run on your own hardware. The other runs only where Google runs it.

The two Flashes, side by side
Both models target the same buyer — high-volume production workloads that cannot afford a flagship per token — but they get there through opposite architectures. DeepSeek V4 Flash's hybrid attention (compressed sparse plus heavily compressed attention) is what makes its 1M context economical enough to sell at these prices; it is the model DeepSeek points legacy `deepseek-chat` and `deepseek-reasoner` endpoints at, retired in July. Gemini 3.7 Flash is Google's algorithmic refinement of its own Flash line, and its edge is serving throughput: independent measurement puts it at roughly 340 tokens/s against DeepSeek V4 Flash's roughly 113, and it reaches that while taking audio and video input that DeepSeek V4 Flash does not.
• Intelligence Index — 56 for Gemini 3.7 Flash versus 50 for DeepSeek V4 Flash, per Artificial Analysis.
• Output speed — ~340 tokens/s versus ~113 tokens/s, both per Artificial Analysis.
• Context window — 1M tokens for both; DeepSeek V4 Flash outputs up to 384K against Gemini 3.7 Flash's 64K.
• Input price — $0.75 (promotional, through 2026) versus $0.14 on DeepSeek's own API.
• Output price — $3.75 (promotional) versus $0.28.
• Weights — proprietary versus MIT-licensed and downloadable.

What six index points actually buy
The six-point Index gap is meaningful but not chasmic, and it sits mostly in areas DeepSeek V4 Flash was not optimized for. Gemini 3.7 Flash's reported agentic-coding numbers (65.3 on DeepSWE v1.1, 43.6 on FrontierCode 1.1 Main, vendor-reported) are well ahead, and its web-development Elo (1588, also vendor-reported) reflects real UI-generation improvements. DeepSeek V4 Flash's own reported figures — 79.0 on SWE-bench Verified, 91.6 on LiveCodeBench, a τ²-Bench score of 95.0 that independent trackers corroborate — hold up well on bounded coding and tool use, and its 1M-context retrieval (78.7 on MRCR, 60.5 on CorpusQA) is competitive with anything in its price class. The pattern: Gemini 3.7 Flash leads on agentic depth and web work; DeepSeek V4 Flash trades those for raw token economics and a long-context ceiling that no closed workhorse matches.
Open weights change the procurement, not just the price
The MIT license is the part of this comparison no benchmark table captures. DeepSeek V4 Flash can be downloaded and run on your own hardware, fine-tuned on your own data, and used in environments where API calls are not permitted — and the license has no scale-based restrictions, so nothing in your growth changes the terms. The practical cost is operational: serving a 284B MoE model at the speed a production team wants takes real GPU inventory, and for most teams the honest choice is the hosted API. That is where the price gap does its real work: even the hosted route, at $0.14 / $0.28, is cheap enough to be the default for the long tail of traffic. Gemini 3.7 Flash offers none of the ownership path — what you buy is a reliably managed, fast, multimodal service with a promotional price tag and a January cliff.
The volume bill
Run the same day on both: 20 million input tokens and 4 million output tokens. On DeepSeek V4 Flash's own API that is $2.80 + $1.12, about $3.92. On Gemini 3.7 Flash at the promotional rate that is $15 + $15, about $30 — a 7.6x gap even while Google is running its discount. After January 1, 2027, when Gemini 3.7 Flash reverts to $1.50 / $7.50, the same day bills about $58, fifteen times the DeepSeek figure. The speed advantage partially offsets this on interactive workloads — faster tokens mean fewer concurrent requests — but for batch and background work the open model wins the bill without contest. The two models are not even aiming at the same cost curve, which is the entire point.

Who should build on which
Choose DeepSeek V4 Flash when cost per token is the binding constraint, when you want long outputs up to 384K, when you need the option of self-hosting or air-gapped deployment, or when you are building something whose margins will not survive flagship-priced tokens. Choose Gemini 3.7 Flash when you need speed in an interactive loop, multimodal input, Google's managed reliability, or the stronger agentic-coding results Google reports — and when you are comfortable that the promotional price is temporary. They are not rivals in the way two flagships are rivals; they are two different procurement strategies wearing the same name. The routing layer is where the strategies meet: DeepSeek V4 Flash is live on OrcaRouter at DeepSeek's list price with 0% markup, and the failover pattern fits this pairing naturally — a rule that runs the bulk of traffic on V4 Flash and escalates the hard subset to a stronger closed model, with Gemini 3.7 Flash reachable through Google's own Gemini API on the other leg of the same router.
The honest read
If you can self-host or you are purely cost-driven, DeepSeek V4 Flash is the better model to build on, and the license makes it a decision you never have to revisit. If you need Google's serving speed, multimodal input, or the reported agentic-coding edge, Gemini 3.7 Flash is the better service while the promotion lasts — with the January price doubling already on the calendar. Two models, one name, and the market gets to see which philosophy the next year of traffic votes for.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
