
Gemini 3.7 Flash vs Grok 4.6: Google's Price War vs SpaceXAI's Self-Checking Agent
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0259Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0258Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0166Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
Three weeks after Gemini 3.6 Flash largely flopped, Google halved the price and shipped Gemini 3.7 Flash — a move that most coverage read as a direct shot at Grok 4.6, the agent-tuned flagship SpaceXAI had shipped the day before at $2 and $6. The timing is the story: two frontier-adjacent models, one day apart, both aimed at the same buyer — a team that wants agentic coding without paying frontier prices — and priced on opposite philosophies. Google is buying the response with a five-month promotion. SpaceXAI is selling trust per task at a flat list price.
Grok 4.6, released August 12, 2026, is SpaceXAI's refinement of its 1.5-trillion-parameter Grok 4.5 foundation, built for long-running agents that verify their own subtask completion, with a 500K context window, text-and-image input, and a $2/$6 rate that doubles at 200K+ input tokens. Gemini 3.7 Flash, released August 13, 2026, is Google's algorithmic refinement of Gemini 3.6 Flash — 1M context, multimodal input including video and audio, and a promotional $0.75/$3.75 that doubles to $1.50/$7.50 on January 1, 2027. On the independent Artificial Analysis Intelligence Index, Grok 4.6 sits at 61, a clear step above early readings of about 56 for Gemini 3.7 Flash.
The price war, one line per dimension
• Price per 1M input — Gemini 3.7 Flash $0.75 (intro) vs Grok 4.6 $2.00. Google undercuts by more than 60%.
• Price per 1M output — Gemini 3.7 Flash $3.75 (intro) vs Grok 4.6 $6.00.
• After January 1, 2027 — Gemini doubles to $1.50/$7.50 vs Grok unchanged at $2/$6.
• Context window — Gemini 3.7 Flash 1M vs Grok 4.6 500K.
• Long-input billing — Gemini charges flat across its window; Grok doubles to $4/$12 at or above 200K input.
• Thinking tokens — Gemini bills them as output; Grok's cost is captured in an independently measured ~$0.84 per task.
The one-line read: Gemini 3.7 Flash is cheaper on every line of the sticker today, and most of that advantage evaporates on January 1. Grok 4.6 is the same price in August as it will be in March.

Where Gemini 3.7 Flash's money went
Google's own figures report large jumps for Gemini 3.7 Flash over 3.6 Flash: 65.3% on DeepSWE v1.1 (up from 49.0%), 43.6% on FrontierCode 1.1 Main (up from 34.4%), 85.8% on Terminal-bench 2.1 (up from 78.0%), and a WebDev Arena Elo of 1588 (up from 1538). These are vendor-reported, unreproduced numbers — the kind of figure that measures a trajectory, not a cross-vendor verdict. But the trajectory is the point: the fixes landed three weeks after a launch that reviewers had broadly written off, which tells you Google treats the Flash tier as the battleground, not the flagship.
The other thing Google's promotion does is compress the buying window. At $0.75/$3.75 the model is cheap enough to test in volume; at $1.50/$7.50 it is still competitive but no longer a price-war weapon. Any production team adopting it in the intro window should be re-evaluating its economics in December, not assuming the number moves with them.
Where Grok 4.6's money went
Grok 4.6's pitch is different: not cheaper tokens, but fewer wasted ones. It was trained to self-verify — to check that a subtask is actually done before moving on — and independent measurement backs the effect: Artificial Analysis puts its cost at roughly $0.84 per task, with a GDPVal-AA v2 of 1,753 and an Intelligence Index of 61. Those are third-party figures, and they are the strongest number in this matchup, because they capture the thing per-token price misses: a model that takes fewer turns to finish a job is cheaper per finished job even at a higher token price.
The trade-off is where the cheaper model has room. Grok 4.6's 500K context is half of Gemini 3.7 Flash's 1M, and its $2/$6 rate doubles above 200K input — so long-context, high-input workloads are exactly where Grok 4.6's list price stops being competitive with the promotional Gemini rate.
The per-task economics
At intro pricing, Gemini 3.7 Flash's blended rate at an 80/20 input/output split is roughly $1.35 per million tokens; Grok 4.6's sticker at the same split is $2.80 per million before its long-input penalty. On paper the Google model looks like the volume winner by a mile. The complication is that Gemini's thinking tokens bill as output, its benchmark claims are one week old, and its per-task efficiency has not been independently measured yet — while Grok 4.6's has. Cheap tokens plus wasted turns can cost more than expensive tokens that finish, and this is a matchup where one side has data on that question and the other does not.
Betting on the new model without betting the pipeline
Grok 4.6 is on OrcaRouter at SpaceXAI's $2/$6 list price with 0% markup, including the 2x step above 200K, which the model page pulls live from the routing layer. Gemini 3.7 Flash is one day old and on Google's own API; its predecessor Gemini 3.6 Flash is on OrcaRouter at Google's list price.


If the question this comparison is really asking is "should I bet my agent pipeline on the new cheap model," the routing layer is the hedge: a failover rule that runs the agent on Grok 4.6 today and switches to Gemini 3.7 Flash the day it lands, or a fusion panel where both answer and a judge reconciles them — the kind of setup the routing DSL exists for, and one that does not require picking a side before the evidence is in.
The honest read
If you want the cheapest token today and you trust the three-week turnaround, Gemini 3.7 Flash is the aggressive bet — the price war is real, the benchmark gains are reported, and the intro window is a genuine discount. If you want a model whose agent efficiency is independently measured, whose price does not move in January, and whose self-verifying loop is designed for long-running work, Grok 4.6 is the conservative pick, and the AA data gives it an edge on cost per finished task. One side is a sale with a countdown; the other is a price with a track record. Both are defensible — which is what makes this a real comparison rather than a formality.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
