A title card for the Gemini 3.7 Flash versus Grok 4.6 comparison, showing a 50%-off price tag with a lightning bolt on the Gemini 3.7 Flash side facing a shield with a checkmark on the Grok 4.6 side.
Guides & Insights

Gemini 3.7 Flash vs Grok 4.6: Google's Price War vs SpaceXAI's Self-Checking Agent

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Three weeks after Gemini 3.6 Flash largely flopped, Go​ogle halved the price and shipped Gemini 3.7 Flash — a move that most coverage read as a direct shot at Grok 4.6, the agent-tuned flagship SpaceXAI had shipped the day before at $2 and $6. The timing is the story: two frontier-adjacent models, one day apart, both aimed at the same buyer — a team that wants agentic coding without paying frontier prices — and priced on opposite philosophies. Go​ogle is buying the response with a five-month promotion. SpaceXAI is selling trust per task at a flat list price.

Grok 4.6, released August 12, 2026, is SpaceXAI's refinement of its 1.5-trillion-parameter Grok 4.5 foundation, built for long-running agents that verify their own subtask completion, with a 500K context window, text-and-image input, and a $2/$6 rate that doubles at 200K+ input tokens. Gemini 3.7 Flash, released August 13, 2026, is Go​ogle's algorithmic refinement of Gemini 3.6 Flash — 1M context, multimodal input including video and audio, and a promotional $0.75/$3.75 that doubles to $1.50/$7.50 on January 1, 2027. On the independent Artificial Analysis Intelligence Index, Grok 4.6 sits at 61, a clear step above early readings of about 56 for Gemini 3.7 Flash.

The price war, one line per dimension

Price per 1M input — Gemini 3.7 Flash $0.75 (intro) vs Grok 4.6 $2.00. Go​ogle undercuts by more than 60%.

Price per 1M output — Gemini 3.7 Flash $3.75 (intro) vs Grok 4.6 $6.00.

After January 1, 2027Gemini doubles to $1.50/$7.50 vs Grok unchanged at $2/$6.

Context window — Gemini 3.7 Flash 1M vs Grok 4.6 500K.

Long-input billing — Gemini charges flat across its window; Grok doubles to $4/$12 at or above 200K input.

Thinking tokens — Gemini bills them as output; Grok's cost is captured in an independently measured ~$0.84 per task.

The one-line read: Gemini 3.7 Flash is cheaper on every line of the sticker today, and most of that advantage evaporates on January 1. Grok 4.6 is the same price in August as it will be in March.

A comparison scoreboard for Gemini 3.7 Flash and Grok 4.6: Gemini at AA Index ~56, $0.75/$3.75 intro pricing, 1M context, Google-reported GDPVal-AA v2 of 1,525 and per-task cost not yet measured, versus Grok 4.6 at AA Index 61, $2/$6, 500K context, GDPVal-AA v2 of 1,753 and about $0.84 per task per Artificial Analysis, both proprietary.

Where Gemini 3.7 Flash's money went

Go​ogle's own figures report large jumps for Gemini 3.7 Flash over 3.6 Flash: 65.3% on DeepSWE v1.1 (up from 49.0%), 43.6% on FrontierCode 1.1 Main (up from 34.4%), 85.8% on Terminal-bench 2.1 (up from 78.0%), and a WebDev Arena Elo of 1588 (up from 1538). These are vendor-reported, unreproduced numbers — the kind of figure that measures a trajectory, not a cross-vendor verdict. But the trajectory is the point: the fixes landed three weeks after a launch that reviewers had broadly written off, which tells you Go​ogle treats the Flash tier as the battleground, not the flagship.

The other thing Go​ogle's promotion does is compress the buying window. At $0.75/$3.75 the model is cheap enough to test in volume; at $1.50/$7.50 it is still competitive but no longer a price-war weapon. Any production team adopting it in the intro window should be re-evaluating its economics in December, not assuming the number moves with them.

Where Grok 4.6's money went

Grok 4.6's pitch is different: not cheaper tokens, but fewer wasted ones. It was trained to self-verify — to check that a subtask is actually done before moving on — and independent measurement backs the effect: Artificial Analysis puts its cost at roughly $0.84 per task, with a GDPVal-AA v2 of 1,753 and an Intelligence Index of 61. Those are third-party figures, and they are the strongest number in this matchup, because they capture the thing per-token price misses: a model that takes fewer turns to finish a job is cheaper per finished job even at a higher token price.

The trade-off is where the cheaper model has room. Grok 4.6's 500K context is half of Gemini 3.7 Flash's 1M, and its $2/$6 rate doubles above 200K input — so long-context, high-input workloads are exactly where Grok 4.6's list price stops being competitive with the promotional Gemini rate.

The per-task economics

At intro pricing, Gemini 3.7 Flash's blended rate at an 80/20 input/output split is roughly $1.35 per million tokens; Grok 4.6's sticker at the same split is $2.80 per million before its long-input penalty. On paper the Go​ogle model looks like the volume winner by a mile. The complication is that Gemini's thinking tokens bill as output, its benchmark claims are one week old, and its per-task efficiency has not been independently measured yet — while Grok 4.6's has. Cheap tokens plus wasted turns can cost more than expensive tokens that finish, and this is a matchup where one side has data on that question and the other does not.

Betting on the new model without betting the pipeline

Grok 4.6 is on OrcaRouter at SpaceXAI's $2/$6 list price with 0% markup, including the 2x step above 200K, which the model page pulls live from the routing layer. Gemini 3.7 Flash is one day old and on Go​ogle's own API; its predecessor Gemini 3.6 Flash is on OrcaRouter at Go​ogle's list price.

The OrcaRouter model page for grok/grok-4.6, showing the New and Featured badges, the $2.00 per million input and $6.00 per million output pricing, a 500K-token context window, and the released August 12, 2026 label.The OrcaRouter model page for google/gemini-3.6-flash, showing the $1.50 per million input and $7.50 per million output pricing and the 1M-token context window.

If the question this comparison is really asking is "should I bet my agent pipeline on the new cheap model," the routing layer is the hedge: a failover rule that runs the agent on Grok 4.6 today and switches to Gemini 3.7 Flash the day it lands, or a fusion panel where both answer and a judge reconciles them — the kind of setup the routing DSL exists for, and one that does not require picking a side before the evidence is in.

The honest read

If you want the cheapest token today and you trust the three-week turnaround, Gemini 3.7 Flash is the aggressive bet — the price war is real, the benchmark gains are reported, and the intro window is a genuine discount. If you want a model whose agent efficiency is independently measured, whose price does not move in January, and whose self-verifying loop is designed for long-running work, Grok 4.6 is the conservative pick, and the AA data gives it an edge on cost per finished task. One side is a sale with a countdown; the other is a price with a track record. Both are defensible — which is what makes this a real comparison rather than a formality.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube