
Grok 4.7 vs Gemini 3.1 Pro: The Leaderboard Moved, the Model Didn't
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
On February 19, 2026, Artificial Analysis published an article headlined "Gemini 3.1 Pro Preview: The new leader in AI." It led six of ten evaluations and sat four points clear of Claude Opus 4.6. Read that same model's page today and the number is 30, ranked sixty-ninth of 200. The vendor changed nothing. What changed is the ruler: Artificial Analysis revised its Intelligence Index to v4.3.2 on September 19, 2026, re-scoring every model in the catalogue against a different evaluation set, and a February leader became a September mid-fielder while the model itself sat untouched in preview. That is the backdrop for a comparison against Grok 4.7 — released by the vendor on September 21, 2026 — and it is the reason this matchup is less about which model is smarter than about which of these two you can still rely on being the same thing next month.
What each of these actually is
Gemini 3.1 Pro is Google's newest Pro-tier model and it has never reached general availability. It shipped as gemini-3.1-pro-preview on February 19, 2026, and Google's deprecation table still lists it with "no shutdown date announced" — a preview that has now been in preview for seven months while Google shipped a Flash ladder underneath it: 3.5 Flash in May, 3.6 in July, 3.7 in August, 3.8 in September. There is no GA Pro model newer than Gemini 3 Pro. It carries a 1,048,576-token input window and a 65,536-token output ceiling, takes text, images, video, audio and PDF in with text out, and exposes a thinking-level control at low, medium and high with no budget parameter and no minimal setting. Weights are closed.
Grok 4.7 is a day old and closed-weight as well. It holds a 500k-token context — half of Gemini's — and takes text and images in, so it cannot ingest video or audio at all, which is the single sharpest capability difference in this comparison. It lists at $2.00 per million input tokens and $6.00 per million output tokens below 200k prompt tokens, stepping to $4.00 and $12.00 above that line, with cached reads at $0.50 and $1.00; xAI also lists a faster variant at roughly double the output speed and double the price. No maximum output figure is published for it in the documentation checked here.
The ruler moved. That is the story.
Both models are currently scored on Artificial Analysis Intelligence Index v4.3.2, which replaced Terminal-Bench 2.1 with Terminal-Bench 4.0 and tau3-Banking with AutomationBench-AA, and re-anchored the GDPval-AA Elo scale. Every model's number moved when it landed. Gemini 3.1 Pro's moved further than most, because its February strengths were concentrated in evaluations that the new index either retired or reweighted. Reading the two columns side by side today:
• Composite — Grok 4.7 (xhigh): 46 on Artificial Analysis Intelligence Index v4.3.2. Gemini 3.1 Pro Preview: 30 on the same revision, ranked #69 of 200.
• Long context — Grok 4.7: 500k window, AA-LCR v1.1 77%. Gemini 3.1 Pro: 1M window — twice as large — and a 65K output ceiling that is roughly half of what a long-context agentic session usually needs.
• Input modality — Grok 4.7: text and image. Gemini 3.1 Pro: text, image, video, audio and PDF. Not close, and not a benchmark gap — a capability gap.
• Speed — Grok 4.7: no published measurement yet. Gemini 3.1 Pro: 124.4 output tokens per second, ranked #39 of 200, with a 63.62-second time to first token through Google AI Studio.
• Price, standard tier — Grok 4.7: $2.00 in / $6.00 out per 1M. Gemini 3.1 Pro: $2.00 in / $12.00 out. Identical input, twice the output.
• Price, long-context tier — Grok 4.7: doubles to $4.00 / $12.00 at 200k prompt tokens. Gemini 3.1 Pro: steps to $4.00 / $18.00 at the same 200k boundary. Same input, 50% more output.
• Maturity — Grok 4.7: one day old, vendor benchmarks largely unreproduced. Gemini 3.1 Pro: seven months in market, but still a preview build with no GA commitment and no shutdown date.

The preview problem is a real one now
A seven-month preview would be an oddity rather than a risk if nothing had gone wrong with it. Something has. From September 18, 2026, users began reporting that Gemini 3.1 Pro had vanished from the AI Studio model selector, and as of this writing there has been no Google staff response, no deprecation notice and no stated cause — an unresolved availability incident rather than a confirmed retirement. Set that beside the announcement history: Google said at I/O in May 2026 that Gemini 3.5 Pro was "rolling out next month," and press reporting in late August said that model had been scrapped because internal candidates "weren't sufficiently better than the Flash series." Google's own Pro page still advertises "3.5 Pro coming soon," which is the contradiction in one screen. Whatever the resolution, the practical reading is that Gemini 3.1 Pro is a preview endpoint with no GA date, no successor named, and a current availability question mark.
Grok 4.7's risk is the mirror image and smaller in kind. It is one day old, so its numbers are thin — Artificial Analysis has published no speed or latency measurement, and xAI's own benchmark suite for it has not been independently reproduced. What is not in question is that the model is generally available, priced, documented, and stable in identity. A model whose identity is stable and whose numbers are unproven is a much easier thing to build on than a model whose numbers are published and whose availability is not.
What the split costs, in practice
Assume you are choosing for a document pipeline. On 100k input tokens and 8k output, Grok 4.7 costs $0.20 in and $0.048 out — about a quarter. Gemini 3.1 Pro costs $0.20 in and $0.096 out — about thirty cents. The two are within twenty per cent of each other, and if your documents are scans, PDFs, audio or video, Gemini is the only one of the two that can read them at all. That is not a benchmark argument; it ends the comparison for that workload.
Now assume you are choosing for a long-context agent. At 300k input and 8k output, Grok 4.7 costs $1.20 in and $0.096 out, about $1.30. Gemini 3.1 Pro costs $1.20 in and $0.144 out, about $1.35. Effectively identical — and here Gemini's 1M window and 65K output ceiling become the constraint that decides it, because a 65K cap is where long agentic runs fall over. Neither model is the cheap option in this matchup, and neither is expensive by frontier standards.

Routing around a moving target
OrcaRouter carries Gemini 3.1 Pro, listed on its own model page at the provider's rates — $2.00 in and $12.00 out, with the 1M window and a 65k maximum output — because we pass provider list pricing through at 0% markup. That is not a marketing line in a case like this one: Google has an unpriced-duration preview build whose availability wobbled in September, and the useful thing a router does is keep the integration stable while the thing behind it changes. Automatic failover means a preview endpoint that stops answering becomes a routing event rather than an outage, and the routing DSL lets you compose a multimodal model with a text-only one in a single call — which is exactly the shape this pair suggests, Gemini reading the video and Grok doing the reasoning over its transcript. Grok 4.7 is not in the OrcaRouter catalog as of this writing; calling it today means going through xAI's own API and the third-party platforms that carry it.
The pattern worth building: keep Gemini 3.1 Pro for anything that has to ingest video, audio or PDF, and treat its preview status as an operational risk to route around rather than a reason to avoid it. Put Grok 4.7 on the text-and-image work where its lower output price and higher composite score apply, and where the model you integrate today is the model you will still be calling in six months. The one thing not to do is read the sixty-ninth-place ranking as a verdict on the model — it is a verdict on a benchmark revision, and the same revision that demoted Gemini 3.1 Pro also demoted dozens of models Google never touched.
The call
On the current index, Grok 4.7 leads Gemini 3.1 Pro by sixteen points, and on output price it is half the cost in the standard tier and a third cheaper in the long-context tier. If your workload is text and images, that is enough to decide it. If your workload is video, audio or scanned documents, Gemini 3.1 Pro is the only candidate of the two and the comparison ends there — you are buying a capability, not a score. The caveat that should follow both of them is about provenance rather than performance: one is a seven-month-old preview with no GA date and a September availability incident behind it, the other is a day old with a headline number and no independent reproduction of anything else. Neither is a settled choice. Both are worth routing rather than committing to.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
