
What Is Gemini 3.1 Pro? Google's Pro Tier Has Not Moved in Seven Months
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Gemini 3.1 Pro is the only live Pro-tier model in its own family. It was announced on 19 February 2026, and seven months later the vendor's own model documentation still files it under “Preview” at the endpoint gemini-3.1-pro-preview — no general availability, no shutdown date, no successor shipped. In the same seven months the vendor pushed five Flash releases out of the door, and the newest of them, Gemini 3.8 Flash, now scores eleven points higher on the independent index most people quote. That gap is the reason a page explaining what Gemini 3.1 Pro is has to open with what it is not.
If you just want the two-sentence version: Gemini 3.1 Pro is a closed-weights, natively multimodal reasoning model made by Google DeepMind, which reads text, images, video, audio and PDFs and writes text. You call it through Google's own APIs or through a router that carries it, it takes up to 1,048,576 tokens of input, and it costs $2.00 per million input tokens and $12.00 per million output tokens below a 200,000-token prompt. It holds up as a long-context multimodal workhorse. It is no longer the model you pick because it is the smartest one available.
Where it sits in its own family
This is the part that most write-ups skip, and it is the part that matters most. The Gemini 3 family is not one ladder with a Pro at the top. It is two ladders that stopped being comparable.
The Flash ladder is the one that moved. Google's model documentation, read on 23 September 2026, lists Gemini 3.8 Flash, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as new stable releases, with Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash and Gemini 3.5 Flash-Lite still stable behind them. Gemini 3.8 Flash landed on 2 September 2026 — the third Flash release in six weeks.
The Pro ladder is the one that did not. There is no Gemini 3.5 Pro, 3.6 Pro, 3.7 Pro or 3.8 Pro in Google's documentation. Gemini 3 Pro, the model Gemini 3.1 Pro was built to replace, is listed as shut down. Gemini 2.5 Pro survives, but access is restricted to prior users. That leaves Gemini 3.1 Pro as the only Pro-tier Gemini 3 endpoint you can call today — not because it won a race, but because nobody else entered it.
Google's stated plan in February was that the preview would last long enough to “validate these updates” and advance areas such as ambitious agentic workflows before the model went “generally available soon.” Those are Google's words in the February announcement. Seven months on, that has not happened, and the reporting around the delay is worth stating plainly as reporting rather than fact: SemiAnalysis concluded the Gemini 3.5 Pro effort had been quietly cancelled, and the Wall Street Journal, as cited by secondary outlets, reported that the internal 3.5 Pro candidate was dropped because it could not surpass the Flash series. Google's own line is that 3.5 Pro remains in restricted partner testing with no confirmed month. We could not find a Google statement confirming or denying the cancellation, and we are not going to pick a side for them.
The practical consequence is a naming trap. If you go looking for “the newest Gemini Pro,” the freshest Pro-tier model you will find is from February, while the version numbers on the Flash side have climbed to 3.8. Version numbers here do not rank capability across tiers; they are two separate release trains.
The specs that matter
Every figure below is read from Google's own Gemini API documentation on 23 September 2026, unless it is labelled otherwise.
• Input context — 1,048,576 tokens, the same one-million-token window Google has shipped since the 2.5 generation. Output is capped separately at 65,536 tokens, which is a ceiling on the answer, not on the conversation.
• Input modalities — text, image, video, audio and PDF. Output modalities — text only. No audio generation, no image generation, no Live API, so this is not the model for a realtime voice agent or an image pipeline. Google ships separate Gemini 3.x Flash variants for those jobs.
• Reasoning — supported, with thinking-level control. Thinking tokens are billed at the output rate, which is why the output price is the number that dominates your bill.
• Tooling — caching, code execution, function calling, Search grounding, Maps grounding, structured outputs, URL context, the Batch API, flex inference and priority inference are all listed as supported. File search is listed as supported in AI Studio only.
• Licence and weights — proprietary. There are no open weights, no downloadable checkpoint, and no self-hosting path. Every route to this model is a hosted API.
• Endpoints — gemini-3.1-pro-preview is the model code. Google also lists a gemini-3.1-pro-preview-customtools variant. A preview endpoint can be revised under you without a version bump, which is the standard reason to pin exact model IDs in production rather than aliases.
• Documented update date — the model page carries “Last update: February 2026.” We could not find a knowledge cutoff date stated on Google's own model page; third-party listings say January 2025, and we are flagging that as unconfirmed rather than repeating it as fact.
What it costs
Gemini 3.1 Pro has a price cliff, and it is the kind of detail that turns a cheap-looking estimate into a doubled invoice. Google's published developer pricing splits on prompt size rather than on model version.
• Standard tier, prompts up to 200,000 tokens — $2.00 per million input tokens, $12.00 per million output tokens, with thinking tokens counted as output.
• Standard tier, prompts over 200,000 tokens — $4.00 input and $18.00 output. Both rates double at the same threshold. This is the trap: a long-context workload pays twice the headline rate precisely because it is using the feature it bought the model for.
• Context caching — $0.20 or $0.40 per million cached tokens depending on which side of the 200,000 threshold you are on, plus cache storage billed at $4.50 per million tokens per hour.
• Batch and flex tiers — $1.00 input and $6.00 output under 200,000 tokens, $2.00 and $9.00 above it. Roughly half price for work that can wait.
• Priority tier — $3.60 input and $21.60 output under the threshold, $7.20 and $32.40 above it.
• Free tier — not available for this model. There is no free rate limit to prototype against.
Two things are worth knowing about how reliable those numbers are. Google's own pricing page is the authority here and we read it directly, so the figures above are current as of 23 September 2026. But third-party listings of this model are genuinely inconsistent — one tracker lists a $2.50 / $15.00 rate, another a $4.50 blended rate — and the tiering is the reason. If a page quotes you a single number for Gemini 3.1 Pro without telling you which prompt-size band it applies to, that number is not usable for a real estimate.
What it is measurably good at, and measurably not
Start with what Google claims, then keep it separated from what independent evaluators measured, because the two have diverged sharply since February.
Google's own announcement carries exactly one benchmark figure: a “verified score of 77.1%” on ARC-AGI-2, which Google describes as more than double the reasoning performance of Gemini 3 Pro. That is the vendor's number, read from Google's own post, and it is the only one we could confirm there. Launch coverage circulated a much larger set — GPQA Diamond, SWE-bench Verified, Terminal-Bench 2.0, LiveCodeBench, Humanity’s Last Exam — but those figures do not appear in Google’s own announcement, and the secondary reports disagree with each other on the ones we checked (GPQA Diamond appears as both 94.1 and 94.3 depending on who you read). We are not printing them as established results.
The independent picture is where the story turns. On the Artificial Analysis Intelligence Index version live on 23 September 2026, Gemini 3.1 Pro Preview scores 30 and ranks 81st of 212 models in that class. The median model in the class scores 25, so it sits above average — and the leader on the same chart, Claude Opus 5 at its maximum reasoning setting, scores 58. The index itself is a versioned ruler rather than a fixed one: Artificial Analysis revised it after the February launch, and its own revision notes record new evaluations joining the composite. Gemini 3.1 Pro was reported as the outright leader on the index as it stood at its February launch, at 57 points. That is not the same ruler as the 30, and we are not going to put the two numbers side by side as if they traced a decline.
What the current mapping does show is a genuine profile, and it is a lopsided one:
• Intelligence — 30 on the index, rank 81 of 212. Above median, well off the frontier.
• Speed — 116.5 output tokens per second, rank 38 of 212. Notably fast for a reasoning model.
• Cost — $0.67 to run one Intelligence Index task, rank 31 of 212. That is cheap per unit of work, because the model is economical with thinking tokens.
• Verbosity — 67M output tokens to complete the index, against a class median of 88M. Rank 39 of 212; Artificial Analysis calls it “fairly concise.” For comparison, Gemini 3.8 Flash at high reasoning burns 170M tokens on the same suite and costs $1.24 per task — more than Gemini 3.1 Pro, despite cheaper per-token pricing.
So the honest read is not “the old model is beaten everywhere.” Gemini 3.1 Pro is fast and token-efficient, and on a cost-per-task basis it beats its own newer sibling. What it loses is the top end of capability, and on any workload where quality is the binding constraint that is the only axis that counts.

The September availability question
There is one live development around this model, and it is a question rather than a release. Since 18 September 2026, users on Google's own AI Studio developer forum have been reporting that Gemini 3.1 Pro disappeared from the AI Studio model selector — the thread title dates the change to 2026-09-18 — with paid users saying they were moved onto Gemini 3.8 Flash and, from there, onto 3.7 and 3.6 Flash. No Google employee had replied in the thread we read, there is no deprecation notice, and no cause has been stated. The model was still listed and still callable through the Gemini API and through our own catalogue when we checked on 23 September 2026.
We are flagging this as unresolved user reports on Google's own forum, not as a retirement and not as an outage. But if your production path depends on this model specifically, it is a reason to have a fallback configured rather than to assume a preview endpoint is a fixture.
Who should pick it, and who should not
Pick Gemini 3.1 Pro if your problem is long and multimodal. The one-million-token window with native video, audio and PDF input is still the reason to be here, and nothing about the model’s age changes that. Concretely: whole-repository analysis, auditing hours of recorded meetings or footage, extracting structured data out of mixed-media document sets, or any pipeline where a single prompt legitimately runs past 200,000 tokens and the alternative is a retrieval layer you would have to build and evaluate. Its token efficiency is a real second argument — it finishes a task suite in 67M output tokens where Gemini 3.8 Flash spends 170M, so a per-task bill can come out lower even though the per-token rate looks higher.
Pick something else in three cases. If you want Google’s current best answer to a hard reasoning or long-horizon coding question, Gemini 3.8 Flash outscores Gemini 3.1 Pro on the current independent index, is generally available rather than preview, costs $0.75 per million input and $3.75 per million output through 31 December 2026, and has a free tier that Gemini 3.1 Pro does not. It burns more thinking tokens to get there — that is the trade, and it is a real one. If you need the outright frontier, that is not a Google model right now: on the same index chart, Claude Opus 5 at maximum reasoning leads at 58 against Gemini 3.1 Pro’s 30. And if your workload is latency- or cost-sensitive and shallow — classification, extraction, routing, short summarisation — a Flash-Lite tier or a small open-weights model will do the job for a fraction of $2.00 per million input tokens, and paying Pro rates for it is just waste.
What we would not do is pick Gemini 3.1 Pro because it is the newest Pro from Google. It is not new, and as of this month Google has no shipped successor to it.
Calling it without betting your stack on a preview endpoint
The awkward part of a seven-month preview is that it is good enough to build on and not stable enough to assume. Two mitigations are worth naming.
Pin the exact model code rather than an alias. A preview endpoint can be revised in place, so gemini-3.1-pro-preview in a config file is safer than a floating “latest Pro” reference, and it makes a silent change visible in a diff.
Keep the routing above the model. On OrcaRouter, Gemini 3.1 Pro is google/gemini-3.1-pro-preview, served from the same key and the same endpoint as every other model we carry, including a native /v1beta/models/{model}:generateContent shape for code written against the Gemini SDK and an OpenAI-compatible /v1/chat/completions for everything else. Because pricing is pass-through with 0% markup, the $2.00 / $12.00 on Google's rate card is the number you pay here, and a vendor price change is live the same day rather than at the next billing cycle. The reason that matters for a preview model specifically is failover: a configured fallback chain means the day Gemini 3.1 Pro is unavailable, your retry lands on the next model in the chain before the response starts, rather than after your users notice. That is the difference between a preview you can ship on and a preview you can only demo.
Our own catalogue entry for the model carries the specs as we serve them — one-million-token context, 65,536-token maximum output, audio/file/image/text/video input, text output, and a median time-to-first-token of 10.00 seconds over the trailing seven days, which is worth knowing before you assume this is a chat-latency model. It is not.

What is still open
Three things we could not close, and would rather state than paper over. Google has not said whether Gemini 3.1 Pro will ever reach general availability, and there is no published shutdown date either way. The AI Studio reports are unexplained five days on, with no vendor response we could find. And the successor question is answerable only in the negative: Google’s Pro page has advertised an upcoming Pro release for months, the reporting says the internal candidate was dropped for failing to beat the Flash line, and as of 23 September 2026 nothing has shipped. If a Gemini 3.5 Pro or a Gemini 4 Pro arrives, the position of Gemini 3.1 Pro in this family changes for the first time since February — and that, not its original launch, is the event worth watching.

Gemini 3.1 Pro is on OrcaRouter at Google's list price, behind the same key as everything else we carry. See the Gemini 3.1 Pro model page for live pricing, context limits and a fallback chain you can configure before the next preview revision lands.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
