
Qwen 4 Max vs Gemini 3.1 Pro: the unannounced flagship against the preview that never ended
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Most comparisons pit a shipped product against a shipped product. This one is unusual: Qwen 4 Max was previewed at the Yunqi conference in Hangzhou on September 22, 2026 with no release date, no price and no weights, and Gemini 3.1 Pro has been in preview since February 19, 2026 and is still in preview — seven months later, with no general-availability date announced and no shutdown date on its deprecation table. Neither model in this matchup is a settled product. That is not a reason to skip the comparison. It is the comparison, and it is the one thing about it that is genuinely informative.
Seven months of preview
A preview label is supposed to be a short-lived state. Gemini 3.1 Pro has now held it through three Flash releases from the same lab, including Gemini 3.8 Flash on September 2, 2026, which Google describes as its most intelligent Flash model. On the independent composites that model now scores above 3.1 Pro. Google's Pro line has no newer member: there is no Gemini 3.8 Pro, and the 3.5 Pro generation was delayed and reportedly set aside. The practical effect is that the model carrying the Pro name is no longer the top-scoring model from its own vendor.
Google's own limitation text on the catalogue entry advises planning for preview deprecation, which is the honest way to read the status. A preview with no GA date and no shutdown date is a model you can build on and cannot schedule around.
The Qwen side is a mirror image with the sign flipped. Qwen 4 Max is not in a long preview; it is pre-preview, with nothing published at all. The two failure modes are opposite: Gemini 3.1 Pro is a real endpoint whose long-term existence is unspecified, and Qwen 4 Max is a specified roadmap item with no endpoint.

The comparison, and the long-context figure that decides it
Gemini 3.1 Pro's specification is public and specific. Qwen 4 Max's is empty. What is worth dwelling on is one row of the Gemini column, because it is the sort of number that gets quoted and should not be:
• Status — Gemini 3.1 Pro in preview since February 19, 2026, versus Qwen 4 Max announced September 22, 2026 with no date
• Input price — Gemini 3.1 Pro $2.00 per 1M tokens up to a 200K prompt and $4.00 above it, versus Qwen 4 Max not published
• Output price — Gemini 3.1 Pro $12.00 per 1M up to 200K and $18.00 above, versus Qwen 4 Max not published
• Cached input — Gemini 3.1 Pro $0.20 per 1M on a cache read, versus Qwen 4 Max not published
• Context — Gemini 3.1 Pro 1,048,576 tokens, versus Qwen 4 Max not published
• Output ceiling — Gemini 3.1 Pro 65,536 tokens, versus Qwen 4 Max not published
• Modality — Gemini 3.1 Pro text, image, video, audio and PDF input with text output, versus Qwen 4 Max not published
• Reasoning control — Gemini 3.1 Pro low, medium and high thinking levels, versus Qwen 4 Max not published
• Long-context retrieval — Gemini 3.1 Pro vendor-reported MRCR v2 at 84.9 percent averaged across eight needles at 128K, but 26.3 percent at the one-million-token pointwise setting, versus Qwen 4 Max nothing reported
• Vendor-reported scores — Gemini 3.1 Pro SWE-bench Verified 80.6 percent, GPQA Diamond 94.3 percent, ARC-AGI-2 77.1 percent, Humanity's Last Exam 44.4 percent without tools, versus Qwen 4 Max nothing reported
• Independent score — Gemini 3.1 Pro 30 on Artificial Analysis Intelligence Index v4.3.2, versus Qwen 4 Max no independent evaluation exists
That long-context row is the one to carry away. A 1,048,576-token context window is a marketing number; 26.3 percent retrieval at the one-million-token setting is what the same vendor reports when it measures the far end of that window. Both figures come from Google, which makes the gap a statement about the model rather than about the evaluator. If your workload depends on finding a fact buried near the end of a million-token prompt, the headline context figure tells you nothing useful and this row tells you almost everything.
The index moved, the model did not
Gemini 3.1 Pro launched on February 19, 2026 at 57 on the Artificial Analysis Intelligence Index, first in the field. Under index revision v4.3.2, published September 19, 2026, it reads 30. Nothing about the model changed between those two dates. The ruler was rebuilt, and it was rebuilt four days before Alibaba's Qwen 4 announcement.
That coincidence is worth more than the numbers themselves. Three of the four models in this batch of comparisons — Gemini 3.1 Pro, Claude Opus 5 and the Qwen 3.8 flagship — have independent scores that moved under the same revision, and in each case the movement was large enough to reverse a ranking. Any comparison written this month that quotes a single index value without naming the revision is comparing scores taken with different rulers, and the error will not be visible to a reader.
For Qwen 4 Max the consequence is that the target keeps moving. When Alibaba eventually publishes, the score to beat on this index will be whatever the current revision says on that day, and the revision has changed at least once in the last week.

What the operational numbers say, including the uncomfortable one
Gemini 3.1 Pro is routable on OrcaRouter as google/gemini-3.1-pro-preview, carrying the Flagship and Featured badges with no pre-release banner, at Google's list price of $2.00 and $12.00 per million tokens with 0% markup. The routing itself is honest — the rate on the page is the rate Google publishes. The operational figures over the seven days to September 22 are also worth printing rather than skipping: a p50 time-to-first-token of 10.00 seconds, an output speed around 632 tokens per second, and an error rate of 77.5 percent across the window. That error rate is not a footnote. For a preview-tier model with no GA date, it is the single most decision-relevant number on the page, and a comparison that quotes the price without it is selling rather than informing.
The same catalogue carries close to 200 models on one OpenAI-compatible endpoint with automatic failover and a routing DSL, which is the mechanism that makes an error rate like that survivable rather than fatal: a degraded provider path can be failed over instead of taken down. The Qwen 3.8 family — Qwen3.8-Max, Qwen3.8-Flash and Qwen3.8-27B — sits on the same endpoint as Gemini 3.1 Pro Preview, so the substitution this article is really about, the one you can make today, is a routing-policy change rather than a migration project. Qwen 4 Max is not on the catalogue, and no Qwen 4 tier is; when one ships, that is where it would appear.
The decision, such as it is
Neither model here is a product you can commit to, and the honest advice is to treat both accordingly. Gemini 3.1 Pro is callable today at a published price with a published context window and a public evaluation trail, including the retrieval weakness at the far end of that window; it is also a preview whose vendor tells you to plan for deprecation, running at an error rate that would be unacceptable for a production dependency. Qwen 4 Max has none of those problems because it has none of those properties.
If you need a model now, the choice is not between these two. It is between Gemini 3.1 Pro and the shipping Qwen flagship, and on the evidence available that is a decision about long-context retrieval quality versus a $2.00 input rate with published weights. If you can wait, watch two things rather than one: a Qwen 4 Max model card, and a general-availability date for Gemini 3.1 Pro. Whichever arrives first tells you which of these two roadmaps Alibaba and Google are actually prioritising.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
