
Gemini 3.1 Pro vs Qwen3.8-27B: A Seven-Month Preview vs a Checkpoint You Can Download
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
On September 18, 2026, users on the vendor's own AI developer forum began reporting that Gemini 3.1 Pro had vanished from the AI Studio model selector — paid accounts included, cache clears and incognito windows notwithstanding. The thread asking the vendor to say whether it was "a bug or intentional" was still live on September 21. There is no deprecation notice, no migration path, and no staff reply. Meanwhile Qwen3.8-27B, which the lab open-sourced on August 14 under Apache 2.0, cannot be taken away from anyone who downloaded it. That asymmetry — not the benchmark gap, which is real but modest — is what this matchup actually turns on, and it is why the two models are not really competing for the same slot even though the comparison tables put them side by side.
What follows is the head-to-head on the numbers that exist, with each one labelled: Artificial Analysis figures from captures taken on September 23, 2026, Alibaba's own model card, Google's launch post, and our own routing telemetry, kept separate throughout. Where nothing has been measured, this page says so rather than guessing.
What happened this week, and what it does and does not mean
The reports are specific and they come from Google's own forum rather than a competitor's blog. A thread titled "Gemini 3.1 Pro missing from AI Studio since 2026-09-18 — bug or intentional?" describes paying users losing the model with no announcement, Google One support offering unrelated troubleshooting steps, and the Google status page showing nothing wrong. A second thread, "Gemini 3.1 Pro Preview model not available in AI Studio to build App," collects the same complaint, including from a user whose coding assistant had been built on the preview for months and who says Flash "makes slop" on their codebase.
Three honest caveats. It is a developer-console disappearance, not a confirmed API outage: Gemini 3.1 Pro is still listed and routable on our own catalogue, where it moved 94.7M tokens in the last seven days. It is user-reported — Google has said nothing, so nobody outside Google can separate "quiet deprecation" from "capacity problem" from "console bug." And the timing is not flattering given that this model has been in preview since February 19, 2026 — seven months, with no GA date published, while Google shipped a ladder of Flash models underneath it. Separately, GitHub Copilot retired Gemini 3.1 Pro on September 1, 2026 with 30 days' notice, which is a different and unrelated event, but it points the same way: a preview endpoint with no GA commitment is a dependency you are allowed to be nervous about.
One number from our side is worth putting next to that context, because we cannot explain it and would rather say so. Our seven-day telemetry on the Gemini 3.1 Pro catalogue entry shows an 80.5% error rate; the neighbouring models we checked the same morning — Qwen3.8-Max, Claude Fable 5.1 — sit at 5.9% and 6.9%. We do not know whether that is the same problem the forum is reporting, a bad week for one upstream provider, or an instrumentation artifact. Treat it as a reason to run a canary before you commit, not as a verdict on the model.
Same ruler, two different scores
This is the part most comparison pages get wrong, so it is worth doing carefully. Artificial Analysis scores both of these models on Intelligence Index v4.3.2. We captured both pages on September 23, 2026, and both carry the same revision — so unlike most cross-model index claims, this one is same-snapshot and the gap is real.
• Qwen3.8-27B (xhigh) — Intelligence Index 33.7
• Gemini 3.1 Pro Preview — Intelligence Index 29.7
Qwen leads by four points on the current ruler. The harder question is what that means, because the ruler itself has moved at least three times this year. In February, Artificial Analysis published an article calling Gemini 3.1 Pro Preview "the new leader in AI," four points clear of Claude Opus 4.6, and press coverage of the time reported a score of 57 on what was then Index v4.0. On v4.3.2 it is 29.7. Artificial Analysis announced v4.3 on September 7, 2026, four days after v4.2, and the revision swapped Terminal-Bench 2.1 for the much harder 66-task Terminal-Bench 4.0 and replaced τ³-Banking with AutomationBench-AA. Gemini 3.1 Pro's February strengths sat in evaluations the new index retired or reweighted. So: the model did not get worse, the exam did. But if you are choosing a model today, today's number is the one you can actually act on, and today's number favours Qwen.
Where the two genuinely diverge
Aggregate scores hide the shape of the difference, and the shape is the useful part. Every figure below is from the same September 23 capture of Artificial Analysis's v4.3.2 pages for both models, on the same revision.
• Reasoning and knowledge — Gemini leads clearly. GPQA Diamond 94.1 vs 90.5; Humanity's Last Exam 47.0 vs 33.9; SciCode 58.7 vs 46.6; CritPt 17.7 vs 5.4.
• Real-world occupational tasks — Qwen leads, and by a lot. GDPval normalized 45.4% vs 13.8%.
• Agentic banking and transactions — Qwen leads. τ-Banking 48.0 vs Gemini's 21.4.
• Agentic tool use on a general benchmark — Gemini leads where Qwen was not measured. τ²-Bench 95.6 for Gemini; no Qwen3.8-27B figure exists.
• Terminal work — close, and both bad on the new benchmark. Terminal-Bench 2.1: Qwen 79.8, Gemini 73.8. Terminal-Bench 4.0: Qwen 5.6%, Gemini 4.0% — both single digits.
• Calibration — the sharpest divergence on either page. On AA-Omniscience, Gemini scores 31.9 with 54.9% accuracy and a 50.9% hallucination rate; Qwen scores −10.0 with 15.6% accuracy and a 30.3% hallucination rate. Gemini knows more and makes it up more than half the time; Qwen frequently declines to answer and hallucinates on roughly a third of what it does answer.
• Long context — dead even on long-context recall, 82.0 each. On multimodal long-context recall, Qwen 21.7% vs Gemini 15.6%.
Read the "not measured" entries as what they are. Gemini has no GDPval-AA normalized figure published against Qwen's 45.4%, and Qwen has no τ²-Bench, IFBench, Terminal-Bench Hard or ITBench-SRE figure against Gemini's 77.1, 53.8 and 30.3. Every one of those blanks is a place a summary table would have quietly planted a winner.

The divergence that decides budgets: same cost to run the index
Qwen3.8-27B is dramatically cheaper per token. On the market figures Artificial Analysis publishes, it is $0.50 per million input and $3.00 per million output with an 80% cache discount and a blended 7:2:1 rate of $0.47. Gemini 3.1 Pro Preview is $2.00 and $12.00 — four times the input price and four times the output price — with a 90% cache discount and a blended rate of $1.74.
Then the number nobody prints: Artificial Analysis's own accounting puts the cost of running the entire Intelligence Index at $1,310.21 for Gemini 3.1 Pro Preview and $1,335.92 for Qwen3.8-27B. Qwen is not cheaper to run. It is very slightly more expensive, because it spends longer getting there. Of Qwen's output bill, $512.36 was reasoning tokens against $82.51 of actual answer; of Gemini's, $691.82 was reasoning against $113.08 of answer. Per-token price is a rate, not a bill.
Two more pricing facts that comparison tables drop. First, Gemini 3.1 Pro's price is tiered by the size of each request's input: $2.00/$12.00 up to 200K input tokens, then $4.00/$18.00 above it, with cache reads doubling from $0.20 to $0.40. The million-token context window is real, but the last 800,000 tokens of it cost double. Second, our own catalogue entry for Qwen3.8-27B — served block-FP8, vision tower at full precision — lists $0.40 input and $4.21 output, which is cheaper on input and dearer on output than the market figure above, and publishes no cached-token rate at all. Different providers, different split. Check the rate you will actually be billed at.
Both models are reachable through one OrcaRouter key at provider list price with 0% markup, so a vendor price change lands on our side the same day rather than after a repricing cycle — which matters more than usual when one of the two has a tier boundary sitting at 200K tokens.
Latency: the fastest first token belongs to the slower model
This is the most misleading pair of numbers in the whole comparison, and it is the one a reader is most likely to act on wrongly.
• Time to first chunk — Qwen 4.04s vs Gemini 28.37s. Qwen looks seven times faster.
• Time to the actual answer — Gemini 28.37s vs Qwen 52.52s. Gemini wins by roughly 1.8x, because Qwen spends 48.48 seconds thinking between the first chunk and the first answer token.
• Output speed — Gemini 116.5 tokens per second vs Qwen 41.3. Gemini is 2.8x faster once it starts writing, against a comparator median Artificial Analysis puts at 95.4 tokens per second. Qwen's own page calls it "notably slow."
If your product streams a first token to a user, Qwen's headline latency is a lie you will tell yourself once. If your product waits for a complete answer, Gemini is the faster model. Both figures are Artificial Analysis measurements; both were taken at Qwen's xhigh effort setting, which is the model card's default.
That default is a live complaint among practitioners, and the model card half-admits it. Qwen3.8-27B exposes a reasoning_effort parameter with three levels — xhigh (the default, "for complex tasks demanding thorough analysis"), medium, and low. Community write-ups through September report xhigh producing very long thinking phases and recommend dropping to medium or low and capping the reasoning budget. Alibaba's own card cautions that in multi-turn agentic work, lowering effort "does not always reduce overall task completion time" — which is a candid thing for a model card to say and a warning that the obvious fix is not always the fix.
Context: 1M vs 262K, and what actually fits
Gemini 3.1 Pro carries a 1,000,000-token context window with a 65,536-token maximum output. Qwen3.8-27B carries 262,144 tokens natively, extensible to 1,000,000 with YaRN scaling, with separate budgets recommended inside it: reasoning content up to 262,144 and a final response up to 131,072.
Two honest footnotes. Artificial Analysis's specification table lists Qwen's window as 256,000 while its own FAQ text says 260K and the model card says 262,144 — those are rounding differences on the same number, but quote 262,144 because that is what the card says. And the extension to 1M is a scaling technique, not the same thing as a native million-token window: long-context recall at 1M on a YaRN-extended model is not what the 82.0 AA-LCR figure measures. Nobody has published a 1M-context recall score for Qwen3.8-27B. Treat Gemini's window as the one with the measured behaviour at full length and Qwen's as the one you should test yourself before designing around it — and remember that past 200K input tokens, Gemini's price doubles.
Agentic work and tool calling
This is where the matchup is genuinely unsettled, and where a manufactured winner would be easy and wrong. Qwen3.8-27B has measured agentic scores Gemini lacks, and vice versa.
Qwen's case rests on a strong independent number and a strong vendor one. The independent number is τ-Banking at 48.0 against Gemini's 21.4 — more than double, on the same revision, on a benchmark about multi-step tool use inside a banking environment. The vendor number is Alibaba's own card, which lists OSWorld-Verified 84.3, WebArena-Verified 64.8, AndroidWorld 81.9, CoWorkBench 70.7, JobBench 33.4 and Agents' Last Exam 20.4 Pass@1. Those are Alibaba's evaluations, unrerun by anyone else, and they sit in exactly the categories where the independently measured GDPval result (45.4% normalized vs 13.8%) points the same direction.
Gemini's case is that it has the one high absolute agentic number in the comparison — τ²-Bench 95.6 — and Qwen has no τ²-Bench score at all. It also has IFBench 77.1 for instruction following, another blank on Qwen's side. And it has a seven-month production record that a five-week-old open-weights release does not.
The tiebreaker is not a score, it is your topology. Qwen's agentic advantage is measured on benchmarks where the model operates inside a large, pre-existing software system it did not design. Gemini's is measured on benchmarks where the model has to follow instructions precisely for a long time. Those are different skills and both are real.

Coding
Coding splits the same way, which is why "which is better at code" has no clean answer here. Gemini leads the scientific-computing end — SciCode 58.7 vs 46.6 — and its AA Coding Index of 68.8 on our catalogue page is within a rounding error of Qwen's 68.1, well inside any confidence interval worth quoting. On agentic terminal work the two are close on the older benchmark, Qwen 79.8 vs Gemini 73.8 on Terminal-Bench 2.1, and indistinguishable on the newer one, where both collapse to single digits.
Alibaba's card claims much more — SWE-bench Pro 61.7, LiveCodeBench v6 90.3, QwenSWEBench 79.0 — but those are vendor-reported, and the card footnotes that SWE-bench Pro and QwenSWEBench were run in the Claude Code harness, which is not the harness a competitor's number would have used. The safe statement is that on the one coding benchmark both labs' models were measured on by a third party, the two are separated by four points on Terminal-Bench 2.1 and by 1.6 points on Terminal-Bench 4.0, and both are bad at the harder one.
Open weights versus a preview endpoint
This is the axis on which the two are not comparable at all, and it is the one with the largest practical consequence.
Qwen3.8-27B is Apache 2.0, 27B dense parameters, 64 layers, with a hybrid of Gated DeltaNet linear attention and full attention at roughly a 3:1 ratio — the design that makes a 262K window affordable to serve. It runs. Quantised to 4-bit it fits in roughly 17GB, which is one consumer GPU; a whole compression ecosystem grew up around it within five weeks, from Unsloth's Dynamic V3 quants, including a 1-bit build that Unsloth says runs in 8GB at about 77% of original accuracy, to PrismML's Ternary Bonsai 2 27B in mid-September. You can fine-tune it, you can audit it, and you can run it on a laptop with the network unplugged.
Gemini 3.1 Pro is a closed preview endpoint, seven months old, with no GA date, a model card that explicitly warns previews may lack the stability, availability and support of a stable release, and — as of this week — a developer console it appears to have quietly left. You cannot download it, you cannot pin a version, and you cannot fail over to yourself.
If you want the honest routing answer rather than a verdict: the existence of the open checkpoint is what makes the Gemini dependency survivable. Put Qwen3.8-27B behind a local or self-hosted path as the fallback, keep Gemini 3.1 Pro on the primary path for the reasoning-heavy work it is measurably better at, and put both behind one endpoint with automatic failover so that a silent disappearance from someone else's console is an incident your routing layer absorbs rather than an outage your users see. That is not a pitch for a product — it is the only configuration in which this week's news does not cost you a weekend.
Pick Gemini 3.1 Pro if
• Your hardest work is knowledge-intensive reasoning rather than long-horizon tool use — it leads GPQA Diamond 94.1 to 90.5, Humanity's Last Exam 47.0 to 33.9, SciCode 58.7 to 46.6 and CritPt 17.7 to 5.4.
• You need genuinely long inputs. A measured million-token window, with recall behaviour tested at length, beats a 262K window extended by technique — provided you can live with the price doubling past 200K input tokens.
• Time-to-complete-answer matters more than time-to-first-token, and you want 116.5 tokens per second rather than 41.3.
• You need broad multimodality today: audio, file, image, text and video input against Qwen's text, image and video.
• You are willing to accept preview risk, and you have a fallback you control.
Pick Qwen3.8-27B if
• Your agents operate inside large existing systems — τ-Banking 48.0 to 21.4 and GDPval 45.4% to 13.8% normalized are the two biggest single-model advantages anywhere in this comparison.
• You need an answer to be honest rather than confident. A 30.3% hallucination rate against Gemini's 50.9% is a different kind of product.
• Licensing or data residency rules out a closed API, or you need to fine-tune on your own data.
• You want a per-token bill a quarter of Gemini's and accept that you will spend the difference on thinking tokens — the index costs the same to run on both.
• You are prepared to tune reasoning_effort down from its xhigh default rather than shipping it as-is.
What we could not verify
For the record, and because a comparison page is only as good as its blanks: no independent evaluation has been published for Qwen3.8-27B at reduced reasoning effort, so its speed-versus-quality trade at medium and low is unmeasured. No long-context recall score exists for the YaRN-extended 1M configuration. Gemini 3.1 Pro has no published GDPval-AA normalized score, and Qwen3.8-27B has none for τ²-Bench, IFBench or ITBench-SRE. And nobody — including Google — has said why Gemini 3.1 Pro left the AI Studio selector on September 18. We will update this page when any of those change.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
