
Intern-S2 vs Gemini 3.1 Pro: The Multimodal Matchup InternLM Actually Measured
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Most of the matchups in this series require stitching together two vendors' claims. This one does not. Intern-S2, the 397-billion-parameter Apache-2.0 multimodal model InternLM released today, and Gemini 3.1 Pro, Google's flagship reasoning model, both appear as measured columns in the same benchmark table — which makes this the fairest head-to-head in the set and, not coincidentally, the one where the open model has the most to answer for. Gemini 3.1 Pro reads text, images, audio and video, holds a 1M-token context, and has been torn apart by independent labs and production traffic since it entered preview on February 19, 2026. Intern-S2 reads text, images and time series, was uploaded to Hugging Face this morning, and has no third-party score of any kind. On the rows where both were measured, Google's model wins more of them than not.
Both read images. That is where the similarity stops.
The word "multimodal" makes these models sound like peers. They are not peers, and the difference is in what each one was built to look at.
Intern-S2's vision path was trained to consume raw pages of scientific literature — figures, equations, structural diagrams, layout — as a single shared representation rather than parsing text out first. Its multimodal benchmarks are scientific: biological microscopy VQA, remote sensing, scientific chart question answering. It also accepts time-series data and can produce forecasts, which is an input type almost no general flagship handles at all.
Gemini 3.1 Pro's multimodality is general-purpose and broader in kind: text, image, audio, and video, in one model, aimed at the entire range of production tasks where you point a model at a file and ask a question. A meeting recording, a UI screenshot, a chart in a PDF, a video clip — all fair game, all with six months of public evaluation behind them.
If your input is a scientific figure, Intern-S2 was purpose-built for it and has the vendor numbers to argue its case. If your input is anything else, Gemini 3.1 Pro takes audio and video and Intern-S2 does not take either. That resolves a large share of "which should I use" questions before price or benchmarks enter the picture, and anyone who tells you the two are interchangeable has not read the input list.
What the shared table shows, read honestly
Because InternLM benchmarked both, this is one of the few places where a direct comparison is legitimate rather than assembled. The caveat is unchanged and worth stating once: every Intern-S2 figure is vendor-reported, run by InternLM on its own harness through OpenCompass, VLMEvalKit and AgentCompass, with no independent reproduction. Gemini's figures in that table are also from InternLM's run of the comparison, though Gemini has an extensive independent record outside it.
• Multimodal knowledge — Gemini 3.1 Pro 83.99 on MMMU Pro vs Intern-S2 81.68.
• Scientific chart QA — Gemini 3.1 Pro 71.18 on ChartQAPro vs Intern-S2 71.12. A dead heat to two decimals.
• Remote sensing — Gemini 3.1 Pro 54.27 on XLRS-Bench vs Intern-S2 51.38.
• Microscopy VQA — Gemini 3.1 Pro 71.02 on MicroVQA vs Intern-S2 69.67.
• Scientific multimodal tasks — Intern-S2 60.74 on SFE vs Gemini 3.1 Pro 59.57. Intern-S2's one multimodal win.
• General knowledge — Gemini 3.1 Pro 91.00 on MMLU Pro vs Intern-S2 89.77.
• High-school maths — Gemini 3.1 Pro 94.70 on HMMT-2026 vs Intern-S2 93.56.
• Scientific literature reasoning — Intern-S2 55.71 on Biology-Instructions vs Gemini 3.1 Pro 13.87, and 53.95 vs 38.84 on Mol-Instructions.
• Science olympiad problems — Intern-S2 87.00 on FrontierScience-Olympiad vs Gemini 3.1 Pro 79.00.
• Terminal mastery — Gemini 3.1 Pro 73.80 on TerminalBench 2.1 vs Intern-S2 64.04.
The honest reading: Gemini 3.1 Pro wins the broad multimodal rows by narrow margins, Intern-S2 wins the deep scientific-literature rows by enormous ones, and neither result is surprising given what each was trained on. The narrowness of the multimodal losses is arguably the more interesting finding — a same-day open checkpoint landing within two points of a six-month-old closed flagship on MMMU Pro and ChartQAPro is a stronger result for InternLM than any of its blowout science numbers.
Context, output and the preview question
The long-input story splits cleanly. Gemini 3.1 Pro holds a 1M-token context window with a 64K output ceiling — enough for a long document set and a substantial answer. Intern-S2 is evaluated at up to 256K tokens for text reasoning and 64K for multimodal, and does not publish an output ceiling; treat its usable window as smaller than Gemini's until someone measures it independently.
There is one operational caveat on the Google side that belongs in any honest comparison. Gemini 3.1 Pro is still a preview model, not a GA release. Preview models carry lower reliability expectations, and a 7-day error-rate panel on a model page is a standing reminder to route around instability rather than build a critical path on top of it. Intern-S2 is a full Apache-2.0 release with weights you can pin — which, counterintuitively, makes the brand-new open model the more operationally stable of the two in the sense that matters most: nobody can change it under you.
• Context — Intern-S2 256K text / 64K multimodal evaluated vs Gemini 3.1 Pro 1M tokens.
• Inputs — Intern-S2 text, image, time series vs Gemini 3.1 Pro text, image, audio, video.
• Weights — Intern-S2 Apache-2.0, downloadable vs Gemini 3.1 Pro closed.
• Maturity — Intern-S2 hours old, no independent score vs Gemini 3.1 Pro preview since February 2026, heavily and independently tested.
• Price — Intern-S2 no public rate card; self-host or official Intern API vs Gemini 3.1 Pro $2.00 input / $12.00 output per million, rising to $4.00/$18.00 above 200K input tokens.

Price and where each one lives
Gemini 3.1 Pro bills $2.00 per million input tokens and $12.00 per million output tokens on the standard tier, stepping to $4.00/$18.00 once a request crosses 200K input tokens — a long-context premium that bites exactly when you use the 1M window that justifies choosing it. It is routable on OrcaRouter at that same list price, passed through with 0% markup, on a key that also carries 200-plus other models; the pass-through matters here because a Google price move lands on our side the same day rather than after a repricing cycle. The routing DSL is the useful part for this particular matchup: you can send image-and-video work to Gemini 3.1 Pro and scientific-document batches to a self-hosted Intern-S2, and keep both behind one endpoint without a second contract or a code change.
Intern-S2 has no published API price. Self-hosting the Apache-2.0 weights is the route most teams will actually take, and at 403B parameters it is a real hardware commitment. The official Intern API exists for teams that would rather not run GPUs, but without a public rate card the cost comparison to Gemini's $2/$12 is not one you can make on paper — you have to ask for quota and do the arithmetic yourself.

The call
Choose Gemini 3.1 Pro if you need a general multimodal model that takes audio and video, holds a million tokens, has months of independent validation, and can be rented by the token today. It is the better-evidenced, more broadly capable model, and the preview badge is a reliability caveat rather than a capability one.
Choose Intern-S2 if your multimodal input is scientific and your data cannot leave your infrastructure. It is the only one of the two that reads raw literature pages natively, forecasts from time-series signals, and can be fine-tuned on proprietary experimental data under a permissive licence. Accept that you are the first person to run it, and that the benchmark table arguing for it was written by the people who made it.

Neither model wins this outright, and the table proves it better than a verdict could. One is a generalist that happens to be good at science; the other is a scientific instrument that happens to be competitive on general multimodal work — and on a same-day open release, "competitive" is the more surprising result.
To hold both halves of this comparison without a second contract: one API key covering 200-plus models puts the closed multimodal flagship and every alternative behind one endpoint, so you can route image-and-video work one way and document batches the other.
