A generated hero card comparing Intern-S2 and Gemini 3.1 Pro, showing a scientific figure page and a chart feeding a multimodal model on the left and four input icons for image, audio, video and text around a closed API card on the right, labelled with the contrast between Apache-2.0 scientific multimodality and a closed 1M-context general multimodal flagship, with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Intern-S2 vs Gemini 3.1 Pro: The Multimodal Matchup InternLM Actually Measured

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Most of the matchups in this series require stitching together two vendors' claims. This one does not. Intern-S2, the 397-billion-parameter Apache-2.0 multimodal model InternLM released today, and Gemini 3.1 Pro, Goog​le's flagship reasoning model, both appear as measured columns in the same benchmark table — which makes this the fairest head-to-head in the set and, not coincidentally, the one where the open model has the most to answer for. Gemini 3.1 Pro reads text, images, audio and video, holds a 1M-token context, and has been torn apart by independent labs and production traffic since it entered preview on February 19, 2026. Intern-S2 reads text, images and time series, was uploaded to Hugging Face this morning, and has no third-party score of any kind. On the rows where both were measured, Goog​le's model wins more of them than not.

Both read images. That is where the similarity stops.

The word "multimodal" makes these models sound like peers. They are not peers, and the difference is in what each one was built to look at.

Intern-S2's vision path was trained to consume raw pages of scientific literature — figures, equations, structural diagrams, layout — as a single shared representation rather than parsing text out first. Its multimodal benchmarks are scientific: biological microscopy VQA, remote sensing, scientific chart question answering. It also accepts time-series data and can produce forecasts, which is an input type almost no general flagship handles at all.

Gemini 3.1 Pro's multimodality is general-purpose and broader in kind: text, image, audio, and video, in one model, aimed at the entire range of production tasks where you point a model at a file and ask a question. A meeting recording, a UI screenshot, a chart in a PDF, a video clip — all fair game, all with six months of public evaluation behind them.

If your input is a scientific figure, Intern-S2 was purpose-built for it and has the vendor numbers to argue its case. If your input is anything else, Gemini 3.1 Pro takes audio and video and Intern-S2 does not take either. That resolves a large share of "which should I use" questions before price or benchmarks enter the picture, and anyone who tells you the two are interchangeable has not read the input list.

What the shared table shows, read honestly

Because InternLM benchmarked both, this is one of the few places where a direct comparison is legitimate rather than assembled. The caveat is unchanged and worth stating once: every Intern-S2 figure is vendor-reported, run by InternLM on its own harness through OpenCompass, VLMEvalKit and AgentCompass, with no independent reproduction. Gemin​i's figures in that table are also from InternLM's run of the comparison, though Gemin​i has an extensive independent record outside it.

• Multimodal knowledge — Gemini 3.1 Pro 83.99 on MMMU Pro vs Intern-S2 81.68.

• Scientific chart QA — Gemini 3.1 Pro 71.18 on ChartQAPro vs Intern-S2 71.12. A dead heat to two decimals.

• Remote sensing — Gemini 3.1 Pro 54.27 on XLRS-Bench vs Intern-S2 51.38.

• Microscopy VQA — Gemini 3.1 Pro 71.02 on MicroVQA vs Intern-S2 69.67.

• Scientific multimodal tasks — Intern-S2 60.74 on SFE vs Gemini 3.1 Pro 59.57. Intern-S2's one multimodal win.

• General knowledge — Gemini 3.1 Pro 91.00 on MMLU Pro vs Intern-S2 89.77.

• High-school maths — Gemini 3.1 Pro 94.70 on HMMT-2026 vs Intern-S2 93.56.

• Scientific literature reasoning — Intern-S2 55.71 on Biology-Instructions vs Gemini 3.1 Pro 13.87, and 53.95 vs 38.84 on Mol-Instructions.

• Science olympiad problems — Intern-S2 87.00 on FrontierScience-Olympiad vs Gemini 3.1 Pro 79.00.

• Terminal mastery — Gemini 3.1 Pro 73.80 on TerminalBench 2.1 vs Intern-S2 64.04.

The honest reading: Gemini 3.1 Pro wins the broad multimodal rows by narrow margins, Intern-S2 wins the deep scientific-literature rows by enormous ones, and neither result is surprising given what each was trained on. The narrowness of the multimodal losses is arguably the more interesting finding — a same-day open checkpoint landing within two points of a six-month-old closed flagship on MMMU Pro and ChartQAPro is a stronger result for InternLM than any of its blowout science numbers.

Context, output and the preview question

The long-input story splits cleanly. Gemini 3.1 Pro holds a 1M-token context window with a 64K output ceiling — enough for a long document set and a substantial answer. Intern-S2 is evaluated at up to 256K tokens for text reasoning and 64K for multimodal, and does not publish an output ceiling; treat its usable window as smaller than Gemin​i's until someone measures it independently.

There is one operational caveat on the Goog​le side that belongs in any honest comparison. Gemini 3.1 Pro is still a preview model, not a GA release. Preview models carry lower reliability expectations, and a 7-day error-rate panel on a model page is a standing reminder to route around instability rather than build a critical path on top of it. Intern-S2 is a full Apache-2.0 release with weights you can pin — which, counterintuitively, makes the brand-new open model the more operationally stable of the two in the sense that matters most: nobody can change it under you.

• Context — Intern-S2 256K text / 64K multimodal evaluated vs Gemini 3.1 Pro 1M tokens.

• Inputs — Intern-S2 text, image, time series vs Gemini 3.1 Pro text, image, audio, video.

• Weights — Intern-S2 Apache-2.0, downloadable vs Gemini 3.1 Pro closed.

• Maturity — Intern-S2 hours old, no independent score vs Gemini 3.1 Pro preview since February 2026, heavily and independently tested.

• Price — Intern-S2 no public rate card; self-host or official Intern API vs Gemini 3.1 Pro $2.00 input / $12.00 output per million, rising to $4.00/$18.00 above 200K input tokens.

A generated two-column scoreboard titled 'Intern-S2 vs Gemini 3.1 Pro — the scoreboard.' Left column Intern-S2-397B lists Weights Apache-2.0, Context 256K text / 64K multimodal, Inputs text image time series, Independent score none yet, Price self-host or Intern API, MMMU Pro 81.68. Right column Gemini 3.1 Pro lists Weights closed, Context 1M in / 64K out, Inputs text image audio video, Independent score 57 at launch on the then-current scale, Price $2.00 / $12.00 per 1M, MMMU Pro 83.99. Footer reads 'Intern-S2 figures are InternLM's own, unreproduced; both MMMU Pro rows as measured in InternLM's comparison table. Gemini 3.1 Pro figures per Google and independent trackers.'

Price and where each one lives

Gemini 3.1 Pro bills $2.00 per million input tokens and $12.00 per million output tokens on the standard tier, stepping to $4.00/$18.00 once a request crosses 200K input tokens — a long-context premium that bites exactly when you use the 1M window that justifies choosing it. It is routable on OrcaRouter at that same list price, passed through with 0% markup, on a key that also carries 200-plus other models; the pass-through matters here because a Goog​le price move lands on our side the same day rather than after a repricing cycle. The routing DSL is the useful part for this particular matchup: you can send image-and-video work to Gemini 3.1 Pro and scientific-document batches to a self-hosted Intern-S2, and keep both behind one endpoint without a second contract or a code change.

Intern-S2 has no published API price. Self-hosting the Apache-2.0 weights is the route most teams will actually take, and at 403B parameters it is a real hardware commitment. The official Intern API exists for teams that would rather not run GPUs, but without a public rate card the cost comparison to Gemin​i's $2/$12 is not one you can make on paper — you have to ask for quota and do the arithmetic yourself.

A screenshot of the Hugging Face model card for internlm/Intern-S2, captured September 13, 2026, showing the Apache-2.0 licence badge, the tags image-text-to-text and qwen3_5_moe, the model size of 403B parameters in BF16/F32, the Intern-S2-397B heading, an inference providers panel reading 'This model isn't deployed by any Inference Provider,' and the collection listing nine items.

The call

Choose Gemini 3.1 Pro if you need a general multimodal model that takes audio and video, holds a million tokens, has months of independent validation, and can be rented by the token today. It is the better-evidenced, more broadly capable model, and the preview badge is a reliability caveat rather than a capability one.

Choose Intern-S2 if your multimodal input is scientific and your data cannot leave your infrastructure. It is the only one of the two that reads raw literature pages natively, forecasts from time-series signals, and can be fine-tuned on proprietary experimental data under a permissive licence. Accept that you are the first person to run it, and that the benchmark table arguing for it was written by the people who made it.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro at orcarouter.ai/models/google/gemini-3.1-pro-preview, captured September 13, 2026, showing the model tagged FLAGSHIP and FEATURED by Google dated 2026-02-19 with Vision, Audio, Tools, JSON and Reasoning tags, a 1M-token context window and 65K max output, and pricing tiles reading input $2.00 and output $12.00 per 1M tokens.

Neither model wins this outright, and the table proves it better than a verdict could. One is a generalist that happens to be good at science; the other is a scientific instrument that happens to be competitive on general multimodal work — and on a same-day open release, "competitive" is the more surprising result.

To hold both halves of this comparison without a second contract: one API key covering 200-plus models puts the closed multimodal flagship and every alternative behind one endpoint, so you can route image-and-video work one way and document batches the other.