A generated hero card comparing Intern-S2-397B and Kimi K3, showing on the left a scientific literature page with a molecular diagram feeding a 403-billion-parameter open-weights card, and on the right a long document ribbon feeding a 2.8-trillion-parameter flagship with a 1M context meter, with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Intern-S2-397B vs Kimi K3: The 397B Science Model and the 2.8T Reasoning Giant

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Intern-S2-397B and Kimi K3 sit on opposite sides of the open-weights line that will define the rest of 2026: the downloadable scientific specialist versus the independently proven long-context generalist. Intern-S2-397B is the 403-billion-parameter scientific multimodal model that InternLM released under Apache-2.0 on September 13, 2026 — no rate card, no independent score, and a vision path trained on raw pages of scientific literature. Kimi K3 is Moonshot AI's 2.8-trillion-parameter mixture-of-experts flagship that launched in July 2026 at $3.00 per million input tokens and $15.00 per million output tokens, opened its weights on July 27 under a Modified MIT licence, and has been through six weeks of independent evaluation. The unusual thing about this matchup is that both models have the same long-horizon ambition — keep working through a big, messy task — and resolve it in opposite directions.

Both ambitious, opposite mechanisms

The ambition is shared: each model wants to be the thing that keeps going when the task is long. They achieve it differently, and the difference is architectural.

Kimi K3's answer is context. A 1,048,576-token window on both input and output means a whole repository, a whole data room, or a fifty-step agent loop can live in one conversation. Moonshot's Mooncake inference stack reportedly sustains cache-hit rates above 90% on coding workloads, which is how a $15-per-million output price becomes survivable in practice. Kimi K3 exposes a top-level reasoning_effort control and accepts no sampling parameters, reasons at maximum depth by default, and expects you to replay prior reasoning_content and tool_calls verbatim in multi-turn tool loops or quality degrades.

Intern-S2-397B's answer is specialization plus weight-level control. Its context — 256K text reasoning and 64K multimodal, both figures from the model card — is a different category of window, enough for a substantial paper or a long research session but not for a million-token working set. What it offers instead is a vision path that reads scientific literature natively — figures, layout, symbolic notation and prose together — and a time-series input that produces forecasts, alongside thinking that is on by default and switchable per request. If your long task is scientific, the model was trained for exactly it; if your long task is a codebase, this is not the model with the million-token window for it.

• Context — Kimi K3: 1,048,576 tokens in and out vs Intern-S2-397B: 256K text / 64K multimodal.

• Scale — Kimi K3: 2.8T total, ~104B active per token vs Intern-S2-397B: 403B total, 512 experts / 10 active.

• Control surface — Kimi K3: reasoning_effort dial, no sampling parameters vs Intern-S2-397B: enable_thinking toggle plus full weight-level control under Apache-2.0.

• Inputs — Kimi K3: text, image vs Intern-S2-397B: text, image, time series.

• Weights — Kimi K3: open under Modified MIT (July 27) vs Intern-S2-397B: open under Apache-2.0 (September 13).

• Independent evidence — Kimi K3: six weeks of third-party scores vs Intern-S2-397B: none yet.

The evidence asymmetry is the real difference

Kimi K3 has been through the independent wringer and come out with scores you can build on. Artificial Analysis places it in the mid-40s on its current Intelligence Index, with a GPQA Diamond in the low 90s, an HLE in the mid-40s, an independently measured long-context recall of 88.7, and a coding score in the mid-70s. It holds the top spot on the Frontend Code Arena (Elo in the 1670s) and posted 91.2 on BrowseComp, the highest figure on that long-horizon retrieval benchmark — though Moonshot itself cautions that K3 "still trails the most powerful proprietary models," and several of its launch-day rows remain vendor-reported.

Intern-S2-397B's evidence is its own launch table, run through OpenCompass, VLMEvalKit and AgentCompass, published hours before you read this. Every number is the vendor's own and unreproduced: MMLU Pro 89.77, MMMU Pro 81.68, HMMT-2026 93.56, SWE-bench-Pro 68.54, and TerminalBench 2.1 at 64.04 — a figure the card shows losing to an older closed generation by a wide margin. Its scientific rows are the numbers that make a specialist's case: 55.71 on Biology-Instructions, 63.28 on SciReasoner 1.5, 53.95 on Mol-Instructions, 62.35 on MolecularIQ. The two evidence bases do not touch, which is why the honest framing of this matchup is not "who wins" but "which kind of evidence does your workload need."

A generated two-column scoreboard titled 'Intern-S2-397B vs Kimi K3 — the scoreboard.' Left column 'Intern-S2-397B': Weights Apache-2.0 open; Scale 403B total MoE; Context 256K text / 64K multimodal; Inputs text, image, time series; Independent score none yet; Biology-Instructions 55.71 vendor-reported. Right column 'Kimi K3': Weights Modified MIT open; Scale 2.8T total, ~104B active; Context 1M tokens in and out; Inputs text, image; Independent score AA Index mid-40s, long-context recall 88.7; Price .00 / 5.00 per 1M. Footer reads 'Intern-S2-397B figures are InternLM's own, unreproduced; Kimi K3 figures per Artificial Analysis and Moonshot AI.'

What each costs, open-weights edition

Both models are open weights, which changes the cost question from the usual metered-versus-free story into a comparison of two hosting propositions. Kimi K3 has a published API price — $3.00 / $15.00 per million tokens at Moonshot's list, passed through on OrcaRouter at 0% markup, with cache-read pricing on top — and it is also downloadable: a 96-shard open-weights release that needs supernode-class hardware, 64 or more accelerators, to serve. The practical audience for a self-hosted K3 is teams with datacenter budgets. Intern-S2-397B has no published API price at all; the model card points at the official Intern API, and the real economics are the self-host ones: roughly 807 GB of weights across 188 shards in BF16, a serious multi-GPU node, with an FP8 build cutting the footprint by about half.

Both models are callable through the same seam if you route: Kimi K3 is on OrcaRouter at Moonshot's list price with 0% markup, while Intern-S2-397B is not routed — a brand-new Apache-2.0 checkpoint, your path to it being the vendor's own API or a self-hosted deployment. The routing value in this matchup is keeping the long-context generalist traffic on the key you already have and the scientific batch work on the model you run, with automatic failover between the two. The DSL makes that boundary a configuration rather than a second integration.

A screenshot of the Hugging Face model card for internlm/Intern-S2, captured September 13, 2026, showing the Apache-2.0 licence, the description of Intern-S2-397B as a 403-billion-parameter multimodal foundation model for scientific intelligence and long-horizon agents, and the note that text reasoning benchmarks use a 256K inference length while multimodal benchmarks use 64K.

The verdict

Choose Kimi K3 when the job is long-context generalist work — whole-repository coding, sustained agent loops, knowledge work that needs a million tokens of working memory — and you want an independent scoreboard behind it. It is the stronger-evidenced model in every sense, its open weights are real (Modified MIT, July 27), and if you are already operating supernode-class infrastructure the per-token economics with caching are favorable.

Choose Intern-S2-397B when the task is scientific: figure-dense papers, molecular diagrams, scanned literature, time-series signals, and data that has to stay inside your perimeter or weights you want to fine-tune on proprietary results. Apache-2.0 makes that possible, no rented API does, and the scientific rows in the card are a different species from what a generalist was ever tuned for. Budget for the hardware, run your own evaluation, and treat every published number about it as a claim until someone outside InternLM reproduces one.

A screenshot of the OrcaRouter model page for Kimi K3 at orcarouter.ai/models/kimi/kimi-k3, captured September 13, 2026, showing the model listing with its provider MoonshotAI, 1,048,576-token context window, and .00 / 5.00 per-million-token pricing displayed alongside the surrounding catalogue.