A generated hero card comparing Intern-S2-397B and Qwen3.8-Max, showing on the left an open scientific checkpoint with a figure-page input and an Apache-2.0 badge, and on the right a metered flagship endpoint with a rate-card tile reading $2 / $6 and a 1M-token context meter, with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Intern-S2-397B vs Qwen3.8-Max: Two Chinese Flagships, Opposite Economic Wagers

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Intern-S2-397B and Qwen3.8-Max are the two most consequential Chinese flagship launches of early September 2026, and they have made opposite economic wagers. Intern-S2-397B, the 403-billion-parameter scientific multimodal model that Shanghai AI Laboratory's InternLM team released under Apache-2.0 on September 13, bets that open weights are the way to win scientific and agentic workloads: no per-token price, weights on Hugging Face, and your own hardware as the meter. Qwen3.8-Max, Ali​baba's 2.4-trillion-parameter flagship that went GA on August 3 and was refreshed on September 2, bets on a priced API at $2.00 per million input tokens and $6.00 per million output tokens across a 1M-token context, with open weights arriving a week after GA and an independent Artificial Analysis score already on the board. The matchup between them is less about whose benchmarks win and more about which bet — ownership or a metered endpoint — matches the shape of your workload.

Two flagships, two wagers

Qwen3.8-Max is Ali​baba's highest-capability tier in the Qw​en line, a sparse mixture-of-experts model with 2.4 trillion total parameters and about 95 billion activated per token across 512 experts with 10 active plus a shared one. It carries a 1M-token context window (983,616 tokens with thinking enabled in the hosted API), a 131,072-token output ceiling, text, image and video input, and it is served over an OpenAI-compatible endpoint. Its defining pricing move is a flat rate: the same $2.00 / $6.00 whether your prompt is 5,000 tokens or 900,000, with cached input priced lower — no long-context surcharge, which is exactly the lever most competitors still charge for. Ali​baba positions it directly against GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro in its own migration guide.

Intern-S2-397B is a different shape of flagship. It is a mixture-of-experts model with 60 layers and 512 experts (10 active per token), a 256K text reasoning context and 64K multimodal, an input surface of text, image and time series, and a pre-training paradigm that reads raw pages of scientific literature — figures, layout, symbolic notation and prose — rather than parsing text out first. Its reinforcement learning spans more than twenty scientific domains, and it was trained for long-horizon agent work in sandboxed environments. Its pricing is not a rate card but an absence of one: self-host the Apache-2.0 weights, or request access to the official Intern API.

• Price — Intern-S2-397B: no rate card; self-host or Intern API vs Qwen3.8-Max: $2.00 / $6.00 per 1M flat, cached input lower.

• Scale — Intern-S2-397B: 403B total, 512 experts / 10 active vs Qwen3.8-Max: 2.4T total, 95B active, 512 experts / 10 + 1 shared.

• Context — Intern-S2-397B: 256K text / 64K multimodal vs Qwen3.8-Max: 1M tokens, flat-rate.

• Inputs — Intern-S2-397B: text, image, time series vs Qwen3.8-Max: text, image, video.

• Weights — Intern-S2-397B: Apache-2.0, open on release day vs Qwen3.8-Max: custom Qwen3.8-Max licence, opened August 12 after GA.

• Independent score — Intern-S2-397B: none yet vs Qwen3.8-Max: AA Intelligence Index 40, GPQA Diamond 92.7.

Both tables are vendor tables, and only one has a referee

Both models launched with vendor-run benchmark tables, which makes the sourcing discipline identical and the track record different. Qwen3.8-Max's independent numbers have since accumulated: Artificial Analysis puts it at 40 on its current Intelligence Index with a GPQA Diamond of 92.7, an HLE of 43.0 and an independent coding score of 71.8, at a cost of about $2.67 per index task on the current scale. Those are measured numbers from a third-party tracker — the first real referee either of these two flagships has had.

Intern-S2-397B's evidence is its own card, run through OpenCompass, VLMEvalKit and AgentCompass, published alongside the weights: MMLU Pro 89.77, MMMU Pro 81.68, HMMT-2026 93.56, SWE-bench-Pro 68.54, and TerminalBench 2.1 at 64.04, a figure the card itself shows losing to an older closed generation. Its scientific rows are the reason the specialist case exists: Biology-Instructions 55.71, SciReasoner 1.5 63.28, Mol-Instructions 53.95, MolecularIQ 62.35. Every one of those is InternLM's own, unreproduced. Put side by side, the two evidence bases tell the same story from opposite directions: one model is a specialist with respectable general rows and no outside measurement, the other is a generalist with an independent score and no scientific tuning.

A generated two-column scoreboard titled 'Intern-S2-397B vs Qwen3.8-Max — the scoreboard.' Left column 'Intern-S2-397B': Weights Apache-2.0 open; Scale 403B total MoE; Context 256K text / 64K multimodal; Inputs text, image, time series; Independent score none yet; Price self-host or Intern API. Right column 'Qwen3.8-Max': Weights custom Qwen3.8-Max licence, opened Aug 12; Scale 2.4T total, 95B active; Context 1M flat-rate tokens; Inputs text, image, video; Independent score AA Index 40, GPQA 92.7; Price $2.00 / $6.00 per 1M. Footer reads 'Intern-S2-397B figures are InternLM's own, unreproduced; Qwen3.8-Max figures per Artificial Analysis and Alibaba.'

What the money shapes actually buy

The flat $2 / $6 meter is Qwen3.8-Max's signature move, and it is worth spelling out what it buys. A repo-analysis agent that pushes 200,000 input tokens of code context through in a single call pays the same per-token rate as a short chat — no long-prompt surcharge — and with caching, a stable code context drops the effective input price dramatically on repeat traffic. The model is also callable as open weights under Ali​baba's custom licence, which permits commercial use with attribution plus scale-based gates above 100 million MAU or $20 million monthly revenue, and a separate agreement for large model-as-a-service businesses.

Intern-S2-397B's economics are the inverse: no meter at all, but a capital expenditure. The weights are roughly 807 GB in BF16 across 188 shards — a serious multi-GPU node — with an FP8 build cutting the footprint by about half. If you already own that capacity, the marginal cost of the next scientific query approaches electricity, and that is the wager: ownership makes the specialist cheap at the margin. If you do not own the hardware, the "free" model is a pilot, and the official Intern API is the only metered alternative.

The two flagships can share one seam if you route. Qwen3.8-Max is on OrcaRouter at Ali​baba's list price passed through at 0% markup, so the flat $2 / $6 rate and any future price cut land here the same day. Intern-S2-397B is not routed — it is a brand-new checkpoint, and your path to it is the vendor's own API or a self-hosted deployment. Send the general, video-capable, long-context traffic to the metered model on the key you already have; keep the scientific batch work on the model you run; let automatic failover cover you when either side's endpoint is unhealthy. The boundary is a routing rule rather than a second integration.

A screenshot of the Hugging Face model card for internlm/Intern-S2, captured September 13, 2026, showing the Apache-2.0 licence, the description of Intern-S2-397B as a 403-billion-parameter multimodal foundation model for scientific intelligence and long-horizon agents, and the feature blocks covering scientific-literature pre-training and multi-domain scientific reinforcement learning.

The verdict

Choose Qwen3.8-Max for general reasoning and video-capable production work where a flat rate across a million tokens beats anything its rivals charge, and where an independent scoreboard reduces the risk of building on it. Its weights are open enough to matter under a licence most teams are inside the gates of, and its $2 / $6 meter is the most aggressive price on the board among frontier flagships.

Choose Intern-S2-397B for scientific workloads — figure-dense papers, molecular diagrams, time-series signals — where the input is multimodal in a way a generalist was never tuned for, and where data residency or fine-tuning on proprietary results makes Apache-2.0 the decisive property. Budget for the hardware, run your own evaluation, and treat its numbers as claims until an independent referee appears.

The honest summary is that these two flagships are not really substitutes. One is an owned scientific instrument with respectable general ability; the other is a rented, independently scored generalist with a flat rate and a passing interest in science. Teams that need both should treat the boundary as a routing decision — and should re-run their own evals on Intern-S2-397B before believing any number on either vendor's slide.

A screenshot of the OrcaRouter model page for Qwen3.8-Max at orcarouter.ai/models/qwen/qwen3.8-max, captured September 13, 2026, showing the model listing with its provider Qwen, 1M-token context window, and $2.00 / $6.00 per-million-token pricing displayed alongside the surrounding catalogue.