
Intern-S2 vs GPT-5.6 Sol: The Free Checkpoint and the $30 Flagship
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
The price line in this matchup looks absurd at first glance. Intern-S2, a 397-billion-parameter multimodal model that InternLM released under Apache-2.0 today, costs nothing to download and nothing per query once you own the hardware to run it. GPT-5.6 Sol, OpenAI's flagship reasoning tier, lists at $5.00 per million input tokens and $30.00 per million output tokens on OpenAI's base tier — though the promotional rate OpenAI has been running since late July, which our own model page currently shows, is $4.00 in and $20.00 out. That is a real gap and it is the first thing anyone notices — and it is also the least useful fact in the comparison, because the $30 model and the free one are not competing for the same work. GPT-5.6 Sol is the tier OpenAI points at when the hardest reasoning your application will ever attempt has to come back correct. Intern-S2 is a scientific instrument with a vision path trained on raw literature pages and a forecast head for time-series signals. The question is not which is cheaper. It is which one your workload actually needs, and whether the "free" one is free.
What the vendor table does and does not tell you
InternLM benchmarked Intern-S2 against GPT-5.5, not GPT-5.6 Sol. That is worth stating before anything else, because every "Intern-S2 beats GPT" claim in circulation is a comparison against the previous OpenAI generation, and any page that presents it as a Sol result is extrapolating. The GPT-5.6 family also spans three tiers — Sol for the hardest work, Terra for balanced production, Luna for high volume — and OpenAI has been actively repricing them, so a Sol figure quoted from a September article may already be stale.
What the vendor's comparison does establish, with the standing caveat that every Intern-S2 figure is InternLM's own and unreproduced:
• Molecular structure reasoning — Intern-S2 62.35 on MolecularIQ vs GPT-5.5's 76.41. Intern-S2 loses.
• Scientific tool use — Intern-S2 66.76 on TOMG-Bench vs GPT-5.5's 69.89. A narrow loss.
• Bio-molecular instruction — Intern-S2 53.95 vs 40.49. A clear Intern-S2 win.
• High-school maths — Intern-S2 93.56 on HMMT-2026 vs GPT-5.5's 97.06. Intern-S2 loses.
• General knowledge — Intern-S2 89.77 on MMLU Pro vs 88.20. A narrow Intern-S2 win.
• Multimodal knowledge — level at 81.68 on MMMU Pro.
• Terminal mastery — GPT-5.5 79.40 on TerminalBench 2.1 vs Intern-S2 64.04. A fifteen-point deficit.
Even against a one-generation-older OpenAI model, Intern-S2 loses more general rows than it wins and wins on the scientific ones. Against Sol — a tier explicitly positioned above GPT-5.5 for the hardest reasoning — the general gap almost certainly widens. Nothing in the evidence supports the idea that the free checkpoint is competitive with the paid flagship on general capability, and the honest framing is that Intern-S2 is a specialist that is respectable in general and excellent in its lane.
"Free" is a capital expenditure, not a discount
Intern-S2's zero licence cost is real, but it is not the same as free inference, and this is where the comparison gets practical. A 403-billion-parameter checkpoint needs a serious multi-GPU node to serve. If you already own that capacity and it is not fully utilized, the marginal cost of the next scientific query is close to electricity — which is a genuinely powerful economic position, and the reason open weights exist. If you do not own that capacity, "free" means buying or renting hardware, and the arithmetic changes completely.
GPT-5.6 Sol's economics are the inverse. There is no hardware to buy and no model to operate. You pay the metered rate — $5.00 per million input and $30.00 per million output on the base tier, or the $4.00/$20.00 promotional rate currently live on our model page — and the model arrives with a 1.05M-token context window, a 128K output ceiling, vision, tools, and the accumulated reliability of a flagship that has been in production since its API launched on July 9, 2026. The Ultrafast preview, which runs Sol on wafer-scale hardware at up to 750 output tokens per second, pushes the same model toward latency-critical work, though it is limited to a small group of customers and has no published price. What you are buying with the $30 is not just capability — it is the absence of an operations team.
• Weights — Intern-S2 Apache-2.0, downloadable vs GPT-5.6 Sol closed, API-only.
• Cost shape — Intern-S2 capital expenditure on a 403B checkpoint, or metered official Intern API vs GPT-5.6 Sol $5.00 / $30.00 per million list, $4.00 / $20.00 promotional.
• Context / output — Intern-S2 up to 256K text, 64K multimodal evaluated; output unpublished vs GPT-5.6 Sol ~1.05M context, 128K output.
• Inputs — Intern-S2 text, image, time series vs GPT-5.6 Sol text and image.
• Independent score — Intern-S2 none yet vs GPT-5.6 Sol top-tier flagship with a long public evaluation record.
• Speed tier — Intern-S2 hardware-bound vs GPT-5.6 Sol Ultrafast preview at up to 750 output tokens per second, vendor-stated.

Where the free checkpoint genuinely wins
There are three situations where Intern-S2 beats Sol without a benchmark being involved.
The first is data residency. GPT-5.6 Sol is a closed API — every request leaves your infrastructure and travels through OpenAI's systems under OpenAI's content policy. Intern-S2 under Apache-2.0 can run air-gapped inside a regulated environment, on data that legally cannot be sent to a third-party endpoint. For a pharmaceutical lab, a hospital system, or a defense contractor, that property is not a feature; it is the entire procurement decision, and no amount of Sol capability substitutes for it.
The second is modality. Intern-S2 accepts time-series input and produces forecasts; it also reads raw scientific literature pages as a native visual representation rather than as extracted text. Sol takes text and images. If your data is a seismograph trace, a load curve, or a scanned figure-dense paper, Intern-S2 has a path that Sol does not, and the vendor's Biology-Instructions row (55.71 against GPT-5.5's 10.52) is the number that reflects it.
The third is fine-tuning. You can take Intern-S2's weights and train on your own proprietary experimental data, then keep that model forever. You cannot do that with Sol at any price. For a team whose competitive advantage is a specialized dataset, an adaptable open checkpoint can be worth more than a stronger frozen model.

Running both without choosing
The teams that get the most out of this pairing do not treat it as a fork. They route the general, latency-sensitive, high-stakes reasoning to GPT-5.6 Sol — which sits on OrcaRouter at OpenAI's list price with 0% markup, on the same key as 200-plus other models, so a Sol price change passes through the same day and a bad provider hour is covered by automatic failover — and they keep the scientific batch work on Intern-S2 wherever they run it. The routing DSL makes that boundary explicit: hard general reasoning goes one way, document and signal analysis goes the other, and the split is a configuration rather than a second integration.
Intern-S2 is not on OrcaRouter; it is a same-day Apache-2.0 checkpoint, and your route to it is the vendor's own API or a self-hosted deployment. Sol is on the key. Between the two, the practical architecture for a science-heavy application is one endpoint for the frontier calls and one deployment for the specialist — with the option to move either side later.

The verdict
If you need the hardest reasoning your application will attempt, want it in one API call, and are willing to pay the metered flagship rate for it, GPT-5.6 Sol is the answer and a same-day open checkpoint does not change that. Nothing in InternLM's own table shows Intern-S2 matching a current OpenAI flagship on general capability; the comparison was run against the previous generation and the open model still lost most of the general rows.
If your workload is scientific, your data cannot leave your infrastructure, or you need to fine-tune on proprietary experimental results, Intern-S2 is the model that can do those things at all — and its Apache-2.0 licence means the version you validate is the version you keep. Budget for the hardware, expect to run your own evaluation, and treat every published number about it as a claim until someone outside InternLM reproduces one.
For the frontier calls alongside whatever you self-host: 200-plus other models keeps the closed flagship on one key at provider list price, so the specialist deployment and the metered model can be swapped by routing rule.
