A generated hero card comparing Intern-S2 and GPT-5.6 Sol, showing a scientific figure page and signal waveform feeding an open-weights model on the left and a closed frontier API endpoint with a speed gauge on the right, labelled Apache-2.0 free to download and self-hosted on your own GPUs against $5.00 / $30.00 per 1M tokens and 1.05M context / 128K output, with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Intern-S2 vs GPT-5.6 Sol: The Free Checkpoint and the $30 Flagship

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The price line in this matchup looks absurd at first glance. Intern-S2, a 397-billion-parameter multimodal model that InternLM released under Apache-2.0 today, costs nothing to download and nothing per query once you own the hardware to run it. GPT-5.6 Sol, Open​AI's flagship reasoning tier, lists at $5.00 per million input tokens and $30.00 per million output tokens on Open​AI's base tier — though the promotional rate Open​AI has been running since late July, which our own model page currently shows, is $4.00 in and $20.00 out. That is a real gap and it is the first thing anyone notices — and it is also the least useful fact in the comparison, because the $30 model and the free one are not competing for the same work. GPT-5.6 Sol is the tier Open​AI points at when the hardest reasoning your application will ever attempt has to come back correct. Intern-S2 is a scientific instrument with a vision path trained on raw literature pages and a forecast head for time-series signals. The question is not which is cheaper. It is which one your workload actually needs, and whether the "free" one is free.

What the vendor table does and does not tell you

InternLM benchmarked Intern-S2 against GPT-5.5, not GPT-5.6 Sol. That is worth stating before anything else, because every "Intern-S2 beats G​PT" claim in circulation is a comparison against the previous Open​AI generation, and any page that presents it as a Sol result is extrapolating. The GPT-5.​6 family also spans three tiers — Sol for the hardest work, Terra for balanced production, Luna for high volume — and Open​AI has been actively repricing them, so a Sol figure quoted from a September article may already be stale.

What the vendor's comparison does establish, with the standing caveat that every Intern-S2 figure is InternLM's own and unreproduced:

• Molecular structure reasoning — Intern-S2 62.35 on MolecularIQ vs GPT-5.5's 76.41. Intern-S2 loses.

• Scientific tool use — Intern-S2 66.76 on TOMG-Bench vs GPT-5.5's 69.89. A narrow loss.

• Bio-molecular instruction — Intern-S2 53.95 vs 40.49. A clear Intern-S2 win.

• High-school maths — Intern-S2 93.56 on HMMT-2026 vs GPT-5.5's 97.06. Intern-S2 loses.

• General knowledge — Intern-S2 89.77 on MMLU Pro vs 88.20. A narrow Intern-S2 win.

• Multimodal knowledge — level at 81.68 on MMMU Pro.

• Terminal mastery — GPT-5.5 79.40 on TerminalBench 2.1 vs Intern-S2 64.04. A fifteen-point deficit.

Even against a one-generation-older Open​AI model, Intern-S2 loses more general rows than it wins and wins on the scientific ones. Against Sol — a tier explicitly positioned above GPT-5.5 for the hardest reasoning — the general gap almost certainly widens. Nothing in the evidence supports the idea that the free checkpoint is competitive with the paid flagship on general capability, and the honest framing is that Intern-S2 is a specialist that is respectable in general and excellent in its lane.

"Free" is a capital expenditure, not a discount

Intern-S2's zero licence cost is real, but it is not the same as free inference, and this is where the comparison gets practical. A 403-billion-parameter checkpoint needs a serious multi-GPU node to serve. If you already own that capacity and it is not fully utilized, the marginal cost of the next scientific query is close to electricity — which is a genuinely powerful economic position, and the reason open weights exist. If you do not own that capacity, "free" means buying or renting hardware, and the arithmetic changes completely.

GPT-5.6 Sol's economics are the inverse. There is no hardware to buy and no model to operate. You pay the metered rate — $5.00 per million input and $30.00 per million output on the base tier, or the $4.00/$20.00 promotional rate currently live on our model page — and the model arrives with a 1.05M-token context window, a 128K output ceiling, vision, tools, and the accumulated reliability of a flagship that has been in production since its API launched on July 9, 2026. The Ultrafast preview, which runs Sol on wafer-scale hardware at up to 750 output tokens per second, pushes the same model toward latency-critical work, though it is limited to a small group of customers and has no published price. What you are buying with the $30 is not just capability — it is the absence of an operations team.

• Weights — Intern-S2 Apache-2.0, downloadable vs GPT-5.6 Sol closed, API-only.

• Cost shape — Intern-S2 capital expenditure on a 403B checkpoint, or metered official Intern API vs GPT-5.6 Sol $5.00 / $30.00 per million list, $4.00 / $20.00 promotional.

• Context / output — Intern-S2 up to 256K text, 64K multimodal evaluated; output unpublished vs GPT-5.6 Sol ~1.05M context, 128K output.

• Inputs — Intern-S2 text, image, time series vs GPT-5.6 Sol text and image.

• Independent score — Intern-S2 none yet vs GPT-5.6 Sol top-tier flagship with a long public evaluation record.

• Speed tier — Intern-S2 hardware-bound vs GPT-5.6 Sol Ultrafast preview at up to 750 output tokens per second, vendor-stated.

A generated two-column scoreboard titled 'Intern-S2 vs GPT-5.6 Sol — the scoreboard.' Left column Intern-S2-397B lists Weights Apache-2.0, Cost self-host capex or Intern API, Context 256K text / 64K multimodal, Inputs text image time series, Independent score none yet, TerminalBench 2.1 64.04. Right column GPT-5.6 Sol lists Weights closed API only, Cost $5.00 / $30.00 list and $4.00 / $20.00 promo, Context 1.05M in / 128K out, Inputs text and image, Independent score flagship tier widely tested, TerminalBench 2.1 79.40. Footer reads 'Intern-S2 figures are InternLM's own and were measured against GPT-5.5, unreproduced; the TerminalBench row uses GPT-5.5 as the OpenAI reference. GPT-5.6 Sol pricing per OpenAI.'

Where the free checkpoint genuinely wins

There are three situations where Intern-S2 beats Sol without a benchmark being involved.

The first is data residency. GPT-5.6 Sol is a closed API — every request leaves your infrastructure and travels through Open​AI's systems under Open​AI's content policy. Intern-S2 under Apache-2.0 can run air-gapped inside a regulated environment, on data that legally cannot be sent to a third-party endpoint. For a pharmaceutical lab, a hospital system, or a defense contractor, that property is not a feature; it is the entire procurement decision, and no amount of Sol capability substitutes for it.

The second is modality. Intern-S2 accepts time-series input and produces forecasts; it also reads raw scientific literature pages as a native visual representation rather than as extracted text. Sol takes text and images. If your data is a seismograph trace, a load curve, or a scanned figure-dense paper, Intern-S2 has a path that Sol does not, and the vendor's Biology-Instructions row (55.71 against GPT-5.5's 10.52) is the number that reflects it.

The third is fine-tuning. You can take Intern-S2's weights and train on your own proprietary experimental data, then keep that model forever. You cannot do that with Sol at any price. For a team whose competitive advantage is a specialized dataset, an adaptable open checkpoint can be worth more than a stronger frozen model.

A screenshot of the Hugging Face model card for internlm/Intern-S2, captured September 13, 2026, showing the Apache-2.0 licence badge, the tags image-text-to-text and qwen3_5_moe, the model size of 403B parameters in BF16/F32, the Intern-S2-397B heading, an inference providers panel reading 'This model isn't deployed by any Inference Provider,' and the collection listing nine items.

Running both without choosing

The teams that get the most out of this pairing do not treat it as a fork. They route the general, latency-sensitive, high-stakes reasoning to GPT-5.6 Sol — which sits on OrcaRouter at Open​AI's list price with 0% markup, on the same key as 200-plus other models, so a Sol price change passes through the same day and a bad provider hour is covered by automatic failover — and they keep the scientific batch work on Intern-S2 wherever they run it. The routing DSL makes that boundary explicit: hard general reasoning goes one way, document and signal analysis goes the other, and the split is a configuration rather than a second integration.

Intern-S2 is not on OrcaRouter; it is a same-day Apache-2.0 checkpoint, and your route to it is the vendor's own API or a self-hosted deployment. Sol is on the key. Between the two, the practical architecture for a science-heavy application is one endpoint for the frontier calls and one deployment for the specialist — with the option to move either side later.

A screenshot of the OrcaRouter model page for GPT-5.6 Sol at orcarouter.ai/models/openai/gpt-5.6-sol, captured September 13, 2026, showing the model tagged FEATURED by OpenAI dated 2026-07-09 with Vision, Tools, JSON and Reasoning tags, a 1M-token context window and 128K max output, and pricing tiles reading input $4.00 and output $20.00 per 1M tokens with 202.3 million tokens of traffic in seven days.

The verdict

If you need the hardest reasoning your application will attempt, want it in one API call, and are willing to pay the metered flagship rate for it, GPT-5.6 Sol is the answer and a same-day open checkpoint does not change that. Nothing in InternLM's own table shows Intern-S2 matching a current Open​AI flagship on general capability; the comparison was run against the previous generation and the open model still lost most of the general rows.

If your workload is scientific, your data cannot leave your infrastructure, or you need to fine-tune on proprietary experimental results, Intern-S2 is the model that can do those things at all — and its Apache-2.0 licence means the version you validate is the version you keep. Budget for the hardware, expect to run your own evaluation, and treat every published number about it as a claim until someone outside InternLM reproduces one.

For the frontier calls alongside whatever you self-host: 200-plus other models keeps the closed flagship on one key at provider list price, so the specialist deployment and the metered model can be swapped by routing rule.