
Laguna S 2.1: Poolside's Open-Weight Coding Model That Punches Far Above Its Weight (and Price)
- qwenNEWQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaNEWOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicNEWAnthropic: Claude Opus 52026-07-2461Intelligence78Coding
- googleNEWGoogle: Gemini 3.6 Flash2026-07-2150Intelligence69Coding
- googleNEWGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaNEWMeta: Muse Spark 1.12026-07-1651Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1557Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0951Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0955Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0959Intelligence77Coding
- grokxAI: Grok 4.52026-07-0854Intelligence72Coding
- tencentTencent: Hy32026-07-0641Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3053Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
- z-aiZ.ai: GLM 5.22026-06-1651Intelligence69Coding60Math
- kimiMoonshotAI: Kimi K2.7 Code2026-06-1242Intelligence61Coding61Math
- anthropicAnthropic: Claude Fable 52026-06-0960Intelligence77Coding
- qwenQwen: Qwen3.7 Plus2026-06-0139Intelligence56Coding59Math
On July 21, 2026, Poolside released Laguna S 2.1 — a 118-billion-parameter open-weight coding model with only about 8 billion active parameters per token — and pitched it as "the West's most capable open-weight model" and its answer to DeepSeek and Qwen. The headline isn't that it beats every closed frontier system; it's that it beats models many times its size on agentic coding benchmarks while costing $0.10 / $0.20 per million tokens. This guide covers what Laguna S 2.1 actually is, its thinking mode, what's independently versus vendor-reported, what it costs, how to run it, and who should use it.
Every figure below is labeled by source. Laguna's headline scores are Poolside-reported (from its launch materials and Hugging Face card) unless a named third party is cited; treat them as vendor-run until independent labs confirm. Benchmarks and prices move — verify before you commit budget.
TL;DR. Laguna S 2.1 is an open-weight, MoE coding model (118B total / ~8B active) with up to 1M context and a "thinking / no-thinking" toggle, priced at just $0.10 input / $0.20 output per 1M tokens — a fraction of most rivals. Poolside reports 78.5% on SWE-Bench Multilingual (leading open disclosed-size models), 70.2% on Terminal-Bench 2.1, and 40.4% on DeepSWE v1.1 (versus DeepSeek-V4-Pro-Max's 9.0%) — at roughly one-sixth the active parameters. It is the best open-weight coding-per-dollar option we've seen, but closed frontier models like Claude Fable 5, Claude Opus 5, and Kimi K3 still lead on several absolute benchmarks. If you want a cheap, self-hostable coder that outperforms its weight class, Laguna S 2.1 is the story; if you need the outright top coding accuracy, a frontier flagship still wins.
Key takeaways
• Best open-weight coding per dollar. $0.10/$0.20 per 1M tokens for a model that leads open disclosed-size models on SWE-Bench Multilingual (78.5%).
• Punches above its weight. ~8B active parameters, yet matches or beats far larger models on Terminal-Bench 2.1 and SWE-Bench Pro.
• Open-weight and self-hostable. Weights on Hugging Face (poolside/Laguna-S-2.1) — you can run it in your own environment.
• Thinking mode drives the scores. Max thinking lifts DeepSWE from 16.5% to 40.4% — powerful, but it spends reasoning tokens, so budget for it.
• Not the outright #1. Closed frontier flagships (Fable 5, Opus 5, Kimi K3) still lead several absolute coding benchmarks; Laguna's claim is weight-class and price, not top of the board.
A note on reading this brief: everything attributed to Poolside is a vendor claim from launch materials; where a benchmark is independent we name the source. Where an independent general-intelligence score isn't published, we say so rather than guess — Laguna is a coding specialist, not a general-purpose leaderboard topper.
What Laguna S 2.1 actually is
Laguna S 2.1 is Poolside's open-weight, agentic coding model. Architecturally it's a mixture-of-experts: 118B parameters total but only about 8B activated per token, which is why it's cheap to serve relative to its capability. It supports up to a 1M-token context and two operating modes — a fast "no-thinking" path and a "thinking" path that spends extra reasoning tokens for harder problems. It's built for coding and agentic work: tool use, function calling, multi-step tasks. Poolside also ships a smaller sibling, Laguna XS 2.1 (33B total / ~3B active), at an even lower $0.06 / $0.12.
Crucially, the weights are open and published on Hugging Face, so unlike a closed API model you can download Laguna S 2.1, run it in your own infrastructure, fine-tune it, and pin a version. That combination — frontier-adjacent coding, tiny active-parameter count, open weights, and a rock-bottom price — is the whole pitch, and it's why Poolside frames it as the West's answer to the strong open models from DeepSeek and Qwen.
What's new: the benchmarks
The launch numbers are all about coding. On SWE-Bench Multilingual, Poolside reports 78.5%, which it says leads open models of disclosed size. On Terminal-Bench 2.1 it scores 70.2%. On DeepSWE v1.1 it reaches 40.4% — and Poolside's most striking comparison is that this beats DeepSeek-V4-Pro-Max's 9.0% on the same test at roughly one-sixth the active parameters. On Terminal-Bench 2.1 and SWE-Bench Pro, Poolside says Laguna S 2.1 matches or exceeds models several times its size, including DeepSeek V4 Flash, NVIDIA's Nemotron 3 Ultra, and Thinking Machines' Inkling.
The honest framing Poolside itself uses is about weight class, not outright supremacy: closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks. So the right way to read Laguna S 2.1 is "the most capability you can get from an open, ~8B-active, dirt-cheap model," not "the best coder period."

Thinking mode: where the performance comes from
Laguna S 2.1's headline scores lean heavily on its "max thinking" mode. The clearest example is DeepSWE v1.1, where the score jumps from 16.5% without thinking to 40.4% with max thinking — most of the model's benchmark strength comes from letting it reason. Terminal-Bench 2.1 shows the same pattern (60.4% → 70.2%). This matters for cost and latency: thinking mode spends extra reasoning tokens, so while the per-token price is tiny, a hard task run at max thinking uses more tokens (and takes longer) than the no-thinking path. In practice you'd run routine completions in no-thinking mode for speed and cost, and switch to thinking only for the hard, multi-step problems where the accuracy jump is worth the tokens — a coding-specific echo of the effort dials appearing on frontier models.
The benchmark deep-dive: what these coding tests measure
SWE-Bench Multilingual extends the well-known SWE-bench (resolving real GitHub issues) across multiple programming languages, so 78.5% is a strong signal of practical, cross-language bug-fixing — and "leads open disclosed-size models" is a meaningful, if vendor-stated, claim. Terminal-Bench 2.1 tests whether a model can operate a real terminal to complete tasks end to end, which is closer to what an autonomous coding agent actually does than a single-shot code completion. DeepSWE v1.1 is a harder agentic software-engineering suite; the 40.4% (vs DeepSeek-V4-Pro-Max's 9.0%) is Laguna's flashiest number precisely because it's a large gap against a much bigger model.
The caveat that keeps this honest: these are Poolside-run results as of launch, and there is no published Artificial Analysis Intelligence Index or independent SWE-bench Verified figure for Laguna S 2.1 yet. So treat the leadership claims as vendor-reported until a third party confirms them, and remember Laguna is a coding specialist — it is not positioned as a general-purpose reasoning leader, and you shouldn't expect it to top a general-intelligence index.

What it costs in practice
This is where Laguna S 2.1 is genuinely disruptive. At $0.10 input / $0.20 output per 1M tokens, a representative coding call of 15,000 input and 3,000 output tokens costs 15k × $0.10/1M + 3k × $0.20/1M = $0.0015 + $0.0006 = about $0.0021 — roughly a fifth of a cent. The same call on a frontier model like Claude Opus 5 ($5/$25) is about $0.15 — nearly 70× more. Even against cheap rivals, Laguna undercuts: it's positioned below DeepSeek's pricing. The asterisk is thinking mode — a hard task at max thinking may generate several times more output tokens, so a "$0.0006 output" call could become a few cents. But even inflated, the cost is a rounding error next to closed frontier models. For high-volume coding automation — CI bots, bulk refactors, agent fleets — the economics are hard to argue with, especially since you can also self-host to remove per-token cost entirely.
How to access and run Laguna S 2.1
Because the weights are open, you have two paths. You can call it via hosted providers (it appears on aggregators like OpenRouter) at the $0.10/$0.20 rate, which is the fastest way to try it. Or you can download the weights from Hugging Face (poolside/Laguna-S-2.1) and self-host — keeping code and data in your environment, fine-tuning on your repositories, quantizing for your hardware, and capping cost at your infrastructure. The XS variant (33B-A3B) is the option when you want to run something even smaller and cheaper locally. Either way, it exposes tool use and function calling, so it drops into agentic coding harnesses.
Who it's for: three scenarios
1. High-volume coding automation on a budget
If you run CI fix-it bots, bulk migrations, or large agent fleets where token cost compounds, Laguna S 2.1's price is transformative — run no-thinking for routine work and thinking for the hard cases. This is its strongest use case.
2. Teams that must self-host code models
For organizations that can't send proprietary code to a closed API, an open-weight coder that leads its size class is exactly what's been missing. Laguna S 2.1 (or XS for smaller hardware) gives you frontier-adjacent coding you fully control.
3. Cost-sensitive startups building coding products
If your product margins depend on inference cost, Laguna lets you offer AI coding features at a fraction of frontier prices — reserving a frontier flagship only for the hardest requests your users hit.

When not to use it
Don't reach for Laguna S 2.1 when you need the outright best coding accuracy and can pay for it — closed frontier models (Fable 5's 95.0% SWE-bench Verified, Opus 5's #1 general index, Kimi K3's #1 Frontend Coding Arena) still lead several benchmarks. Don't use it for general-purpose reasoning, multimodal, or non-coding tasks — it's a coding specialist without a general-intelligence pedigree. And don't assume the headline numbers hold at no-thinking mode; if you disable thinking for cost, benchmark the lower scores you'll actually get.
Decision checklist
• Choose Laguna S 2.1 if: you want the cheapest capable open-weight coder, you must self-host, or you run high-volume coding automation where cost dominates.
• Choose a frontier flagship instead if: you need the outright top coding accuracy, general reasoning, or multimodal — and can afford it.
• Run both if: default routine coding to Laguna and escalate the hardest tasks to a frontier model, behind one endpoint.
FAQ
How much does Laguna S 2.1 cost?
$0.10 per 1M input tokens and $0.20 per 1M output — with a smaller Laguna XS 2.1 at $0.06 / $0.12. It's one of the cheapest capable coding models available, and being open-weight you can self-host to remove per-token cost. [OpenRouter/BenchLM]
Is Laguna S 2.1 open source?
It's open-weight: the model weights are published on Hugging Face (poolside/Laguna-S-2.1), so you can download, run, and fine-tune it. Check Poolside's license for your specific use.
Is it better than DeepSeek or Qwen for coding?
Poolside positions it as the West's answer to them and reports beating DeepSeek-V4-Pro-Max on DeepSWE (40.4% vs 9.0%) at far fewer active parameters. On open, disclosed-size coding benchmarks it leads per Poolside — but these are vendor-run, so validate on your own repositories.
Does it beat the frontier models?
Not outright. Closed flagships like Claude Fable 5, Claude Opus 5, and Kimi K3 still lead several absolute benchmarks. Laguna's claim is best-in-weight-class and best-per-dollar, not best overall.
What is thinking mode?
A toggle between a fast no-thinking path and a max-thinking path that spends extra reasoning tokens for higher accuracy (e.g., DeepSWE 16.5% → 40.4%). Use no-thinking for routine work, thinking for hard tasks.
What's the context window?
Up to 1M tokens, so large codebases and long agent transcripts fit.
Should I use S or XS?
Laguna S 2.1 (118B/8B) for the strongest coding; Laguna XS 2.1 (33B/3B) when you want something smaller and cheaper to self-host or serve at very high volume.
Bottom line
Laguna S 2.1 is the most compelling open-weight coding model of its size to date: a 118B/8B MoE with up to 1M context and a thinking toggle, priced at a remarkable $0.10 / $0.20, that Poolside says leads open disclosed-size models on SWE-Bench Multilingual (78.5%) and beats far larger rivals on agentic coding tests. It is not the outright best coder — closed frontier flagships still lead several benchmarks — and its headline scores lean on thinking mode, which spends tokens. But for cheap, self-hostable, high-volume coding, nothing in its weight class competes. Route Laguna S 2.1 alongside every frontier flagship through one OpenAI-compatible endpoint at OrcaRouter and send it the bulk of your coding traffic, escalating only the hardest tasks.
