A generated title card reading 'FrogNano-4B-2609 vs Qwen 3.8' with the subtitle 'a 4B agent on your own GPU against a 2.4T hosted flagship', three chips reading '9.32 GB download', '$2 in / $6 out per million' and 'up to 150 steps per task', a footer reading 'FrogNano figures per Microsoft; Qwen3.8-Max pricing per Alibaba. Nothing here is independent', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

FrogNano-4B-2609 vs Qwen 3.8: What a Coding Agent Costs When the Weights Are Yours

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

microsoft/FrogNano-4B-2609 is a four-billion-class repository agent derived from Qwen3.5-4B that you download and serve yourself. Qwen3.8-Max is the hosted flagship — a 2.4-trillion-parameter sparse mixture-of-experts model with roughly 95 billion parameters active per token, a one-million-token multimodal window, and a rate card of $2.00 per million input tokens and $6.00 per million output — reachable with an API key since 3 August 2026. One of them costs a GPU and an afternoon. The other costs nothing until you use it, and then charges by the token on the longest trajectories in common use. That trade, and not the benchmark table, is the real matchup.

The FrogNano figures here come from Microsoft's model card and technical report; the Qwen3.8-Max figures come from Alibaba's published pricing and the model record on our own catalogue, with the independent scores labelled as Artificial Analysis figures wherever they appear. Note the naming carefully: Qwen3.8 the open-weights release has been announced but not shipped, while Qwen3.8-Max is live and priced today. Everything callable in this article is the latter.

The arithmetic that decides the comparison

Start with what a single task costs, because it is not obvious and it is not small.

FrogNano's evaluations ran every trajectory with a 150-step budget inside roughly 131K combined tokens, at up to 8,192 generated tokens per assistant turn, and a 10,800-second ceiling. Multiply the two published limits and you get a ceiling of about 1.23 million generated tokens for one worst-case trajectory — which at Qwen3.8-Max's $6.00 per million output tokens is roughly $7.38 for a single task that ran to the end of its budget. That is arithmetic on two published numbers, not a measurement: real trajectories stop early, the average is far below the ceiling, and tool results and shell output are billed at the cheaper input rate rather than the output rate. But it establishes the shape of the thing. A coding agent does not spend a few hundred tokens on a question; it spends a long, reasoning-heavy conversation with itself, and every token of that conversation is generated.

Now put the other side of the ledger next to it. FrogNano-4B-2609 is 9.32 GB of BF16 weights in two shards, downloadable without a gate. Serving it wants SGLang with a Qwen3 reasoning parser and a Qwen3 coder tool-call parser, plus an isolated sandbox for the five tools. After that the marginal cost of a trajectory is electricity, and the memory it occupies is whatever your context grows to rather than what the weights weigh.

• Cost model — FrogNano-4B-2609 is fixed hardware cost and near-zero marginal cost. Qwen3.8-Max is zero fixed cost and per-token billing, with a $2.00 / $6.00 split that front-loads nothing and back-loads everything at the output rate.

• What the money buys — a 4B-class agent on your own silicon, against a frontier-scale model you never have to capacity-plan for.

• Break-even — it is reached when your monthly trajectory volume times the average per-trajectory token bill exceeds the cost of the GPU you would otherwise rent. For a team running a handful of repository tasks a day, that line is a long way off and the hosted model wins on convenience alone.

Two models that do not do the same job

The cost comparison is only fair if the outputs are comparable, and here they are not, which is the part most spec sheets bury.

FrogNano-4B-2609 was post-trained exclusively on repository-level software engineering. Its whole interface is a five-tool loop — Read, Write, Edit, Glob and Bash — executed by the Leaf harness in an isolated sandbox, iterating until the model stops calling. Microsoft's card states its supported language as English and its validated programming domain as Python, and says outright that image and video components in the checkpoint were never post-trained and are not supported. It does not reason about your architecture, answer questions about your business, or read a screenshot.

Qwen3.8-Max is a general model. Alibaba's catalogue record gives it text, image and video input with text output and a 1,000,000-token context window. It reasons, writes, reads documents and holds a conversation, and its function calling is a feature of a general model rather than the entire point of the checkpoint. Its Artificial Analysis figures, read from the record on 2 September 2026, are an Intelligence Index of 45.4 and a Coding Index of 76.2.

• Input — FrogNano-4B-2609 is text only. Qwen3.8-Max takes text, images and video.

• Context — roughly 131K combined tokens for FrogNano's evaluated configuration, against 1,000,000 for Qwen3.8-Max. That is a factor of seven and a half, and it matters most for exactly the repository-sized inputs a coding agent gets.

• Output per turn — FrogNano's validated configuration caps a single assistant turn at 8,192 tokens. Qwen3.8-Max publishes no equivalent per-response ceiling.

• Output format — FrogNano emits structured Leaf tool calls and needs a harness to execute them. Qwen3.8-Max returns prose or a function call to whatever client you already have.

• Independent scores — none for FrogNano-4B-2609 beyond its own card; the Qwen3.8-Max figures above are third-party, from Artificial Analysis.

• Parameters — approximately 4.66B for FrogNano by its card's own description, against 2.4 trillion total and roughly 95B active for Qwen3.8-Max.

A screenshot of the OrcaRouter model page for Qwen3.8 Max showing the breadcrumb Home, Models, Qwen; the model name and the qwen/qwen3.8-max model id; the FEATURED badge with Vision, Tools, JSON and Reasoning capability chips; the 1M-token context, text, image and video input and text output row; the $2.00 and $6.00 per-million price cards with the p50 TTFT figure; the byline 'by Qwen, 2026-08-03'; the on-page section list for Code samples, Pricing, Performance, Public benchmarks, Community buzz, How it compares and FAQ; and the opening of the long description describing Qwen3.8-Max as Alibaba's newest flagship.

Where the 4B actually earns its place

There is one axis on which a 4B agent beats a frontier model outright, and it is not accuracy.

FrogNano's paper starts from Qwen3.5-4B, which scored 39.4% on SWE-bench Verified through the Leaf harness as a plain general model. Five reinforcement-learning iterations on roughly 1,500 synthetic tasks took it to 61.5%, with the same paper's appendix reporting the per-iteration progression as 48.2%, 53.4%, 58.3%, 58.6% and 61.6%. The method — generating tasks calibrated to the current policy's learnability frontier, with no distillation from a stronger model — is the contribution, and the reason a research group without frontier-model budget would care.

Read what that implies for the comparison. Microsoft's card is candid that FrogNano is not uniformly better than the model it came from: its parallel tool-call rate is 1.71%, because later iterations lost the ability to fire concurrent calls, and the team added a consolidation stage to try to recover it. A model that solves more issues while becoming worse at one specific behaviour is a real engineering artefact with visible seams, and its card says so rather than burying it.

And the SWE-bench number is a closed loop. The base-model baseline, the harness and the final score all come from the same lab, and the same paper's two tables give two different ladders for the same experiment. Qwen3.8-Max's coding figure, by contrast, was produced by Artificial Analysis rather than by Alibaba — different kind of evidence, different kind of confidence, and the two should never be subtracted from each other.

The 4B you cannot quite plan around yet

Two things stand between FrogNano-4B-2609 and a production slot, and both are in its own documentation rather than in any reviewer's opinion.

The first is hardware. The card gives the BF16 weight requirement as about 9.3 GB and then adds that memory "increases substantially as the combined context approaches approximately 131K tokens", before stating that the exact minimum GPU model, VRAM configuration, runtime versions and performance profile "must still be validated before release" — a sentence that appears on a model already downloadable. Nobody has published a VRAM figure for the evaluated configuration, and the card does not contain one.

The second is safety posture. Microsoft's own alignment note says the agent-specific post-training used no safety-preference, refusal, harmful-content or adversarial datasets, that it optimised for functional correctness and regression avoidance instead, and that the model "should not be considered independently safety-aligned for unrestricted autonomous deployment". The card closes by naming the concrete risks: patches may be incorrect or insecure despite passing their tests, performance is sensitive to the quality of those tests, and every candidate patch needs qualified human review and independent regression and security testing before use. A general hosted model carries none of those specific caveats, and it also will not run a shell command in your repository.

One key for whichever side you pick

The useful thing about running this comparison in practice is that only one half of it is a download. Qwen3.8-Max, and the general models you would benchmark a self-hosted agent against, are all reachable through one OpenAI-compatible endpoint on OrcaRouter at 0% markup — provider list price passed through, so a vendor price change lands on our side the same day, which is the number that actually moves when a coding agent's output-token bill is the thing you are trying to control. That also makes the failover question simple: if the hosted half of your pipeline degrades, the retry is automatic and the experiment continues.

FrogNano-4B-2609 is not one of our routes and there is no date for it. The weights are yours to serve, and the card is honest that the serving configuration has not been fully validated. What we can do is make the other side of the comparison cost nothing to set up, so the question you end up answering is the real one — whether a 61.5% vendor-reported SWE-bench agent on your own GPU beats a metered frontier model on your own tasks, at your own volume.

A generated two-column scoreboard titled 'FrogNano-4B-2609 vs Qwen3.8-Max'. The left column reads Cost your GPU, electricity; Output tool calls and patches; Context about 131K evaluated; Input text only; Published score 61.5% SWE-bench Verified, vendor, unreproduced; Turn cap 8,192 tokens. The right column reads Cost $2.00 in / $6.00 out per million; Output prose or function call; Context 1,000,000 tokens; Input text, image, video; Published score AA Intelligence 45.4 and AA Coding 76.2, third-party; Turn cap not published. A footer reads 'FrogNano figures per Microsoft; Qwen3.8-Max scores per Artificial Analysis. Different benchmarks; not comparable.'

Which one to reach for

Reach for Qwen3.8-Max when the work is general, when someone else's operations team should be carrying the capacity, when the input might contain an image or a long document, or when you need an answer today rather than a serving stack by Friday. Its third-party scores mean you can predict roughly what you are buying, and its per-token bill is a variable cost you can see before you commit. For a team running a modest number of repository tasks a month, the arithmetic is not close.

Reach for FrogNano-4B-2609 when the volume is high enough that per-token billing on long trajectories is the constraint, when the data cannot leave your infrastructure, or when the research question — can a small agent be trained to repository-level competence without a teacher — is the thing you are actually testing. Go in knowing three things: the 61.5% is a lab's own measurement on a lab-maintained harness, nobody has published a VRAM figure for the evaluated configuration, and the card itself says the serving setup is not yet validated.

What makes the pairing worth writing about is that they share an ancestor two generations back. FrogNano descends from Qwen3.5-4B — a general model from the same company whose 2.4-trillion-parameter flagship sits on the other side of this page. Microsoft took the small one, spent five RL iterations on synthetic repositories, and got a specialist. Alibaba went the other way and scaled. Both answers are published, and the gap between what the download costs and what the API costs is the honest way to choose between them.

A screenshot of the microsoft/FrogNano GitHub repository README showing the description that FrogNano evaluates coding agents with the Leaf harness in isolated Kubernetes sandboxes against OpenAI-compatible endpoints and five tools Read, Write, Edit, Glob and Bash; the requirements list naming a Kubernetes cluster with pod and network-policy permissions and an OpenAI-compatible endpoint with reasoning and tool-call parsers configured for Qwen3.5; the frognano-eval run commands; and the benchmark table listing SWE-bench Verified at 500 tasks with an 8,192-token cap, SWE-bench Pro at 731 with 32,000, Terminal-Bench 2.0 Verified at 89 with 32,000 and PatchEval Verified at 230 with 8,192.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily