
FrogNano-4B-2609: Microsoft Shipped a 4B Coding Agent to Hugging Face and Never Announced It
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 41 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 105 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The repository went up on 17 September 2026 with the initial commit and the checkpoint itself landed four minutes later. Then nothing. No tweet from Microsoft AI, no entry on the Microsoft Research blog, no launch page. Seven weeks later microsoft/FrogNano-4B-2609 — a four-billion-class coding agent built on Qwen3.5-4B — still has no announcement behind it, and that silence is the single most important fact about this model's publication. What it does have is a model card that claims a paper and a harness, a GitHub repository that turns out to be the harness and not the paper, and a createdAt of 17 September sitting under a card that says "Release date 22-SEP-2026". Both dates are improvised; neither is wrong; they disagree with each other.
What is actually published
Everything below is checkable on the hub right now, and every figure in this piece comes from Microsoft's own card or from the byte counts in the repository itself. Nothing has been reproduced by a third party — there is no Artificial Analysis entry, no arena rating, and no independent evaluation of FrogNano anywhere.
• Weights — one public repository under the Microsoft organisation, ungated, MIT-tagged on the hub, 15 files, no download gate and no agreement to sign.
• Weights on disk — 9.32 GB across two safetensors shards, 59.6 billion bytes of repository data including the optimizer-free file set, 738 tensors in the index.
• Architecture — Qwen3_5ForConditionalGeneration, the dense 32-layer hybrid stack inherited from Qwen3.5-4B, with a 24-layer vision tower that the card says was inherited and never post-trained.
• The card's own parameter field — "500M-5B". That is a band, not a number, and the download is the more precise document.
• License — the card's front matter says MIT. The card's own body says Apache License 2.0, and the GitHub harness really does ship an MIT LICENSE file. Those two statements are about different artefacts in the same release.
• Announcement — none. Not on the Microsoft Research blog, not on the team's own research site, not in the Hugging Face organisation feed as anything other than a repository listing.
That last line is the one that should set a reader's posture for the rest of this piece. A quietly uploaded checkpoint is not a lesser artefact than an announced one — its card is often longer, because nobody is editing it down for a press release. But it has not been through the one filter an announced release gets for free: other people looking at it.

The scores, and who produced every one of them
Microsoft reports FrogNano reaching 61.5% on SWE-bench Verified, 37.6% on SWE-bench Pro, 31.1% on Terminal-Bench 2.0, and 47.3% on PatchEval-Verified, all as Avg@3 resolution rates measured through the Leaf harness. The same card gives the starting point: Qwen3.5-4B scored 39.4% on SWE-bench Verified under the identical harness and budget. Five reinforcement-learning iterations took it through 49.1%, 53.1%, 56.7%, 59.1% and 61.5% — 22.1 points over the base model, which Microsoft's own framing rounds to about 56% relative improvement.
Two things about that ladder are worth pausing on, because they are places where a summary can go wrong rather than places where the vendor did.
The first is the Iter 5 discrepancy. The card's headline table says 61.5% for the final iteration; the paper's own efficiency appendix reports the same five-checkpoint progression as 48.2%, 53.4%, 58.3%, 58.6% and 61.6%. Only the last number is close, and only the last number is the one anybody quotes. Treat the final figure as the stable result and the intermediate rungs as measurement of different things, because under a different aggregation they plainly are.
The second is that FrogNano is not uniformly better than the model it started from on every axis. Its published parallel tool-call rate is 1.71%. The card is candid that later iterations lost the ability to fire several tool calls in a single turn, and that the team's consolidation work exists partly to recover it. A model that solves more issues while issuing almost no concurrent calls is a real engineering trade, not a footnote.

The harness is the product, and the repository is the harness
FrogNano does not run on its own. It emits structured calls to five tools — Read, Write, Edit, Glob and Bash — and something has to execute them in an isolated environment and hand back the output. That something is Leaf, and here the trail does something mildly unusual.
The model card's "additional related assets" row links a technical report at aka.ms/frognano-tech-report and a harness at github.com/microsoft/FrogNano. The aka.ms link does not point at a PDF; it redirects straight to the arXiv abstract page, which is where the actual report lives. The GitHub repository, meanwhile, contains no training code, no RL recipe and no checkpoints. Its own README describes an evaluation harness that runs coding agents in Kubernetes sandboxes against an OpenAI-compatible endpoint, and points back at the arXiv paper for the training story. So the card's two asset links are a pointer to the paper and a pointer to the reply of the paper, and the name is a shared one.
This matters for anyone planning to use the model, for a practical reason. The documented serving recipe is SGLang with --reasoning-parser qwen3 and --tool-call-parser qwen3_coder, and the harness wants an endpoint that already speaks tool calls and reasoning. Get the parsing configuration slightly wrong and the model produces text where the harness expects JSON, which read from the outside looks exactly like a bad model rather than a bad configuration. The card is explicit that any matching score requires a matching checkpoint, tokenizer, serving configuration, task images and evaluation protocol — that is the vendor telling you the harness is half the result.
Also worth stating plainly for anyone sizing hardware: a 9.32 GB checkpoint at the card's evaluated context is not a 9.32 GB inference problem. The evaluations ran at roughly 131K combined tokens with a 150-step budget. Memory scales with the context, not with the weights, and the card says the minimum GPU configuration still "must be validated before release" — a phrase that appears on a model that is already downloadable.
What the training actually did, in one paragraph
The paper's contribution is not the model, it is the loop. TaskPilot generates candidate software-engineering tasks from real repository snapshots, runs rollouts from the current checkpoint against them, keeps the ones near the edge of what the policy can sometimes solve, and discards candidates that are permanently trivial or permanently impossible. The accepted set trains the next checkpoint. That next checkpoint then calibrates the following round of task generation, so the task distribution moves as the policy moves. Roughly 1,500 validated environments in total. No distillation: the card states outright that agent-specific post-training used no stronger-model solution trajectories, actions, reasoning traces or patch targets. Stronger models write tasks; they do not demonstrate answers.
Microsoft ran that loop for five iterations and reports an 8.7-point gain before any consolidation, with a log-length penalty added mid-run because reasoning traces were growing faster than they were improving.
The model card is unusually direct about where the whole approach is fragile. Training data are Python-heavy and primarily English; the card's stated supported natural-language set is English and nothing else, with the base model's broader multilingual coverage explicitly not claimed. Performance is sensitive to the harness and to the quality of the tests. Generated patches "may be incorrect or insecure despite passing available tests." The model card for the 4B closes that paragraph with the sentence every reader should carry: not to be used without qualified human review and independent regression and security testing. That is a vendor statement about a vendor artefact, and it is a more useful sentence than any of the scoreboard rows above it.
Where this leaves a buyer, and where a router fits
FrogNano is not a model you call. There is no first-party API, no hosted endpoint, and no serverless image. It is a checkpoint, and using it means either standing up your own SGLang deployment behind a Kubernetes-hosted sandbox, or evaluating it as a component inside an agent stack you already operate. That is the whole adoption path today, and no announcement would have changed its shape.
What a routing layer can honestly do here is a narrower thing than it first sounds. If the plan is to compare a self-hosted FrogNano against a hosted coding model inside your own harness, the hosted half is the half that benefits from being behind one key instead of a second contract: OrcaRouter carries 200-plus models behind a single OpenAI-compatible endpoint at 0% markup, meaning provider list prices pass through untouched and a vendor price change is live on our side the same day. That matters for a bake-off whose outcome is decided by cost per resolved issue rather than by a single benchmark number. We do not host FrogNano and there is no date for it; the weights and the serving work are yours.

The three things that would settle it
First, an independent reproduction. Every number in this article — 61.5, 37.6, 31.1, 47.3 — was produced by the lab that trained the model, on a harness the same lab maintains, against a base-model figure the same lab measured. That is a complete and internally consistent picture, and it is also a closed loop. The useful test is whether someone with no stake reproduces the 39.4-to-61.5 jump on the same 500 tasks without Microsoft's calibration work on the task set.
Second, someone has to check the harness claim end to end. The repository is public, which is more than many releases manage, but it is the evaluation rig. If the five tools and the sandbox isolation are genuinely the performance mechanism, a third party running the published configuration should land near the published numbers. If they do not, the gap is the real finding.
Third, and cheapest to answer: Microsoft should say whether this is a product or a paper artefact. The card's distribution section treats the weights, the card and the harness as the deliverable, which reads like publication rather than launch. "Release date 22-SEP-2026" reads like launch. An announced model would have resolved that ambiguity in the first sentence of a blog post, and there is no blog post.
Until then the correct posture is the one the release itself implies. FrogNano-4B-2609 is a real, downloadable, MIT-and-Apache-and-maybe-not checkpoint with an unusually thorough card, a public evaluation harness, a genuine methodological idea, and a scoreboard that has never been touched by anyone without a Microsoft address. For a lab, that is a complete and interesting artefact. For a team about to put an agent in front of a repository at three in the morning, it is a lead worth following and not yet a decision worth making.
