
FrogNano-4B-2609 vs Gemma 4 12B: 61.5% on SWE-bench and Nobody Has Checked It
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 41 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 105 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Microsoft's FrogNano-4B-2609 is a four-billion-class coding agent derived from Qwen3.5-4B that reports 61.5% on SWE-bench Verified. Gemma 4 12B is DeepMind's 11.95-billion-parameter encoder-free multimodal model, and its card reports 72.0% on LiveCodeBench v6. Those are the two numbers a buyer will put side by side, and putting them side by side is the mistake. They come from different benchmarks, different harnesses, different labs, and — more importantly — only one of them has a published evaluation methodology that a stranger has read. Everything else about this matchup is a question of what each model actually is, and the answer is that they are not the same category of thing at all.
The FrogNano figures below come from Microsoft's model card and its technical report, both public since September and neither reproduced by anyone outside the lab. The Gemma 4 12B figures come from Google's model card, which is a vendor document too — the difference is that Gemma's numbers have had twenty weeks and roughly 2.6 million downloads of scrutiny behind them, and FrogNano has had two weeks and a download counter that has not cleared single digits.
What each model is for
This is the part that a spec sheet hides. FrogNano-4B-2609 is not a general model that happens to be good at code. It is a checkpoint whose entire post-training was repository-level software engineering, and which is designed to be driven by a harness rather than talked to.
Its five tools are Read, Write, Edit, Glob and Bash. It emits calls to them, a sandbox executes them and returns output, and the loop runs until the model stops asking. The evaluated configuration is 150 interaction steps inside roughly 131K combined tokens, and it produced a candidate patch per task. Microsoft's card is explicit that the model is for English-language, Python-heavy repositories with reproducible environments and executable test suites, and that image and video components inherited from the base model were never post-trained and are not supported.
Gemma 4 12B is the opposite shape. It is a 48-layer unified model with a 256K context and no separate vision or audio encoder — raw image patches and audio waveforms project straight into the single decoder's embedding space through lightweight linear layers. It takes text, image and audio in and writes text out. It has a thinking mode triggered by a token in the system prompt, native function calling, and Google's card makes a point of listing agentic workflows among its intended uses.
So the real question is not which one is smarter. It is whether you want a specialist component that only functions inside a rig you have to operate, or a general model that does a decent job at many things and can be called by anything.

Line by line, where they actually differ
• Parameters — FrogNano-4B-2609 publishes a band, "500M-5B", on a card whose own description reads approximately 4.66 billion; the download settles it at 9.32 GB of BF16 weights. Gemma 4 12B is 11.95B, stated once and never hedged. The ratio is roughly 2.6 to one, and it shows up directly in what you have to rent.
• Context — FrogNano's evaluated configuration is about 131K combined tokens, with reasoning and tool output sharing the budget. Gemma 4 12B carries 256,000 tokens, and on its own long-context row scores 43.4% on MRCR v2 with eight needles at 128K. Nearly double the window, and the second number is a published measurement rather than a capability claim.
• Modality — FrogNano-4B-2609 is text in, text out, notwithstanding the vision tower sitting in its checkpoint. Gemma 4 12B is text, image and audio in.
• Output ceiling — the validated FrogNano configuration allows 8,192 generated tokens per assistant turn, and its card notes the RL training configuration used that same cap. Gemma 4 12B publishes no equivalent single-turn ceiling; the constraint there is the 256K window.
• Agentic interface — FrogNano emits structured Leaf calls and needs a harness to execute them. Gemma 4 12B ships native function calling in its instruction-tuned form and is callable directly.
• License — Gemma 4 12B is Apache 2.0 under Google's Gemma 4 terms, cleanly stated. FrogNano's card is MIT in its front matter and Apache 2.0 in its own body, which is the kind of ambiguity that ends up in a legal review rather than a footnote.
• Numbers — FrogNano reports 61.5% SWE-bench Verified, 37.6% SWE-bench Pro, 31.1% Terminal-Bench 2.0 and 47.3% PatchEval-Verified. Gemma 4 12B reports 77.2% MMLU Pro, 72.0% LiveCodeBench v6, a Codeforces ELO of 1659, 78.8% GPQA Diamond and 69.0% on Tau2. Google's card does not publish a SWE-bench row, and Microsoft's card does not publish LiveCodeBench. There is no overlap between the two scoreboards at all.
That last bullet is the one that should end the comparison-by-number instinct. Not one benchmark appears on both cards. A reader who lines 61.5 up against 72.0 has compared a repository-resolution rate against a competitive-programming pass rate, which is roughly like comparing a marathon time to a hundred-metre time because both are in seconds.
Why Google's silence on SWE-bench is not a red flag
The tempting conclusion from the bullet list is that FrogNano has a real SWE-bench number and Gemma does not, so FrogNano wins the coding question. That does not follow, for a reason worth being precise about.
Gemma 4 12B reports Tau2 at 69.0% — an agentic tool-use benchmark over three runs — alongside its function-calling support and its explicit positioning for agentic workflows. A model that scores 69.0 on Tau2 is not a model that cannot do agentic work; it is a model whose vendor chose to report tool-use performance rather than repository resolution. Google also does not report SWE-bench rows for the 26B or the 31B, which is a family-wide editorial choice rather than a 12B weakness.
Meanwhile the counter-intuitive result in Microsoft's own paper is one a Gemma buyer should read carefully. FrogNano starts from Qwen3.5-4B, a general model, which scored 39.4% on SWE-bench Verified through the Leaf harness. Five rounds of reinforcement learning on roughly 1,500 synthetic tasks moved it to 61.5%. The lesson is not "small models cannot code" — it is that a general 4B checkpoint at 39.4% was already solving more than a third of a hard, human-validated issue set when somebody ran it inside a competent agent loop. Gemma 4 12B is three times that size and twenty weeks more mature. Nobody has run it through Leaf, and until somebody does, "Gemma is not a repo agent" is an assumption rather than a finding.
The harness tax, which the scoreboards hide completely
Here is the practical asymmetry, and it is the one that decides the purchase.
Gemma 4 12B is a model. Download 11.95B of weights, point Transformers, vLLM or SGLang at them, send text, get text. Google publishes the serving path, the license is unambiguous, and 2.6 million downloads' worth of people have already hit the rough edges and documented them. If the task is "summarize this stack trace", "read this screenshot and tell me what is broken", or "call this function with these arguments", it works today.
FrogNano-4B-2609 is a model plus a harness, and Microsoft's own README makes the harness non-optional. The repository at github.com/microsoft/FrogNano is the Leaf evaluation rig, not the training code — it needs a Kubernetes cluster, an existing namespace, permissions to manage pods and network policies, pull access to benchmark container images, and an OpenAI-compatible endpoint configured with the right reasoning and tool-call parsers. Microsoft's card states the matching condition bluntly: identical scores require matching checkpoint, tokenizer, serving configuration, task images and evaluation protocol. The half of the release that generates the score is the half the card points at a GitHub link.
There is real value in that, and it is worth naming rather than dismissing. FrogNano's contribution is a task-synthesis loop that regenerates training problems against the current policy's ability — the paper's claim is that a small agent can be trained to competitive level on synthetic tasks with no distillation at all, and that opens a path for anyone who cannot afford a frontier teacher. Its API is public. Its methods are reproducible in principle. That is a stronger contribution than another incremental checkpoint, and it is also more work than most teams signing up for.

If you are choosing one, choose by the job
Take Gemma 4 12B if the work mixes modalities, if anything needs to see an image or hear audio, if the context has to hold a large file tree and a long transcript at the same time, or if you want a model that any framework can load without a second system behind it. Its 69.0 Tau2 score and native function calling make it a defensible choice for agentic pipelines, and nobody has to take a lab's word for it — the model has been evaluated by thousands of people and the results are not a secret.
Take FrogNano-4B-2609 if you already run a sandboxed repository agent and you want a 4B-class checkpoint to drop into it, and if a 9.32 GB download that fits on modest hardware is the constraint driving the decision. Be clear-eyed about the sequencing: you are not buying a score, you are taking a bet on a method, and the first thing you will discover is whether your polling loop and your tool-call parser look enough like Leaf. Some teams will get 61.5%-adjacent behaviour on the third afternoon and some will spend two weeks discovering their evaluation set was easier than SWE-bench Verified, and nobody outside the lab can yet tell you which you are.
If the decision is genuinely close, the cheap experiment is to run both inside the same harness on your own issues rather than debating published numbers that share no benchmark. OrcaRouter carries 200-plus models behind one OpenAI-compatible key at 0% markup, so provider list prices pass through untouched — which makes it possible to put a hosted general model and the rest of your candidate set behind a single endpoint and one billing line while you decide. FrogNano is not one of our routes and there is no date for it; the weights are a download you host yourself. What we can remove is the friction on the other side of the comparison: swapping candidate models in your own harness without a second contract, a second SDK, or a second set of credentials.

The honest summary is that the 61.5 number is real work by the lab that made it, and that it is currently the only reason to prefer FrogNano on coding. When somebody outside Microsoft reproduces 61.5 on the same 500 tasks — or fails to — this matchup gets a genuine answer. Until then the safer default for most teams is the 11.95B model with the clean license, the published long-context measurement, and twenty weeks of other people's mistakes already absorbed into its documentation.
