A generated title card reading 'FrogNano-4B-2609 vs Gemma 4 12B' with the subtitle 'a 4B repository agent against an 11.95B multimodal generalist', three chips reading '61.5% SWE-bench, vendor-reported', '72.0% LiveCodeBench, vendor-reported' and 'no shared benchmark', a footer reading 'Different benchmarks with no overlap; both scoreboards vendor-reported', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

FrogNano-4B-2609 vs Gemma 4 12B: 61.5% on SWE-bench and Nobody Has Checked It

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Microsoft's FrogNano-4B-2609 is a four-billion-class coding agent derived from Qwen3.5-4B that reports 61.5% on SWE-bench Verified. Gemma 4 12B is DeepMind's 11.95-billion-parameter encoder-free multimodal model, and its card reports 72.0% on LiveCodeBench v6. Those are the two numbers a buyer will put side by side, and putting them side by side is the mistake. They come from different benchmarks, different harnesses, different labs, and — more importantly — only one of them has a published evaluation methodology that a stranger has read. Everything else about this matchup is a question of what each model actually is, and the answer is that they are not the same category of thing at all.

The FrogNano figures below come from Microsoft's model card and its technical report, both public since September and neither reproduced by anyone outside the lab. The Gemma 4 12B figures come from Google's model card, which is a vendor document too — the difference is that Gemma's numbers have had twenty weeks and roughly 2.6 million downloads of scrutiny behind them, and FrogNano has had two weeks and a download counter that has not cleared single digits.

What each model is for

This is the part that a spec sheet hides. FrogNano-4B-2609 is not a general model that happens to be good at code. It is a checkpoint whose entire post-training was repository-level software engineering, and which is designed to be driven by a harness rather than talked to.

Its five tools are Read, Write, Edit, Glob and Bash. It emits calls to them, a sandbox executes them and returns output, and the loop runs until the model stops asking. The evaluated configuration is 150 interaction steps inside roughly 131K combined tokens, and it produced a candidate patch per task. Microsoft's card is explicit that the model is for English-language, Python-heavy repositories with reproducible environments and executable test suites, and that image and video components inherited from the base model were never post-trained and are not supported.

Gemma 4 12B is the opposite shape. It is a 48-layer unified model with a 256K context and no separate vision or audio encoder — raw image patches and audio waveforms project straight into the single decoder's embedding space through lightweight linear layers. It takes text, image and audio in and writes text out. It has a thinking mode triggered by a token in the system prompt, native function calling, and Google's card makes a point of listing agentic workflows among its intended uses.

So the real question is not which one is smarter. It is whether you want a specialist component that only functions inside a rig you have to operate, or a general model that does a decent job at many things and can be called by anything.

A screenshot of the google/gemma-4-12B-it model card on Hugging Face showing the Gemma 4 12B Unified heading, the Any-to-Any, Transformers, Safetensors and image-text-to-text tags, the Apache 2.0 license line credited to Google DeepMind, and the model overview explaining that text, audio, image and video inputs are handled without separate encoders.

Line by line, where they actually differ

• Parameters — FrogNano-4B-2609 publishes a band, "500M-5B", on a card whose own description reads approximately 4.66 billion; the download settles it at 9.32 GB of BF16 weights. Gemma 4 12B is 11.95B, stated once and never hedged. The ratio is roughly 2.6 to one, and it shows up directly in what you have to rent.

• Context — FrogNano's evaluated configuration is about 131K combined tokens, with reasoning and tool output sharing the budget. Gemma 4 12B carries 256,000 tokens, and on its own long-context row scores 43.4% on MRCR v2 with eight needles at 128K. Nearly double the window, and the second number is a published measurement rather than a capability claim.

• Modality — FrogNano-4B-2609 is text in, text out, notwithstanding the vision tower sitting in its checkpoint. Gemma 4 12B is text, image and audio in.

• Output ceiling — the validated FrogNano configuration allows 8,192 generated tokens per assistant turn, and its card notes the RL training configuration used that same cap. Gemma 4 12B publishes no equivalent single-turn ceiling; the constraint there is the 256K window.

• Agentic interface — FrogNano emits structured Leaf calls and needs a harness to execute them. Gemma 4 12B ships native function calling in its instruction-tuned form and is callable directly.

• License — Gemma 4 12B is Apache 2.0 under Google's Gemma 4 terms, cleanly stated. FrogNano's card is MIT in its front matter and Apache 2.0 in its own body, which is the kind of ambiguity that ends up in a legal review rather than a footnote.

• Numbers — FrogNano reports 61.5% SWE-bench Verified, 37.6% SWE-bench Pro, 31.1% Terminal-Bench 2.0 and 47.3% PatchEval-Verified. Gemma 4 12B reports 77.2% MMLU Pro, 72.0% LiveCodeBench v6, a Codeforces ELO of 1659, 78.8% GPQA Diamond and 69.0% on Tau2. Google's card does not publish a SWE-bench row, and Microsoft's card does not publish LiveCodeBench. There is no overlap between the two scoreboards at all.

That last bullet is the one that should end the comparison-by-number instinct. Not one benchmark appears on both cards. A reader who lines 61.5 up against 72.0 has compared a repository-resolution rate against a competitive-programming pass rate, which is roughly like comparing a marathon time to a hundred-metre time because both are in seconds.

Why Google's silence on SWE-bench is not a red flag

The tempting conclusion from the bullet list is that FrogNano has a real SWE-bench number and Gemma does not, so FrogNano wins the coding question. That does not follow, for a reason worth being precise about.

Gemma 4 12B reports Tau2 at 69.0% — an agentic tool-use benchmark over three runs — alongside its function-calling support and its explicit positioning for agentic workflows. A model that scores 69.0 on Tau2 is not a model that cannot do agentic work; it is a model whose vendor chose to report tool-use performance rather than repository resolution. Google also does not report SWE-bench rows for the 26B or the 31B, which is a family-wide editorial choice rather than a 12B weakness.

Meanwhile the counter-intuitive result in Microsoft's own paper is one a Gemma buyer should read carefully. FrogNano starts from Qwen3.5-4B, a general model, which scored 39.4% on SWE-bench Verified through the Leaf harness. Five rounds of reinforcement learning on roughly 1,500 synthetic tasks moved it to 61.5%. The lesson is not "small models cannot code" — it is that a general 4B checkpoint at 39.4% was already solving more than a third of a hard, human-validated issue set when somebody ran it inside a competent agent loop. Gemma 4 12B is three times that size and twenty weeks more mature. Nobody has run it through Leaf, and until somebody does, "Gemma is not a repo agent" is an assumption rather than a finding.

The harness tax, which the scoreboards hide completely

Here is the practical asymmetry, and it is the one that decides the purchase.

Gemma 4 12B is a model. Download 11.95B of weights, point Transformers, vLLM or SGLang at them, send text, get text. Google publishes the serving path, the license is unambiguous, and 2.6 million downloads' worth of people have already hit the rough edges and documented them. If the task is "summarize this stack trace", "read this screenshot and tell me what is broken", or "call this function with these arguments", it works today.

FrogNano-4B-2609 is a model plus a harness, and Microsoft's own README makes the harness non-optional. The repository at github.com/microsoft/FrogNano is the Leaf evaluation rig, not the training code — it needs a Kubernetes cluster, an existing namespace, permissions to manage pods and network policies, pull access to benchmark container images, and an OpenAI-compatible endpoint configured with the right reasoning and tool-call parsers. Microsoft's card states the matching condition bluntly: identical scores require matching checkpoint, tokenizer, serving configuration, task images and evaluation protocol. The half of the release that generates the score is the half the card points at a GitHub link.

There is real value in that, and it is worth naming rather than dismissing. FrogNano's contribution is a task-synthesis loop that regenerates training problems against the current policy's ability — the paper's claim is that a small agent can be trained to competitive level on synthetic tasks with no distillation at all, and that opens a path for anyone who cannot afford a frontier teacher. Its API is public. Its methods are reproducible in principle. That is a stronger contribution than another incremental checkpoint, and it is also more work than most teams signing up for.

A generated two-column scoreboard titled 'FrogNano-4B-2609 vs Gemma 4 12B'. The left column reads Parameters approx 4.66B with the card saying 500M-5B, Context about 131K evaluated, Input text only, Harness required Leaf with 5 tools, SWE-bench Verified 61.5% vendor, and Independent evals none. The right column reads Parameters 11.95B, Context 256,000 tokens, Input text, image, audio, Harness none required, LiveCodeBench v6 72.0% vendor, and Independent evals widely run. A footer reads 'Different benchmarks, no overlap; both scoreboards vendor-reported.'

If you are choosing one, choose by the job

Take Gemma 4 12B if the work mixes modalities, if anything needs to see an image or hear audio, if the context has to hold a large file tree and a long transcript at the same time, or if you want a model that any framework can load without a second system behind it. Its 69.0 Tau2 score and native function calling make it a defensible choice for agentic pipelines, and nobody has to take a lab's word for it — the model has been evaluated by thousands of people and the results are not a secret.

Take FrogNano-4B-2609 if you already run a sandboxed repository agent and you want a 4B-class checkpoint to drop into it, and if a 9.32 GB download that fits on modest hardware is the constraint driving the decision. Be clear-eyed about the sequencing: you are not buying a score, you are taking a bet on a method, and the first thing you will discover is whether your polling loop and your tool-call parser look enough like Leaf. Some teams will get 61.5%-adjacent behaviour on the third afternoon and some will spend two weeks discovering their evaluation set was easier than SWE-bench Verified, and nobody outside the lab can yet tell you which you are.

If the decision is genuinely close, the cheap experiment is to run both inside the same harness on your own issues rather than debating published numbers that share no benchmark. OrcaRouter carries 200-plus models behind one OpenAI-compatible key at 0% markup, so provider list prices pass through untouched — which makes it possible to put a hosted general model and the rest of your candidate set behind a single endpoint and one billing line while you decide. FrogNano is not one of our routes and there is no date for it; the weights are a download you host yourself. What we can remove is the friction on the other side of the comparison: swapping candidate models in your own harness without a second contract, a second SDK, or a second set of credentials.

A generated timeline card headed 'Two releases, twenty weeks apart' showing an upper bar labelled Gemma 4 12B uploaded 23 May 2026 with 2.6 million downloads across the family, and a lower bar labelled FrogNano-4B-2609 uploaded 17 September 2026 with no announcement, with a span marker of 20 weeks between them and a footer reading 'Download figures from Hugging Face on 3 October 2026.'

The honest summary is that the 61.5 number is real work by the lab that made it, and that it is currently the only reason to prefer FrogNano on coding. When somebody outside Microsoft reproduces 61.5 on the same 500 tasks — or fails to — this matchup gets a genuine answer. Until then the safer default for most teams is the 11.95B model with the clean license, the published long-context measurement, and twenty weeks of other people's mistakes already absorbed into its documentation.