A generated title card reading 'FrogNano-4B-2609 vs Intern-Decision-4B' with the subtitle 'the same Qwen3.5-4B ancestor, two opposite outputs', three chips reading 'tool calls and patches', 'one symbol per field' and 'neither independently scored', a footer reading 'Both scoreboards are vendor-reported; no third party has reproduced either', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

FrogNano-4B-2609 vs Intern Decision 4B: One Base Model, Two Entirely Different Bets

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two labs took the same checkpoint, Qwen3.5-4B, and pushed it in directions that do not resemble each other. microsoft/FrogNano-4B-2609 was post-trained with reinforcement learning on roughly 1,500 synthetic software-engineering environments until it could emit structured tool calls and multi-file patches through a five-tool harness. internlm/Intern-Decision-4B was fine-tuned to emit no text at all — it takes a state plus a schema of named questions and returns a calibrated probability for each option in a single forward pass. One writes code. The other refuses to write anything and returns a distribution. Comparing them on quality is meaningless; comparing them on what they cost you to run and what they can be trusted with is the honest exercise.

Both shipped without announcement, which is the other thing they share. Neither has an Artificial Analysis entry, an arena rating, or a single independent evaluation. Every figure below — the SWE-bench ladder on one side, the Brier and ECE calibration rows on the other — was produced by the lab that trained the model, on that lab's own harness, and has never met a test set it did not come from.

The fork happened before either model existed

The base checkpoint is a dense 32-layer hybrid Gated DeltaNet and gated-attention model, and both derivatives inherit its skeleton. After that the two post-training pipelines have nothing in common.

Microsoft's pipeline is a closed loop. TaskPilot generates candidate repository tasks from real snapshots, runs rollouts from the current checkpoint to find ones the policy sometimes solves, keeps those, and trains. Five iterations, each calibrated against the previous iteration's policy, ending at 61.5% on SWE-bench Verified and 37.6% on SWE-bench Pro. No distillation — the card states that no stronger-model trajectories, actions or reasoning traces were used as targets.

InternLM's pipeline is a single supervised objective over a fixed task shape. The model is given a system prompt, a state, a decision schema, and a complete assistant JSON skeleton with one placeholder per field. It runs one causal forward pass, reads the logits at each placeholder position, and takes a softmax over only that field's permitted candidate symbols. InternLM's card says it outright: this path "does not call generate() or sample free-form text."

That difference is not a scale difference. It is a difference in what the artefact even is. FrogNano-4B-2609 is an agent that can be wrong in a hundred interesting ways and be corrected by tests. Intern-Decision-4B is a scorer whose failure mode is a confident number.

Two contracts, and neither one is "send a prompt"

FrogNano's contract is a loop. The model emits calls to Read, Write, Edit, Glob and Bash; the Leaf harness executes them inside an isolated repository sandbox and returns the output; the loop continues until the model stops calling. Evaluations ran at 150 steps inside roughly 131K combined tokens, with up to 8,192 generated tokens per assistant turn. Serving needs SGLang with a Qwen3 reasoning parser and a Qwen3 coder tool-call parser — get that wrong and the model produces prose where the harness expects JSON, which from the outside looks exactly like a broken model.

Intern-Decision-4B's contract is one shot with a hard edge. Questions and options keep their order. Options are mapped onto single-token symbols — A through Z, then a through z, then 0 through 9, which is the arithmetic reason a question maxes out at 62 options. The wrapper's ceiling is 8,192 tokens and InternLM's card is explicit that longer inputs are rejected without truncation. Three question types are supported: choice with an ordered criteria map, score with a list or numeric-keyed map, and noul, a binary. Up to eight images are allowed and their tokens count against the same 8,192.

• Output shape — FrogNano-4B-2609 emits tool calls, reasoning text and a patch. Intern-Decision-4B emits one symbol per field and nothing else.

• Determinism — FrogNano samples at temperature 0.6 across three seeds, so it is probabilistic by design. Intern-Decision-4B's argmax is fixed by the forward pass; the fitted temperature moves only its confidence.

• Failure surface — FrogNano can hallucinate an API, over-broaden an edit, or pass tests while introducing a vulnerability, all of which its card names. Intern-Decision-4B cannot hallucinate an answer, because it can only pick from options you supplied — it can only be miscalibrated.

• Input ceiling — roughly 131K combined tokens for FrogNano in its evaluated configuration, against a hard 8,192 for Intern-Decision-4B, which is a factor of sixteen and there is no workaround.

• Cost shape — FrogNano bills by the trajectory, and trajectories are long. Intern-Decision-4B has no output tokens to bill at all.

• Parameters — the FrogNano card gives a "500M-5B" band and describes approximately 4.66B, with a 9.32 GB BF16 download; Intern-Decision-4B is stated at 4.54B with a 612 MB vision tower and a 54 MB projector alongside its language shards.

The number that decides this pairing

InternLM's card puts Intern-Decision-4B at 44.16 ms mean latency and 44.03 ms median on a single RTX 4090 over the local Hugging Face path, with a P95 of 44.60 ms. Read that spread again: median to tail is well under a millisecond, because a fixed-shape forward pass has almost nothing to vary. That is the whole argument for the architecture.

FrogNano's corresponding figure is not in milliseconds. Its evaluations allowed a 150-step budget with a 10,800-second agent ceiling per task. Those are not comparable units and should never be put in the same sentence as a latency claim — but they do tell you the shape of the trade. One model answers a decision in forty-four milliseconds and the other spends minutes of tool calls chasing a patch. If your problem is "route this ticket into one of six queues and tell me how sure you are", the first one is not a cheaper version of the second, it is a different machine.

Both labs measured themselves carefully, and both sets of measurements should be read as intentions rather than results. InternLM reports a seven-set average of 90.02 with a Brier score of 0.347 and an expected calibration error of 0.065, fitted at a temperature of 1.99241824 on 1,728 designated calibration cases. Microsoft reports a five-iteration climb from 39.4% to 61.5% on SWE-bench Verified, and the same paper's appendix reports that same progression as 48.2%, 53.4%, 58.3%, 58.6% and 61.6% — a reminder that even inside one lab, the same fact measured two ways produces two ladders.

A generated two-column scoreboard titled 'FrogNano-4B-2609 vs Intern-Decision-4B'. The left column reads Output tool calls and patch, Latency minutes per trajectory, Context about 131K evaluated, Independent evals none, Base Qwen3.5-4B, and Score 61.5% SWE-bench Verified vendor. The right column reads Output one symbol per field, Latency 44.16 ms mean, Context 8,192 tokens rejected above, Independent evals none, Base Qwen3.5-4B, and Score 90.02 seven-set average vendor. A footer reads 'Both scoreboards vendor-reported; no third party has reproduced either.'

The tell is in what each card chooses to warn you about

Both cards are unusually honest, and the honesty points in opposite directions — which is the most useful thing in this comparison.

Microsoft's known-limitations section reads like a deployment warning. English and Python only. Performance sensitive to harness and test quality. Patches that may be incorrect or insecure despite passing their tests. The card's own sentence is that FrogNano "should not be considered independently safety-aligned for unrestricted autonomous deployment", and it names the specific gap: the agent-specific post-training used no safety-preference, refusal or adversarial data, because it optimised for functional correctness and regression avoidance instead. It also volunteers the 1.71% parallel tool-call rate, meaning the model almost never fires two tools in one turn despite the harness allowing it — a real capability regression from the base model that the team added a consolidation stage to try to recover.

InternLM's card warns about the opposite class of thing. There is no hallucination surface to warn about, so the caveats are about inputs: the 8,192 rejection, the 62-option ceiling, the fact that the underlying text tower declares a 262,144-token position limit that the released wrapper refuses to use. Its risk is that a confident number is trusted further than it deserves to be, and the card is candid that the calibration narrows a gap with the Jev baseline without closing it on every slice.

Read together they describe two different trust postures. FrogNano needs supervision because it acts. Intern-Decision-4B needs an audit because it scores — and a number that moves a production decision is exactly the kind of output nobody thinks to test.

Where a router fits, and an honest note on both

Neither model is a hosted endpoint. FrogNano ships as weights plus a Kubernetes evaluation harness; Intern-Decision-4B ships as a Python class you import after downloading four shards. For a team that wants to try either in front of something real, the safe pattern is the same for both and it is not glamorous: put the experimental component behind a fallback so that a bad trajectory or a miscalibrated day costs a retry rather than an incident.

That pattern is what a routing layer is for. OrcaRouter runs 200-plus models behind one OpenAI-compatible key at 0% markup — provider list price passed through, so a vendor price change lands on our side the same day — and its failover sits in front of the general model you fall back to, not in front of these two. To be plain: we do not carry FrogNano-4B-2609 or Intern-Decision-4B, and there is no date for either. What a single endpoint buys you here is that the comparison itself gets cheap — one credential, one billing line, and no second integration each time you swap the general model you are benchmarking the specialist against.

A screenshot of the internlm/Intern-Decision-4B model card on Hugging Face showing the model heading and the Qwen3.5-4B base-model line, the five-step explanation of how inference works (single-token symbol mapping, one causal forward pass, a softmax over each field's allowed symbols and the checkpoint's probability calibration), the Benchmark results table listing Intern-Decision-0.8B, Intern-Decision-2B and Intern-Decision-4B against the Jev, Laya, SemIf, Kev and JevK5 comparison rows, and the inference latency table with the 4B row at 44.16 ms mean, 44.03 ms median and 44.60 ms P95.

Which one, and when

Take Intern-Decision-4B when the answer already exists in your prompt and you need it chosen, with a confidence attached, at a rate of thousands per second. Fixed-taxonomy routing, rubric scoring, schema-constrained extraction, moderation against a closed label set. The 44-millisecond forward pass and the zero output-token bill are the product, and the calibration is the thing to audit before you trust it.

Take FrogNano-4B-2609 when the answer does not exist yet and has to be discovered by reading a repository and running its tests. That is a minutes-long, sandboxed, reviewable process, and the 61.5% on SWE-bench Verified — vendor-reported, unreproduced — is the current best evidence that a 4B checkpoint can do it at all.

What should not pass either way is the habit of comparing their scores. A 90.02 seven-set average and a 61.5 SWE-bench Verified rate are measurements of two different questions, produced by two labs, on two harnesses, and neither number has been through a third party's hands. They share an ancestor and nothing else.

A screenshot of the file listing for the microsoft/FrogNano-4B-2609 repository on Hugging Face showing the model header with its size and the Safetensors, qwen3_5 and license:mit tags, the two safetensors shards model-00000-of-00002 and model-00001-of-00002, the config, tokenizer, vocab, merges, chat template and preprocessor files, and the commit column listing 'Update README.md', 'Upload FrogNano 4B SWE checkpoint' and 'Update FrogNano 4B README.md' across two contributors and four commits.

The one genuinely useful prediction from the shared ancestor: because both models start from the same 4B checkpoint, the difference between them is almost entirely in the post-training, which means the interesting question for anyone building on Qwen3.5-4B is which of these two reward structures transfers to their own domain. A synthetic task loop that calibrates against its own policy, or a fixed-shape scoring head with a fitted temperature. Those are two recipes, both published, both from labs that have not yet let anyone else run them. That is a rare situation and worth watching rather than buying into.