A generated title card for the comparison of Decision 3.0 and Intern-Decision-4B, subtitled 'same Qwen3.5-4B base, two different answers', with chips reading '26 Sept vs 10 Oct', 'video vs images only', 'Brier published vs not', and a footer reading 'Decision 3.0 figures are vLLM-SR's own; Intern-Decision-4B figures are InternLM's own; none independently reproduced.' The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Decision 3.0 vs Intern-Decision-4B: Two Teams Fine-Tuned the Same Model and Disagreed About Everything Else

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put d3-mini, the 4B member of Decision 3.0, next to Intern-Decision-4B and the first thing you notice is not a difference. Both are fine-tunes of the same base checkpoint, Qwen3.5-4B. Both are listed at 4.54 billion parameters. Both take a state, a schema of named questions and a set of candidate answers, and return a calibrated probability per candidate without generating a token. Both are Apache-2.0. Both were released without an announcement — InternLM uploaded three checkpoints in forty seconds on 26 September 2026, and vLLM-SR's Decision 3.0 family landed on Hugging Face on 10 October 2026 with the news carried only by the vLLM project's X account.

Everything after that is a disagreement. They disagree about how to read a probability out of the model, about whether video counts as input, about how long a request may be, and — most sharply — about whether the team is willing to publish the number that says its own confidence is trustworthy. This piece is about those four disagreements and what each one costs you, not about which model is "better", because the two are not measured on the same scale and cannot be ranked against each other without doing work neither vendor has done.

The coincidence worth understanding first

Two labs picking the same 4B backbone within a fortnight is not entirely surprising — Qwen3.5-4B is a reasonable base for a structured-output model, and both teams clearly reached for it because it is small enough to run cheaply and strong enough to read instructions. What is surprising is that they arrived at the same parameter count to four significant figures. That tells you the fine-tuning preserved the architecture, and that neither added a separate vision tower big enough to move the total. Both fold their multimodal capability into the same weights.

The interesting part is the readout. Both models answer questions the same way in principle — score candidate answers rather than generate them — and completely differently in mechanism:

• Intern-Decision-4B — maps every option to a single token symbol (A–Z, then a–z, then 0–9), renders a JSON skeleton into the prompt with a placeholder per field, and runs one causal forward pass, reading logits at the position immediately before each placeholder. The mechanism is documented step by step on its model card, including the exact softmax and temperature step.

• d3-mini — ships a custom architecture in modeling_d3.py with a separate readout head in its own readout.safetensors, a decision_config.json declaring noncausal_full_attention and last-token pooling, and a token-to-code mapping defined in the config rather than described in prose.

Neither approach is obviously better. The InternLM route has the advantage that it runs on a stock Hugging Face model class with a documented numerical sequence — you can check its work. The vLLM-SR route has the advantage that the readout is a trained head rather than a projection of an existing token embedding, which is a freer fit, and the cost is that you must trust_remote_code=True and run their code to do anything at all.

The limits each one publishes

This is where a real preference starts to form, because one card is far more specific than the other about where it stops working.

• Input ceiling — Intern-Decision-4B: 8,192 tokens by default and requests over it are rejected outright, never truncated, with the ceiling set by a constructor argument. d3-mini: max_length is null in the shipped config and no token budget appears anywhere on the card.

• Questions per request — Intern-Decision-4B: 1 to 16, with a stated maximum of 62 options in any one question. d3-mini: no stated limit; the card says only that questions are answered together, each from its own forward pass.

• Images — Intern-Decision-4B: up to eight per request, ordered by a list you supply, with the checkpoint processor handling resize and token expansion. d3-mini: multiple per request as paths, URLs, PIL images or base64 data URLs, each read at up to 1.6 megapixels, every question seeing every image.

• Video — Intern-Decision-4B: none. d3-mini: multiple videos, read at 2 frames per second, capped at 32 frames spread across the clip and 0.2 megapixels per frame.

The input ceiling is the line that will decide this for most people. An 8,192-token budget shared between state, question instructions and candidate descriptions is a genuine constraint on the document-grading and long-context routing work these models are sold for, and InternLM deserves credit for saying so plainly rather than leaving it to be discovered. vLLM-SR leaving it unstated is the opposite: not a hidden limit, but an unknown one, and no amount of reading the repository settles it.

Calibration is the real split

Every decision model makes the same promise: the number it returns is a probability, and thresholds set against it mean something. Almost none of them prove it. This is where the two releases diverge most, and the divergence runs in the opposite direction from the one you would guess from release dates.

Intern-Decision-4B publishes, on its own card, a Brier score of 0.347 and an expected calibration error of 0.065 across its seven-benchmark average, a fitted temperature of 1.99241824 derived by NLL minimisation on 1,728 designated calibration cases with 1,693 separate validation cases, an explicit statement that test-suite labels were not used to select that temperature, and a 96-case diagnostic showing its calibration moving from 0.628 Brier / 0.213 ECE before temperature scaling to 0.550 / 0.089 after. It also states the default and says the calibration is per-checkpoint, so using another size with this module will not match.

Decision 3.0 publishes an accuracy index and a coverage claim — every one of 140,178 public requests answered, none unsupported — and no calibration figure at all. No Brier score, no ECE, no stated temperature, on any of the six checkpoints. temperature in d3's shipped decision_config.json is 1.0, which is the identity and may or may not be the fitted value; the file does not say.

Read the two index headlines side by side and the asymmetry gets worse. d3-mini's card reports a Jev Decision Index 0.3 public-suite score of 54.90, described as measured with the official kit on the released weights, while the comparison rows on the same board are described as live board data. Intern-Decision-4B reports a 90.02 average across its own seven benchmarks. Those two numbers are not on the same scale, they do not use the same tasks, and putting them in one sentence as a comparison would be dishonest. What is comparable is the disclosure: one card tells you how wrong its confidences are, and the other does not know or will not say.

A generated two-column scoreboard headed 'Decision 3.0 d3-mini vs Intern-Decision-4B - the scoreboard'. The left column for d3-mini reads: base model Qwen3.5-4B fine-tuned, 4.54B parameters; input ceiling not stated on the card; video input yes, up to 32 frames at 2 fps; calibration figures none published; reported score Jev Decision Index 0.3 public suite 54.90; latency median 17.5 ms text and 96.2 ms image. The right column for Intern-Decision-4B reads: base model Qwen3.5-4B fine-tuned, 4.54B parameters; input ceiling 8,192 tokens, rejected not truncated; video input none, images only up to eight; calibration Brier 0.347, ECE 0.065 and temperature 1.99241824; reported score 90.02 seven-benchmark average; latency mean 44.16 ms and median 44.03 ms on one RTX 4090. A footer reads 'd3-mini figures are vLLM-SR's own; Intern-Decision-4B figures are InternLM's own; neither is independently reproduced.'

Latency, and why the two sets of milliseconds are not comparable either

Both cards publish per-request latency, and taking them at face value would be a mistake for the same reason the accuracy figures are not comparable.

• Intern-Decision-4B — mean 44.16 ms, median 44.03 ms, p95 44.60 ms, measured on a single RTX 4090 through the local Hugging Face path, described as workload and hardware dependent.

• d3-mini — median 17.5 ms for text, 96.2 ms with an image, 371.5 ms with a ten-second video, on one AMD Instinct MI325X, one request at a time.

Two things make these incomparable. The first is the hardware and the software path: a 4090 against an MI325X, a stock Hugging Face forward pass against a custom attention implementation with masked-layer kernels available through flash-linear-attention. The second is the workload: InternLM's number is described as per-query end-to-end on an unstated mix, and vLLM-SR's is broken out by input modality, so the text-only comparison is the only like-for-like line and even that crosses two GPU vendors.

The number to take from both cards is not the ranking, it is the shape. A decision model is called repeatedly inside one workflow — a support record might need a destination, a refund check, an escalation decision and a priority score, four questions, and a batch of 128 records turns that into 512 decisions. At that volume, 17 ms and 44 ms both disappear next to whatever the generative model downstream costs. The modality-dependent figures are the ones to watch, because an image or a video request is between five and twenty times the cost of a text one on d3-mini's own numbers, and if your decision is being made on a screenshot you have imported a cost profile most decision-model deployments do not have.

Which one to actually pick

If the decision you need to make depends on a video, there is no contest and no analysis required: Decision 3.0 reads video and Intern-Decision-4B does not. That is the entire answer for anything involving screen recordings, camera clips or frame sequences, and it is the capability gap that justifies the newer family's existence on its own.

If your inputs are text and occasional images, the choice turns on two things and neither of them is the leaderboard.

Take Intern-Decision-4B when you need to reason about thresholds. It is the only one of the two that tells you whether a 0.9 means nine times out of ten, it names its temperature, it says which cases that temperature was fitted on, and it documents its inference as a short numbered procedure you can reimplement against a stock model class. For a scorer sitting in front of an automated action, that is the property that matters, and it is rarer than accuracy points.

Take Decision 3.0 when you need the range or the modalities. Six checkpoints from 0.59B to 26.09B means the same request format can be served by a 6.7 ms edge model and a 27B one, and the family shares one interface, so moving between tiers is a configuration change rather than a rewrite. The catch is that you are trusting an unstated input budget and an unaudited calibration claim, and the largest model in the family is the one whose index number the vendor measured on its own harness.

Neither is a safe default today. d3's most-downloaded checkpoint has been on Hugging Face for about a day; Intern-Decision-4B has been up for two weeks and drawn roughly 3,200 downloads and 83 likes, which is attention but not production traffic. Both are cheap enough to test and neither has a third-party evaluation behind it. If you are putting a scorer in front of something that spends money, the right move is to run both on your own labelled cases and compare the calibration curves, not the index rows.

A screenshot of the Intern-Decision-4B model card on Hugging Face. The header shows 83 likes, 1.34k followers and tags for Image-Text-to-Text, Transformers, Safetensors, qwen3_5, decision-making, multimodal, structured-prediction and conversational under an Apache-2.0 licence, with the model size listed as 5B parameters in F32 or BF16 and the base model pinned to Qwen/Qwen3.5-4B. A section titled 'How inference works' lists five numbered steps: map each question's options to single-token symbols A to Z then a to z then 0 to 9; render the state, decision schema and a JSON skeleton with one decision placeholder per field; run one causal forward pass and read logits immediately before each placeholder; take a softmax over the allowed candidate-symbol logits and apply the checkpoint's probability calibration; and map symbols back to the original option values. It states that the API never calls generate() and samples no free-form text. A benchmark results table below carries Jevbench Easy, Jevbench Original, Jevbench Hard, Typed Decision, ToolACE, AG News and WildJailBreak columns for the Jev, Laya, Semif and Kev rows. The sidebar reports 3,179 downloads last month.

Where a router fits, honestly

OrcaRouter does not carry either of these models. Decision 3.0's checkpoints are a local Python inference path in a Hugging Face repository with no HTTP endpoint published, and Intern-Decision-4B ships as a DecisionEngine class you instantiate yourself. Neither is something we could route today, and no part of this article should be read as an availability claim.

What we do carry is the hosted end of the same family. typesafe/jev-1.13 is in our catalogue, served over POST /v1/systemone — the same state-and-named-questions contract both open models above implement — at $0.042 per million input tokens with no completion charge, since it never generates one. It sits alongside 200-plus other models, and that is the practical point for anyone comparing these two: the scorer is the cheap part of the loop and the model that acts on the decision is the expensive part. Routing both through one key, with automatic failover when a provider wobbles and provider list price passed through at 0% markup, means evaluating a decision model does not require signing a second contract or rewriting the call site when you switch backends. If you are mid-evaluation — which is where both of these models are, today — that is the part worth setting up before you commit to either.

A screenshot of the OrcaRouter model page for Jev 1.13. The header reads 'Jev 1.13', by TypeSafe, dated 2026-09-24, tagged NEW, with a specification panel reading 65K tokens of context, text input, text output and a p50 time-to-first-token of 176 ms, and the endpoint listed as /v1/systemone. The description says it is TypeSafe's structured decision and evaluation model, given a state and a set of named questions (noul / choice / score), returning a structured answer for each, served via POST /v1/systemone, non-streaming, up to about 64K input tokens, text in and structured JSON out. The metric strip reads input /bin/bash.04 per 1M tokens, no output price, p50 TTFT 176 ms, p95 TTFT 423 ms and 59.3M tokens of traffic over 7 days. Buttons read 'Get the Jev 1.13 API' and 'Try in playground', and a code sample shows a POST to https://api.orcarouter.ai/v1/systemone with the model typesafe/jev-1.13 and a state plus noul, choice and score questions.

The open question

The two cards disagree about what a model author owes a reader, and that disagreement is more interesting than the models. InternLM published a temperature and the cases it was fitted on, then published the diagnostic showing how much calibration improved. vLLM-SR published file hashes, pinned base revisions, a stated hardware target, a coverage claim — real provenance work — and no calibration number at all.

The test of which release matures is not which one wins a board. It is whether the next Decision checkpoint ships with a Brier score on it, and whether InternLM's next upload reaches video. Both are visible from the outside, both are cheap to check, and neither has happened yet.