A generated title card for Microsoft-Decision-1 vs Gemma 4 12B subtitled 'a scorer with no output tokens against a writer with no closed set', with chips reading probabilities versus prose, 32,768-token context versus 256,000, text only versus text image audio and video, and hosted API versus open weights.
Guides & Insights

Microsoft-Decision-1 vs Gemma 4 12B: One Picks From Your List, One Writes You a New One

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Microsoft-Decision-1 and Gemma 4 12B side by side and the difference that decides everything is not size, accuracy or licence — it is what happens after the model has read your input. Gemma 4 12B, DeepMind's 11.96-billion-parameter encoder-free multimodal model released on 3 June 2026 under Apache 2.0, reads text, images, audio or video and then writes you an answer, token by token, across a 256,000-token window and more than 140 languages. Microsoft-Decision-1 went generally available on Microsoft Foundry on October 8, 2026 as a text-only scorer post-trained on Qwen3.5-9B: you hand it a state and a question whose answers you have already enumerated, and it returns a calibrated probability for each option in a single pass over up to 32,768 tokens, generating nothing. One model tells you which of your options is right. The other tells you what the options should have been — and bills you for every word of that.

Framing this as a contest would be a category error, and it is worth saying so before the spec lines rather than after them. These two are not substitutes; they sit at opposite ends of a workflow. The useful question is which end you are actually building, because the answer determines whether you are buying a judge or a writer.

The bill is where the difference stops being philosophical

Gemma 4 12B is billed the way every generative model is: input tokens and output tokens, both. When you ask it for structured JSON, every character of that JSON is generated, sampled and paid for — and the more nested your schema, the longer the answer and the larger that line item. That is not a defect of the model. It is what a decoder does, and it is precisely the line item that decision heads exist to delete.

Microsoft-Decision-1 has no output side to bill. It runs a single forward pass over your state, your question and your option set, reads a score for each candidate, and returns. The consequence compounds with volume rather than with prompt size: a pipeline that labels ten thousand records with Gemma 4 12B produces ten thousand JSON documents, and the same pipeline on a decision scorer produces ten thousand prompts and nothing else. Whether the absolute saving is large depends entirely on your volume and the fatness of your schema, and anyone with a per-token budget can settle it in a spreadsheet. The direction is not in dispute.

There is a second, smaller asymmetry: Microsoft does not print a rate on the Microsoft-Decision-1 model page. Pricing links out to Microsoft's own pricing surface, so the per-decision cost is a thing you read off Azure rather than off the card. What the page does establish is that 0% of the call is output tokens, and that batch inference is disabled — you cannot amortize a bulk scoring run through a batch channel the way you would with a generative model.

A two-column generated scoreboard titled Microsoft-Decision-1 vs Gemma 4 12B. Left column Microsoft-Decision-1 rows read: parameters 9B Qwen3.5 base post-trained; context 32,768 tokens; output probabilities with zero output tokens; modalities in text only; weights hosted, not distributed. Right column Gemma 4 12B rows read: parameters 11.96B encoder-free multimodal; context 256,000 tokens; output generated text billed per token; modalities in text, image, audio, video; weights Apache 2.0 and self-serve. A footer line reads that the Gemma 4 12B figures are Google-reported and Microsoft-Decision-1 has published no benchmark figures.

Same input, two different answers

Give both models a customer support ticket and a question. Microsoft-Decision-1 returns something like 0.83 for "billing", 0.11 for "technical", 0.04 for "cancellation", 0.02 for "cannot tell". You have a nameable label, a number you can threshold against, and an abstention path — Microsoft documents that a "cannot tell" style option is supported when the supplied evidence is insufficient, which is what makes escalation logic work. What you do not get is a reason.

Give the ticket to Gemma 4 12B and you get a paragraph. It can reason about the request with configurable thinking, call functions, read a scanned invoice attached to the ticket, transcribe the voicemail stapled to it, and answer in the customer's own language. Every one of those is something Decision-1 cannot do at any accuracy: its input surface is text and nothing else, and it is explicitly not designed for open-ended question answering, conversation, translation or summarization.

None of Gemma 4 12B's output is a probability. Ask it for a routing decision with confidence and you will get prose containing a number, and that number is whatever the sampler produced rather than a softmax over the option set you supplied. If a downstream threshold depends on that number meaning something, you have a calibration project on your hands, not a scorer.

Whether your question has a closed set is the whole decision

This is the test to apply before anything else, and it takes about ten minutes with a whiteboard.

• If your answer is drawn from a list you can enumerate in advance — a label, a rating, a yes/no, a rubric score, a route — Microsoft-Decision-1 is the correct instrument, and Gemma 4 12B is a wasteful and less reliable way to pick from that list. You would be paying a decoder to guess at a number you already bounded.

• If your answer is not a closed set — summarise this, explain that, write the reply, describe the chart — Microsoft-Decision-1 cannot help at any price. It has no path to prose, no rationale field, and no mechanism for producing anything you did not enumerate as an option.

• If it is both, you have a two-model pipeline rather than a choice, and the interesting design question becomes which one owns the escalation threshold. The scorer sets it; the writer fills in what happens above and below it.

Where Gemma 4 12B is simply a different sport

Outside the decision framing, Gemma 4 12B is not ahead on a line — it is playing a different game, and pretending otherwise would be dishonest.

A screenshot of the google/gemma-4-12B-it model card on Hugging Face, showing the Gemma 4 banner, the Hugging Face, GitHub, launch blog, documentation and technical report links, an Apache 2.0 licence, the note that the card covers the Gemma 4 12B Unified model, and the opening paragraphs describing a multimodal family handling text, image, video and audio with a context window of up to 256K tokens and multilingual support in over 140 languages.

• Modalities — Microsoft-Decision-1 is text-only in and numeric out. Gemma 4 12B takes text, image, audio and video natively through an encoder-free design that projects raw patches and waveforms straight into embedding space, with interleaved input in any order.

• Context — 32,768 tokens for Microsoft-Decision-1 against 256,000 for Gemma 4 12B, which scores 43.4% on MRCR v2 eight-needle at 128k on Google's own card.

• Language — 25 supported languages for Microsoft-Decision-1, with Microsoft warning that coverage, quality and calibration "may vary by language" and naming lower-resource non-English as a weak spot; 140+ for Gemma 4 12B, with an evaluated MMMLU figure of 83.4 for this size.

• Published performance — Microsoft-Decision-1's Benchmarks tab carries a methodology and no figures; Google publishes 77.2% MMLU Pro, 77.5% AIME 2026 without tools, 72.0% LiveCodeBench v6, 78.8% GPQA Diamond, 69.1% MMMU Pro and 79.7% MATH-Vision for Gemma 4 12B.

• Distribution — Microsoft-Decision-1 is a Foundry-only hosted API with no weights, no download and no fine-tuning path. Gemma 4 12B is Apache 2.0 weights that shipped into a day-one ecosystem of LM Studio, Ollama, llama.cpp, MLX, vLLM, SGLang and Unsloth.

• Runtime — decision scoring is one pass with no decoding loop, so latency scales with the prompt rather than with the answer length; Gemma 4 12B generates token by token, and Google's launch materials place it at 16 GB of VRAM or unified memory.

The benchmark row is the one worth pausing on, in the direction that flatters Microsoft least. Gemma 4 12B's numbers are vendor-reported too — but they have had months of third-party use behind them, in public harnesses, on hardware other people own. Microsoft's entry in this category is the newest and the least verifiable of the two despite coming from the larger company.

What each costs you to actually call

Gemma 4 12B in its 12B form is not on our catalogue — if that is the size you want, you fetch 11.96 billion parameters of BF16 and serve it yourself. What our catalogue does carry is the rest of the Gemma 4 line, and it is inexpensive: Gemma 4 26B A4B at $0.06 per million input tokens and $0.33 per million output, and Gemma 4 31B at $0.13 and $0.38, both listed with 262,144-token contexts. Those are vendor list prices passed through with no markup added on our side, behind one OpenAI-compatible key alongside more than 200 other models. If your problem turns out to be general reasoning rather than a closed-set decision, one of those two sizes is the cheap thing to measure against a self-operated scorer — and if you would rather not commit to a single writer, automatic failover keeps the generation leg of a scoring loop alive when one provider degrades.

A screenshot of OrcaRouter's own model page for google/gemma-4-31b-it, showing the Google breadcrumb, the model name Gemma 4 31B, a release date of 2026-04-02, the description of it as a 30.78B dense multimodal model with text and image input, a 256K token context window and configurable thinking mode, list pricing of $0.13 per million input tokens and $0.38 per million output tokens, and a P50 time to first token figure.

Microsoft-Decision-1 is a Foundry deployment you provision yourself, and it is not available through us. Nothing in this article is an availability claim for it, on either side of the call.

Bottom line

Microsoft-Decision-1 went generally available on Microsoft Foundry on October 8, 2026: a text-only, 32,768-token scorer on a Qwen3.5-9B base that returns calibrated probabilities over options you enumerate, with zero output tokens, no weights and no published benchmark figures. Gemma 4 12B is an 11.96B encoder-free multimodal generalist released on 3 June 2026 under Apache 2.0 that reasons, reads images and audio and writes answers across 256,000 tokens. Pick by the shape of your answer, not by the parameter counts: a closed set belongs to the scorer, an open one belongs to the writer, and paying a decoder to select from a list you already wrote is the most common way teams spend ten times what a decision costs.

What our catalogue does carry is the rest of the Gemma 4 line, and it is inexpensive: Gemma 4 26B A4B at $0.06 per million input tokens and $0.33 per million output , and Gemma 4 31B at $0.13 and $0.38.