A generated title card reading 'Clef vs Gemma 4 12B', subtitled 'a 64K-token decision model against a 256K-token generalist', with chips reading '65,536 vs 256,000 context' and 'probabilities vs prose'. The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Clef vs Gemma 4 12B: A 64K-Token Decision Appliance Against a 256K-Token Generalist

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most useful number in a comparison between Clef and Gemma 4 12B is not a benchmark score. It is 65,536 against 256,000 — the context windows. Cloudflare/clef is a 27-billion-parameter multimodal decision model, post-trained from Qwen3.8-27B, that reads a state and a schema of typed questions and returns a probability for every allowed answer in a single non-autoregressive pass. Gemma 4 12B — the Unified checkpoint, google/gemma-4-12B-it — is the vendor's 11.95-billion-parameter encoder-free any-to-any model, released in the summer of 2026, which generates text over a 256,000-token window in more than 140 languages. Put a folder of contracts in front of both and the decision model runs out of room four times sooner, because its ceiling is a documented 65,536 tokens while the generalist's is 256,000. That asymmetry is not a footnote. For a whole class of routing work — long documents, pasted threads, multi-page forms — it decides the choice before accuracy is even discussed.

So this is not a page about which model is better. It is a page about the two axes on which they genuinely differ: how much state each can hold, and what shape of answer each is physically able to produce. Everything else is either a vendor number that nobody has reproduced or a comparison that does not mean what it looks like.

What each one is, in one paragraph each

Clef is Cloudflare's 27B decision model, published under Apache-2.0 with the weights on Hugging Face and a hosted endpoint on Workers AI. Cloudflare froze the Qwen3.8-27B backbone including its vision encoder, attached a joint schema head plus rank-256 low-rank adapters, and trained it with label-smoothed cross-entropy for valid schema answers alongside a Brier loss to sharpen calibration. You supply state — text, JSON, images or video frames — and one to 64 named questions, each typed noul (true/false, returning the probability of true), choice (2 to 26 named options) or score (2 to 26 ordered levels). What comes back is a distribution per question and nothing else. There is no text-generation path at all.

Gemma 4 12B is a generalist that generates text. It is encoder-free: instead of separate vision or audio towers, raw image patches and waveforms are projected straight into the language model's embedding space through lightweight linear layers, which is what lets it take text, image, video and audio without a per-modality pipeline. It has configurable thinking modes, native function calling, a 262K vocabulary and a hybrid attention scheme that interleaves 1,024-token sliding-window attention with full global attention. It is Apache-2.0 with the vendor's own usage terms alongside, and it is the checkpoint most people mean when they say "run this size locally."

One thing to state plainly before going further: neither model is a route on OrcaRouter. Our catalogue returns a 404 for cloudflare/clef and for the 12B checkpoint in the same family, and nothing in this piece should be read as an availability claim from us. What we do carry from the Gemma family are the two larger sizes, and their prices appear further down.

A two-column comparison scoreboard titled 'Clef vs Gemma 4 12B — the scoreboard'. The left column 'Clef' reads: Parameters: 27B multimodal; Context: 65,536 tokens; Output: probability per option, no text; Input: text, JSON, images, video; Latency: 209.3 ms median, H200; Evidence: vendor's own Decision Index. The right column 'Gemma 4 12B' reads: Parameters: 11.95B dense, encoder-free; Context: 256,000 tokens; Output: generated text; Input: text, image, video, audio; Latency: about 2.4 s to first chunk; Evidence: vendor card plus Artificial Analysis. A footer reads 'Clef figures unaudited; Gemma figures per Google's card and Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.

The option list is an interface, and it has a failure mode

The structural difference between these two is what happens to input that does not fit.

• Output — Clef returns a probability per declared option, and only those options. Gemma 4 12B returns generated text.

• Refusal — Clef has no "none of the above" unless you write one into the criteria yourself. Gemma 4 12B can describe a case that fits nothing you asked about.

• Failure mode — Clef's error is silent: a confident pick inside your schema, no exception raised, no fallback branch triggered. The generalist's error is usually loud: prose that does not parse.

• Testability — a fixed option set makes the component reproducible and cheap to audit. Free text makes it flexible and expensive to validate.

Clef's own numbers for that refusal problem are the weakest figures it publishes. On CLINC150 with out-of-scope handling — intent detection where one valid answer is "this belongs to no category" — the 27B scores 97.4 macro-F1, which is the best in Cloudflare's table and the reason to care about the 27B over its own 9B sibling, which scores 66.8 on the same benchmark. A thirty-point gap between two models built from the same recipe says that out-of-scope behaviour is the thing post-training scale buys, and it is the thing most routing pipelines quietly depend on. The mitigation Cloudflare documents is a second question in the schema — add a noul field asking whether the state fits at all, then threshold it — which works, and which is your job rather than the model's.

Where the numbers come from matters more than what they say

Every Clef figure in this article comes from Cloudflare: its model card, its launch post, or its Decision Index run. The suite itself is Cloudflare's own, and the leaderboard that hosts the scores is Cloudflare's own. That is a supplier measuring its product on hardware it controls against a rival whose API it deliberately mirrors. No third party has reproduced any of it, and this piece is not going to pretend otherwise.

The Gemma 4 12B column is a different kind of evidence. the vendor's own card publishes MMLU Pro 77.2%, GPQA Diamond 78.8%, AIME 2026 without tools at 77.5% and LiveCodeBench v6 at 72.0%. Artificial Analysis, an independent tracker, indexes the reasoning configuration at an Intelligence Index of 14.18 with a listed $0.10 / $0.30 per million tokens and a measured median output speed of about 113 tokens per second — and its own page notes those figures are workload- and hardware-dependent.

The honest reading is that the two columns are not comparable at all. One is a vendor's decision-quality suite run by the vendor; the other is a mix of vendor benchmarks and an independent evaluation of a general model. Quoting a 97.4 next to a 77.2 as though they were columns of the same table would be the single most misleading thing this page could do.

Latency is the one place where the two can at least be discussed in the same units, and even there the caveats stack. Cloudflare reports Clef at 209.3 ms median and 238.6 ms p95, measured on its own infrastructure. Artificial Analysis measures the 12B's time to first chunk at about 2.4 seconds in a reasoning configuration, which includes a thinking budget the decision model does not have. A 209 ms non-autoregressive pass against a 2.4-second generation start is a real architectural difference — one forward pass versus a decoding loop — and it is the strongest argument Clef has for a hot path. It is also, at 27B, a model that needs H200-class hardware to hit that figure, and Cloudflare charges $0.24 per million input tokens to run it for you.

A headless capture of the google/gemma-4-12B-it model card on Hugging Face, showing the model title, the Any-to-Any and Transformers and Safetensors tags, the gemma4_unified image-text-to-text architecture line, the Apache 2.0 licence line credited to Google DeepMind, and the note that the card covers the Gemma 4 12B Unified model.

What 64K of context actually buys you

Context is where this comparison turns practical and where the generalist wins on paper by a wide margin.

• Clef — 65,536 tokens, per the Workers AI model page and the launch post; the bundled encoding helper defaults to max_length=16,384, which is a default rather than a ceiling, but it is the default you will get if you use the loader as shipped.

• Gemma 4 12B — 256,000 tokens, per the vendor's card, with 1,024-token sliding-window attention to keep the long-context cost down.

• Images — Clef scores images and video frames jointly with the text state through the inherited vision encoder. Gemma 4 12B takes text, image, video and audio through its encoder-free path.

• Languages — Clef's card makes no multilingual claim; Gemma 4 12B is documented across more than 140 languages.

For a ticket, an invoice or a rendered screenshot, 64K is generous and the difference is academic. For a 200-page contract you want classified field by field, it is the whole argument: the generalist can hold the document, and the decision model cannot. Clef's answer to that is chunking, and chunking a document before you decide about it means you have already made a decision the model was supposed to make. That is a real limit and it deserves to be stated as one.

What you would actually pay, and where

Neither model is on our catalogue, so the price question splits three ways.

• Clef — $0.24 per million input tokens on Cloudflare's Workers AI, or self-hosted from Apache-2.0 weights; Cloudflare's published latency was measured on a single H200 with torch 2.11 and transformers 5.10.2, and the community quantisations that appeared within two days suggest it can be pushed smaller with some accuracy cost nobody has measured.

• Gemma 4 12B — no hosted route here at all. You fetch roughly 24 GB of BF16 weights and serve them yourself, or reach for a third-party platform. No per-token price exists that we can quote.

• The routed substitute — Gemma 4 26B A4B, a mixture-of-experts with 25.2B total and 3.8B active parameters, lists on OrcaRouter at $0.06 per million input and $0.33 per million output tokens with a 262,144-token context. Gemma 4 31B, the dense multimodal sibling with configurable thinking and native function calling, lists at $0.13 and $0.38. Both are on one key with provider list price passed through at no per-token markup and automatic failover across providers.

That last line is the honest version of our involvement in this comparison. If what you need is a generalist that reads a long document and writes a sentence, the cheap route to one is a Gemma 4 size we actually serve, and it costs less per million tokens than self-hosting a 12B on rented GPUs. If what you need is a bounded decision in the hot path, Clef's 209 ms median is a real architectural property, and the closest thing to it on our catalogue is TypeSafe's Jev 1.13 — the model Cloudflare benchmarked against and copied the wire format from — at $0.042 per million input tokens over a 65,536-token context, served through the same POST /v1/systemone shape. Because the two are API-compatible behind a thin adapter, trying Clef against the incumbent is an experiment rather than a migration, and the experiment stays cheap when both sides are reachable through one endpoint.

A headless capture of the OrcaRouter model page for google/gemma-4-31b-it, showing the Google breadcrumb, the Gemma 4 31B title, the import date, a p50 TTFT latency figure, and the code-sample and benchmark navigation tabs.

The decision, stated as a decision

Pick Clef when the answers are enumerable, the state fits in 64K, the decision sits in a latency-sensitive path, and you are willing to own the residual class yourself. The 27B's 97.4 on out-of-scope intent detection is why you would pay for it over the 9B sibling, and its fixed output shape is what makes the component auditable without a parser.

Pick Gemma 4 12B when the input is long, when the modality is audio or video, when the answer needs to be prose, or when the categories are not known yet — because "triage this pile into whatever groupings make sense" is a request a decision model cannot accept, not one it handles badly. It is a generalist with an independent evaluation behind it, and for open-ended work that is worth more than a vendor table.

And be sceptical of both figure sets for the reason this article keeps repeating: Cloudflare has not been independently reproduced on anything, and that card is the vendor's own benchmark table. The one genuinely interesting number in the pair — whether a 27B schema head really does hold its 97.4 out-of-scope score on data it was not trained against — is exactly the number neither vendor can settle and a third party could.