
Clef vs Qwen3.8-27B: One Backbone, Two Output Contracts
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 41 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Clef cannot write a sentence, and it is built out of a model that can. Cloudflare/clef is a 27-billion-parameter multimodal decision model: you hand it a state and a schema of typed questions, and it returns a probability for every allowed option without producing a token of prose. The backbone underneath it is Qwen3.8-27B, an open-weights dense vision-language model, frozen in place with a small extra head attached. That makes this one of the few model-versus-model pages where the two sides are not rivals at all — they are a checkpoint and its derivative, sharing 55 gigabytes of weights and disagreeing about whether the answer is something you read or something you compute with.
The weights for both went public before the words did. Cloudflare created the Cloudflare/clef repository on Hugging Face at 21:15 UTC on 30 September 2026, under Apache-2.0, and published the announcement post — "Introducing Clef: our open-source decision models, and new RL fine-tuning platform" — the following afternoon. For roughly eighteen hours a 55-gigabyte model was downloadable and running on Cloudflare's own edge GPUs before its vendor said anything about it. Within three days it had an Ollama library page behind Ollama 0.35.1, which is where this piece found it, and NVFP4, FP8, EXL3, INT8, Q4 and Q8 builds from a dozen different uploaders.
What is knowable from the repository is unusually complete, and it is worth separating that from what is not. Knowable: the architecture, the base checkpoint, the licence, a reference implementation that runs, and a full evaluation matrix. Not knowable: whether any of those numbers survive a rerun by someone who is not Cloudflare. Every score attributed to Clef below comes from the vendor's own run of a suite the vendor hosts. That asymmetry against a model with a month of independent use behind it is the honest spine of this comparison, and the rest of the piece is about where it actually bites.
The two repositories, side by side
Start with the download, because the sameness is the point. Clef's safetensors come to 54.97 GB across twelve shards. The parent's come to 55.56 GB across eighteen. Clef keeps Qwen3.8-27B's vision encoder — the processor, the vision tower, the whole multimodal front end — so nothing was stripped to make it smaller. What Cloudflare added is a joint schema head of 256 MB: four layers, 1,024 wide, sixteen heads, a 4,096-wide feedforward, two routing layers reading the backbone's final hidden states. The training recipe is a frozen backbone, that head, and rank-256 low-rank adapters, with label-smoothed cross-entropy for valid schema answers paired against a Brier loss to sharpen calibration.
Then the divergence, which is entirely in what the model is allowed to emit.
• Parameters — Clef 27B, post-trained from Qwen/Qwen3.8-27B. Qwen3.8-27B 27B dense, 64 layers, 5,120 hidden, 248,320-token vocabulary, 16 × (3 × Gated DeltaNet → FFN) interleaved with Gated Attention.
• Output — Clef: one logit per allowed option, softmaxed per question, no generation path at all. Qwen3.8-27B: tokens, tool calls, and a reasoning block that is on by default.
• Question shapes — Clef takes 1 to 64 typed questions per request in three forms: noul (probability of true), choice (2 to 255 named options), score (2 to 10 ordered levels, returning a probability-weighted value that can land between them). Qwen3.8-27B has no such contract; you get prose and you write the parser.
• Licence — Apache-2.0 on both, which is why the third-party quantisation happened in a weekend.
• Cost of the frozen half — Clef's head is 256 MB against 54.97 GB of backbone. Almost the entire download is the parent. If you already have Qwen3.8-27B on disk, you are downloading it a second time to get a different opinion about how to answer.

What the re-heading costs, in two numbers
Cloudflare's evaluation is the Decision Index 0.2.1 suite, run internally, and it is a table about decision models — six columns: Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B and Laya. Qwen3.8-27B has no row in it, and Clef has no row on the Artificial Analysis Intelligence Index. So there is no shared scoreboard and this piece will not invent one.
There is, however, one benchmark both sides have been put through, and the gap is large enough to be the most decision-relevant number in the release. On GPQA Diamond — graduate-level science questions, multiple choice — Cloudflare measures Clef at 48.0. The parent sits at 90.5 on Artificial Analysis's harness and 89.2 on the vendor's own model card. That is a forty-point cliff on general knowledge, and it is not a mystery: Clef answers by scoring a fixed set of options it was handed, and general-knowledge multiple choice is about as far from "route this ticket" as the same weights can be pointed. Treat the comparison as indicative rather than controlled — two harnesses, two labs, neither the vendor's nor the tracker's. But the direction is not in doubt, and it is the price of the contract.
The same pattern shows in the rest of Clef's general-purpose cells. MMLU-Pro 65.9, BBH 73.7, CLINC150 with out-of-scope handling 97.4 for the 27B against Jev's 89.3 — that last one is the rare cell where the derivative beats the older decision model outright, and it matters because refusing to classify is the hard half of routing. Where Clef is strong it is strong exactly where a house-trained classifier should be: BANKING77 intent macro-F1 at 94.2, BFCL case-exact at 98.5, API-Bank at 91.9. Where it is weak, a text-only decision model from a competing lab — TypeSafe's Jev — beats it by thirty points on GPQA Diamond while billing $0.042 per million input tokens on our own catalogue, which is a useful reminder that "27B" is not itself an argument.
What the parent keeps is everything the Decision Index does not ask about. On its own card and in the AA-sourced rows our catalogue carries: Terminal-Bench 2.1 at 73.0, SWE-bench Pro 61.7, LiveCodeBench v6 90.3, OSWorld-Verified 84.3, DeepSWE 1.1 at 42.2, AA Coding 68.1 at rank 35 of 138 sampled models. None of those have a Clef row, because Clef cannot attempt them. It has no tool-calling path, no long-horizon planning, no way to emit the patch.
What one decision costs, and why the output line is the whole argument
The Workers AI model page lists exactly one unit price for Clef: $0.24 per million input tokens. There is no output line, and the reason is in the responses. Ollama's own documented example — a support ticket routed across three questions — comes back with "usage": {"input_tokens": 1204, "output_tokens": 3}. Three output tokens for three questions. Whether Cloudflare bills those three tokens is nearly irrelevant: the output cost of a decision is not a rate, it is a count of questions.
Put that against the parent's rate card on our own catalogue, which is $0.33 per million input tokens and $2.40 per million output with a 262,144-token window. Same 1,204-token state, same single routing decision:
• Clef — 1,204 input tokens at $0.24 per million is about $0.00029 per decision, and about $289 per million decisions. The number barely moves when the answer gets harder, only when the state gets longer.
• Qwen3.8-27B with thinking disabled — the same state costs about $0.00040 in input, plus roughly ten output tokens at $2.40 per million. Call it $0.00042, or $421 per million. Clef is about 1.5× cheaper, which is not the story anyone is telling.
• Qwen3.8-27B with thinking left on — which is the default on this checkpoint. If the model spends 400 tokens reasoning about which of four buckets a password-reset email belongs in, that is another $0.00096 and the decision lands near $0.0014, or $1,357 per million. Those 400 tokens are an assumption, not a measurement, and the multiple moves linearly with it.
So the honest version of the pricing argument is narrower than it first looks. Against a generalist running with thinking off, Clef wins on dollars by a factor of about one and a half — real, but not transformative. Against the same model in its default configuration, the gap is a function of verbosity, and verbosity is exactly what a language model has and a scorers-over-options design does not have at all. That is the structural claim: Clef's cost is bounded by the length of the state, and the parent's is bounded by how much the parent decides to say.

The one number that is not close
Cloudflare reports median request latency of 209.3 ms for Clef and p95 of 238.6 ms, measured across the 43 benchmarks in its run, on its own edge GPUs. Nobody outside Cloudflare has reproduced it, and the vendor's infrastructure is the reason the number is as low as it is — this model is designed to sit on the same network as the caller.
For the parent, we can do better than a vendor number, because it is one of our routes. Over the seven-day window ending 2 October 2026, OrcaRouter's own playground measured Qwen3.8-27B at a p50 of 1,929 ms and 235.8 output tokens per second, on 136.3 million tokens of traffic in that window. The daily series is the interesting part: p95 sits at exactly 10,000 ms on six of the seven days, which reads like a ceiling rather than a measurement, and the p50 swings between 1,751 ms and 4,296 ms depending on the day. Treat the median as roughly nine times Clef's, and treat the tail as uninformative until someone publishes a real distribution.
That gap is not a tuning difference and it will not close. A prefill-only pass plus a parallel scoring head finishes in one sweep; an autoregressive model finishes when it runs out of things to say. The vendor absolute figures for Clef deserve the usual discount, but the ordering is a property of the architecture, not of Cloudflare's hardware.
Context is quoted three ways for one of them
Both models take text, JSON, images and video. The windows are not the same, and one of them is quoted inconsistently depending on where you read.
• Clef, hosted — 65,536 tokens, per the Workers AI model page and the announcement post. That is the number that binds an API caller, and Cloudflare's own framing is comparative: twice the state that Jev's 32K window holds.
• Clef, on disk — the released config.json declares max_position_embeddings of 262,144, because the frozen backbone is the parent's and still carries the parent's position ceiling. The bundled encode_record helper defaults to max_length=16,384, with max_state_tokens available to bound the state separately. None of those three numbers is wrong; they answer different questions, and a runtime listing that advertises 256K is reading the config, not the endpoint.
• Qwen3.8-27B — 262,144 natively, extensible to 1,000,000 with the right serving configuration, per the vendor's card and our catalogue row. Four times the hosted Clef window at minimum.
The practical consequence is smaller than the ratio suggests. Clef consumes a state — a ticket, an invoice, a screenshot — not a conversation, and one of its selling points is that you can hand it a large flattened payload and ask sixty-four questions about it without paying for the payload again per question. The parent needs the window because it has to hold the history, the tool results and its own reasoning. Different requirements, different ceilings.
Both of them see the same pixels
The vision encoder is inherited, so this is not a case of one model being multimodal and the other not. Clef accepts up to four images per request, base64-encoded PNG, JPEG or WebP, shared by every question in the payload, and video frames can be added as frame arrays — the receipt example in the card asks a yes/no question about whether a total is legible, which is the design in miniature. Qwen3.8-27B is a native vision-language model with OmniDocBench 1.5 at 91.1, RealWorldQA at 85.9 and CharXiv at 78.8 on the vendor's card.
What differs is what you get back. Clef gives you a field-by-field decision with a probability attached, which is a typed answer about a document. The parent gives you a description of the document, which you then parse. For a large class of document work — check this form, locate this clause, is this total above the threshold — the typed answer is the entire job, and the description is overhead you pay for and then throw away.
Where the line actually falls
• Take Clef — the answers are enumerable in advance and you would rather not write a parser. Routing, triage, gating, policy checks, document field extraction. Closed sets, 2 to 255 options per question, up to 64 questions scored together.
• Take Clef — the decision is in a hot path and 209 ms against 1,929 ms changes what the pipeline can do. This is the single strongest reason to move, and it is a latency reason, not an accuracy one.
• Take Clef — you want determinism. An option you did not declare cannot come back, because there is no generation step to produce it.
• Keep Qwen3.8-27B — the answer is not a choice from a list: summarising, drafting, reasoning through a novel problem, writing code, driving tools over a long horizon. Clef cannot do any of it, and it will not tell you so.
• Keep Qwen3.8-27B — you need the model to say "none of these." A decision model scores the options it was given. Hand it a billing/technical/sales schema and a privacy request, and it returns the least-bad of those three rather than a refusal. The parent surfaces the question you did not think to ask.
• Keep Qwen3.8-27B — context or long-document work. Four times the hosted window, before the million-token extension.
Read that list as a pipeline rather than a verdict, because that is what it is. The architecture makes the relationship literal: Clef is the parent's prefill pass with a different answer mechanism on the end. A sensible design puts Clef in front — decide the route, the priority, the escalation flag — and keeps Qwen3.8-27B behind it to act on the decision. The two are complements that happen to share a download.
Where to run each of them
Clef is hosted on Cloudflare's Workers AI at the $0.24 per million input figure above, and the Apache-2.0 weights are yours to run locally. The Ollama library listing is the newest surface: ollama pull clef needs Ollama 0.35.1 or later, because the model speaks through the /v1/systemone endpoint that Ollama added for decision models. Watch the tags, not the headline — the default and clef:27b tags are 18 GB, which is a four-bit build of a 55-gigabyte model; clef:27b-q8_0 is 30 GB, clef:27b-nvfp4 is 18 GB, clef:27b-mxfp8 is 31 GB, and clef:27b-mlx-bf16 is the full 55 GB for Apple silicon. TypeSafe's own Python SDK will talk to it if you point it at a local Ollama and supply any API key, which the runtime ignores. One caveat that will bite in production: confidence measures how concentrated the probabilities are, not the chance the answer is correct, and ties resolve in your declared option order.
The parent is the easier one to put behind an API today. Qwen3.8-27B is routable on OrcaRouter at the $0.33 and $2.40 per million rates used above, with the 262,144-token window, and it is one of 200-plus models reachable through a single OpenAI-compatible key. Because provider list price is passed through at 0% markup, a vendor price cut on either side of this comparison is live on our routes the same day it lands.
Clef itself is not one of our routes — our public catalogue returns nothing for it, and nothing here should be read as an availability claim from us. What we do carry is the model that makes the decision pattern reachable without a Cloudflare account: TypeSafe's Jev 1.13 is a listed route at $0.042 per million input tokens with a 65,536-token context, measured at a p50 of 148 ms across the same seven-day window, and it speaks the identical POST /v1/systemone request shape. Since Clef is API-compatible with Jev on purpose, a pipeline built against either one is a config change away from the other — which is the case a routing layer exists to make free. Keep both reachable, measure on your own labels, and let the winner be a data point rather than an architectural commitment. If a decision model ends up gating an agent, the model that acts on the decision still has to be reachable, and that half is already behind the same key.

What would change this read
Three things, in ascending order of how much they would move it.
A third party needs to rerun the Decision Index. The suite is public and Cloudflare describes the evals as reproducible, so this is achievable — and the CLINC150 out-of-scope cell is the one to watch, because a 97.4 for the 27B is either a genuine advantage over Jev's 89.3 on the hardest half of routing or an artefact of a house-built harness. Nobody outside Cloudflare knows yet.
Someone needs to score Clef and Qwen3.8-27B on the same classification task with the same labels. That is the only comparison that would settle whether the re-heading is a quality trade or a quality upgrade for closed-set decisions, and neither vendor has an incentive to run it. Until then the forty-point GPQA gap stands as the best available evidence about what was removed, and it is evidence about the wrong task.
And Clef's latency needs a p95 under concurrency. The 209 ms median was measured by the vendor, on hardware the vendor controls, for a model whose entire pitch is that it sits in a hot path. The ordering against an autoregressive parent is architectural and will survive. The absolute milliseconds are the number to distrust, and they are the number the whole business case rests on.
Qwen3.8-27B is one of 200-plus models reachable through a single OpenAI-compatible key — Qwen3.8-27B.
