A generated title card reading 'Clef', subtitled 'Cloudflare's decision models that never write a token', with two chips reading 'weights 30 Sept 2026' and 'blog 1 Oct 2026'. The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Clef Copies Jev's API on Purpose — and That Is the Interesting Part

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Clef is a 27-billion-parameter multimodal model from Cloudflare that answers typed questions and nothing else. You hand it a state — text, JSON, an image, a video frame — plus a schema of up to 64 named questions, and it returns a probability for every allowed option in one forward pass. No generated text, no parser, no decoding loop. What makes Cloudflare/clef worth reading about is not that design, which it borrowed, but the wire format it chose: Cloudflare built Clef to speak the request and response body of Jev System One, the decision model TypeSafe shipped first, served at the same POST /v1/systemone path. Interoperability here is the whole commercial argument. If your pipeline already asks a decision model questions in that shape, Clef is a URL change and a model string, not a migration.

That is an unusual thing for a release post to bury, and Cloudflare does not bury it — the model card states flatly that the API is "fully compatible with Jev and SystemOne," and the blog opens by crediting TypeSafe with putting the category on the map before positioning Clef as the second-generation version of an internal experiment. So the honest read of this release has three parts: what the compatibility decision buys you, what Cloudflare's own numbers say about where Clef wins and loses, and what remains unverified because every figure in that table was produced by the vendor that ships the model.

The category came from somewhere else

Cloudflare's post is candid about the lineage. The idea of a model that returns bounded structured outputs rather than prose is credited to TypeSafe's Jev System One. Cloudflare's own route in was an earlier demo that bent a diffusion model into emitting deterministic probabilities by exposing logprobs, plus community work by Matt Mastracci on making the inference engine handle that pattern. Clef builds on that concept with a different backbone, a purpose-built scoring head, and — the part that matters commercially — a request body that matches the incumbent's.

When a vendor both copies a rival's wire format and benchmarks itself against the rival's index, the compatibility is not a footnote to the model, it is the product strategy. Swapping a decision model behind an existing endpoint is the cheapest possible way for a new entrant to be tried, and the cheapest possible experiment for a team that already has one of these wired up.

What actually shipped, and what it costs

• Parameters — 27B, built by freezing Qwen3.8-27B with its vision encoder and training a joint schema head plus rank-256 low-rank adapters on top.

• Sibling — Clef-Flash, the same recipe at 9B from Qwen3.5-9B. Separate article, separate trade-offs.

• Context — 65,536 tokens, per the Workers AI model page and the blog post. The bundled encoding helper defaults to max_length=16,384, a default rather than a ceiling, but the default you get if you use the loader as shipped.

• Question types — noul (true/false, returning the probability of true), choice (2 to 26 named options), score (2 to 26 ordered levels). One to 64 questions per request.

• Price — $0.24 per million input tokens on Cloudflare's Workers AI, or self-hosted from Apache-2.0 weights that measure roughly 55 GB across twelve shards.

• Licence — Apache 2.0, following the Qwen3.8-27B base checkpoint.

A single-column scoreboard titled 'Clef — the scoreboard', with six rows: Parameters: 27B, vision encoder kept; Trained from: Qwen3.8-27B, head plus rank-256 adapters; Context: 65,536 tokens; Output: a probability per option, no generated text; Latency: 209.3 ms median, 238.6 ms p95; Licence: Apache 2.0. A footer reads 'Cloudflare figures, self-run on the Jev Decision Index, unreproduced.' The OrcaRouter logo is composited in the bottom-right corner.

Where it wins, and where it does not

Cloudflare's Decision Index run is the source for everything in this section, and the leaderboard hosting it is Cloudflare's. Nothing below has been independently reproduced. Read it as a supplier's best case, and notice how uneven it is.

The wins cluster where a house-trained classifier should win. Clef takes BANKING77 intent macro-F1 at 94.2 against Jev's 79.7, and CLINC150 with out-of-scope handling at 97.4 against 89.3 — a thirty-point gap that is the most decision-relevant cell in the table, because refusing to classify is the hard half of routing. It also takes CRUXEval at 86.7, CLadder at 94.0, and the home-appliance simulator at 83.0.

The losses are just as real and belong in the same paragraph. On general knowledge and hard reasoning the older Jev wins outright: GPQA Diamond 78.3 to 48.0, MMLU-Pro 82.7 to 65.9, BBH 92.9 to 73.7. Those are the three biggest gaps in the table and they all point the same way. Clef is also not the best model in its own comparison on the security-domain PhishNChips set, where it lands at 79.6 and a competing entry scores higher.

On the four end-to-end workflow evaluations, drawn from TypeSafe's own public eval set and scored against consensus labels, the two are close enough to be a coin toss. Clef takes invoice processing on exact actions at 64.7 to 61.8 and the security-incident set 62.9 to 61.7; Jev takes agent-trace observability at 71.6 against 68.5; customer service is a 76.3-to-76.0 draw that means nothing at that margin.

Latency is where the gap is structural rather than marginal. Cloudflare reports 209.3 ms median and 238.6 ms p95 for Clef, against 524.1 and 536.0 for Jev. That ordering falls out of the non-autoregressive design rather than out of a benchmark's quirks, which is why it is the number most likely to survive contact with a real deployment. The absolute milliseconds were measured on Cloudflare's own hardware and mean little on their own.

The most convincing datapoint in the release is not a benchmark row. Cloudflare says its threat-intelligence team handed a domain to Clef alongside its browser-rendering service and got a category verdict in 2.2 seconds, against 4.7 seconds for its own fastest general LLM on the same workflow, with more classifications returned. Internal anecdote, not a study — but it is the kind of story that explains why anyone would build this rather than prompt a general model.

A headless capture of Cloudflare's blog post announcing Clef, showing the Cloudflare site header, the breadcrumb 'Blog', the banner date 'October 1, 2026', the headline 'Introducing Clef: our open-source decision models, and new RL fine-tuning platform', and the three named authors.

Two things the API reference tells you, and one it does not

The documented traps are small and worth reading before you port anything. The confidence field measures how concentrated the probabilities are, not the probability that the answer is correct — a well-calibrated model can return a low-confidence correct answer and a high-confidence wrong one. And when two options tie, the answer follows the model's option order, so put your preferred option first. Anyone porting a prompt-based classifier without reading those two sentences will misread their own logs.

The larger constraint is not documented anywhere because it is structural. Clef has no text-generation path at all, which means it cannot draft a reply, summarise a thread, call a tool, or hold a conversation, and it will not invent a category you did not declare. The whole design assumes you can enumerate the allowed answers in advance. That quietly excludes a large class of problems, and no amount of benchmark performance changes it.

What the compatibility argument is worth

This is where the release stops being a model review and becomes an architecture question. If Clef speaks Jev's request shape, then the model behind your decision endpoint is a config value, and the sensible way to adopt it is side by side rather than as a replacement — run both against your own traffic, keep the one that measures better on your data, and let the switch stay cheap.

Clef itself is not on our route list; the OrcaRouter catalogue returns a 404 for it, and nothing here should be read as an availability claim from us. What we do carry is the incumbent it was built to displace. TypeSafe's Jev 1.13 is a listed route at $0.042 per million input tokens over a 65,536-token context, served through the same POST /v1/systemone body, so an A/B test between the two is a model string rather than a rewrite. Having both behind one key — 200-plus models, provider list price passed through at no per-token markup, automatic failover across providers — is what keeps that test running after the week someone decides to care about it, instead of decaying into a stale comment in a config file. The general model you keep around to act on whatever the decision model concludes is reachable from that same key rather than a second contract.

A headless capture of the OrcaRouter model page for typesafe/jev-1.13, showing the TypeSafe breadcrumb, the model title, the description naming the noul, choice and score question types served over POST /v1/systemone, a code sample set to model 'typesafe/jev-1.13', and the performance strip reading p50 TTFT 148 ms.

What would change this read

Someone outside Cloudflare needs to rerun the Decision Index. The suite is public and the evals are described as reproducible, so a third-party run is the difference between a vendor's table and a fact — and the cells that matter most are the ones pointing in opposite directions. A 97.4 on out-of-scope intent detection next to a 48.0 on GPQA Diamond is not a smooth capability curve; it describes a model that is very good at the classification shapes in its training distribution and much weaker at reasoning it has not seen. If that holds under independent testing, it is the most operationally important fact about Clef and it is not in the highlights.

Production traffic also needs to look like the latency table. A 209 ms median measured on one H200 says nothing about p95 under concurrency, and the design's whole selling point is that it sits in a hot path.

And the fine-tuning platform needs to produce something. Cloudflare's post describes an RL loop — AI Gateway to capture a dataset from your own traffic, Workers AI to generate rollouts, Containers as the scoring sandbox, Trainer to update weights, then redeploy onto Workers AI — in the same post that announces the models, without a customer result attached. A fine-tuned Clef that beats the base on a task the base was never trained for would be far stronger evidence for the category than any row in the table.

Until then the fair summary is this: Cloudflare shipped a real, permissively-licensed, genuinely fast decision model and made it trivially swappable with the one that came first. That is a better story than most open releases offer, and it is still, so far, entirely the vendor's own account of itself.