Cloudflare Launches Clef and Clef-flash: First In-House Models Return Typed Probabilities
When TypeSafe launched Jev on September 15, 2026, it opened a new category: decision models that return typed probabilities instead of free text. Two weeks later, Cloudflare answered with Clef and Clef-flash — but almost every number circulating about the new models comes from the vendors themselves and has yet to be independently reproduced, and Clef costs nearly six times as much per token as Jev.
What Was Actually Launched
Cloudflare released Clef and Clef-flash on October 1, 2026, during the company's Birthday Week. According to MarkTechPost, these are the first models trained by the company's Workers AI team, and they are not chatbots: each model reads an input state and a schema of typed questions, and returns a probability for each allowed answer — no free text (MarkTechPost).
Both models are open-weight under the Apache 2.0 license and compatible with TypeSafe's Jev API. Clef uses the same "System One" API as Jev, which according to MarkTechPost means that switching from Jev in practice involves changing the endpoint and model name. In a market where switching costs could otherwise have been significant, the shared specification reduces friction for developers who want to compare the models.
The same day, AWS released its own open-source Strands Decider 2B — a sign of how quickly the category is filling up.
How the Models Work
Technically, Clef is built on specially trained, frozen versions of Qwen3.8-27B, while Clef-flash uses Qwen3.5-9B, according to Cloudflare via The Register (The Register). "Frozen" here means the backbone is locked as a fixed base; post-training adds the decision behavior on top of it.
The inference pattern may be the most interesting part: the Qwen backbone runs, according to Cloudflare, a pure prefill pass — no token generation — and the allowed answers are scored in parallel. This mechanically explains why latency can be low: the model does not need to generate text sequence by sequence, it only has to evaluate a fixed set of options. The models also handle images and video as input, in addition to text.
The models are hosted on Workers AI, so developers can run them without their own infrastructure — while the Apache 2.0 weights also allow self-hosting.
The Numbers This All Hinges On — and Who Supplied Them
Cloudflare published a shortlist of benchmarks from Decision Index 0.2.1 in which Clef led on 7 of 10. The most prominently featured figure is BANKING77 (macro-F1): 94.20 for Clef versus 79.74 for Jev. But Jev leads on the heaviest reasoning benchmarks: GPQA Diamond 78.3 versus 48.0, MMLU-Pro 82.7 versus 65.9, and BBH 92.9 versus 73.7 (all figures vendor-reported, via MarkTechPost).
On TypeSafe's own workflow evaluations, the picture is far more even than a "Jev killer" narrative would suggest. Clef won three of four areas by small margins: invoice handling 64.7 versus 61.8, customer service 76.3 versus 76.0, and security incidents 62.9 versus 61.7. Jev leads on agent trace observability, 71.6 versus 68.5 — these figures are also vendor-reported.
The crucial caveat: MarkTechPost notes that all the figures were supplied by the vendors without independent replication, and The Register writes that Cloudflare's self-evaluations have not yet been reproduced for ranking on the official Decision Index. For now, then, we do not know how Clef scores when someone else runs the tests.
The Latency and Pricing Picture — With Caveats
On speed, Cloudflare's internal run reported a median latency of 38.8 ms for Clef-flash versus 524.1 ms for Jev (Cryptobriefing). In an internal threat-intelligence example, Clef was said to have fetched and classified a web page in 2.2 seconds versus 4.7 seconds for a general model. TechBooky stresses that these are company-driven comparisons, not a guarantee that every customer will see the same results (TechBooky).
Two factors complicate the "faster than Jev" story. First, the latency comparison rests on different hardware assumptions than TypeSafe's hosted service. Second, The Register notes that Clef is slightly slower than other open models on the same ranking — so Cloudflare beats Jev in its own run, but is not the fastest in its class.
On pricing, Clef sits at $0.24 per million input tokens and Clef-flash at $0.09, with no billing for output tokens — logically enough, since the models do not generate text. That is nearly six times the price of Jev, which costs $0.042 per million tokens (The Register). Cloudflare is thus betting that lower latency and higher task-specific accuracy justify the price; the market will decide.
One detail is worth noting: the context window is reported differently across the coverage. TechBooky and Cryptobriefing give 64k for Clef, while Jev's limit is given as 32k in some sources and up to 64k in others, depending on how the state plus the longest question is counted. The disagreement cannot be resolved with the available documentation, and should remain open until TypeSafe's and Cloudflare's own specifications are in place.
A Category Filling Up at Record Speed
The timeline shows how fast the market is moving. TypeSafe launched Jev on September 15, 2026. Open alternatives such as Kev-9B and Laya followed shortly after. On October 1 — the same day as Clef — AWS released Strands Decider 2B, an open-source decision model that, according to TechCrunch, is inspired by Jev (TechCrunch). According to TechCrunch's coverage of the launch, the model has around 2 billion parameters, built on Qwen3.5-2B, and AWS distinguished engineer Marc Brooker has described the use case as a structured, low-latency "decider for a workflow step" (TechCrunch).
Two structural features distinguish this wave from earlier model races. First, several players share the same API specification (System One), which makes switching costs low — change the endpoint and model name, and you are up and running. Second, several of the models are open-weight, so you can self-host and compare on your own data. Over time, that could push both prices and claims toward verifiability.
The category logic itself makes sense: many agent and automation workflows do not need a generative language model at all. They need a structured choice — should the case be escalated, approved, flagged? — with calibrated probabilities, low latency, and predictable cost. Decision models are built for exactly that, and the prefill-only architecture Cloudflare describes for Clef shows how much can be gained by dropping text generation entirely.
The Fine-Tuning Platform: a Direction, Not a Product
Cloudflare tied the launch to a reinforcement learning (RL) platform where customers are to be able to fine-tune the decision models on their own workflows. But it is worth being precise: initially, this is carried out by forward-deployed engineers working hands-on with customers. A self-service platform is only promised for later. It is a stated direction — not an available feature — and it should be assessed as such by anyone considering putting Clef into production.
What Remains to Be Answered
Several questions remain open. Will Cloudflare's numbers hold when independent parties run them against the official Decision Index? The latency and accuracy map has so far been drawn by the interested party itself. Second: TypeSafe has not published Jev's underlying architecture, so comparisons of approach are one-sided — we know how Clef is built, but not Jev. Third: will the price premium or the latency win? Clef is roughly six times more expensive than Jev per token, but faster in Cloudflare's own run; for workflows where decisions are the bottleneck, latency may weigh more heavily than token price. Fourth: will the shared System One API harden into a real standard, or will the category fragment as soon as a major player deviates?
What can be said with certainty: two weeks after the category opened, there are at least four competing models, several of them with open weights, a shared API that makes switching cheap — and a pile of self-reported numbers that no one but the vendors has yet confirmed. For developers, the most sensible advice for now is to test on their own data and their own workflows, and let the official Decision Index rankings — when they come — settle the disputes.

