Cloudflare launches Clef and Clef Flash for fast agent decisions
Cloudflare releases two open decision models for Workers AI, with typed probability outputs, visual inputs and different speed and quality tradeoffs.

Cloudflare released Clef and Clef Flash on October 1, adding two decision models to Workers AI and publishing their weights under the Apache 2.0 license. The dated release notice identifies them as the first models trained by the Workers AI team.
The models evaluate content against questions with defined answers. They return probabilities that software can use to select an option, estimate urgency or assign a score. That puts them in the part of an agent workflow that decides where work should go, before another model or application performs it.
Two sizes with different costs
The Clef model page lists 27 billion parameters, a context window of 65,536 tokens and a hosted price of $0.24 per million input tokens. Its supported inputs include text, structured data and visual material.
The smaller Clef Flash model has 9 billion parameters and the same context window. Its listed input rate is $0.09 per million tokens. Cloudflare positions it for applications where response time matters most.
The hosted interface accepts up to 64 questions in a request. Developers can ask for a yes or no probability, choose among named options or score against an ordered rubric. The image field accepts up to four embedded images. Remote image URLs are not supported, and long text is truncated to fit the token limit.
Speed results come with tradeoffs
Cloudflareโs model card reports median request latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef Flash, compared with 524.1 milliseconds for Jev. These are measurements from Cloudflareโs internal evaluation, rather than an independent test of a customer deployment.
The quality results also vary by task. On the CLINC150 plus out of scope intent benchmark, Clef scored 97.4 and Flash 66.8 on macro F1. On GPQA Diamond, their accuracy was 48.0 percent and 51.0 percent, while Jev reached 78.3 percent in the same reported table. The faster option does not lead on every workload.
For a support routing system, the useful test is how often the model sends a difficult or ambiguous ticket to the correct destination. Teams should check those errors alongside latency and cost, and decide when a low confidence result needs human review.
Customization begins with assisted projects
In the launch article, Cloudflare explains that Clef uses Qwen 3.8 27B while Flash uses Qwen 3.5 9B. Each scores the permitted answers after reading the input, without generating an intermediate written response.
Cloudflare also introduced a service for adapting Clef to specific workloads with its engineering team. A platform that lets customers train and redeploy models themselves is described as a later step. The immediate offer is an assisted engagement for design partners, alongside the hosted models and downloadable weights.
For another approach to bounded choices, see our coverage of Strands Decider 2B and its limits for local experiments.
Original AI generated conceptual illustration of a distributed decision model choosing among bounded options



