Source-led article
Cloudflare Launches Clef Decision Models, Aiming to Cut Humans Out of AI Agent Loops

Cloudflare has released two new decision models, Clef and Clef-flash, designed to let AI agents make structured decisions without generating text responses. The company says the models can handle classification tasks so quickly that humans no longer need to be in the loop for routine agentic decisions.
The models are built on Alibaba’s Qwen architecture and are available under the Apache 2.0 license. Cloudflare positions them as a direct competitor to TypeSafe AI’s Jev model, with speed as the primary differentiator. Clef-flash delivers classifications in about 39 milliseconds median latency, while Clef takes around 209 milliseconds. By comparison, Jev requires just over 524 milliseconds for the same tasks, according to Cloudflare’s benchmarks.
What decision models do differently
Decision models occupy a niche between large language models and traditional classifiers. A large language model can reason and call tools, but its outputs vary and it can be slow. Traditional classifiers are fast but need retraining for every new category. Clef returns a brief classification with probability scores instead of a long text response.
For example, given a customer support message, Clef assesses its urgency and identifies the team that should handle it. Downstream code can use those results to route a ticket, trigger an escalation, or hand the case to a human. Cloudflare states that “a human does not necessarily need to be in the loop for agentic decisions anymore.” Agents can “programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.”
The name Clef comes from music notation, where a clef assigns pitches to the lines of a staff. The phonetic resemblance to Jev is likely intentional.
Technical details and benchmarks
Clef is based on Qwen3.8-27B and Clef-flash on the smaller Qwen3.5-9B. Cloudflare leaves the base models unchanged during training and uses its own synthetic data to train extra components. These components derive answer options and probabilities from the models’ internal computations.
Cloudflare uses its own variant of Reinforcement Learning for Calibrated Decisions (RLCD), the same training method TypeSafe used for Jev. RLCD trains models to answer multiple questions about an input in a single call, aiming to assign probabilities that match how often the answers are actually correct.
Across 43 benchmarks, Clef and Clef-flash are faster than all relevant competing decision models, according to Cloudflare. The company’s threat intelligence team is already testing Clef to classify websites. In one example, it assigns a domain a 95 percent probability of being a fashion website and 85 percent of being an online store. The probability of it being a phishing site is under one percent. Fetching, rendering, and classifying the site took 2.2 seconds, compared to 4.7 seconds for Cloudflare’s fastest general-purpose language model, which returned only two categories.
Clef can also process images, while Jev is limited to text so far. Its 64,000-token context window holds twice as much input as Jev’s.
Datos clave
| Model | Median Latency | Base Model | Context Window |
|---|---|---|---|
| Clef | ~209 ms | Qwen3.8-27B | 64,000 tokens |
| Clef-flash | ~39 ms | Qwen3.5-9B | 64,000 tokens |
| Jev (TypeSafe) | ~524 ms | undisclosed | 32,000 tokens |
Availability and customisation
Both models run on Cloudflare’s Workers AI platform and are available on Hugging Face under the Apache-2.0 license. The API is fully compatible with Jev’s, so customers can switch easily.
Alongside the launch, Cloudflare is rolling out a reinforcement learning service that lets customers tailor Clef to their own tasks. A team of forward deployed engineers will initially handle fine-tuning with customers, with a self-service platform planned for later. Customers can build a dataset by logging requests through AI Gateway, then evaluate those requests in containers that serve as an RL sandbox. A new trainer component will let them deploy the fine-tuned model on Workers AI. To run custom models, Cloudflare uses technology from Replicate, which it acquired in late 2025.
Cloudflare says it plans to use Clef internally to review abuse reports, sort support requests, and distinguish useful bots from harmful ones.
Context and competition
The decision model approach was popularised by TypeSafe AI, the startup founded by former OpenAI researcher Diogo Almeida, which introduced Jev in mid-September. TypeSafe markets Jev as a model “without hallucinations,” though that only guarantees it stays within predefined answer options and does not prevent it from choosing incorrectly.
In late September, OpenAI followed with a Decisions API built on GPT-6 Luna that also accepts context as text or images. Cloudflare primarily runs a global network for content delivery, DNS, and security services and has attracted little attention for its own AI models. Its recent AI headlines have focused on giving website owners control over access, such as letting site operators block or allow AI bots based on their purpose in July 2025.
For Indian enterprises and startups building AI agents, the availability of fast, open-source decision models on Cloudflare’s edge infrastructure could simplify deployment of automated customer support, content moderation, and threat classification workflows. The Apache 2.0 license and compatibility with existing APIs reduce vendor lock-in concerns.
Source: The Decoder – https://the-decoder.com/cloudflare-says-its-new-clef-model-means-humans-no-longer-need-to-be-in-the-loop-for-ai-agents/