Cloudflare's Clef targets the agent decision layer

Cloudflare has released Clef and Clef-flash, a pair of decision models built for AI agents and aimed squarely at TypeSafe AI's Jev. The company's central pitch is speed. Instead of producing a long natural-language response, Clef returns a short classification with probabilities attached. Those probabilities give downstream code something concrete to act on, whether that means routing a support ticket, triggering an escalation, or handing a case to a human operator.

Cloudflare's framing is blunt: with this kind of model, "a human does not necessarily need to be in the loop for agentic decisions anymore." In its description, agents can gather context programmatically, make decisions, and take action on tasks, while still deferring to a person when the situation calls for it. That is a significant claim about autonomy, and it places Cloudflare in competition with a model that has already established itself as a frontrunner in the category.

What a decision model actually does

The term "decision model" describes a system that sits between general-purpose language models and conventional classifiers. It does not write essays or hold conversations. Instead, it answers a set of predefined questions about an input and returns a probability for each possible answer.

A practical example from Cloudflare involves a customer support message. Given such a message, Clef can assess how urgent it is and identify which team should handle it. The downstream system reads those answers and acts: routing the ticket to the right queue, escalating it, or passing it to a human. The model never needs to compose a reply itself.

The name comes from music, where a clef assigns pitches to the lines of a staff. Cloudflare uses the analogy to suggest that a decision model sets the framework for the actions that follow. The phonetic resemblance to Jev is, in all likelihood, not accidental.

Speed is the headline number

Cloudflare is leaning heavily on latency as its differentiator. Across 43 benchmarks, the company says Clef and the smaller Clef-flash are faster than all relevant competing decision models. The reported median response times are roughly 39 milliseconds for Clef-flash and about 209 milliseconds for Clef. By comparison, Cloudflare puts Jev at just over 524 milliseconds.

Diagram showing a support ticket about API errors and two questions about team and urgency flowing into a decision model that outputs 100 percent for the technical team and 100 percent for urgency in parallel.
A decision model answers multiple predefined questions about an input at the same time, returning a probability for each answer option. | Image: Cloudflare
  • Clef-flash: about 39 milliseconds median latency.
  • Clef: about 209 milliseconds median latency.
  • Jev: just over 524 milliseconds, according to Cloudflare's comparison.

Those are self-reported figures, and Cloudflare has an obvious interest in presenting them favorably. Even so, the gap it describes is large enough to matter in systems where many small decisions happen in sequence. A tenfold difference in response time can determine whether an agent feels responsive or sluggish, and it can shape how many steps a pipeline can afford to take before a user notices a delay.

Cloudflare also notes that both models run directly on its own infrastructure. The company says this lets them benefit from proximity to edge data centers, which could reduce the distance data travels before a decision comes back.

Built on Qwen, compatible with Jev

Both Clef and Clef-flash are built on Qwen, according to Cloudflare, and both support text and images. That multimodal capability matters for agent workflows that need to interpret screenshots, scanned documents, or other visual inputs alongside text.

Just as important for adoption, Cloudflare has kept the API fully compatible with Jev. Customers already using TypeSafe AI's model can switch without rewriting their integrations. That compatibility is a classic competitive move: it lowers the cost of trying an alternative and makes the speed comparison the main variable in the decision.

Filling the gap between language models and classifiers

Cloudflare positions Clef in a niche that has been awkward to fill. Large language models can reason about a problem and call tools, but their outputs vary from run to run, and they can be slow. Traditional classifiers are fast, but they typically need retraining whenever a new category is introduced. A decision model aims to take the speed and predictability of a classifier while keeping the flexibility that agentic systems need.

That distinction is useful when thinking about where this technology fits:

  • Language models are general-purpose reasoners whose outputs are flexible but variable and relatively slow.
  • Traditional classifiers are fast and consistent but rigid, requiring retraining for every new category.
  • Decision models return probabilities over predefined options, aiming for speed, structure, and enough flexibility for agent pipelines.

For teams building agents, the appeal is straightforward. An agent that must decide repeatedly whether to escalate, route, tag, or pause does not need a full generative response each time. It needs a fast signal it can trust enough to act on.

Scatter plot with Decision Index score on the y-axis and median latency in milliseconds on the x-axis. Clef scores 61.2 points at about 210 milliseconds, Clef-flash scores 57.1 points at about 40 milliseconds, and Jev scores 57.9 points at about 525 milliseconds, alongside open models like AutoJev-27B and Kev 9B.
In Cloudflare's self-reported numbers, Clef delivers the highest decision quality while Clef-flash nearly matches Jev's accuracy at a fraction of the latency. | Image: Cloudflare

Humans in the loop, or on the bench

The most provocative part of Cloudflare's messaging is the suggestion that human oversight is no longer a mandatory step for every agentic decision. Cloudflare describes a setup in which agents gather context, decide, and act programmatically, while retaining the option to defer to a human when needed.

That phrase, "when needed," is doing a lot of work. A decision model that returns probabilities can support automatic action, but the threshold for acting without a person is a design choice made by whoever builds the system. A high-confidence routing decision may be safe to automate; an escalation involving a frustrated customer, a safety issue, or a financial commitment may warrant review even when the model is confident. The technology makes faster automation possible, but it does not by itself settle where the line should be drawn.

Probability outputs are also not the same as guarantees. A classification can be wrong, and a fast wrong answer can propagate through a pipeline faster than a slow one. The value of a decision model therefore depends not only on latency but on calibration and on how gracefully the surrounding system handles uncertainty.

What to watch as Clef rolls out

Cloudflare's claims are specific and measurable, which makes them testable. Independent benchmarks will matter, particularly on the 43 benchmarks Cloudflare cites and on real-world workloads that differ from vendor test suites. The Jev comparison will also be scrutinized, since TypeSafe AI's model is the incumbent that Clef is explicitly courting.

Several questions will shape whether Clef becomes a default choice for agent builders:

  • Do independent tests reproduce the latency advantage, especially under production load?
  • How well do the models handle edge cases, ambiguous inputs, and categories that were not anticipated?
  • Does edge proximity translate into meaningful gains for globally distributed applications?
  • How easily can teams move existing Jev integrations over, given the compatible API?

Cloudflare has framed Clef as the moment agents stop needing a human in the loop for every decision. The more accurate reading may be that it makes that choice cheaper and faster to exercise. Whether autonomy expands will depend less on the model and more on the guardrails that teams build around it.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com