Cloudflare has released Clef and Clef-flash, two open-source decision models designed to enable AI agents to make fast, predictable choices. Unlike traditional large language models that generate text and tool calls fluidly, decision models receive information, evaluate options, and return typed responses with probabilities. In a customer support scenario, for example, a decision model can classify whether a request is urgent, route it to the right team, and determine when to escalate to a human agent.
The models are available on Workers AI and Hugging Face under the Apache 2.0 license, and Clef maintains API compatibility with competing solutions. Clef accepts images, offers a 64,000-token context window compared to 32,000 for Jamba, and delivers the structured classification that agent workflows require.
In Cloudflare's published latency benchmarks, Clef achieved a median of 209.3 ms, while Clef-flash reduced that to 38.8 ms. Jamba reached 524.1 ms, and while Laya was fastest at 5.8 ms, the company notes that speed came with considerable quality trade-offs. Cloudflare built Clef on the Qwen model foundation using an architecture that evaluates valid options in parallel instead of generating responses token by token.
Beyond raw performance, Cloudflare introduced a fine-tuning service powered by Reinforcement Learning to adapt the model for specific tasks: classifying support requests, analyzing Trust & Safety cases, and detecting bots. Initial training happens with expert guidance; a self-service platform is planned for later. The vision positions Clef in the critical path of agent workflows, handling structured decisions while a traditional LLM handles open-ended reasoning and task execution.