Cloudflare released two open-source "decision models," Clef and Clef-flash, hosted on its Workers AI platform, alongside a new reinforcement learning (RL) fine-tuning product.
Unlike large language models, which generate open-ended text, decision models are built to produce bounded, typed, probability-scored outputs for a fixed set of questions, the kind of output a workflow can act on directly without parsing free text. Cloudflare's example: pass in a support message and ask if it's urgent, which team should own it, and how severe the impact is, and get back structured, probability-weighted answers your code can use to route the ticket or escalate to a human.
Cloudflare says Clef currently leads the Jev Decision Index benchmark and is fully API-compatible with Typesafe AI's Jev models, so existing integrations can swap in Clef without rework. Both models are released on Hugging Face under an Apache 2.0 license for local use, and are also hosted on Workers AI for immediate use via API.
On speed, Cloudflare reports that using Clef to classify a website domain (fetching, rendering, and categorizing it) took 2.2 seconds and returned more category labels, versus 4.7 seconds for its general-purpose LLM gpt-oss-120b on the same task, which returned only two classifications. Clef also adds a vision encoder for classifying images, a capability Jev does not currently have, and a 64k context window versus Jev's 32k.
Alongside the models, Cloudflare is launching an RL fine-tuning service. Initially this runs as a hands-on engagement with Cloudflare's forward-deployed engineer (FDE) team; a self-serve version is planned later, letting customers capture their own request data via Cloudflare's AI Gateway, generate training rollouts on Workers AI, and redeploy a fine-tuned Clef model on Cloudflare's infrastructure. Cloudflare states it does not read, store, or train on customer requests unless they opt into this fine-tuning product.
Top comments (0)