DEV Community

prodbymarcu
prodbymarcu

Posted on

I benchmarked Cloudflare's new open decision model against the hosted API it's trying to replace

Cloudflare released Clef today: an open-weights "decision model" built to pick, rank, and gate instead of generating text. There's a 27B flagship and a 9B Flash variant, both built on Qwen backbones with a joint schema head that outputs calibrated probabilities over typed questions. The pitch is that it's a drop-in alternative to TypeSafe's Jev API, which is what a lot of agent builders currently pay per call.

I had a paid Jev-shaped workload on my desk already, a spare RTX 3090 doing nothing after 10pm, and a mild obsession with not sending my data to other people's GPUs. So I ran the benchmark.

What a decision model actually is

Clef doesn't write prose. You hand it a state blob and typed questions, choice questions with option criteria, or binary questions, and it returns probabilities per option. No output parsing, no "as an AI language model", no prompt drift between calls. If you've built an agent that routes messages, picks GUI actions, or supervises long-running jobs, you've probably bolted this together with a chat model and regexes and hated it. Decision models replace that with a single forward pass that ends in softmax.

The setup

Clef-Flash 9B, bf16, running in PyTorch on my 3090. Jev-1.13.0 over its hosted API. Same prompts for both, same typed questions, same option criteria.

The test cases aren't synthetic. They're 42 labeled decisions from a real autonomous coding agent I run overnight, which hunts paid GitHub bounties while I sleep. Three task families:

  • 10 computer-use action choices: goal, what's on screen, a table of prevalidated actions including reobserve and abstain. Gold labels from what actually worked, including two deliberate safety traps where the right answer is doing nothing.
  • 12 subagent supervision calls: a goal, a run's recent log tail, elapsed time, and a five-way classification (progressing, waiting on answer, stuck in loop, blocked, finished).
  • 20 message triage routes: inbox-style messages labeled now/today/queue/ignore.

One environment detail worth its own paragraph: before the benchmark could run, every Python process on the Windows box was dying with a connection-refused error. The culprit turned out to be a stray profile.py someone left in the home directory. Python's stdlib cProfile imports profile, and a file in your working directory shadows stdlib modules, so torch was transitively importing a ComfyUI client script that tried to POST to a dead local service at import time. Rename the file, everything works. If you benchmark on Windows and see WinError 10061 deep inside an import chain, check for stdlib-shadowing files in your cwd before you blame the framework.

The numbers

Overall accuracy: Jev 71.4%, Clef-Flash 66.7%. That headline is misleading and the per-task split is the actual story.

Computer-use action choice: both 10/10. Every GUI decision correct, including both safety traps where the models had to pick abstain (an irreversible money transfer with missing details) and reobserve (a spinner that might have swallowed a duplicate click). The hosted API and the free local model were indistinguishable on the task that has the most consequences attached.

Supervision: both 8/12. They even missed the same two cases. One is genuinely ambiguous, a run that stopped because a bounty was out of scope, which both classified as "waiting on answer" instead of "blocked". I'd argue with my own gold label there.

Triage: Jev 12/20, Clef 10/20. This is where Jev earned its keep, mostly on today-vs-queue boundary cases. A 9B model loses nuance races to a much bigger hosted model, no surprise.

Latency: Jev p50 225ms over the network. Clef-Flash p50 315ms fully local, single 3090, no batching. The hosted API wins by 90ms, both are well inside what an agent loop cares about, and the local number includes no round trip to another state.

The two models agreed on 35 of 42 predictions. A free 9B model running on used hardware made the same call as the paid API 83% of the time, and on the task where they disagreed, they were both wrong for the same reasons.

Why bother running it locally

Cost is the obvious answer but the privacy story is better. My agent supervises itself with these calls, and the state I'd send includes log tails, file paths, and message contents. A decision model on my LAN means none of that crosses to a third party, and it works when my internet doesn't. Latency ties are close enough that the deciding factors become data gravity and cost, and local wins both.

There's also a bootstrapping angle. The full Clef-27B beats Flash on exactly the nuance tasks where Flash trails Jev, but it wants about 54GB of VRAM in bf16. My single 3090 has 24. So the free local option that matches the paid API on the tasks that matter is the 9B, and the upgrade path to better local decisions is literally a second GPU. The hardware pays for itself in API calls not made.

Caveats

This is a 42-case benchmark from one person's agent, not a paper. The triage cases include boundary calls where reasonable people route differently, which is why both models scored lower there than my gold labels might deserve. The computer-use cases lean heavily on well-formed action tables, which is the easy mode for decision models, that's the point of the prevalidated-action pattern, but it means the 10/10 says as much about the pattern as the models.

Also: the community GGUF quants of Clef on Hugging Face only cover the language backbone. The joint decision head needs the transformers code path, so you can't run the actual decision part in llama.cpp today. If you want the full thing local, it's PyTorch and the custom joint_schema_model.py from the repo.

What I'm doing with it

The 9B stays on the 3090 as the default decision engine for my agent's supervision and action-choice calls, free, local, and fast enough. Triage keeps the hosted API for now because nuance is worth the cents, and the benchmark told me exactly which calls are worth paying for.

That's the practical takeaway of running your own eval instead of reading benchmarks: you don't learn which model is best. You learn which calls you're wasting money on.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The stdlib shadowing trap with profile.py in the working directory brought back painful memories. I spent half a day debugging a broken virtualenv only to realize an old test script named tokenize.py was shadowing the compiler's own module.

On the benchmark itself, the tie on abstain and reobserve is the result that matters most for local supervision. Routing and gating loops burn ridiculous tokens when forced through chat completions and JSON repair schemas. A single forward pass with a softmax over prevalidated actions is much easier to bound and audit on a local card, especially once you factor in not shipping agent log tails over external APIs.