<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: prodbymarcu</title>
    <description>The latest articles on DEV Community by prodbymarcu (@prodbymarcu).</description>
    <link>https://dev.to/prodbymarcu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155609%2Fbca687dc-ae23-403a-920d-7ea2b30b8ee0.png</url>
      <title>DEV Community: prodbymarcu</title>
      <link>https://dev.to/prodbymarcu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/prodbymarcu"/>
    <language>en</language>
    <item>
      <title>I benchmarked Cloudflare's new open decision model against the hosted API it's trying to replace</title>
      <dc:creator>prodbymarcu</dc:creator>
      <pubDate>Thu, 01 Oct 2026 20:45:14 +0000</pubDate>
      <link>https://dev.to/prodbymarcu/i-benchmarked-cloudflares-new-open-decision-model-against-the-hosted-api-its-trying-to-replace-2ded</link>
      <guid>https://dev.to/prodbymarcu/i-benchmarked-cloudflares-new-open-decision-model-against-the-hosted-api-its-trying-to-replace-2ded</guid>
      <description>&lt;p&gt;Cloudflare released Clef today: an open-weights "decision model" built to pick, rank, and gate instead of generating text. There's a 27B flagship and a 9B Flash variant, both built on Qwen backbones with a joint schema head that outputs calibrated probabilities over typed questions. The pitch is that it's a drop-in alternative to TypeSafe's Jev API, which is what a lot of agent builders currently pay per call.&lt;/p&gt;

&lt;p&gt;I had a paid Jev-shaped workload on my desk already, a spare RTX 3090 doing nothing after 10pm, and a mild obsession with not sending my data to other people's GPUs. So I ran the benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a decision model actually is
&lt;/h2&gt;

&lt;p&gt;Clef doesn't write prose. You hand it a state blob and typed questions, choice questions with option criteria, or binary questions, and it returns probabilities per option. No output parsing, no "as an AI language model", no prompt drift between calls. If you've built an agent that routes messages, picks GUI actions, or supervises long-running jobs, you've probably bolted this together with a chat model and regexes and hated it. Decision models replace that with a single forward pass that ends in softmax.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Clef-Flash 9B, bf16, running in PyTorch on my 3090. Jev-1.13.0 over its hosted API. Same prompts for both, same typed questions, same option criteria.&lt;/p&gt;

&lt;p&gt;The test cases aren't synthetic. They're 42 labeled decisions from a real autonomous coding agent I run overnight, which hunts paid GitHub bounties while I sleep. Three task families:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;10 computer-use action choices: goal, what's on screen, a table of prevalidated actions including &lt;code&gt;reobserve&lt;/code&gt; and &lt;code&gt;abstain&lt;/code&gt;. Gold labels from what actually worked, including two deliberate safety traps where the right answer is doing nothing.&lt;/li&gt;
&lt;li&gt;12 subagent supervision calls: a goal, a run's recent log tail, elapsed time, and a five-way classification (progressing, waiting on answer, stuck in loop, blocked, finished).&lt;/li&gt;
&lt;li&gt;20 message triage routes: inbox-style messages labeled now/today/queue/ignore.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One environment detail worth its own paragraph: before the benchmark could run, every Python process on the Windows box was dying with a connection-refused error. The culprit turned out to be a stray &lt;code&gt;profile.py&lt;/code&gt; someone left in the home directory. Python's stdlib &lt;code&gt;cProfile&lt;/code&gt; imports &lt;code&gt;profile&lt;/code&gt;, and a file in your working directory shadows stdlib modules, so torch was transitively importing a ComfyUI client script that tried to POST to a dead local service at import time. Rename the file, everything works. If you benchmark on Windows and see &lt;code&gt;WinError 10061&lt;/code&gt; deep inside an import chain, check for stdlib-shadowing files in your cwd before you blame the framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;Overall accuracy: Jev 71.4%, Clef-Flash 66.7%. That headline is misleading and the per-task split is the actual story.&lt;/p&gt;

&lt;p&gt;Computer-use action choice: both 10/10. Every GUI decision correct, including both safety traps where the models had to pick &lt;code&gt;abstain&lt;/code&gt; (an irreversible money transfer with missing details) and &lt;code&gt;reobserve&lt;/code&gt; (a spinner that might have swallowed a duplicate click). The hosted API and the free local model were indistinguishable on the task that has the most consequences attached.&lt;/p&gt;

&lt;p&gt;Supervision: both 8/12. They even missed the same two cases. One is genuinely ambiguous, a run that stopped because a bounty was out of scope, which both classified as "waiting on answer" instead of "blocked". I'd argue with my own gold label there.&lt;/p&gt;

&lt;p&gt;Triage: Jev 12/20, Clef 10/20. This is where Jev earned its keep, mostly on today-vs-queue boundary cases. A 9B model loses nuance races to a much bigger hosted model, no surprise.&lt;/p&gt;

&lt;p&gt;Latency: Jev p50 225ms over the network. Clef-Flash p50 315ms fully local, single 3090, no batching. The hosted API wins by 90ms, both are well inside what an agent loop cares about, and the local number includes no round trip to another state.&lt;/p&gt;

&lt;p&gt;The two models agreed on 35 of 42 predictions. A free 9B model running on used hardware made the same call as the paid API 83% of the time, and on the task where they disagreed, they were both wrong for the same reasons.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother running it locally
&lt;/h2&gt;

&lt;p&gt;Cost is the obvious answer but the privacy story is better. My agent supervises itself with these calls, and the state I'd send includes log tails, file paths, and message contents. A decision model on my LAN means none of that crosses to a third party, and it works when my internet doesn't. Latency ties are close enough that the deciding factors become data gravity and cost, and local wins both.&lt;/p&gt;

&lt;p&gt;There's also a bootstrapping angle. The full Clef-27B beats Flash on exactly the nuance tasks where Flash trails Jev, but it wants about 54GB of VRAM in bf16. My single 3090 has 24. So the free local option that matches the paid API on the tasks that matter is the 9B, and the upgrade path to better local decisions is literally a second GPU. The hardware pays for itself in API calls not made.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;p&gt;This is a 42-case benchmark from one person's agent, not a paper. The triage cases include boundary calls where reasonable people route differently, which is why both models scored lower there than my gold labels might deserve. The computer-use cases lean heavily on well-formed action tables, which is the easy mode for decision models, that's the point of the prevalidated-action pattern, but it means the 10/10 says as much about the pattern as the models.&lt;/p&gt;

&lt;p&gt;Also: the community GGUF quants of Clef on Hugging Face only cover the language backbone. The joint decision head needs the transformers code path, so you can't run the actual decision part in llama.cpp today. If you want the full thing local, it's PyTorch and the custom &lt;code&gt;joint_schema_model.py&lt;/code&gt; from the repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm doing with it
&lt;/h2&gt;

&lt;p&gt;The 9B stays on the 3090 as the default decision engine for my agent's supervision and action-choice calls, free, local, and fast enough. Triage keeps the hosted API for now because nuance is worth the cents, and the benchmark told me exactly which calls are worth paying for.&lt;/p&gt;

&lt;p&gt;That's the practical takeaway of running your own eval instead of reading benchmarks: you don't learn which model is best. You learn which calls you're wasting money on.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cloudflare</category>
      <category>llm</category>
    </item>
    <item>
      <title>My local LLM fixed a bug overnight. The "success" was the bug.</title>
      <dc:creator>prodbymarcu</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:25:13 +0000</pubDate>
      <link>https://dev.to/prodbymarcu/my-local-llm-fixed-a-bug-overnight-the-success-was-the-bug-6dj</link>
      <guid>https://dev.to/prodbymarcu/my-local-llm-fixed-a-bug-overnight-the-success-was-the-bug-6dj</guid>
      <description>&lt;p&gt;I run a small autonomous coding agent on a secondhand RTX 3090 in my apartment. Last week I pointed it at paid GitHub bounties and let it run unsupervised, thinking I'd wake up to pull requests. I did wake up to pull requests. One of them was a lie, and it taught me more about agent pipelines than a week of everything going right.&lt;/p&gt;

&lt;p&gt;The model is Qwen3 27B, quantized, served locally with llama.cpp. It is not a frontier model. That's the whole point. Frontier models can write correct diffs most of the time. A 27B model will happily produce something that looks like a diff, fails to apply, and tells no one. Building for that gap is the actual engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The pipeline is boring on purpose. A Python runner loops over bounties, clones the repo, installs dependencies, runs the test suite to get a baseline, then asks the model to fix the issue in up to six rounds. Each round: generate a patch, apply it, run tests, feed failures back. If tests pass, save the patch for human review.&lt;/p&gt;

&lt;p&gt;The naive version of "apply the patch" is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;applied&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;diff&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...)&lt;/span&gt;
&lt;span class="nf"&gt;run_tests&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That has a flaw that took me an embarrassing amount of logging to see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The false green
&lt;/h2&gt;

&lt;p&gt;Round one on a real bounty came back green in 40 seconds. One round, tests pass, done. I checked the diff file and it was empty. Zero bytes.&lt;/p&gt;

&lt;p&gt;What happened: the model produced a diff with fake hunk headers. &lt;code&gt;git apply&lt;/code&gt; rejected it with "corrupt patch", which my code caught. It logged the failure and moved to the next round. But the test suite runs on the repo either way, and the repo was in its original state. Tests passed because the original code was fine. My pipeline saw exit code 0 and declared victory over a patch that never existed.&lt;/p&gt;

&lt;p&gt;I'd built a machine that manufactures green checkmarks out of nothing. If I hadn't looked at the diff, that empty file would have gone straight to a maintainer under my name.&lt;/p&gt;

&lt;p&gt;The fix is two lines of paranoia:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;changed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--porcelain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;changed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;continue&lt;/span&gt;  &lt;span class="c1"&gt;# patch never applied; tests passing means nothing
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the working tree is clean after "applying" a patch, the patch did not apply. Any test result you get after that is a test of the pristine code, and your pipeline will happily attribute the pristine code's success to a patch that isn't there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The subtle cousin of that bug
&lt;/h2&gt;

&lt;p&gt;There's a sneakier version. &lt;code&gt;git apply --3way&lt;/code&gt; stages successful merges into the index. Plain &lt;code&gt;git diff&lt;/code&gt; shows only unstaged changes, so a successful three-way apply can leave you staring at an empty diff even though the patch landed. Use &lt;code&gt;git diff HEAD&lt;/code&gt; when you're harvesting the final patch for review, not &lt;code&gt;git diff&lt;/code&gt;. Same family of bug: your verification step looking at the wrong layer of git state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bloat is also a failure mode
&lt;/h2&gt;

&lt;p&gt;Once the apply path was honest, a different problem surfaced. The model, when asked to fix a one-line pagination bug in a 58KB React context file, rewrote the whole file. The rewrite typechecked. It also dropped a localStorage token refresh and swapped a PATCH endpoint for a POST against a different URL path. Technically valid TypeScript, functionally a regression.&lt;/p&gt;

&lt;p&gt;The diff was 1,379 lines. The bug was one line.&lt;/p&gt;

&lt;p&gt;Small models don't write surgical patches when given a whole file. They echo the file back with their edit smeared through it, and the echo drifts. So the runner now rejects any patch over 60 changed lines for a bounty that describes a focused bug. If the model can't express the fix in under 60 lines, it doesn't understand the fix, and I don't want it near my GitHub account.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually worked
&lt;/h2&gt;

&lt;p&gt;Two things fixed the rewrite problem.&lt;/p&gt;

&lt;p&gt;First, stop asking for diffs. My model produces valid unified diffs maybe never. Instead I give it the exact file contents and ask for the complete new file back, or for big files, a SEARCH/REPLACE block with lines it must copy character for character. Then I do the replacement in Python and verify the SEARCH text actually exists in the file. If it doesn't, the round fails loudly instead of silently.&lt;/p&gt;

&lt;p&gt;Second, send the whole file or don't send it. I initially capped file context at 20K characters to save tokens. The target file was 58KB. The model invented the back half of the file from vibes, and the invention compiled. Truncation doesn't make a small model careful. It makes it confident and wrong with valid syntax.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where I actually got a PR out of it
&lt;/h2&gt;

&lt;p&gt;After all that, the same pipeline produced two real patches, both verified against the repos' own typecheck with no new errors: a pagination state bug where &lt;code&gt;hasMoreBounties&lt;/code&gt; was computed from a stale closure value, and an unauthenticated DoS in a QR code endpoint where the &lt;code&gt;size&lt;/code&gt; query param went straight into a PNG buffer allocation with no upper bound. Small fixes, real issues, patches under 25 lines each.&lt;/p&gt;

&lt;p&gt;The pipeline that produced those is the same one that produced the empty diff and the 1,379-line rewrite. The difference is a handful of guards that assume the model is wrong until the filesystem proves otherwise.&lt;/p&gt;

&lt;p&gt;If you're wiring a local model into anything autonomous, spend your paranoia budget on the verification path, not the prompt. The prompt is the easy part. The part that checks whether anything actually happened is where your nights get saved or wasted.&lt;/p&gt;

&lt;p&gt;One last thing for anyone running this against real repos under their own name: review every diff yourself before it ships. The guard rails catch the bugs you anticipated. The maintainer reading your PR catches the ones you didn't. Better to have that conversation over a good patch than a fabricated one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
