DEV Community

Cover image for Supercharge Your LLM Router with a Decision Model: Clef on Cloudflare Workers
Alexey Sazhin
Alexey Sazhin

Posted on

Supercharge Your LLM Router with a Decision Model: Clef on Cloudflare Workers

In part 1 we built a classification router in about 50 lines on a Cloudflare Worker. It classifies each prompt and routes it through AI Gateway: coding questions to Claude Sonnet, everything else to a cheap Workers AI model. The classifier was a small LLM, Llama 4 Scout, with a one-line system prompt: "Reply with only that single word."

Since then, a new kind of model has appeared for exactly this job: decision models. TypeSafe introduced Jev, and on October 1 Cloudflare published Clef and Clef-flash on Workers AI. A decision model doesn't write an answer. It takes your content and a typed question and returns the choice directly, with no text to generate token by token. For a router that only needs one word, coding or simple, that looks like a perfect fit.

So in this post we'll bolt Clef and Clef-flash onto the existing router as drop-in classifiers, then test them against Llama on the same labelled prompts: which one routes each prompt to the right model?

Disclosure: I work at Cloudflare. This is a personal side project; opinions and measurements are my own.

Make the classifier testable

Before changing the classifier, I wanted a way to score it. Two small additions made that possible.

1. A /classify endpoint. POST /classify runs only the classifier and returns its decision as JSON. It doesn't call Claude, so testing costs nothing beyond the classifier itself:

curl https://prompt.demolabs.fyi/classify \
  -H "x-classifier: clef-flash" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "What is the capital of Portugal?"}]}'


{"task":"simple","raw":"simple","confidence":0.9117,
"fallback":false,"classifier":"clef-flash",
"model":"@cf/cloudflare/clef-flash","ms":734}
Enter fullscreen mode Exit fullscreen mode

2. A benchmark script. scripts/bench.mjs sends 20 labelled prompts (10 coding, 10 general, a few of them long) through /classify N times and reports accuracy against the labels, listing every prompt it got wrong:

npm run bench -- \
  --url https://prompt-router.<your-subdomain>.workers.dev \
  --classifier llama-4-scout,clef --runs 10
Enter fullscreen mode Exit fullscreen mode

The x-classifier request header picks the classifier per request, so one deployment can compare several models on identical prompts.

Example of the benchmark script's output:

 npm run bench -- --url https://prompt.demolabs.fyi --classifier llama-4-scout,clef,clef-flash --runs 10

Endpoint: https://prompt.demolabs.fyi/classify
Prompts: 20 × 10 runs, concurrency 1, warmup 2

=== llama-4-scout ===
  20/20 requests
  model: @cf/meta/llama-4-scout-17b-16e-instruct
  latency (ms)               n   min   p50   p90   p95   p99   max  mean
  server classify           20   111   129   267   801  4117  4117   372
  client round-trip         20   141   164   293   829  4142  4142   402
  accuracy: 75.0% (15/20)
  ✗ expected simple, got coding (raw: "coding") — Translate 'Good morning, how are you?' into German.
  ✗ expected simple, got coding (raw: "coding") — Summarize the plot of Hamlet in two sentences.
  ✗ expected simple, got coding (raw: "coding") — I have a job interview tomorrow for a solutions architect ro
  ✗ expected coding, got simple (raw: "**coding**\n\n## Explanation\n\nIn") — Explain the difference between useEffect and useLayoutEffect
  ✗ expected coding, got simple (raw: "To classify the user based on ") — Write a Dockerfile for a Node 20 Express app that runs as a
  saved → \Projects\prompt-router\bench-results\llama-4-scout-2026-10-08T21-56-12-168Z.json

=== clef ===
<...>
Enter fullscreen mode Exit fullscreen mode

What the LLM classifier got wrong

The first run of Llama 4 Scout scored 75% accuracy: one prompt in four went to the wrong model. The failures came in two kinds.

It ignored "one word". For some coding prompts, Llama answered like a chatbot:

// "Explain useEffect vs useLayoutEffect"
raw: "**coding**\n\n## Explanation\n\nIn..."

// "Write a Dockerfile for a Node 20 app"
raw: "To classify the user based on ..."
Enter fullscreen mode Exit fullscreen mode

My parser expected exactly coding, so both went to the cheap model: the two prompts that most needed Sonnet. It also means the router spent its time generating a paragraph nobody reads.

It was confidently wrong. Three general prompts came back as a clean coding:

  • "Translate 'Good morning, how are you?' into German."
  • "Summarize the plot of Hamlet in two sentences."
  • "I have a job interview tomorrow for a solutions architect role…"

Those aren't expensive mistakes (they just cost Sonnet prices), but they repeated on every run.

Patching it only goes so far

The obvious fix is a safer default: only an explicit simple goes to the cheap model, and anything else (rambling, empty, unexpected) goes to the strong one.

const isSimple = raw.trim().toLowerCase() === "simple";
const task: Task = isSimple ? "simple" : "coding";
Enter fullscreen mode Exit fullscreen mode

That lifted Llama to 85%: the rambling answers now route correctly. But the three confident mistakes stayed, on every run. A safe default can catch a model that is unsure or off-format. It can't catch a model that is sure and wrong.

Decision models: answers, not text

A chat model is asked to write a classification and hopes it complies. Clef is a decision model: you give it the content and typed questions, and it returns typed answers. There is no generated text, so there is nothing to ramble and nothing to parse.

A request has two parts:

  • state: the content to judge (here, the user's prompt).
  • questions: up to 64 named questions, each one of three types: noul (yes/no, returns a probability), choice (pick one option from a set), or score (place on an ordered scale).

For routing, one choice question is enough. The options are object keys, and each value describes when that option applies:

const CLEF_PROMPT = "Which model tier should handle this user request?";
const CLEF_CRITERIA = {
  coding:
    "Writing, debugging, reviewing or explaining code, " +
    "scripts, SQL, regex, configs or developer tooling",
  simple:
    "General knowledge, writing, translation, advice " +
    "and everything else",
};
Enter fullscreen mode Exit fullscreen mode

The answer comes back as the chosen key, a probability per option and a confidence between 0 and 1. It can only be coding or simple, by construction. There is also Clef-flash, a smaller, faster variant with the same API.

The criteria descriptions do the work that my system prompt failed at. "Translation" and "writing" are listed under simple explicitly, which is exactly where Llama went wrong.

The swap

To compare models fairly, the classifier became a registry. Every entry takes the prompt and returns the same small shape, { raw, confidence? }, and parsing into a route happens in one place:

interface ClassifierOutput {
  raw: string;
  confidence?: number; // only classifiers that report one set it
}

interface Classifier {
  model: string;
  run(env: Env, prompt: string): Promise<ClassifierOutput>;
}
Enter fullscreen mode Exit fullscreen mode

The Clef entry is about 15 lines:

"clef": {
  model: "@cf/cloudflare/clef",
  async run(env, prompt) {
    const out = (await env.AI.run(this.model, {
      model: "clef", // must match the model ID ("clef-flash" for Flash)
      state: prompt,
      questions: {
        classify: {
          type: "choice",
          instructions: CLEF_PROMPT,
          criteria: CLEF_CRITERIA,
        },
      },
    })) as ClefResult;

    return {
      raw: out.answers?.classify?.choice ?? "",
      confidence: out.answers?.classify?.confidence,
    };
  },
},
Enter fullscreen mode Exit fullscreen mode

Two gotchas I hit. The text goes in state, not in a messages array; leave it out and Clef has nothing to judge. And the body's model field has to match the model ID: calling @cf/cloudflare/clef-flash with model: "clef" throws, which surfaces as a bare error code: 1101 unless you catch it.

When Clef isn't sure, pay for Sonnet

The two possible mistakes don't cost the same. A coding question sent to the cheap model gets a bad answer, which users notice. A simple question sent to Sonnet costs a little more and is still answered well. So low confidence routes to the strong model:

const CONFIDENCE_THRESHOLD = 0.7;

const lowConfidence =
  confidence !== undefined && confidence < CONFIDENCE_THRESHOLD;
const isSimple = raw.trim().toLowerCase() === "simple";

const task: Task = isSimple && !lowConfidence ? "simple" : "coding";
Enter fullscreen mode Exit fullscreen mode

The result carries a fallback flag, so the logs show how often the safety net fires. If it's often, the threshold is too strict and you're paying Sonnet prices for easy prompts. Llama reports no confidence, so for it the rule reduces to the safe default from earlier.

The router picks the classifier from the x-classifier header, then the CLASSIFIER var, then a default in code. That's what made the side-by-side benchmark possible without redeploying.

Results

Both Clef models routed all 200 test requests correctly. The patched Llama still misrouted 3 of the 20 prompts, every time.

Classifier Accuracy Prompts misrouted (of 20)
Llama 4 Scout, strict parsing 75% (150/200) 5
Llama 4 Scout, safe default 85% (170/200) 3
Clef 100% (200/200) 0
Clef-flash 100% (200/200) 0

How I tested: 20 labelled prompts × 10 runs = 200 requests per classifier, sent through /classify:
npm run bench -- --url https://prompt.demolabs.fyi --classifier llama-4-scout,clef,clef-flash --runs 20
The mistakes were the same prompts on every run, so this is about which prompts each model gets wrong, not random noise.

A few things stand out:

  • The criteria descriptions matter. Llama's three remaining mistakes were translation, a summary and interview prep. All three are named under simple in Clef's criteria, and Clef got them right.
  • Clef vs Clef-flash was a tie for this task. Both scored 100%. For a two-way coding/simple split either works; a harder, multi-category task is where the bigger model should earn its keep.
  • 20 prompts is a small test set. 100% here means "no mistakes on these 20, repeated 10 times", not "never wrong". The confidence fallback is there for the prompts I didn't think of.

Takeaways

  • Test the classifier, not just the router. Routing looked fine in manual tests. A labelled prompt set and a small bench script showed one prompt in four going to the wrong model.
  • Don't ask a text generator for a label. "Reply with one word" is a request, not a guarantee. A decision model can only return one of the options you defined.
  • Make "unsure" fail towards quality. Only an explicit simple goes cheap. Everything else, including low confidence, goes to the strong model.
  • To be fair to Llama: Clef also got a better prompt, with a description per option. A tighter Llama prompt would likely fix some of its confident mistakes, but not the off-format answers.

Try it

The code is on GitHub: palermo-777/prompt-router. The part 1 version is still at the part-1 tag.

git clone -b part-2 https://github.com/palermo-777/prompt-router.git
cd prompt-router && npm install

# copy the config, then fill in your gateway details
cp wrangler.example.jsonc wrangler.jsonc

npx wrangler secret put AI_GATEWAY_TOKEN
npm run deploy

npm run bench -- \
  --url https://prompt-router.<your-subdomain>.workers.dev \
  --classifier llama-4-scout,clef,clef-flash
Enter fullscreen mode Exit fullscreen mode

The Worker doesn't authenticate its callers, so anyone who finds the URL can spend your Anthropic credits. Put Cloudflare Access or a bearer-token check in front of it before you share it.

What's next

  • More routes. Clef takes up to 64 questions per call, so reasoning, creative or needs-tools can be added as options or as separate questions, all answered in the same call.
  • A harder test set. Borderline prompts ("explain what an API is", "write a cover letter for a developer job") are where confidence and the threshold actually matter.

If you have prompts that trip up your classifier, share them in the comments. I'd like to grow the test set with real-world edge cases.

Top comments (0)