DEV Community

M Hossein
M Hossein

Posted on

Generate, then decide: using Cloudflare Clef as a decision layer

Most LLM pipelines have one model do two jobs: write the answer, then judge whether it's good. Those are different problems. Writing is open-ended. Judging is a bounded choice: supported or not, safe or not, ship or not.

Clef is Cloudflare's model for the second job. You don't prompt it for prose. You give it a state (text, JSON, images) and a schema of typed questions, and it returns a probability for every allowed option of every question. That changes how you build: the decision becomes a structured object your code can check, not a sentence you have to parse.

I built a small proof of concept, Evidence Lab, to see what this looks like in a real pipeline: checking whether a RAG answer is actually supported by its sources.

The problem: right citation, wrong meaning

Suppose a policy says members can return unopened items. An answer that says all customers can return any item can cite exactly the right document and still be wrong. Citation checks pass. The meaning doesn't.

You can't catch this with string matching, and asking the same LLM "is your answer correct?" is weak evidence. What you want is a separate classifier that reads the evidence and the claim and commits to a label.

Clef in 30 seconds

A Clef call has three parts:

  • state: what to evaluate
  • questions: up to 64 typed questions, each with an ID
  • answers: one per question, keyed by the same IDs, with probabilities

There are three question types: noul (a yes/no question returning a probability), choice (pick one of several labelled options) and score (an ordered scale). Here is Cloudflare's own triage example in short form:

const res = await env.AI.run("@cf/cloudflare/clef", {
  model: "clef",
  state: "Checkout has been failing for every customer for the last hour.",
  questions: {
    team: {
      type: "choice",
      instructions: "Which team should handle this request?",
      criteria: {
        billing: "Payments, invoices, and refunds",
        technical: "Outages, errors, and configuration",
        sales: "Plans and upgrades",
      },
    },
  },
});
// res.answers.team -> chosen team + per-option probabilities
Enter fullscreen mode Exit fullscreen mode

Notice what's missing: no output parsing, no "respond only in JSON," no retry when the model chats back. The schema is the contract.

Applying it: a verification gate for RAG

In Evidence Lab, the generator writes the answer and Clef decides whether it can be shown. Application code, not a model, makes the final call.

The diagram follows one question through the pipeline. A corpus is the collection of documents being searched. An evidence pack is the set of excerpts retrieved for this question. Keeping that pack fixed means the answer is checked and repaired against the same material it was written from.

flowchart TD
  Question[Question] --> Retrieve[Retrieve and freeze evidence]
  Retrieve --> Draft[Generate cited answer blocks]
  Draft --> Structure[Check structure and citations]
  Structure -->|Valid| Clef[Clef: check support and global criteria]
  Structure -->|Invalid| Failure[Technical failure]
  Retrieve -. Same evidence .-> Clef
  Clef --> Verdict[Validate complete verdict]
  Clef -->|Malformed response or timeout| Failure
  Verdict -->|Incomplete or mismatched| Failure
  Verdict -->|Checks pass| Gate[Publication gate]
  Verdict -->|Unsupported, repair available| Repair[One repair]
  Repair --> Structure
  Verdict -->|Repair exhausted| Abstain[Abstain]
  Gate -->|Qualified live policy| Response[Release response]
  Gate -->|Shadow or unqualified live policy| Review[Operator trace only]

Here is what each block does:

  1. Question: the pipeline receives the user's question and the selected corpus.

  2. Retrieve and freeze evidence: retrieved excerpts are pinned with their IDs and versions. Generation, verification and repair all use this same pack. Otherwise you'd be verifying against material the generator never saw.

  3. Generate cited answer blocks: the LLM writes structured blocks, each with citation IDs. A block can contain several factual claims. This is the open-ended "generate" half.

  4. Check structure and citations: plain code validates the schema, the citation IDs and any quoted text. Invalid drafts stop here, before spending a Clef call.

  5. Clef: check support and global criteria: this is the "decide" half. The question, the exact draft and the frozen evidence go to Clef as the state, and each block becomes one choice question with four options:

   questions[`block.${i}`] = {
     type: "choice",
     instructions: `Is block ${i} supported by the cited evidence?`,
     criteria: {
       supported: "Every claim follows from the cited excerpts",
       contradicted: "An excerpt says the opposite",
       insufficient_evidence: "The excerpts don't establish the claim",
       conflicting_evidence: "Excerpts disagree on this point",
     },
   };
Enter fullscreen mode Exit fullscreen mode

Three global questions go into the same request:

  • Does the answer stay within the task's scope?
  • Is it internally consistent?
  • Does any retrieved excerpt contradict it, including excerpts the answer didn't cite?

That last check catches an answer that quietly ignored inconvenient evidence. Clef accepts up to 64 questions per call, so the whole verdict comes back in one round trip.

  1. Validate complete verdict: don't trust the verdict blindly. Code requires every expected question to be answered exactly once, then checks labels, probabilities, hashes of the answer and evidence, and the round ID. To pass, every block must be supported and every global check must pass. The probabilities are not calibrated yet, so any threshold on them needs evaluation first.

  2. One repair: a rejected draft gets exactly one rewrite, using the failed check IDs and the same evidence. The new draft goes back through structural checks and Clef.

  3. Abstain: if the repair also fails verification, the system returns an abstention. The draft is never published.

  4. Technical failure: the run stops on any of these: invalid data, missing checks, hash mismatches, timeouts or exhausted budgets. A malformed verdict is treated as a failure, never as a pass. Remote calls reserve their spend before sending, and transport retries have their own separate limits.

  5. Publication gate: passing checks is necessary but not sufficient. A live release also requires a qualified policy that matches the current configuration and implementation. "Qualified" means the release settings are tied to reviewed evaluation evidence.

  6. Release response: once the gate passes, the accepted blocks and their citations are published in the normal answer field.

  7. Operator trace only: in shadow mode, or under a live policy that isn't qualified, drafts and verdicts are kept for inspection. Users see no answer.

The general pattern

Strip out the RAG details and the design is reusable:

  1. An LLM generates (open-ended, creative, expensive to reason about).
  2. Clef decides (bounded labels, probabilities, one call).
  3. Your code enforces (validation, thresholds, retries, release policy).

Where I think this fits beyond RAG (ideas, not things I've tested):

  • Support triage: route, flag urgency and score severity in one call (Cloudflare's own example).
  • Agent guardrails: before an agent runs a tool, ask whether the action matches the user's request and whether it's reversible.
  • Content moderation: multi-label choice questions instead of one fuzzy "is this OK?"
  • Extraction QA: check that a field pulled from a document actually appears in it.
  • Screenshot and UI checks: Clef accepts images, so "does this screen show an error state?" is a valid question.

The rule of thumb: if the decision can be written as a fixed set of options with clear definitions, it belongs in Clef, not in a second free-text prompt.

What this PoC does not prove

I'm being explicit here because this is the part that usually gets skipped:

  • The probabilities aren't calibrated. A 0.9 for supported doesn't mean "right 90% of the time" until it's measured. Thresholds need evaluation data.
  • Quality is unmeasured. I haven't yet run held-out, human-reviewed cases to see how often the gate lets bad answers through or blocks good ones.
  • Verification can't fix retrieval. If the right document was never retrieved, Clef can only say "insufficient evidence."

Evidence Lab is a PoC that shows the mechanics: frozen evidence, structured checks in one call, a single repair, and an explicit release decision, with automated tests for response validation, evidence binding, and failure handling.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

Treating verification as a classification task with explicit probabilities is so much cleaner than asking the generator to grade its own homework in prose. The practical win is being able to set numerical threshold gates in application code instead of hoping a model follows a JSON prompt. Where I have seen this split run into trouble is chunk fragmentation across the evidence pack. If the generator synthesizes a conclusion across two separate excerpts, a classifier evaluating chunks in isolation often flags a hallucination on a claim that was valid across their union. Passing the concatenated pack with clear delimiter tags usually prevents the gate from becoming overly trigger-happy.