DEV Community

Cover image for Cloudflare Launched Clef. Let's Build a Tiny Decision Gate in TypeScript.
Bobby Hall Jr
Bobby Hall Jr

Posted on

Cloudflare Launched Clef. Let's Build a Tiny Decision Gate in TypeScript.

On September 30 I published OpenAI Launched Dots and built a tiny always-on agent.

That agent has rules. Allow. Ask. Deny.

The rules are still handwritten.

On October 1, Cloudflare shipped a model whose whole job is the decision.

Clef.

  • On October 1, Cloudflare released Clef and Clef-flash, the first models trained by the Workers AI team. They are decision models. You pass a state and a set of typed questions. They return a probability for every allowed answer. No generated sentence. No tool-call text to parse.
  • Clef is a 27B model post-trained on a frozen Qwen3.8-27B. Clef-flash is 9B, on Qwen3.5-9B. Weights are Apache 2.0 on Hugging Face. The hosted names are @cf/cloudflare/clef and @cf/cloudflare/clef-flash.
  • Workers AI lists unit pricing at $0.24 per million input tokens for Clef and $0.09 for Clef-flash. Those pages quote input tokens. I am not going to invent an output price.
  • Cloudflare's own latency table, 43 runs, puts median latency at 209.3 ms for Clef, 38.8 ms for Clef-flash, and 524.1 ms for Typesafe's Jev. That is their number. About 2.5x and 13x against Jev at the median. p95 is 238.6 ms, 122.4 ms, and 536.0 ms.
  • In that same table, a Clef model is highest on 7 of 10 decision benchmarks. Jev still leads When2Call (80.97 vs 72.37) and BRIGHT (47.52 vs 45.91). DiffusionGemma Jev leads PhishNChips. On CLINC150 with out-of-scope, Clef scores 97.43 and Clef-flash scores 66.77. The fast model is not the careful one.
  • The API is the System One shape Jev already uses. Three question types: noul (yes or no), choice (one option from a set you wrote), and score (a rung on a rubric). Up to 64 questions. Context window on the model pages is 65,536 tokens. Clef also takes up to four images. Jev, today, is text.

Different product from dots.

Same place in the stack.

Something is true about the world.

A score comes back.

Your code decides whether that score is allowed to move.

So let's build a tiny decision gate.

By the end, you'll run one command:

npx tsx gate.ts
Enter fullscreen mode Exit fullscreen mode

And watch two support messages get scored, routed, or held.

No API key.

No real model.

Just TypeScript.

One honesty note: this is not Clef. Clef is a trained model with a vision encoder, a prefill-only pass, and a scoring head. Cloudflare describes that in the launch post. The scorer below counts keywords and turns the counts into shares with a softmax. The policy around it is the thing worth keeping.

What We Are Building

A decision gate turns support message state into scores, passes those scores through code-owned policy, then chooses act, ask, or skip.

A gate with two jobs.

  1. Score only the answers you already listed.
  2. Turn those scores into act, ask, or skip with thresholds that live in code.

The example is a checkout inbox. The messages are made up.

Real Clef would take natural-language criteria. This toy takes keyword lists, so the lesson runs offline.

Project Setup

You will need Node.js 18 or newer.

mkdir tiny-gate
cd tiny-gate
npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the blocks below, in order, as gate.ts.

Step 1: Score Only Allowed Answers

type Noul = { type: "noul"; yesIf: string[] };
type Choice = { type: "choice"; options: Record<string, string[]> };
type Question = Noul | Choice;

type Answer =
  | { type: "noul"; yes: number }
  | { type: "choice"; pick: string; probs: Record<string, number> };

const round = (n: number) => Math.round(n * 100) / 100;

function hits(text: string, needles: string[]) {
  const hay = text.toLowerCase();
  return needles.filter((needle) => hay.includes(needle)).length;
}

function share(weights: number[]) {
  const lifted = weights.map((weight) => Math.exp(weight));
  const total = lifted.reduce((sum, value) => sum + value, 0);
  return lifted.map((value) => value / total);
}

function decide(question: Question, state: string): Answer {
  if (question.type === "noul") {
    const [yes] = share([hits(state, question.yesIf), 1]);
    return { type: "noul", yes: round(yes) };
  }

  const names = Object.keys(question.options);
  const probs = share(names.map((name) => hits(state, question.options[name])));
  const scored = Object.fromEntries(names.map((name, i) => [name, round(probs[i])]));
  const pick = names.reduce((best, name) => (scored[name] > scored[best] ? name : best));
  return { type: "choice", pick, probs: scored };
}
Enter fullscreen mode Exit fullscreen mode

noul is Cloudflare's name for a yes-or-no question. The yes score is a share between "how many yes-phrases matched" and a prior of 1 for no. If nothing matches, yes lands near 0.27, not 0. That is a prior, not confidence.

choice can only return a key you put in options. There is no free-text team name for a later regex to miss.

The model, even this fake one, does not pick the action. It returns numbers.

Step 2: Put the Thresholds in Code

type Action = "act" | "ask" | "skip";

function policy(id: string, answer: Answer): { action: Action; reason: string } {
  if (id === "refund" && answer.type === "noul") {
    return answer.yes >= 0.4
      ? { action: "ask", reason: "money stays with a human" }
      : { action: "skip", reason: "no refund was asked" };
  }

  if (answer.type === "noul" && id === "urgent") {
    return answer.yes >= 0.8
      ? { action: "act", reason: `urgent ${answer.yes} clears 0.80` }
      : { action: "skip", reason: `urgent ${answer.yes} stays under 0.80` };
  }

  if (answer.type === "choice" && id === "team") {
    const confidence = answer.probs[answer.pick];
    return confidence >= 0.6
      ? { action: "act", reason: `${answer.pick} ${confidence} clears 0.60` }
      : { action: "ask", reason: `${answer.pick} ${confidence} is too close to call` };
  }

  return { action: "skip", reason: "no rule" };
}

const questions: Record<string, Question> = {
  urgent: { type: "noul", yesIf: ["failing", "every customer", "last hour"] },
  team: {
    type: "choice",
    options: {
      billing: ["refund", "invoice", "payment"],
      technical: ["failing", "every customer"],
      sales: ["plan", "upgrade", "pricing"],
    },
  },
  refund: { type: "noul", yesIf: ["refund"] },
};

const states = [
  "Checkout has been failing for every customer for the last hour. Please refund order 4412.",
  "Love the new checkout page. Can someone send me the enterprise plan?",
];

for (const state of states) {
  console.log(`\nSTATE  ${state}`);
  for (const [id, question] of Object.entries(questions)) {
    const answer = decide(question, state);
    const gate = policy(id, answer);
    const detail =
      answer.type === "noul"
        ? `yes ${answer.yes.toFixed(2)}`
        : `${answer.pick} ${answer.probs[answer.pick].toFixed(2)}`;
    console.log(`${id.padEnd(8)} ${answer.type.padEnd(7)} ${detail.padEnd(18)} ${gate.action.padEnd(4)} ${gate.reason}`);
  }
}
Enter fullscreen mode Exit fullscreen mode

Read the refund branch first.

A refund score of 0.50 does not issue a refund. It asks.

A refund score of 0.27 does not ask either. Nobody asked for a refund. Waking a human for a question the message never raised is how inboxes die.

Urgency can act, because the action on the other side of that score is a route, not a charge.

Team can act only when the winning option clears 0.60. Under that, the gate asks. A 58% sales guess is a guess.

An unknown or close action should cost you a click, not an incident.

That line is the same rule as the dots post. The difference is where the number comes from. There, the rule was a map of tool names. Here, the rule is a threshold on a score.

Run It

npx tsx gate.ts
Enter fullscreen mode Exit fullscreen mode

You should see:

STATE  Checkout has been failing for every customer for the last hour. Please refund order 4412.
urgent   noul    yes 0.88           act  urgent 0.88 clears 0.80
team     choice  technical 0.67     act  technical 0.67 clears 0.60
refund   noul    yes 0.50           ask  money stays with a human

STATE  Love the new checkout page. Can someone send me the enterprise plan?
urgent   noul    yes 0.27           skip urgent 0.27 stays under 0.80
team     choice  sales 0.58         ask  sales 0.58 is too close to call
refund   noul    yes 0.27           skip no refund was asked
Enter fullscreen mode Exit fullscreen mode

The same three questions produce different permissions for a checkout outage and an enterprise plan request. The outage routes and holds the refund; the plan request waits for a person.

Same questions.

Different states.

Different permissions.

The outage gets a technical route. The refund waits. The compliment does not page anyone. The plan question is close enough that a person should see it.

Where It Breaks Down

This is a teaching gate. Here is what a real one still needs.

The Score Is Not a Probability

share is a softmax over keyword hits. Two phrases from the same sentence, "failing" and "every customer," both count. That is double counting, not evidence.

Clef is trained with a Brier loss so the number is closer to a calibrated probability. Cloudflare says so. I have not checked the calibration. Treat every vendor probability as a score until your own labels say otherwise.

The Fast Model Drops the Weird Cases

On CLINC150 with an out-of-scope bucket, Cloudflare's table shows Clef at 97.43 and Clef-flash at 66.77. Jev sits at 89.27.

If you put the 38.8 ms model on "is this even our problem," you will auto-route things that never belonged in the taxonomy. Use flash on the hot path after a narrower question. Use the bigger model, or a human, when the label set can be wrong.

Their Leaderboard Is Theirs

Seven of ten is Cloudflare's count, on Cloudflare's runs, published the day of the launch. Jev still wins the "should you call a tool" benchmark, When2Call. On Typesafe's own workflow suite, Clef wins invoice processing, customer service, and security incidents, and loses agent-trace observability (Jev 71.6, Clef 68.5).

Laya, another model in the latency table, has a 5.8 ms median. It also scores 14.29 on BANKING77, where Clef scores 94.20. Fast and empty is still empty.

Ship the gate. Then put yesterday's scope eval on the trace. A decision model that was never graded on your tickets is a new way to be confidently wrong.

Thresholds Are Product Decisions

Demo thresholds: urgency routes at 0.80, a team routes at 0.60, and a refund always waits for a human.

0.80 and 0.60 are demo numbers. They are not a research result.

Move 0.60 to 0.90 and the technical route starts asking. Move the refund line from ask to act and you have built a machine that spends money because a score cleared a constant you typed on a Thursday.

Write the constant next to the action it unlocks. Review it like a permission, not like a hyperparameter.

The Bigger Idea

The harness post split "what the model wants" from "what the process allows."

The dots post split background mode from assigned work.

This one splits the decision from the generator.

A chat model is a bad policy engine. It can explain a refund, talk itself into a refund, and format the refund as JSON, all in one completion. You then parse the JSON and hope the sentence was wise.

A decision model is narrower. The answers exist before the call. The model distributes a score across them. Your code holds the line between a score and a side effect.

State ──→ scored questions ──→ probabilities
                                      │
                                your thresholds
                                      │
                                act / ask / skip
                                      │
                         the model does not cross this line
Enter fullscreen mode Exit fullscreen mode

Cloudflare's interesting claim is that this scoring pass can sit on the hot path. Their threat-intel note says Clef, paired with Browser Run, fetched, rendered, and classified a domain in 2.2 seconds, against 4.7 seconds for gpt-oss-120b in the same workflow. That includes the fetch. It is one internal example. It is still the right shape. Classify in milliseconds. Spend the large model only when the score says the work is worth it.

They also opened a fine-tune path, first with their forward-deployed team, later as a self-serve RL product. That part is a design-partner form, not a button I would tell you to press today.

I think the pattern is clear anyway. A probability is not a permission.

The generator can propose.

The decision model can score.

The threshold is yours.


Try Roster

I'm building Roster around this idea: AI employees with a lane, tools, and a rule for anything that spends money, sends a message, or is hard to undo. Preparation can move. The consequential step waits.

If the same triage, follow-ups, and "should we even do this?" loops keep eating your week, give them to an AI employee.

Try Roster →

Top comments (0)