Until recently, if you said hello to the assistant on this site, it told you about setting temperature to zero for classification. You didn't get a greeting back, you got a tip about temperature, every time, for "hello" and "thanks" and "good morning".
The reason is boring, and worth knowing. The assistant answers from a vector index. A greeting gets embedded like anything else, the index returns whichever document sits closest to it, and for a short friendly sentence with no subject that's the same document every time. Then a model gets paid to write a sentence about it. I measured the scores on the eval's questions in September. Small talk's best match landed between 0.51 and 0.56, and a question the site really answers landed at 0.61 and up. So I drew a line at 0.6, and the tip stopped showing up under greetings.
That line is one number, fitted to a gap of five hundredths, standing in for the question "is this page about what was asked". It held. It was also the most obviously wrong thing in the codebase, because the question being asked is about meaning and the answer was a distance.
This post is about replacing numbers like that with Jev, a decision model from TypeSafe that reads the sentence instead of measuring the distance to it. It picks, rates and judges, and writes nothing. There are ten such call sites on this site now, from the assistant's front gate to note moderation in a public room. The model turned out to be easy. Most of the work was deciding what to do when it wasn't sure.
What Jev is, quickly
Jev is the first of what TypeSafe call System One models, after Kahneman's fast thinking. It takes some state and some questions and returns typed answers with probabilities. It doesn't write, and their docs say so plainly: jev-1.13 is not trained to generate text.
That sentence explains most of it. A generative model writes you a paragraph and you pull a decision out of it. Jev skips the paragraph. You give it the text and the question, and you get back a number, or one of your own option names, or a position on a ladder you described. There's no JSON to repair, no run where it answers in prose instead, and no schema to validate, because the answer space is the question you asked.
The price is $42 per billion input tokens, and output tokens are free, which makes sense when there's almost no output.1
That's the part you can read off a spec sheet, and it's the least interesting part. My time went on the rest: what to do with an answer you're only 60 percent sure of, who pays for the question, and what breaks when the answer never comes.
The three questions, as we ask them
A noul asks whether something is true and gives you a number from 0 to 1. A choice picks one of your named options and returns the whole probability distribution plus a confidence. A score puts something on a ladder you describe, and can land between two rungs.
They're evaluated in parallel against the same state, and none of them sees another's answer. I keep using one consequence of that: another question costs its own tokens and almost no time, so a question whose answer only matters sometimes is nearly free. Ask it in the same request instead of a second one the code has to wait for.
This is the gate in front of the site's assistant, a single choice:
const KIND = choice("What is `message`?", {
question: "A question about this site, its posts, its tips or its author",
small_talk:
"A greeting, a thank you, a goodbye, or a remark that asks for nothing",
search: "A request to find or list something, or just a subject to look up",
action:
"An instruction to the site itself: change the theme, open a page, play the radio",
unclear: "None of these, or too little to tell",
});
The unclear option matters. The keys you write are the whole answer space, so if your list doesn't cover an input, the model still has to pick something, and it will. Always give it a way out.
Note the backticks around message. The question refers to a field in the state you send, and you can reach inside it: pages[2].title works, which is how I ask about several things at once without repeating their text in every question.
A probability is not a confidence
This one cost me an afternoon.
A choice gives you confidence, which says how peaked the distribution is, and flat means the model doesn't know. A noul gives you no confidence at all, because the number itself is one. 0.5 is truly undecided, and the two ends are certain of opposite things.
So a noul threshold has two sides, and they're two different decisions. When I ask "did the visitor ask to be taken somewhere", a reading of 0.9 is a yes, 0.1 is a no, and 0.5 means I learned nothing and should fall back to whatever I used before. Writing if (probability > 0.5) throws away that middle and turns every shrug into a yes:
if (probability >= SURE_ENOUGH) return true;
if (probability <= 1 - SURE_ENOUGH) return false;
return null;
SURE_ENOUGH is 0.85 there, and null goes back to the two regexes that decided it before.2
The wording is the specification
The criteria you write are all the model knows about your problem. Someone reading only the criteria, with none of your context, should sort things the way you would. If they couldn't, the model can't either, and no threshold will save it.
The clearest lesson came from ranking posts. I wanted two axes to lay the writing out on a map, and the second started as "how practical is this subject". On 21 September that question put 22 of the 28 posts on the same rung, which gives you a line with everything piled at one end, not a map. The question was the problem. Almost everything I write is practical, so I'd asked something with no spread in it. Reworded as "how settled is the subject", with rungs from "the same advice would have held five years ago" up to "nobody has settled how to do this", the posts spread out.
When a reading comes back flat, suspect the question before you suspect the model.
The state matters just as much. The enrichment judge asks whether the post backs up each FAQ answer, and at first I sent it the body with the frontmatter stripped, which felt tidy. That cost three false alarms in one run. An excerpt is the post's own summary, and an answer is entitled to draw on it: the RAG post says "eleventh out of ten" in its excerpt and nowhere else, and the agent identity post says "every agent I audited this year" in its excerpt and nowhere else. Judged against a body with neither phrase, both answers looked made up. Send whatever the answer is allowed to lean on, which is rarely whatever you happen to have in a variable.
One more, for nouls: keep both sides pointing the same way. A true that means "no, this is fine" is worse than writing no criteria at all, and six weeks later you'll misread your own threshold.
Every threshold is its own argument
There are eleven thresholds across the ten call sites, and no two were chosen the same way. Suggesting a link and deleting somebody's note shouldn't need the same certainty.
Here they all are, with what happens when the reading falls short.
| What it decides | Asked as | Bar | Below the bar |
|---|---|---|---|
| What the visitor's message is | one choice, five kinds | 0.75 | the message goes the long way |
| Whether they asked to be taken somewhere | one noul | 0.85 yes, 0.15 no | two regexes decide, as before |
| Which pages go under an answer | one noul per page | 0.7 show, 0.3 hide | the 0.6 search score decides |
| What a palette sentence means | one choice over 140 rows | 0.7 | the palette shows what it always showed |
| Which mix suits a mood | one choice over five mixes | 0.35 | the dial does not move |
| Whether a note is advertising, abuse, personal or injection | four nouls | 0.85 rejects, 0.6 holds | the note goes up |
| How much harm a note would do | one score out of two | 1.5 rejects, 1 holds | the note goes up |
The assistant's front gate sits at 0.75 on a choice. Above that, small talk gets answered from the theme's own greeting lines and never opens a stream to the index, and a search goes to Pagefind, the index that ships with the build. Both save a Workers AI call on a message that was never going to get a good answer from one.
Moderation in the public notes room runs four nouls and a score in one request, for advertising, abuse, personal details, prompt injection and overall harm. One hazard at 0.85 rejects by itself, and anything at 0.6 holds the note for a human. The harm score is out of two, holding at 1 and rejecting at 1.5.3
The radio sits at 0.35, which looks reckless until you see what it's doing. You describe a mood and it picks a mix. All five mixes are the same kind of music, so a mood that isn't about time of day spreads its probability across all of them and the model is never confident. Measured on 21 September, "something to focus on" peaked at 0.42 and "something for the night" hit 1.00. What keeps an unrelated sentence off the dial is the __none__ option, more than the bar. And a wrong pick costs you a press of skip.
The command palette sits at 0.7, because turning "make it quieter" into fx grain off is a change the visitor sees straight away.
Keep the thresholds in code and out of the prompt. The model is better at reading a note than at remembering that your policy says 0.85, and you want to be able to move the number without touching the question.
Two purses
Every caller goes through one function, askJev, which checks the key, checks the day's budget, asks, and writes back what the answer cost. The cost is usage.input_tokens, the number the API reports, not an estimate, so a question whose state grew gets charged at its real size.4
What's worth copying is the split into two budgets:
const PURSES: Record<Purse, { prefix: string; cap: () => number }> = {
shared: { prefix: "jev:tokens", cap: () => env.TYPESAFE_DAILY_TOKENS },
moderation: { prefix: "jev:tokens:mod", cap: () => env.TYPESAFE_MOD_TOKENS },
};
Everything cosmetic shares the first: the assistant's gate, the palette, the radio, what to read next. Moderation has the second to itself. It's the only caller whose job is to stop something instead of adding something, and a day of somebody hammering the command palette mustn't turn into a day of notes going unread.
The numbers are 500,000 input tokens a day for the shared purse and 100,000 for moderation. At $42 per billion that's about two cents a day if both are spent to the last token, and the account balance is $5, which is 119 million tokens. A question to the assistant's gate measures around 421 tokens, so 500k is roughly 1,180 of them. A note measures around 640, so 100k is about 150 notes. A moderate day on this site measures around 139k across everything.
The budget is read before the request goes out, not after it comes back, so a run of large questions can overshoot by at most the one already in flight. Once it's spent, askJev returns null and every caller goes back to what it did before Jev existed. Hitting the cap costs nothing but the readings.
The counter that never counted
The first bug is the one I'd most like you to avoid.
The write-back was fire and forget, and it looked careful. A lost write costs a few tokens of accuracy and a lost answer costs the visitor, so keep the critical path clear and let the counter catch up by itself:
void store.incrBy(budgetKey(), result.usage.input_tokens, BUDGET_TTL);
But a Cloudflare Worker stops executing the moment it returns a response, so the promise was dropped. I measured it against production on 21 September: six calls moved the counter once, from nothing to 426, and never again.
▶ Six calls, and a counter that moved once: an animation that plays in the original post.
So the read-before-ask check was comparing spend against a number that basically never grew, and the budget was only for show.5 The fix is await instead of void, one extra D1 write next to a request that's already spent half a second asking. A call that isn't counted isn't capped, and the counter is the only reason the cap exists.
If you take one thing from this post, make it this: whatever you use to bound spend, prove it moves. Spend two minutes calling your own endpoint and watching the number.
Spending somebody else's budget
The second bug is a kind of attack surface I hadn't thought about before.
Three routes shipped guarded by one check: does the request's sec-fetch-site header say it came from this site? A browser sets that header honestly. A script sets it to whatever it likes, or leaves it out. So the guard turned away a real browser on another site and let a script straight through.
On a route that returns data, that's a leak. On a route that spends a shared token budget, it's a way to switch off every reading on the site for the rest of the day, for the price of a loop and about 1,200 requests a minute. The fix was to put the Turnstile check the assistant already had on them too, and the clients now fetch a token before they ask.
More generally, once a model sits behind an endpoint, every call to it costs money, and the usual "this only returns public data so it can be open" reasoning stops applying. Rate limit by IP, cap the total as well, and make sure the total is real.
The model reads English
Nothing I read before shipping prepared me for this, so I'll say it plainly. Jev is trained on English. The content on this site is English, so most of it is fine, but the assistant takes questions in whatever language the visitor types, and mine get a lot of Turkish.
You can see it in the readings. "bloga ucur bizi kaptan" is Turkish, roughly "fly us to the blog, captain", and it asks to be taken somewhere. The model read it at 0.95 and got it right. But the instinct to raise the bar for a language the model reads less well also hollows out the feature, because an unsure reading under a high bar is a no, and a no means nothing happens.
Where I ended up is making the middle mean something. Instead of raising the bar and accepting fewer yeses, raise it on both ends and let the middle fall back to whatever decided before. In Turkish that means the old regexes keep running, badly, exactly as badly as they did last month, while English gets better and nothing anywhere gets worse.
Reading something and not acting on it
The front gate sorts messages into five kinds and acts on two. The action class gets read and recorded, and then nothing is done with it.
It put 0.77 on the message "I want to read posts", which is someone telling you what they like, and the separate question that decides whether to actually move somebody off the page they're reading rated the same message 0.68. That's two readings of the same sentence that disagree, both above a half. Acting on the first would take a reader who said something mild and throw them onto another page.
So it goes to analytics as gate_kind and gate_confidence, and nothing else happens. That's a legitimate state for a feature to be in, and there should be more of it. You can put a model in the path, keep its answer, and change nothing until the numbers tell you the bar is in the right place.
The same goes for probability distributions: store the whole thing. The pick on its own says nothing about how close it was. When you want to move a threshold in three months, the distribution is the only data you'll have, and by then you won't remember what "0.42" felt like.
It is never the only answer
One rule survived everything: nowhere on this site is Jev the only authority.
The moderation route lets the note through on every failure. No key, empty note, request timed out, budget spent: the note goes up. A public room that stops taking writes because a classifier is down is worse than one that shows a rude line until somebody removes it.6 The last word still belongs to the rate limit and the ownership check, and those are code.
The last piece I built shows it most clearly. Remember MIN_LINK_SCORE = 0.6, the number from the opening? It's still there. What changed is that a reading can overrule it, but only when it's sure:
const verdict =
probability >= WORTH_SHOWING ? true
: probability <= WORTH_HIDING ? false
: null;
▶ A reading that only decides where it is sure: an animation that plays in the original post.
The bars are 0.7 and 0.3. A clear yes rescues a page the distance would have dropped, at 0.56, inside that five-hundredth gap I fitted the number to. A clear no drops the tip that shows up under "hello". Anything in between hands the decision back to 0.6. The model can't make the links worse than they were, because where it has nothing to say, whatever decided before still decides.
I value that more than accuracy. It meant I could ship it without an eval proving it beats the old number, and it means the feature falls back to last month's behaviour instead of to nothing.
What it costs
Four things run at build time on this site, once per change. Ranking all 30 posts against each other for related links costs 48,000 tokens. Placing them on a two-axis map costs 12,600. Judging the enrichment text (the FAQ answers, cover alt text and social descriptions) against what each post actually says costs 111,000 across every post. A full eval run of the assistant, 18 cases, costs 22,500. All four from scratch come to 194,000 tokens, which is 0.8 cents.
On the request path each call is cheaper than that, and the ceiling is the daily cap, not the balance. The cap is there because the failure I care about is somebody finding an endpoint and making the readings stop for everybody else, much more than the cost.
Latency is what you have to plan around, and it's not the 70 to 500 milliseconds in the marketing. That figure is the model. Yours is the model plus your worker's network hop plus whatever else the request is doing. The front gate gets 1.5 seconds, because a visitor is waiting with nothing on screen. The question about moving pages gets 5 seconds, because the reply is already written and another second is free. Links get 2.5, and build scripts get 30.
That gap matters. On 21 September a reading came back at 0.95 in 782 milliseconds from a laptop, and the worker still didn't answer within the two-second budget it had then. It fell back to the patterns and dropped the page change the visitor had asked for. The model was fast enough, and the budget I'd given it wasn't.
Where it stops
It can't count, do arithmetic or compare dates, and their own docs say so. Reading time, prices, ordering and anything involving money stay in code, as they always did.
It isn't a parser. Terminal command parsing, URL matching, file paths: all of that is regex and should stay regex. The places worth replacing are the ones where the code was guessing at meaning and passing a number off as an opinion.
The accuracy claims are the vendor's. I haven't compared it with a small classifier on the same data, and the one independent benchmark I found was somebody's 40 hand-labelled tickets. It's early access, not GA, and there's no free tier, so your balance runs down while you experiment.
And it doesn't remove work, it moves it somewhere else. Instead of tuning one cosine threshold, I now maintain eleven thresholds, two budgets, a fallback for every call site, and a set of questions worded so a model reading only the criteria would answer the way I would. That's more to look after than I had before. It's better to look after, because each piece is about one decision and says what it's for, but anyone selling you the idea that this deletes code is selling you something.
How I measured, so you can argue with it
Token counts are usage.input_tokens as the API reports them, read from the same field the budget is charged against. The per-call sizes come from a smoke script that makes one real request and prints the count.
The 0.51 to 0.56 and 0.61 figures are retrieval scores over the eval's 18 questions in September, from the index as it was that day. The confidence readings quoted for the radio and the front gate are single measurements against production on 21 September, not averages, and I've said so each time instead of dressing them up.
I found the counter bug by calling production six times and reading the D1 row, and that was the whole method. The $42 per billion and the free output tokens are TypeSafe's published prices, and the 119 million figure is the account balance divided by that.
Every threshold in this post is a constant you can read in the repository, and where a number came from judgement instead of measurement, I've tried to say so.
Originally published at zeybek.dev.
-
Context is 64k per request, 32k of it for the state. ↩
-
I'll come back to why that matters more than the threshold. ↩
-
I got those numbers by reading what the model said about notes I'd already judged myself. ↩
-
Output tokens are free, so input is the whole bill. ↩
-
Spend was still bounded, because every route has its own per-IP limit, but the one number meant to cap the lot did nothing. ↩
-
I've argued this before about anything with a model in it, and a decision model doesn't change the argument. ↩
Top comments (0)