After a press release, the contact address of one of my products started filling up with sales pitches: webinar invites, free whitepapers, trade show booths, agencies offering to help. I run more than ten products on my own, so I didn't have time to sort these by hand, and my Gmail filters and regex weren't holding.
This post is about what replaced them: a judgment model that answers yes/no questions about each email, and the rules I put around it. On October 2, 2026, it took the inbox from 92 messages to 2.
Why rules broke
Four kinds of mail were landing in the same inbox:
- replies people had written to us (these need an answer)
- auto-replies from contact forms (we reach out to companies through their website forms)
- pitches, newsletters and event invites
- bounces
The same phrases show up across kinds. "Thank you for your inquiry" appears in auto-replies and in thoughtful human replies. Filtering on "webinar" catches real inquiries that mention one.
Some pitches don't even come from the sender. A common trick, mostly from English-speaking companies, is to write the pitch in a Google Doc and share it with you, so the email arrives from Google. Another is signing you up for a Zoom webinar, so the invite comes from Zoom's shared address. You can't send Google or Zoom to spam without losing the shares and invites you actually need.
A wording rule also hurt us once. We treated "we'll pass this time" as a refusal and stopped contacting that company for good. It says "this time." A regex can't tell "this time" from "never" reliably.
The model: Jev
I used Jev from TypeSafe AI. It's what TypeSafe calls a System One model (the name comes from Kahneman's fast, intuitive System 1). It reads natural language like an LLM does, but it doesn't generate text. You send a state and a set of typed questions, and it returns typed answers with probabilities. No JSON to coax out of a chat model, no parsing.
There are three question types:
| Type | Asks | Returns |
|---|---|---|
choice |
pick one option from a list |
choice, per-option probabilities, confidence
|
score |
place the state on 2 to 10 described levels |
score (can land between levels), per-level probabilities, confidence
|
noul |
yes or no |
noul, the probability of yes (0 to 1) |
A few properties mattered for this job:
- All questions in one request are evaluated against the same state, in parallel and independently. One answer doesn't leak into another.
- The docs say most queries come back in about 100 ms, and adding questions barely changes that.
- The question id is never sent to the model. Only
instructions(and optionalcriteria) are, so the edge cases go in the instructions. - Pricing at the time of writing (Jev 1.13) is $0.042 per million input tokens, output free. Rate limits are 100K tokens/s and 40 requests/s, and you get a 429 past either.
- English is the primary training language. Other languages, Japanese included, work but less accurately. My inbox is mostly Japanese, which turned out to matter for thresholds.
Five questions, one call
Each email becomes one request with five noul questions. Most of the cost is sending the body, so you put all questions in a single call instead of one call per question.
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"model": "jev-latest",
"state": "From: ...\nSubject: ...\n\nBody...",
"questions": {
"human": { "type": "noul", "instructions": "A person wrote this reply individually. Auto-replies, promotions and newsletters don't count." },
"needReply": { "type": "noul", "instructions": "It would be appropriate, socially or commercially, for us to reply to this email." },
"counterSales": { "type": "noul", "instructions": "The sender is trying to sell us their own product, service or seminar." },
"meetingRequest": { "type": "noul", "instructions": "The sender wants to set up a meeting, call or visit." },
"stopRequest": { "type": "noul", "instructions": "The sender clearly asks us not to contact them again. Declining just this one time doesn't count." }
}
}
EOF
That last instruction is the direct fix for the "this time" mistake.
Same input, slightly different answers
The probabilities wobble between calls. In an earlier test on a different task, I sent the same three questions about the same 69 items three times. Average spread was 0.013 to 0.041 per question, the worst was 0.15, and only 2 to 11 of the 69 items got identical values all three times.
So anything near a threshold can flip if you ask again. I ask twice and only act when both answers land on the same side of the line. If they disagree, a person looks at it.
const runs = await askRepeated(state, QUESTIONS, 2);
const pick = (key) => {
const g = agreed(runs.map((r) => gateNoul(r.answers[key], LINES[key])));
return g.act === 'auto' ? (g.value ? 'yes' : 'no') : 'unsure';
};
gateNoul maps a probability to yes, no or "don't decide", and agreed only returns auto when every run reached the same verdict. Both live in a small wrapper I share across my products.
Draw the lines from your own mail
My wrapper's default for a single noul is 0.9 for yes and 0.1 for no. On 26 real Japanese emails, that left nearly everything undecided. The measured values looked like this:
- "should we reply": real replies 0.61 to 0.93, pitches and auto-replies 0.05 to 0.49
- "stop contacting us": the highest of the 26 was 0.19
So each question got its own line:
const LINES = {
human: { yes: 0.5, no: 0.3 },
needReply: { yes: 0.6, no: 0.599 },
counterSales: { yes: 0.6, no: 0.599 },
meetingRequest: { yes: 0.6, no: 0.599 },
stopRequest: { yes: 0.7, no: 0.3 }, // can't be undone, so the bar is higher
};
Two things I learned the hard way. Changing the set of questions in a request moves the numbers for the others, so the lines belong to that set of questions and need re-measuring when you add one. And jev-latest is an alias that moves when a new version ships, so if you tuned thresholds on a version, pin it (jev-1.13.0) and re-measure before moving. The response's model field tells you which version answered, which is worth logging.
What never goes to the model
If code can decide it, code decides it:
- bounces are recognizable from the sender (mailer-daemon, postmaster)
- anything about invoices, payments or contracts stays in the inbox regardless of the model
- if the call fails or returns nothing, the email goes to a person, never to "fine, archive it"
That last one matters most. "Couldn't judge" and "judged as harmless" aren't the same thing.
Result
On October 2, 2026:
- about 70 pitches and invites moved to "no reply needed"
- about 15 form auto-replies filed away
- 3 bounces
- 2 left in the inbox: a Google Workspace invoice and one reply I needed to answer
It now runs every five minutes, with judged message IDs cached so the same email isn't paid for twice.
Summary
- Wording rules break when the same words show up across different kinds of mail, and platform-borne pitches (Google Docs shares, Zoom invites) can't be filtered by sender at all.
- Put every question for one email in one call.
- Ask twice, act only on agreement, send disagreements to a person.
- Measure thresholds on your own data, especially outside English, and raise the bar for anything irreversible.
- Pin the model version you measured against.
TypeSafe AI docs:
https://docs.typesafe.ai/introduction
The longer story, including setup steps, is on my blog:
Top comments (0)