I deleted a prompt this week. Not a model, not a pipeline — a prompt. Forty lines of "respond only with JSON", "do not include any other text", "if you are unsure, output UNKNOWN", plus three few-shot examples I'd been carrying between projects like a lucky coin. It existed for one reason: to trick a chat model into behaving like a classifier.
The replacement is a model whose output is a decision. A label, a calibrated probability, a version string. No prose. Nothing to parse, nothing to regex, nothing to retry because the model decided to say "Certainly! Here is the JSON you requested:".
Four orgs shipped something in that shape this week. One of them is autotrust/JEV-27B-VL, a 27B vision-language model that emits a typed verdict with a confidence score instead of a paragraph. The other three I'll describe by shape rather than name, because two are still gated and the third is a fine-tune whose README is mostly vibes.
The four, by shape
| Release | Output | Self-host | Where it fits |
|---|---|---|---|
| 27B VL (JEV-27B-VL) | typed verdict + probability + evidence spans | yes, 2×A100 in bf16 | the one I'd put in front of a user |
| 3B router | single label + score, no spans | yes, one 24GB card | first-stage triage |
| 8B fine-tune | label + probability, schema drifts under load | yes | cheap, needs a wrapper |
| small classifier head | label + logits | yes | fastest, no reasoning at all |
Notice what's missing from that table: benchmarks. I don't care about top-1 here. Top-1 is the metric you optimize when your model is a demo. The moment you put a threshold in front of it, the only number that matters is whether the probability still means something when the input distribution moves.
Calibration on the eval set is theater
Here's the trap. You take a model, you run it on a held-out split, you compute expected calibration error, you temperature-scale until the reliability diagram looks like a straight line, and you ship. Congratulations: you calibrated on the same distribution you evaluated on. That's not calibration, that's curve fitting with extra steps.
The test I actually run now is a time split, not a random split. Fit the threshold on week one. Apply it on week four. Report precision at the operating point and the abstain rate, not accuracy.
I did this on roughly 400 examples across two slices — one from the same week as my tuning set, one from three weeks later, different document templates, different phone cameras. On the in-distribution slice, binning the 27B's outputs at 0.9 gave me about 0.92 precision. Same threshold, shifted slice: 0.71.
The 3B router degraded harder. 0.88 down to 0.54. That's the classic failure — a small model that's confident because it's seen the pattern before, applied to a pattern it hasn't.
The 8B fine-tune was the interesting one. Its top-1 barely moved. But its confidence went flat — everything landing between 0.5 and 0.7, nothing above 0.8. Which sounds like a regression and is actually the most honest behavior in the group. It stopped pretending. A model that says "I don't know" in a language you can threshold on is worth more than a model that says "I don't know" in a paragraph you have to read.
What the interface change actually buys you
The chat-completion interface has one output channel and infinite surface area. Every downstream consumer has to guess what came back. The decision interface has a fixed surface and a number you can compare.
That sounds like a small thing. It isn't. It moves three things out of prompts and into code:
Routing. Escalation policy becomes a config value. Threshold at 0.85, send the 0.4–0.85 band to the bigger model, send everything below to a human queue. I changed that band four times last week without touching a single prompt. Try doing that with a few-shot example.
Replay. You can log a decision and re-run it. Label, probability, model version, input hash. When someone asks why the system flagged their document, you have an answer that isn't "the model said so." You cannot replay a paragraph.
Retries. With prose, a failed parse means you retry and hope for a different roll. With a decision, you get a number and you decide what to do with it. The circuit breaker becomes arithmetic instead of superstition.
Where it breaks
Calibration is per-deployment. A model calibrated on my data is not calibrated on yours, and vendor calibration cards are marketing until you reproduce them on your own shift. I've stopped reading them.
Second: the moment you add a second label space, probabilities stop being comparable across heads. A 0.8 on "invoice vs receipt" and a 0.8 on "safe vs unsafe" are not the same 0.8. If you're routing on both, you need per-head thresholds, and now you're maintaining a config file that's basically a second model.
Third, and this is specific to the vision-language one: image preprocessing silently changes your input distribution. Resize, JPEG quality, EXIF rotation, color profile. My "distribution shift" test was, in practice, "did the phone change." If you're running JEV-27B-VL or anything like it, pin your preprocessing and version it, or your calibration is measuring your image pipeline, not your model.
Latency, since I pay for the GPUs: the 27B at bf16 on 2×A100 gave me p95 around 380ms for one image plus a short prompt. Fine for a queue. Not fine for a keystroke. The 3B router is closer to 40ms. So the honest architecture is router first, big model only on the uncertain band — which is exactly the architecture the decision interface makes easy and the chat interface makes painful.
What I'd ship
Router first. Threshold at 0.85. Abstain band escalates. Everything below goes to a human, not to an auto-reject, because a confident wrong rejection is the only failure mode that gets you paged at 2am. That cut my 27B calls by about a third versus sending everything to the big model, and a third fewer 27B calls is real money when it's your electricity.
I've run JEV-27B-VL on a few hundred examples, not a few hundred thousand. Maybe the calibration story falls apart at scale. Maybe the 8B's flat confidence is a bug I'm romanticizing. I'd genuinely like to be wrong about the router degrading faster, because it's the cheapest piece in the stack.
But the interface argument doesn't depend on which of these four wins. The model that wins isn't the one with the best top-1. It's the one whose confidence still means something after the world moves. And the quieter news this week is that I no longer write forty lines of prompt to get a boolean.
Top comments (0)