DEV Community

Cover image for When Every Gateway Ships Its Own Decision Model
TuringCorp
TuringCorp

Posted on

When Every Gateway Ships Its Own Decision Model

When Every Gateway Ships Its Own Decision Model

A decision model used to be a product you called. Now it is a thing you host.

On September 15, TypeSafe introduced Jev and the announcement went to the front page of Hacker News. What followed did not look like the usual adoption curve, where a new capability stays a dependency for a year before anyone builds on it. It became a genre in about two weeks.

I pulled the current numbers from the Hacker News search API (hn.algolia.com/api/v1/search?query=%22jev%22&tags=story&numericFilters=points>100) instead of quoting a snapshot, because this table keeps moving. The official announcement post is at 1,989 points and 520 comments. Then the copies and the riffs: Jev in 25 Lines of Python (691 points, Sep 23), Ollaya – Ollama for open-source, Jev-style decision models (614, Sep 25), Kev: Tiny Jev-like family of decision models (462, Sep 21), OpenAI is well positioned to fast-follow Jev (328, Sep 22), and Reverse-engineered Jev-like model (169, Sep 16).

The dates are the interesting column. Reverse-engineered Jev-like model landed on September 16, one day after the announcement — a working imitation of the interface, built without the training recipe, without the weights, without any cooperation from the people who made it. By September 23 the interface fit in twenty-five lines of Python. By September 28 a small model trained at home was answering in about thirty milliseconds. The September 28 clone scored 571 points, more than the earliest family of models managed a week after it appeared. A late entrant beating early ones at the only game the front page scores, attention, is not a saturation curve.

What the fast copies actually prove

The obvious reading is that something got stolen. That reading is wrong, and being precise about why matters.

What those two weeks demonstrated is that the interface is now a known shape. It fits in a sentence: take a program state and typed questions, return a structured choice with a score and a confidence, write no prose. Twenty-five lines of Python is enough to re-implement that, and a weekend is enough to re-train a small open model against it. Nobody needed permission, and nobody needed to be clever.

That is a genuine loss of a moat. If your business is being the only place that offers typed judgment at low latency, that business just ended. It was not undercut. It was generalized.

The part that is easy to miss from inside a company that builds one of these: this is mostly good. An interface many people can implement gets implemented everywhere — in gateways, routers, CI pipelines, and the classifier slot that used to hold a fine-tuned BERT and a lot of maintenance. The copies are not free-riding. They are the original becoming infrastructure.

The layer that gets eaten, and why you should let it

Every one of the clones is a stateless transformation with a clean boundary. Text and a question go in, a structured judgment comes out, and nothing downstream depends on whether that judgment was good — only on whether it was well-formed. Three properties make that shape easy to commoditize. The interface is narrow, so the surface to reproduce is small. The volume is enormous, so the margin per call was always going to be competed toward the cost of compute. And errors are cheap and re-runnable: a wrong classification on a product listing is a bad row you re-run next week and never sign.

When a layer has those three properties, commoditization is not a threat to defend against. It is a service. Judgment becomes a commodity the way TLS certificates and JSON parsing did: something nobody thinks about, priced near zero, available at the edge, on by default.

I will be direct about the self-interest here. We sell judgment. If the only thing we sold were that stateless transformation, this article would be a eulogy. It is not, and the reason is in the next section.

What did not get copied

Go back to the clones and ask a question the point totals do not answer: how do you know whether any of them is right?

You cannot tell from the interface, because the interface has no place to put that information. choice, score, confidence — those are outputs, not evidence. A model that reports 0.9 on everything and a model whose 0.9 means something produce byte-identical JSON for the same input. The interface is silent on the only thing a buyer eventually needs to know.

So two properties stay behind when the shape gets copied, and neither is protected by a clever trick. They are protected by being work.

Verification is the first. It means a record that says: here is the data this number was computed on, here is the counting rule, here is what was excluded and why, and here is what happens when you change the order the candidates are presented in. That record cannot be reverse-engineered from an API, because it is not a property of the API. It is a property of a measurement someone had to run, on a named benchmark, under a stated protocol, and publish even where it looks bad. The thing to check is order sensitivity: does the judgment hold when you swap the two candidates? A confidence number that moves because you re-ordered the same two options is not a confidence number. It is a position in a sequence.

We publish ours, which makes this a claim you can audit rather than a principle you take on faith. Self-run on JudgeBench, 620 judgments, with the 6 that failed on a first verdict disclosed rather than quietly retried, raw accuracy came out at 92.5% against 92.2% for a direct model baseline. That is a tie. We publish it as a tie, and we claim no accuracy advantage over a direct model call, or over anyone else. The bands are where the measurement earns its keep: calls reported at 90% confidence or above were right 99.6% of the time, and calls in the 80–90% band were right 94.0%. On ContextualJudgeBench, self-run under the official protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — consistent accuracy is 67.1% against the benchmark's official reference, with 12 orders excluded after repeated platform failures and that exclusion stated next to the result. The deliberately constructed near-tie splits sit at 46–60%.

That last figure is the least flattering number we have, and it is the one that makes the others mean something. A vendor that publishes only the top row tells you which row it wants read.

None of this is a moat in the patent sense. Any competitor can run the same measurement tomorrow, and if they do, the category gets more legible. What they cannot do is run it retroactively, because the record is a history with dates on it. The number you can produce today is not a substitute for the number someone published before they had a reason to.

Responsibility is the second, and it is the one that moves the economics.

Why the accountable end gets more expensive, not less

Here is where this is not the Jevons argument, and the difference belongs on the record. The Jevons case is about a price mechanism: make a unit of work cheaper, you buy more units, and total spend on the remaining hard cases rises even as the easy ones approach free. That mechanism is real, and it is about volume.

What happens to verification and ownership is a different mechanism. It is about scale — or rather the absence of it.

Verification does not get cheaper when you add customers. The hundredth deployment does not make the reliability curve for the first one any easier to produce. Each materially different deployment — different data, different questions, different cost of being wrong — needs its own curve, and a curve is a measurement, which is someone's afternoon with a benchmark harness and a decision about what to exclude. No version of this batch-processes.

Responsibility fails to scale at all. When a judgment runs automatically and turns out wrong, the question is not "what was the confidence." It is "who decided this was good enough to run on its own, and on what evidence." That has an answer only if a specific party put its name on a specific curve and said: above this line, the machine acts; below it, a person does. The name does not get cheaper when volume goes up. A signature is a transfer of liability, and liability never had a volume discount.

So the two ends move in opposite directions, and not for the reason people usually give. It is not that the high end is more intelligent, or that it has better models — we explicitly do not claim that, and the published accuracy tie is the evidence. It is that the low end's unit of value is a transformation, which scales beautifully, and the high end's unit of value is a record with a name on it, which scales not at all.

That asymmetry has a predictable consequence for the ecosystem the gateways are building. When every gateway ships a decision model, typed judgment becomes a default — in the routing layer, the moderation layer, the triage layer, everywhere a boundary is clean and a re-run is cheap. That is a large amount of value created, and almost none of it will be captured by whoever shipped the model. The model is the floor.

What gets scarce is the thing the clones left out. In a world where every system has a cheap opinion, the scarce good is not a better opinion. It is a judgment that arrives with the evidence behind it and the name in front of it — something a reviewer can read, disagree with, and hold someone to.

What this argument does not claim

Two caveats. First, none of this predicts that unaccountable judgment will fail. It will work fine for a long stretch, and for most of the volume it is the correct choice — which is exactly why the commoditized layer is worth having. The claim is about where the price goes, not about which layer is ethical.

Second, verification is not a permanent moat. Someone can copy our protocol tomorrow, and I would rather they did. What cannot be copied is the fact that a record was published before the argument needed it, with the unflattering rows included, under a name. Publishing it is available to anyone. Doing it early, and continuing after it stops being flattering, is a policy rather than a feature.

The interfaces are already copied, which is what an interface is for. What the copies cannot carry is the part where someone says: this is what we measured, this is what we left out, this is how wrong we are allowed to be, and this is who is answerable when the machine makes the call alone. That part has a name attached, and a name does not batch.

If you are the one who has to put that name on the line — a hard question, two candidate answers, and a record you can be held to — that is the case we built for: Decider.

Top comments (0)