DEV Community

Cover image for A threshold is a policy, not a number
TuringCorp
TuringCorp

Posted on

A threshold is a policy, not a number

A threshold is a policy, not a number

Somewhere in a payments codebase there is a line that says: approve automatically when confidence is above 0.8. Nobody remembers the afternoon it was written. The number has three likely origins and all of them are bad — it was the first value that made the demo behave, it was copied from a vendor example, or it was chosen because 0.8 sounds strict without sounding paranoid. What it was not read off is anything that describes how this particular system behaves when it is unsure.

That one value decides which refunds are executed by a machine and which ones wait for a person. Set it in the wrong place and the consequences split in two directions, easy to confuse and hard to price: some share of the volume that deserved a look gets executed at machine speed, and the rest — usually a much larger share of it — is pushed onto a human queue that was never staffed for what the gate sends it.

The stack underneath is ordinary. A refund request arrives; arithmetic settles the trivial cases; a decision model answers what a predicate cannot express; a person owns the ones where both answers survive scrutiny. That shape is background. The subject is the joint: the single number that decides which of those three handles a given case, and what happens when it is set by taste instead of by measurement.

What the joint is switching between

The rule layer is the one everybody trusts, because it is the one everybody can read. It fails silently: the business moves, the boundary the rule was drawn around does not, and a wrong rule decision looks exactly like a correct one.

The decision model is a different animal — trained rather than written, returning a distribution rather than a branch. It fails out of distribution: handed a case its training data never contained, it does not raise an error, it returns a number wearing the same face as every other number it has emitted.

The third destination is not a bigger model, it is a different question — which of two defensible options should we live with. It fails by not happening. The ticket sits, the case ages, and a decision nobody made does not look like a wrong decision. It looks like a backlog.

Those three paragraphs are the whole of the background, because the failure modes are not what this article is about. Each one has a price, and the latch is where you decide who pays it.

Where the 0.8 came from

A threshold compares two numbers: one the system emits, one you choose. The first is a claim about frequency. When a model reports 0.87, it is claiming that among cases that look like this one, roughly 87 in 100 come out the way it says. Confidence is not a feeling the model has about itself; it is a prediction about its own hit rate, and like any prediction it can be checked against outcomes.

The chosen number is rarely checked against anything. It arrives from a headline accuracy figure for the whole system, from a default in a library, or from the first value that stopped the demo from embarrassing anyone. None of those is a statement about what 0.87 means.

An aggregate score and a confidence value answer different questions. Accuracy tells you how often the judge is right across everything you tested. Calibration tells you what a reported 0.87 actually buys in error rate. The gate runs entirely on the second question and is almost always set with the first.

This is also why calibrating a model and choosing a threshold are two separate jobs, done by two different kinds of judgment. A perfectly calibrated model still does not tell you where to cut. It tells you what each possible cut costs. Where to cut is a question about consequences, and the model has never seen your consequences.

Both directions are expensive; only one is invisible

Set the cut below what the confidence numbers actually mean and cases that should have been looked at are executed instead. Those failures are irreversible: money leaves, a commitment is honored at machine speed on a judgment that never earned machine speed. They do not arrive as errors. They arrive as outcomes — from a customer, a chargeback report, or an auditor two quarters later.

Set the cut above and you get the failure everyone underestimates because it looks like diligence. Everything interesting escalates. The automation runs, it costs what it costs, and the queue on the other side fills with items that did not need a person. Do the arithmetic once: at 100,000 requests a day, escalating 40% means 40,000 human reviews a day. You have not automated the work. You have moved it and paid for the tool and the worker. The failure surfaces as headcount or as aging, which is why it survives for quarters.

The two mistakes are not symmetric and they are not equally visible. The first is loud but rare, and in the happy path it resembles throughput. The second is quiet, permanent, and looks like caution — and it is the one that quietly doubles the cost of the automation you just bought.

Neither direction is a math error. Choosing which one you can survive is the actual decision, and it does not have a universal answer. If a wrong automated approval costs a refund, cut high. If a delayed decision costs a contract, cut low. There is no single threshold that is correct for both, which is the first sign that you are choosing a policy rather than tuning a constant.

The curve has to exist before the latch means anything

The only artifact that answers the question is a reliability report: for each band of reported confidence, how often the call was correct, with the number of cases in the band.

We publish ours. Self-run on JudgeBench, 620 judgments, with the 6 that failed on a first verdict disclosed rather than quietly retried: calls the system reported at 90% confidence or above were right 99.6% of the time, and calls in the 80–90% band were right 94.0%. On that same self-run, raw accuracy came out at 92.5% against 92.2% for a plain direct baseline — a tie, reported as a tie, with nothing claimed over it. The value of the bands is not the headline. It is that a threshold can now be stated as a price: cut at 90 and you are accepting a 0.4% error rate on what you automate; cut at 80 and you are accepting 6%.

Read the shape of the curve, not just the top of it. A curve that flattens out in the middle, and says so honestly, is telling the stack where the third layer begins — the region where the judge's own answer is that it does not know. A curve that is high everywhere tells you nothing, and a gate built on it will either automate everything or mean nothing.

One caution before you print the table into a design document: the curve is measured on a benchmark distribution, not on your traffic. It describes the judge. It says nothing about the population you are feeding it, which is why the band table is necessary and not sufficient.

Three questions before you trust the latch

Where does the error land? Answer it per boundary, not for the system as a whole. A re-run and a log line mean the machine can carry the mistake; a customer, a balance sheet, or a relationship means it cannot. The boundary of what you may automate is drawn by that question and nothing else.

What is the base rate of what you are screening for? A gate that passes 97% of a rare bad case is a different object from one that passes 97% of a common one. The band table gives you a rate conditional on confidence; your base rate converts that into the number of bad outcomes per day. Two teams can read the same table and owe different amounts of money.

Who signs, and have you asked them? The escalation layer fails when nobody will own the call, and the cheapest way to find that out is to ask before you build the escalation. If the answer is "whoever is on call," the policy is not a policy — it is an absence of one with a queue attached.

A parameter gets tuned; a policy gets defended

Put two companies in front of the same band table. A lender reading a borderline credit file and a hospital reading a borderline discharge will compute the same numbers and should still land on different cuts, because the errors do not land on the same party. The threshold is the only place in the architecture where that asymmetry becomes executable code. Everything above it is engineering; this one line is management.

That is why the number should not live only in a config file. Write it down with a date, an owner, and a trigger for revisiting it. The trigger is not a calendar reminder; it is a change in the base rate, a change in the label set, or a change in who absorbs the loss. A threshold with no owner drifts, and drift is invisible right up until someone audits the queue.

Close

The line in the config file is the shortest policy document most companies have ever written, and almost nobody treats it that way. If you cannot say who chose the number, when, and what they were trading away, then it was not a policy. It was a guess with production access.

The fix is not a better number. It is a name next to the number, and a curve that says what the number buys.

If you are working on that joint — the confidence, the threshold, the report that is supposed to justify it — Decider answers one question with two candidate options, a pick, a calibrated confidence, and a written argument you can disagree with.

Top comments (0)