Your classifier returns 0.90.
Your agent sees it and takes the action.
But what does that number mean?
If predictions assigned about 90% confidence are correct about 90% of the time, the score is doing its job. If they are correct half the time, your automation threshold is reading a promise the model hasn't kept.
That's calibration.
It is easy to confuse it with accuracy. A classifier can pick useful labels and still attach misleading probabilities to them. Guo and colleagues studied this distinction in On Calibration of Modern Neural Networks, published at ICML in 2017. Their experiments also found temperature scaling useful on many of the datasets they studied. That is a result from that paper, not a guarantee for your model.
I wanted a tiny check I could read end to end.
So let's build one in TypeScript. We'll group predictions by confidence, compare scores with observed outcomes, and audit what happens when an agent only acts above a threshold.
No API key. No model download. No training loop.
Honesty note: the 12 predictions below are invented. Every percentage in the output comes from that fixture. This is a calibration audit, not a calibrator, not an LLM truth detector, and not evidence about any deployed model. Four examples in a bin are far too few to set an automation policy.
What the score is supposed to mean
Our score estimates the probability that a classifier's chosen label is correct. correct is the observed outcome for that prediction.
That differs from a binary positive-class calibration curve, where the outcome is whether the positive event happened. Both compare predicted probabilities with observed event frequencies, but you must define the event before interpreting the number. The scikit-learn calibration guide explains the positive-class version and reliability diagrams.
A generated sentence saying "I am 90% confident" is not automatically this kind of probability. Neither is a cosine similarity, a tool permission, or an arbitrary score squeezed into the range zero to one.
First define the event. Then collect labels. Then check whether the score predicts it.
The dashed diagonal means predicted confidence matches observed accuracy. The cyan points are our synthetic fixture.
The high-score point is the one to watch: average confidence 0.90, observed accuracy 0.50. In this small fixture, the score is overconfident in that bin by 40 percentage points.
Project setup
Use Node.js 22.20.0 or later in the Node 22 line, with built-in TypeScript type stripping. Save the complete block below as calibration.ts, then run:
node calibration.ts
I tested it on Node.js 22.20.0 in the Local VM. This example needs no npm packages. Type stripping executes it; it does not type-check it.
Build the audit
There are three pieces:
- A labeled fixture with a defined correctness event.
- Equal-width probability bins and their observed accuracy.
- A threshold report that keeps coverage next to selected accuracy.
Here is the whole runnable file:
import assert from 'node:assert/strict';
type Example = { score: number; correct: boolean };
// Synthetic held-out predictions. These are not a real model's probabilities.
const data: Example[] = [
{ score: .2, correct: true }, { score: .2, correct: false },
{ score: .2, correct: false }, { score: .2, correct: false },
{ score: .6, correct: true }, { score: .6, correct: true },
{ score: .6, correct: true }, { score: .6, correct: false },
{ score: .9, correct: true }, { score: .9, correct: false },
{ score: .9, correct: true }, { score: .9, correct: false },
];
for (const x of data) {
if (!Number.isFinite(x.score) || x.score < 0 || x.score > 1)
throw new Error('score must be a finite probability');
}
const mean = (xs: number[]) => xs.reduce((a,b) => a+b,0) / xs.length;
const pct = (x: number) => `${(100*x).toFixed(1)}%`;
let ece = 0;
console.log('Reliability bins: [low, high), with 1 included in the last bin');
for (let i=0; i<5; i++) {
const low = i / 5, high = (i+1) / 5;
const bin = data.filter(x => x.score >= low && (x.score < high || (i===4 && x.score===1)));
if (!bin.length) continue; // Empty bins contribute zero weight.
const confidence = mean(bin.map(x => x.score));
const accuracy = mean(bin.map(x => Number(x.correct)));
const gap = Math.abs(confidence - accuracy);
ece += bin.length / data.length * gap;
console.log(`${low.toFixed(1)}-${high.toFixed(1)}: n=${bin.length}, confidence=${pct(confidence)}, accuracy=${pct(accuracy)}, gap=${pct(gap)}`);
}
assert(Math.abs(ece - .2) < 1e-12);
console.log(`ECE=${pct(ece)} (bin-dependent; descriptive, not a guarantee)`);
console.log('\nAutomation threshold audit');
for (const threshold of [0, .55, .85, .95]) {
const selected = data.filter(x => x.score >= threshold);
const coverage = selected.length / data.length;
const accuracy = selected.length ? pct(mean(selected.map(x => Number(x.correct)))) : 'n/a';
console.log(`threshold=${threshold.toFixed(2)}: selected=${selected.length}/${data.length}, coverage=${pct(coverage)}, accuracy=${accuracy}`);
}
const high = data.filter(x => x.score >= .85);
assert.equal(high.length, 4);
assert.equal(mean(high.map(x => Number(x.correct))), .5);
assert.equal(data.filter(x => x.score >= .95).length, 0);
console.log('\nAll 3 scenario assertions passed.');
Exact output
Reliability bins: [low, high), with 1 included in the last bin
0.2-0.4: n=4, confidence=20.0%, accuracy=25.0%, gap=5.0%
0.6-0.8: n=4, confidence=60.0%, accuracy=75.0%, gap=15.0%
0.8-1.0: n=4, confidence=90.0%, accuracy=50.0%, gap=40.0%
ECE=20.0% (bin-dependent; descriptive, not a guarantee)
Automation threshold audit
threshold=0.00: selected=12/12, coverage=100.0%, accuracy=50.0%
threshold=0.55: selected=8/12, coverage=66.7%, accuracy=62.5%
threshold=0.85: selected=4/12, coverage=33.3%, accuracy=50.0%
threshold=0.95: selected=0/12, coverage=0.0%, accuracy=n/a
All 3 scenario assertions passed.
Read the bins before the summary
ECE is expected calibration error. Here it is the weighted average of the absolute gaps between mean confidence and observed accuracy in the occupied bins.
ECE = sum over bins: (bin count / total count) × |confidence - accuracy|
= (4/12 × 0.05) + (4/12 × 0.15) + (4/12 × 0.40)
= 0.20
The output writes that as 20.0%. Think of it as a 20-percentage-point average gap under this binning scheme. It is not the error rate; this fixture's overall error rate is 50%.
Binning matters. Change the boundaries and you can change ECE. A single aggregate can also hide a bad slice, so keep the counts and individual bin results.
Notice the intervals. They include the lower edge and exclude the upper edge, except that the final bin includes 1. A score on a boundary belongs in one bin, not two. Empty bins contribute zero weight; they tell us nothing about reliability there.
A higher threshold is not automatically safer
At 0.55, the toy selects eight cases and gets five correct: 62.5% selected accuracy.
At 0.85, it selects four and gets two correct: 50.0% selected accuracy.
Raising the threshold selected a worse slice here. The fixture deliberately breaks the assumption that larger scores always imply better outcomes. A monotonic calibration mapping cannot repair that ordering problem by itself.
At 0.95, it selects nothing. Coverage is zero. Accuracy is undefined, so the code prints n/a. Reporting 100% accuracy there would reward doing no work.
For an agent, I want these numbers side by side:
- How many cases did the policy handle?
- How many of those handled cases succeeded?
- What happened to the cases it deferred?
That last question isn't implemented here. A real workflow needs to measure review load, review outcomes, and the cost of errors too.
Calibration describes probability estimates. Permission still comes from your workflow. A calibrated classifier predicting a category correctly does not authorize a refund or prove an entire multi-step task will succeed.
How I would turn this into a real check
Keep data used to train the classifier separate from data used to fit a calibration mapping. Keep the final evaluation separate from both. The scikit-learn guide describes independent calibration data and cross-validation options.
Then choose an automation threshold on validation data before touching the final test. A test set reused to tune the threshold becomes part of the tuning process.
For a real audit, I would also:
- Report sample counts and uncertainty intervals, especially near the action threshold. Our four-case bins are teaching examples.
- Inspect meaningful slices such as task type, language and input source. A good aggregate can conceal a weak subgroup.
- Compare reliability diagrams with proper scoring rules such as Brier score or log loss. Those scores capture more than calibration alone, as the scikit-learn guide explains.
- Recheck after a model change or a distribution shift. Calibration on yesterday's workload does not establish calibration on tomorrow's.
This script does none of that fitting or monitoring. It makes one mistake visible: treating a score as a measured success rate without checking outcomes.
The bigger idea
The model supplies a number.
The labeled outcomes tell you whether the number means what you think it means.
The automation policy decides what you are willing to do with it.
Those are three different responsibilities.
Before an agent uses confidence to act, make it earn the word.
I'm building Roster, AI employees that do real work. This is the kind of engineering question I care about around those workflows: what evidence should let software act, and what should send the work to a person?
What score is your system treating as a promise today?



Top comments (0)