DEV Community

Cover image for A Confidence Score Is Not a Probability: Act, Ask, or Abstain
Raju Dandigam
Raju Dandigam

Posted on

A Confidence Score Is Not a Probability: Act, Ask, or Abstain

An agent emits confidence: 0.91.

It is tempting to read that as “a 91% chance of being correct” and allow an automated action. Unless the score was calibrated and validated for this task and population, that interpretation is unjustified.

The number might be model self-report, a classifier score, normalized log probability, retrieval similarity, or an application heuristic. Those sources are not interchangeable.

Preserve where the score came from

Make provenance part of the type:

type ConfidenceSignal = {
  value: number;
  source:
    | "model_self_report"
    | "classifier"
    | "calibrated_classifier"
    | "retrieval_score"
    | "policy_heuristic";
  modelVersion: string;
  calibrationVersion?: string;
};
Enter fullscreen mode Exit fullscreen mode

This prevents an application from silently treating a model-written 0.9 as a calibrated probability.

Calibration asks an empirical question: among decisions assigned approximately 0.8, how often was the target outcome actually correct? Reliability diagrams and metrics such as expected calibration error can reveal whether the score systematically overstates or understates accuracy.

Calibration is conditional. A score calibrated on billing questions may not remain calibrated on medical questions, another language, a new model, or a changed prompt.

The product decision has three paths

Many systems need more than allow or deny:

type Disposition = "act" | "ask" | "abstain";

type DecisionContext = {
  confidence: ConfidenceSignal;
  impact: "low" | "medium" | "high";
  missingRequiredFacts: boolean;
};

function decide(context: DecisionContext): Disposition {
  if (context.missingRequiredFacts) return "ask";

  if (context.impact === "high") {
    return context.confidence.source === "calibrated_classifier" &&
      context.confidence.value >= 0.98
      ? "ask"
      : "abstain";
  }

  if (context.confidence.value >= 0.85) return "act";
  if (context.confidence.value >= 0.55) return "ask";
  return "abstain";
}
Enter fullscreen mode Exit fullscreen mode

The numbers are illustrative, not recommended thresholds. Real thresholds should come from measured error costs and representative validation data.

Notice that this example never auto-executes a high-impact action. Even a well-calibrated signal only chooses between requesting a controlled confirmation and abstaining. If your product permits automation at that risk level, authorization and effect controls belong in a separate policy layer.

The important design is the disposition:

  • Act: execute a reversible or sufficiently controlled action.
  • Ask: request one fact, confirmation, or human decision that can reduce uncertainty.
  • Abstain: decline the task or route it to a safer system because more conversation will not resolve the risk.

Thresholds should reflect asymmetric harm

A false positive on “show a help article” is different from a false positive on “issue a refund.” Use a loss matrix:

Decision Correct Incorrect
Act User gets immediate value Wrong real-world action
Ask Small amount of friction May prevent a harmful action
Abstain Safe fallback Lost automation opportunity

Choose thresholds to minimize expected harm under the product's constraints, not to maximize raw model accuracy.

For irreversible actions, confidence should rarely be the only control. Require authorization, argument validation, freshness, idempotency, and post-action verification as appropriate.

Evaluate the policy, not one score

Build a dataset containing the situations near each boundary. Record:

type DecisionEvaluation = {
  caseId: string;
  predictedDisposition: Disposition;
  expectedDisposition: Disposition;
  confidence: ConfidenceSignal;
  segment: string;
  actualOutcome?: "correct" | "incorrect" | "unknown";
};
Enter fullscreen mode Exit fullscreen mode

Measure at least:

  • precision of automated actions;
  • unnecessary-ask rate;
  • unsafe-act rate;
  • abstention rate;
  • task success after clarification;
  • calibration by confidence band and segment.

Plot risk against coverage as the action threshold changes. “Coverage” is the share of eligible cases the system acts on; “risk” is the observed error rate among those actions. This exposes the real tradeoff hidden by a single accuracy number. Report intervals and sample counts for each segment—a 100% success rate over three actions is not a stable operating point.

Version the threshold policy independently from the model. A model rollback should not require guessing which decision boundary was active, and a threshold change should be evaluated against the same labeled cases before release. Record modelVersion, calibrationVersion, and policyVersion with every disposition.

Do not evaluate only easy cases far from the thresholds. Mine production failures, disagreements, and near-boundary decisions into a reviewed regression set.

Monitor drift without pretending every outcome is known

Some outcomes are delayed or ambiguous. Keep unknown explicit instead of counting it as failure or success. Define the observation window and update calibration only from outcomes that are valid labels.

Track score distribution and disposition rate by stable segment. A sudden jump in high-confidence actions after a prompt or model change is a release signal even before enough outcome labels arrive.

Confidence is evidence, not authority

A numerical score can help allocate automation, human attention, and clarification. It becomes dangerous when its provenance disappears and the UI presents it as a probability it never was.

Name the source. Measure calibration. Design act, ask, and abstain as product behaviors. Then place confidence inside a policy that also considers impact, missing facts, and system controls.

References

Top comments (1)

Collapse
 
johnnylemonny profile image
𝗝𝗼𝗵𝗻 •

This is an important distinction that many production systems still get wrong. Confidence scores are often used as if they were calibrated probabilities, even when they aren't. The act/ask/abstain approach feels much more practical, especially for applications where the cost of a wrong answer is higher than the cost of asking for clarification. Nice explanation and examples.