I’ve been thinking about a class of AI systems that do not primarily generate text.
Instead, they take some state or evidence and return bounded judgments such as:
is_risky: 0.91
category:
safe: 0.03
suspicious: 0.79
unknown: 0.18
review_priority: 2.6 / 3
Recent systems such as Jev make this pattern particularly visible: instead of asking a language model to generate an explanation or JSON object and then parsing it back into a decision, the model directly produces typed probabilistic judgments.
That seems extremely useful for agent runtimes.
But it raises a question I’m not sure how these systems should be placed inside the control plane.
A tempting architecture
Suppose an agent wants to use a tool.
A probabilistic model inspects the request and returns:
dangerous_operation = 0.94
The obvious implementation is:
if dangerous_operation > 0.8:
reject()
This is simple, fast, and probably useful in many applications.
But I’m increasingly uncomfortable with calling the probabilistic model a verifier.
The model did not prove that the operation was unsafe.
It produced an observation about the operation.
That seems closer to a sensor.
Sensor versus verifier
The architecture I am experimenting with separates four roles:
Artifact / Runtime State
↓
Finding Producer
↓
Finding + confidence
↓
Policy
↓
Decision
↓
Verifier
↓
Enforcer
A probabilistic AI model would live at the first stage:
AI judge
↓
FindingProducer
For example:
Finding:
type: SEMANTIC_RISK
probability: 0.94
model_version: X
evidence_ref: Y
A separate policy layer could then decide what that finding means:
if SEMANTIC_RISK >= 0.90:
require_human_review
The verifier would check whether the policy was applied consistently, whether the correct policy version was used, and whether required preconditions were satisfied.
The enforcer would finally allow, block, quarantine, or escalate the action.
So:
probabilistic judgment
≠
policy decision
≠
verification
≠
enforcement
Why bother separating them?
Imagine the model changes.
Version A says:
risk = 0.92
Version B evaluates the exact same evidence and says:
risk = 0.41
Was the original decision wrong?
Maybe.
But another interpretation is that the historical decision was valid under:
model A
+
policy version 3
+
evidence snapshot E
while a new evaluation would produce a different result under a new model.
That suggests model outputs might need provenance similar to other evidence:
model identity
model version
question / rubric version
input evidence version
timestamp / epoch
confidence
Otherwise, replaying or auditing an agent decision later becomes ambiguous.
Another problem: calibrated does not mean correct
Even if a probabilistic model is well calibrated, a result such as:
0.93
does not mean:
this statement is objectively 93% true
It is still a model judgment.
And even a perfectly type-safe system can confidently choose the wrong member of the allowed output space.
So I currently prefer this interpretation:
A probabilistic AI judge is a semantic observation mechanism, not an authority source.
Its result can influence a decision, but it should not silently become the decision itself.
But perhaps this separation is too strict
There are obvious counterexamples.
Spam filters routinely make automated decisions.
Fraud systems block transactions.
Content moderation models directly suppress material.
Anomaly detectors can trigger circuit breakers.
At some point, a probabilistic judgment clearly does become operational authority.
Maybe the right distinction is not:
AI model may never decide
but instead:
model produces judgment
policy explicitly grants that judgment
a bounded amount of authority
For example:
confidence >= 0.99
AND
effect is reversible
AND
blast radius is low
→ automatic action allowed
otherwise
→ review / escalation
In that architecture, the model still does not define its own authority.
The runtime does.
The question
So I’m curious how others model this boundary.
Should probabilistic AI judges be treated primarily as sensors / finding producers rather than verifiers?
More specifically:
- Should model output ever directly authorize an irreversible action?
- Should model version and rubric version be part of the evidence provenance?
- If a newer model disagrees with an older model, should historical decisions be re-evaluated?
- Where should confidence thresholds live: inside the model interface, inside policy, or inside the application?
- Is there established terminology for separating probabilistic semantic judgment from deterministic policy enforcement?
I’m especially interested in related patterns from:
- runtime assurance
- policy engines
- autonomous systems
- fraud / risk systems
- capability security
- agent runtimes
- formal methods
The architecture that led me to this question is an experimental agent-runtime project called HANDS, but I’m deliberately trying to phrase the problem independently of that implementation.
I would be particularly interested in counterexamples where treating the model as “just a sensor” is the wrong abstraction.
Disclosure: I used ChatGPT to help structure and edit this post. The underlying architecture question and design are from my own ongoing project, and I reviewed the content before publishing.
Top comments (0)