Your agent's hardest call is the one you can't test
There's a moment in almost every agent loop where two options are both defensible and one of them has to win.
Two migration plans. Two refactors. Two diagnoses. Two vendors. Two drafts.
You put both in the prompt, ask for a pick, and get back something confident. You ship it. And you have no idea whether that choice was well-founded or whether the model just happened to prefer the second thing it read.
Here's the part that took me too long to notice.
When the judgement lives in a prompt, it isn't a component — it's a paragraph
A paragraph can't be asserted on. You can't log it as a value, can't regression-test it, can't diff it between model versions, and can't tell a confident pick from a coin flip. When it's wrong you find out in production, and what you learn is "the agent chose B" — not how close the call was.
There's a second, quieter failure hiding in the same place. Ask a model to pick between A and B and report how sure it is, and it will usually do both in one breath. The number you get back was generated to be consistent with the answer it already gave, not measured before it. If you route on that number — "escalate below 70%" — you're routing on a vibe with a percent sign.
That isn't a prompting mistake. It's a structural one: one call can't both make the judgement and measure it.
The reframe: give the agent something to call, not an opinion to have
If a decision matters enough to be tested, it should be a component with a contract — inputs, a typed result, and a way to assert on it.
Concretely: take the judgement out of the prompt and make it a tool call. The agent sends the neutral task plus both options; it gets back which one won, how far apart they were judged, and the reason.
What that buys you, in code
The point isn't the call — it's that the call becomes assertable. With the judgement as a component, the same scenario can go into your test suite like anything else:
# tests/test_escalation.py
def test_close_call_escalates_to_human(decide):
r = decide(
task="Which support reply do we send this customer?",
optionA="Refund immediately, no questions.",
optionB="Explain the policy and offer a partial credit.",
)
# The pick is not the contract. The *shape* of the answer is.
assert r["betterOption"] in ("option_A", "option_B")
assert r["reason"], "a pick with no reason is not reviewable"
if float(r["confidence"].rstrip("%")) < 70:
escalate_to_human(r["reason"]) # your policy, not ours
Two things are worth noting, because they're the whole reason to do this:
- The assertion is on the contract, not on the answer. You're not asserting "the model says B" — you'd be re-recording the answer every model change. You're asserting that a pick came back, that it carried a reason, and that your threshold did its job. That last line is your policy: the service reports how far apart the two options were judged to be; it doesn't tell you what to do about it.
- The reason is a first-class output. It's the part you show a human when the call was close. If you only get a pick, escalation has nothing to hand over.
Wiring it up
Any MCP-capable client can hold this. For a client that takes a remote server URL and a static bearer token, the config is the boring part — which is exactly how it should be:
{
"mcpServers": {
"decider": {
"type": "streamableHttp",
"url": "https://mcp.turingcorp.net/mcp",
"headers": {
"Authorization": "Bearer <your-agent-pass>"
}
}
}
}
Three field-level traps, all of which I've hit:
-
"type"is not decoration. At least one popular client silently falls back to SSE if the type isn't spelled the way that client expects, and SSE against a Streamable-HTTP-only endpoint fails with a405. Copy the value from the client's own docs, not from a blog post (including this one). -
Discovery and calling are separate. Reading the tool list doesn't need a credential; calling does. So a scanner reporting "open — no credentials needed" is reading the wrong half, and you'll meet a
401with aWWW-Authenticateheader on the first real call. No credential belongs in a tool argument — it goes in the header, or it ends up in your transcripts. - Reserve minutes, not milliseconds. A judgement between two options is a long call: budget 180–300 seconds in the client, and note that the timeout is a client setting — there's nothing to pass in the tool call. A 60-second default will cut it off before the answer arrives.
The part most write-ups skip: the call that never came back
A long call can be cut off — by your client, your host, or a network hiccup. The failure mode to design out is the expensive one: silently re-running it. A retry is a second paid call, and the declaration is explicit rather than pretend-safe — the operation is not idempotent, and we say so instead of implying otherwise.
So the recovery path is retrieval, not repetition:
try:
r = decide(task=t, optionA=a, optionB=b) # may be cut off
except TimeoutError:
r = get_result(job_id=j) # read-only, free; same body the call would have returned
When you never received an id at all — the cut-off-before-handshake case — retrieval with no argument lists the ids this credential created in the last seven days. That's the difference between "we lost a call" and "we lost a call and paid for it twice".
Why I'm publishing the numbers instead of a free tier
This is a paid call with no free tier, so the honest thing is to give you evidence instead of a trial.
There are 27 real decisions recorded verbatim — the question, both options, which was preferred, the reported confidence, and the full reason — published at https://github.com/TuringCorp-net/poe-demo-public. In that set the confidence runs 27.3%–88.3%, median 74.0%, none above 90%, because they're everyday close calls rather than easy ones. A service reporting 99% on questions like those would be telling you something false.
The fairest test costs one call: run a decision whose answer you already know, and check whether the reason is one you'd accept from a colleague.
The thing to take away
If a judgement is worth acting on, it's worth being able to test. Move it out of the prompt, give it a contract, assert on the contract, and log the reason — and then your agent's hardest call becomes the one part of the loop you can actually regression-test.
Decider is at https://mcp.turingcorp.net — a decision model for the calls that don't have a right answer.
Top comments (1)
You need to verify your account .
Link is in the profile.