Cross-post of Insights #8 — canonical: https://sheikhwasim.com/insights/boundary-contract-tools-fail-boring/
Most agent failures I see in production don't start in the model. They start at the edge of a tool.
A call times out. An API returns an empty list. A search gives back three results out of thirty. The model gets something shaped like data, so it does what models do: it writes a fluent, confident answer on top of it.
The transcript reads fine. The customer gets the wrong answer. Nobody notices until a human asks why.
I call the missing control a Boundary Contract: every tool the agent can call returns a typed outcome, not just a payload. Ok, empty, partial, timeout, rejected. The agent has to handle each one by name. It can't turn a broken call into the truth.
Clever prose can wait. The boundary has to be boring first.
Quick use case
Situation. A field-service company shipped a scheduling agent. A customer asks for a repair visit. The agent checks the technician calendar, offers the next open slot, and books it. Staging looked sharp.
What broke. On a busy Monday the calendar service started timing out under load. The tool wrapper caught the timeout and returned an empty list, because that's what the original developer thought "safe" meant. The agent read "no slots" and told customers the next opening was three weeks out. Some cancelled. Some called a competitor. The logs showed a successful tool call, a polite reply, and no error anywhere. The dashboard stayed green all morning.
The fix (Boundary Contract).
-
Typed outcomes on every tool. The calendar tool now returns
ok,empty,partial,timeout, orrejected, plus the payload. "Empty" and "timed out" stopped looking the same. -
Input checked before the call. A bad zip code or a missing service type gets
rejectedwith a reason, before it reaches the calendar API and comes back as a mystery. -
A named policy per outcome.
timeoutmeans retry once, then tell the customer "I can't see the calendar right now" and offer a callback. Never "no slots." - Partial results are surfaced. If one region's calendar answered and another didn't, the agent says so instead of quietly offering only what came back.
- The outcome goes in the trace. Every run logs which outcome each tool returned, so "how many timeouts did we hide this week" has a real answer.
Same model. Same prompt. The difference was that a broken call could no longer pass as an answer.
Why "it returned something" isn't enough
A tool payload answers: What came back?
A Boundary Contract answers: What actually happened at the edge, and what is the agent allowed to say about it?
Most stacks only fund the first question. They wrap errors into empty results to "keep the agent from crashing," and the model fills the gap with confidence. That isn't resilience. It's a lie with good grammar.
The Boundary Contract (five layers)
Build it into the tool layer, not into a line of the system prompt that says "be careful with errors."
1. Input contract
Each tool declares what valid input looks like: required fields, allowed values, sane ranges. Invalid input is rejected before the call, with a reason the agent can repeat or fix. A tool that accepts anything will eventually do anything.
2. Outcome types, not just payloads
Every response carries one of a small, fixed set of outcomes. Five is enough for most tools: ok, empty, partial, timeout, rejected. The important part is that empty and failed are never the same value.
3. Time budget, named
Each tool has an explicit timeout and retry rule written down next to it. When the budget runs out, the outcome says timeout. It doesn't silently become empty, and it doesn't hang until the whole run dies.
4. Outcome policy in the orchestrator
For each outcome, decide what the agent may do and say. partial means disclose the gap. timeout means fall back or hand off to a human. rejected means ask a clarifying question. Write these as rules the orchestrator enforces, so the model doesn't improvise them per ticket.
5. Outcome in the trace
Log the outcome for every tool call, not only the payload. Freeze one happy path and one timeout path as Golden Traces, and add "timeout must never produce a confident answer" to your Smoke Evals.
How this maps to what you already have
- 5-layer agent stack. Boundary Contract lives in Tools, enforced by Guardrails, recorded in Traces.
- Receipt Gate. Receipt Gate checks a write landed before the next one fires. Boundary Contract makes sure the agent knows when a read didn't land.
- Permission Envelope. The envelope decides whether a tool may run. The contract decides what its result means.
-
Cheap Twin. A
partialortimeoutoutcome is a good escalate line on the checklist. - Ship Gate. Inject a timeout in canary and prove the agent says "I can't see that right now" instead of making something up.
The test
Pick the tool your agent calls most. Then ask:
If this tool timed out right now, would the agent tell the user something true, and would the trace show it was a timeout?
Run it for real: block the endpoint in staging and read the reply. If the agent still answers confidently, the boundary isn't a contract yet. It's a hope.
Closing
Agents don't need to be clever at the edges. They need to be honest there. Every tool should fail in a way that's boring, typed, and visible, so the clever parts of the system are built on something true.
I'm Wasim Sheikh, AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept, production.
Follow for practical AI architecture that ships.
Connect on LinkedIn: Wasim Sheikh · Site: sheikhwasim.com · Notes: Practical AI Notes · X: @anciwasim
Top comments (0)