DEV Community

Wasim Sheikh
Wasim Sheikh

Posted on Originally published at sheikhwasim.com

Boundary Contract: Make Every Agent Tool Fail Boring

Cross-post of Insights #8 — canonical: https://sheikhwasim.com/insights/boundary-contract-tools-fail-boring/

Most agent failures I see in production don't start in the model. They start at the edge of a tool.

A call times out. An API returns an empty list. A search gives back three results out of thirty. The model gets something shaped like data, so it does what models do: it writes a fluent, confident answer on top of it.

The transcript reads fine. The customer gets the wrong answer. Nobody notices until a human asks why.

I call the missing control a Boundary Contract: every tool the agent can call returns a typed outcome, not just a payload. Ok, empty, partial, timeout, rejected. The agent has to handle each one by name. It can't turn a broken call into the truth.

Clever prose can wait. The boundary has to be boring first.

Quick use case

Situation. A field-service company shipped a scheduling agent. A customer asks for a repair visit. The agent checks the technician calendar, offers the next open slot, and books it. Staging looked sharp.

What broke. On a busy Monday the calendar service started timing out under load. The tool wrapper caught the timeout and returned an empty list, because that's what the original developer thought "safe" meant. The agent read "no slots" and told customers the next opening was three weeks out. Some cancelled. Some called a competitor. The logs showed a successful tool call, a polite reply, and no error anywhere. The dashboard stayed green all morning.

The fix (Boundary Contract).

  1. Typed outcomes on every tool. The calendar tool now returns ok, empty, partial, timeout, or rejected, plus the payload. "Empty" and "timed out" stopped looking the same.
  2. Input checked before the call. A bad zip code or a missing service type gets rejected with a reason, before it reaches the calendar API and comes back as a mystery.
  3. A named policy per outcome. timeout means retry once, then tell the customer "I can't see the calendar right now" and offer a callback. Never "no slots."
  4. Partial results are surfaced. If one region's calendar answered and another didn't, the agent says so instead of quietly offering only what came back.
  5. The outcome goes in the trace. Every run logs which outcome each tool returned, so "how many timeouts did we hide this week" has a real answer.

Same model. Same prompt. The difference was that a broken call could no longer pass as an answer.

Why "it returned something" isn't enough

A tool payload answers: What came back?

A Boundary Contract answers: What actually happened at the edge, and what is the agent allowed to say about it?

Most stacks only fund the first question. They wrap errors into empty results to "keep the agent from crashing," and the model fills the gap with confidence. That isn't resilience. It's a lie with good grammar.

The Boundary Contract (five layers)

Build it into the tool layer, not into a line of the system prompt that says "be careful with errors."

1. Input contract

Each tool declares what valid input looks like: required fields, allowed values, sane ranges. Invalid input is rejected before the call, with a reason the agent can repeat or fix. A tool that accepts anything will eventually do anything.

2. Outcome types, not just payloads

Every response carries one of a small, fixed set of outcomes. Five is enough for most tools: ok, empty, partial, timeout, rejected. The important part is that empty and failed are never the same value.

3. Time budget, named

Each tool has an explicit timeout and retry rule written down next to it. When the budget runs out, the outcome says timeout. It doesn't silently become empty, and it doesn't hang until the whole run dies.

4. Outcome policy in the orchestrator

For each outcome, decide what the agent may do and say. partial means disclose the gap. timeout means fall back or hand off to a human. rejected means ask a clarifying question. Write these as rules the orchestrator enforces, so the model doesn't improvise them per ticket.

5. Outcome in the trace

Log the outcome for every tool call, not only the payload. Freeze one happy path and one timeout path as Golden Traces, and add "timeout must never produce a confident answer" to your Smoke Evals.

How this maps to what you already have

  • 5-layer agent stack. Boundary Contract lives in Tools, enforced by Guardrails, recorded in Traces.
  • Receipt Gate. Receipt Gate checks a write landed before the next one fires. Boundary Contract makes sure the agent knows when a read didn't land.
  • Permission Envelope. The envelope decides whether a tool may run. The contract decides what its result means.
  • Cheap Twin. A partial or timeout outcome is a good escalate line on the checklist.
  • Ship Gate. Inject a timeout in canary and prove the agent says "I can't see that right now" instead of making something up.

The test

Pick the tool your agent calls most. Then ask:

If this tool timed out right now, would the agent tell the user something true, and would the trace show it was a timeout?

Run it for real: block the endpoint in staging and read the reply. If the agent still answers confidently, the boundary isn't a contract yet. It's a hope.

Closing

Agents don't need to be clever at the edges. They need to be honest there. Every tool should fail in a way that's boring, typed, and visible, so the clever parts of the system are built on something true.

I'm Wasim Sheikh, AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept, production.

Follow for practical AI architecture that ships.

Connect on LinkedIn: Wasim Sheikh · Site: sheikhwasim.com · Notes: Practical AI Notes · X: @anciwasim

Top comments (0)