Most agent failures are not one bad tool call. They are three "successful" writes in a row that nobody can reconstruct.
Tool A updates the CRM. Tool B opens a ticket. Tool C emails the customer. Each step returns ok. The dashboard stays green. Hours later support asks why a customer got a resolution note for a case that never closed — and the only answer lives in a token dump no human will read.
I call the missing control a Receipt Gate: after every write, produce a short, human-readable receipt. The next write stays closed until that receipt is skim-worthy. No receipt, no second write. Most cascading messes die right there.
Golden Trace freezes the path that worked. Receipt Gate makes each live step legible before the next one fires.
Quick use case
Situation. A revenue-ops team shipped an "exception closer" agent. On a flagged invoice it could (1) patch the CRM opportunity stage, (2) post a note on the finance ticket, and (3) send the customer a status email. Staging looked sharp: three tools, one happy path, under thirty seconds.
What broke. A model bump plus a "helpful" prompt tweak patched the CRM to Closed-Won while the ticket note still said "pending PO match." The email went out with a warm "resolved" tone. Every tool returned success. Nobody could skim, in ten seconds, what actually changed.
Ops found it when a customer asked for a closed invoice that did not exist. Logs showed tokens, not a receipt chain a human could trust.
The fix (Receipt Gate).
- Require a receipt after every write — object, fields, tool, success/fail in plain language.
- Gate the next write on that receipt — tool B cannot fire until tool A's receipt is present and non-mushy.
- Make receipts skim-first — ten-second read; raw JSON alone is not enough.
- Fail closed on mush — vague "updated record" blocks the chain.
- Keep the chain in the run record — receipts travel with the trace for audits and Golden Trace.
Same three tools. Same model family. Different blast radius. The next "Closed-Won + resolved email" never left the first gate.
Receipt Gates aren't about slowing the agent for sport. They're about refusing to stack silent side effects before a human — or the next step — can see them.
Success responses lie. Receipts don't have to.
A tool status code answers: Did the API accept the call?
A receipt answers: What business state actually changed, in words a teammate can check?
Those are different questions. Most stacks fund the first and hope the second shows up in observability later.
When writes chain, hope is the bug. Later tools inherit a world the first write may have only partly changed — or changed the wrong way while still returning 200.
If the only bridge is the model's short-term memory, you are one polite hallucination away from a customer-facing contradiction.
The Receipt Gate stack (five layers)
Build it like a control on the write path, not a nicer log sink.
1. Write → receipt contract
Every mutating tool emits a structured receipt before the orchestrator continues:
- Object / resource id
- Fields or side effects that changed
- Tool name + schema id
- Inputs that mattered (redacted)
- Outcome in plain language: success, partial, or fail
If the tool cannot say what it did, it is not ready for production chaining.
2. Skim rule (ten seconds)
A receipt only a debugger can parse is not a receipt. Target a human skim: one-sentence summary, concrete change bullets, and an explicit "did not change" when an expected field was skipped.
Mushy language ("synced the account," "handled the exception") fails the gate. Specificity is the product.
3. Next-write latch
The orchestrator holds a latch: the next mutating tool stays closed until the prior receipt passes schema + skim checks. Read-only tools can still run. Irreversible or customer-visible writes wait.
Cascading damage dies at the latch — not in the postmortem.
4. Chain in the run record
Persist receipts in order next to the trace. Receipt[n] can reference Receipt[n-1]. Partial failures keep the chain. Ops get an export without opening the full token stream.
Golden Trace can freeze a successful chain. Smoke evals can assert "no second write without receipt." Ship Gate should refuse agents that mutate twice with nothing skimmable in between.
5. Human or policy break-glass
Sometimes a human must accept a partial receipt and continue. That is fine if it is explicit: named approver, why mush was allowed, and a short TTL on the override (pair with Permission Envelope).
Silent auto-continue on a failed skim check is how the CRM went Closed-Won again.
How this maps to what you already have
- 5-layer agent stack — Receipt Gate lives across Tools + Traces; Guardrails enforce the latch.
- Permission Envelope a write may be authorized and still blocked if it left no receipt.
- Golden Trace — freeze a chain whose receipts still match milestones after a model bump.
- Smoke evals — include "second write with missing/mushy receipt → blocked."
- Ship Gate prove chained writes leave skim-worthy receipts, not only that demos finish.
Architecture without receipts is just faster confusion with better models.
The test
Ask one question before the next multi-write agent ship:
If write #1 returns ok, can a teammate skim a plain-language receipt of what changed and does write #2 stay closed until that receipt exists?
If the answer is no, you are still chaining side effects on status codes and model memory.
Closing
I'm Wasim Sheikh — AI Architect. I build systems teams trust and organizations depend on: not demos, not proofs of concept — production.
Follow for practical AI architecture that ships.
Connect on LinkedIn: Wasim Sheikh
Related: Golden Trace · Permission Envelope · Smoke Evals · Ship Gate · The 5-Layer Agent Stack
Top comments (0)