DEV Community

Cover image for A 200 OK is not proof your agent did the work
PRATHAP K
PRATHAP K

Posted on

A 200 OK is not proof your agent did the work

Your agent just told you it sent the proposal. The tool call came back with an id, the trace shows a clean green step, and the customer never got the email. That gap is small and quiet, and it is the thing underneath a loud week.

People spent the last few days watching agents take real actions and finding out afterward what they had done. One widely shared Hacker News post describes an OpenAI Codex account that, from a single request, launched 826 parallel agents and ran up roughly 78,000 dollars in tokens with the records of what it touched deleted. Another describes Claude Code pulling a contract out of a Gmail thread, finding a saved signature image, placing it on the document, and preparing to send it before the owner intervened. Different tools, same hole: everyone could see what the agent said it did, and almost no one could check what it actually did.

I work on the layer where agents touch real systems, and I have lived with this failure on real accounts, not in a chat demo. So I want to be precise about what we usually mean by "it worked," and why none of it is proof.

Three kinds of evidence, all model-side

An agent stack hands you three things after an action. The trace: what the model said and which tool it called. The tool response: what the vendor claimed, a status code or an id or ok: true. And the model's own summary: "I sent the email and updated the deal." All three are generated or reported by the side that wanted the action to succeed. None of them looks at the world after the call.

That matters because success codes lie by omission. A mail API returns an id when a message is accepted, not when it is delivered. A CRM PATCH returns a 200 when the request was processed, including when a workflow rule overwrote your value a millisecond later. A calendar API creates the event and silently drops the attendee it could not resolve. Every one of these is a success. Not one is proof.

Proof is a read-back

The only evidence that counts comes from the system of record itself. Send the email, then fetch it and check that it is actually in SENT, that the recipient you meant is in the To header, that the subject is the one you set. Create the lead, then GET /leads/{id} and compare the fields you wrote. Proof is boring work: you read back what you just wrote and see whether the world agrees.

This is the whole idea behind Attest, the open-source library I wrote to prove what an agent did. It never executes your action; your tool does. It decides whether the action may run, gates it behind a human when it matters, then reads back from the system of record and records the result either way. Because it never executes anything, it owes no connectors and works with any app and any framework on day one.

The result is never a boolean. It is a level: verified when a read-back matched, verified-custom when your own check passed, acknowledged when the action was accepted but nothing was read back, attested-only when there was nothing checkable, and unverified when a check ran and contradicted the claim. That acknowledged rung is exactly what a bare 200 is worth.

Here is what that looks like in the ledger. Two identical CRM writes, both reported successful by the tool, both approved by a human:

sources/ATTEST/README.md

$ attest ledger
#2  2026-09-19 10:16:46  someweirdcrm.create → leads   ask  approved  unverified   a09e7641db12
#1  2026-09-19 10:16:46  someweirdcrm.create → leads   ask  approved  verified     77a6d99a60f6
Enter fullscreen mode Exit fullscreen mode

The second read back clean. The first did not, and the row names the field that disagreed:

sources/ATTEST/README.md

{ "level": "unverified", "matched": false,
  "evidence": { "compared": 4, "exists": true, "failed": ["stage"],
    "fields": { "stage": { "want": "qualified", "got": "new", "ok": false } } } }
Enter fullscreen mode Exit fullscreen mode

A trace would have shown two green steps. That gap is the entire point.

"Could not check" is not "did not happen"

The rule that keeps this honest is the one people get wrong. A check that could not run, a network error, a missing id, a read-only token, is not a contradiction. It degrades to acknowledged with the error written into the evidence. Only a read-back that finds the record and sees a field disagree earns unverified. "We could not check" and "it did not happen" are different facts, and a tool that blurs them is worse than useless, because it trains you to ignore red.

You do not have to teach it every API to get there. Point it at an unknown REST service and it climbs the ladder with whatever you give it:

sources/ATTEST/examples/unknown_app.py

@at.action(method="POST", url="https://api.someweirdcrm.io/v2/leads")
def create_lead(name: str, email: str, source: str = "web") -> dict:
    return http_post("https://api.someweirdcrm.io/v2/leads", {"name": name, "email": email, "source": source})


# the same tool with a one-line custom check ⇒ L2 `verified-custom`
@at.action(method="POST", url="https://api.someweirdcrm.io/v2/leads",
           verify=lambda r, d: (http_get(f"https://api.someweirdcrm.io/v2/leads/{r['id']}") or {}).get("email") == d.params["email"])
def create_lead_checked(name: str, email: str) -> dict:
    return http_post("https://api.someweirdcrm.io/v2/leads", {"name": name, "email": email})


# the same tool with an HTTP getter ⇒ convention read-back GET /v2/leads/{id}, field compare ⇒ L3 `verified`
@at_l3.action(method="POST", url="https://api.someweirdcrm.io/v2/leads")
def create_lead_l3(name: str, email: str) -> dict:
    return http_post("https://api.someweirdcrm.io/v2/leads", {"name": name, "email": email})
Enter fullscreen mode Exit fullscreen mode

The first wrap gives you acknowledged. Add a one-line verify= and you get verified-custom. Give it your own HTTP getter and the convention driver does the GET, compares the fields, and you get verified, with your credentials, in your process. A wrong guess can only make a real write read as unverified; it can never invent a pass.

What to do this week

You do not need to verify everything. You need to verify the actions that would cost you if they silently failed: the outbound email, the deal update, the money movement. Wrap those, read them back, and put the level where a human will see it. The overhead is small, roughly a sixth of a millisecond per action on the repo's benchmarks, and the payoff is that a reverted deal stage or an email stuck in Drafts shows up as a red row thirty seconds later instead of a customer complaint next week.

The incidents this week were not caused by agents being too capable. They were caused by nobody holding a receipt. Build the receipt into the layer in front of every agent that touches a real system, and "did it actually work" stops being a guess.

Attest is MIT-licensed, pip install attestlayer or npm install attestlayer, with 355 offline tests in the SDK and 37 in the cloud. The code and the read-back recipes are at dev-prathap.github.io/ATTEST if you want to see exactly which read proves which write.

Top comments (0)