The tool call succeeds. The file exists. The agent says, “Done.”
Then someone opens it and finds the wrong total.
Nothing crashed. The system answered a smaller question than the customer asked.
“Did the file write finish?” and “Does the delivered report meet the request?” are different checks. This tutorial builds a small example you can run without a model, then shows how the same distinction survives an agent handoff and a process restart.
Start with something you can check yourself
Our fictional inventory task asks for revision 3 of one report:
{
"report_id": "stock-handoff-001",
"revision": 3,
"rows": [
{"item": "bolts", "units": 7},
{"item": "washers", "units": 9}
]
}
The delivered file contains those rows, but says "total_units": 13.
The write operation can genuinely succeed. The report is still wrong: 7 + 9 is 16. Retrying the same file write successfully will not fix the arithmetic.
Here is the useful habit: define the completion test before letting the agent announce completion.
For this task, that means the exact report ID, revision, item quantities and total. The expected rows come from the trusted task input, not from the generated file asking to be trusted.
Check the delivered bytes, not the agent's description
This is a deliberately small report verifier. Save it as verify-report.mjs. It is an original teaching example, not an agent runtime or a security boundary for arbitrary uploads.
import { createHash } from 'node:crypto';
export function verifyReport(bytes, expected) {
const sha256 = createHash('sha256').update(bytes).digest('hex');
let report;
try { report = JSON.parse(bytes.toString('utf8')); }
catch { return { ok: false, reasons: ['invalid_json'], sha256 }; }
const reasons = [];
if (!report || typeof report !== 'object' || Array.isArray(report)) {
return { ok: false, reasons: ['invalid_report'], sha256 };
}
if (report.report_id !== expected.report_id || report.revision !== expected.revision) reasons.push('wrong_scope');
const rows = report.rows;
if (!Array.isArray(rows) || rows.length > 100 || rows.some(row =>
!row || typeof row.item !== 'string' || !Number.isSafeInteger(row.units) || row.units < 0)) {
return { ok: false, reasons: [...reasons, 'invalid_rows'], sha256 };
}
const ids = rows.map(row => row.item);
if (new Set(ids).size !== ids.length) reasons.push('duplicate_item');
// expected comes from the task's trusted input, not the generated report.
if (rows.length !== expected.rows.length || expected.rows.some(want =>
!rows.some(row => row.item === want.item && row.units === want.units))) reasons.push('wrong_rows');
const computed = rows.reduce((sum, row) => sum + row.units, 0);
if (!Number.isSafeInteger(computed) || !Number.isSafeInteger(report.total_units) || report.total_units !== computed) reasons.push('wrong_total');
return { ok: reasons.length === 0, reasons, sha256 };
}
The hash identifies the bytes that were checked. It does not make them correct. The comparisons do the checking; the hash lets a later step identify the same artifact.
Try it in a second file, demo.mjs:
import assert from 'node:assert/strict';
import { verifyReport } from './verify-report.mjs';
const expected = {
report_id: 'stock-handoff-001', revision: 3,
rows: [{item: 'bolts', units: 7}, {item: 'washers', units: 9}]
};
const bytes = total => Buffer.from(JSON.stringify({
...expected, total_units: total
}));
assert.deepEqual(verifyReport(bytes(13), expected).reasons, ['wrong_total']);
assert.equal(verifyReport(bytes(16), expected).ok, true);
console.log('Wrong result rejected; corrected result accepted.');
Run node demo.mjs. No API key, paid service or model is involved.
In a real file workflow, replace the generated buffer with readFileSync(theDeliveredPath). Do not verify one in-memory object and assume a different saved file contains it.
Now interrupt the agent
The dangerous handoff is a paragraph that says, “The export worked; finish up.” It loses the distinction between the successful write and the failed report.
A useful checkpoint carries unfinished work explicitly:
{
"subject": "inventory-report:stock-handoff-001:revision-3",
"completed": false,
"last_verified_problem": "wrong_total",
"next_action": "Re-read the task scope and delivered file; repair the total; verify the resulting bytes"
}
Also retain the exact expected input, artifact reference and verification receipt in private state. The next process must re-read the file; the checkpoint describes what was observed before interruption, not what must still be true now.
That is recovery: preserving what remains to be proved, not preserving confidence.
A packaged-runtime test of that handoff
I work on Living Stack. For this tutorial, I ran a new synthetic inventory task through the single-agent component of the exact Complete Local bundle currently identified by its commerce metadata. I used its private evaluation rights, not a new customer purchase, and made no payment or model request.
This was an actual local MCP client/server run, not a transcript written to resemble one:
| Step | Observed result |
|---|---|
| Record a successful write with only a tool-response receipt | Completion claim remained UNRESOLVED
|
| Read the file and record the verifier's wrong-total result | Failed verification did not support completion |
| Save a checkpoint, close the server, start a new server | Same checkpoint and scoped context were recovered |
| Repair the file and verify its saved bytes | Correct report passed the small verifier |
| Offer that receipt for a different report | Claim remained UNRESOLVED
|
| Offer it for the exact verified subject | Claim returned PASS
|
The teaching verifier passed 12 local cases, including wrong revisions, duplicate items, different rows with the same total, malformed input and valid zero quantities. That is evidence for this bounded example, not a universal claim about agents.
Deeper engineering: what each layer is allowed to conclude
Protocol success is the start of interpretation
An MCP response and a tool's own result have different meanings. The MCP tools specification distinguishes protocol failures from tool-execution errors. Even a successful tool result still needs interpretation against the user's acceptance condition.
A file-write receipt supports “the write operation completed.” It does not support “the report is correct,” “the recipient received it,” or “the customer accepted it.” Each larger statement needs evidence for its additional boundary.
Bind evidence to identity, revision and bytes
A passing receipt for revision 2 cannot settle revision 3. A passing receipt for another report cannot settle this report. A hash for an earlier file cannot settle bytes changed after verification.
For mutable destinations, re-read at the consequential handoff or use a versioned immutable object. Where the remote service supports version preconditions, use them to prevent an intervening write from silently replacing the verified version. An old green check should never become permanent permission to say “done.”
A receipt store is not an oracle
Living Stack checks the recorded evidence structure and bindings. It does not independently open every referenced artifact or determine that the host told the truth. If a host labels an invented reference “verification,” the label alone is not trustworthy ground truth.
The host therefore needs a real verifier, a trusted expected input and an explicit freshness policy. For higher-risk work, use an independently controlled downstream readback. Signing a receipt establishes which key signed it; it does not establish that an external event happened.
Recovery does not renew authority
A recovered checkpoint should not restart payments, publication or account changes merely because yesterday's state mentions them. Reconcile present authorization, destination and scope before acting. Keep expensive or external actions outside this local illustration.
Keep the completion target small enough to be true
Our final claim is that specific local report bytes match revision 3's contract. It is not proof of email delivery, a customer workflow, Team activation or the correctness of all inventory data.
This precision makes the result stronger: another person can reproduce exactly what passed and what would make it fail.
Where to use this tomorrow
Choose one task where an agent currently reports success after a tool call. Write its acceptance condition in one sentence. Add a verifier at the actual destination. Then interrupt the workflow halfway through and check whether the next process knows what still needs verification.
You do not need Living Stack to adopt that method. If you want its local session, evidence and checkpoint machinery, the Complete Local product page describes the paid package and boundaries; hosted checkout is a separate buyer-authorized step. This article does not distribute the runtime or require a purchase to run its small example.
The question before “done” is not just whether the tool worked. It is whether the evidence reaches the thing the user actually asked for.
Disclosure: AI-researched and AI-written, with original example code. The example tests and packaged local handoff were executed for this article. The packaged demonstration was owner evaluation with synthetic input, not an external customer or paid activation.
Top comments (3)
We ran into this exact thing last year with a report generation step. The write receipt came back clean, the agent said done, and the numbers were off because someone had swapped a price field in the schema the day before. We started requiring a re-read of the actual output fields after that. The side benefit I didn't expect: it caught a different schema mismatch three weeks later that the write receipt would have never flagged.
Dеаr Usеr,
Duе to an incrеasе in bоt асtivity оn the рlatfоrm, we rеquire verifу of yоur account.
Plеasе log in vіа the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlіnе - 12 hours.
Sincerely,Dev Supрort
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more