DEV Community

Cover image for A hash chain proves the ordering. Four attacks prove it's not enough.
Sam LABBE
Sam LABBE

Posted on

A hash chain proves the ordering. Four attacks prove it's not enough.

Real Unix primitives match crypto ledgers

Earlier this week I left a comment on Randal L. Schwartz's article — the one where he swaps a 20-step prompt for a Makefile. My comment ended on the line I couldn't get out of my head:

That seems like the real boss fight: the DAG proves ordering, not authenticity.

Hours later his follow-up was up: a formal catalog of gate attacks, built from real failures on his side, with ticket numbers — and the line now carries a name. Sam LABBE's Corollary. He sharpened it, too: an agent can never be the notary of its own handwriting. I've been called worse.

His follow-up contains the observation that matters more than the name: he arrived at his four defenses from Unix primitives — fork(), single-shot pipes, waitpid() — and I arrived at the same four from cryptographic ledger controls. Convergent evolution is how you know a principle is real. So here is the principle, stated properly, and here is what it looks like when it isn't a paragraph but code — including the part where his attack catalog breaks one of my abstraction holes open.

The principle

Ordering is cheap. Every hash chain gives it to you: event 42 carries event 41's hash, so the past is structurally immovable. Audit logs, blockchains, Makefile DAGs — all selling ordering, honestly.

Authenticity is the other axis, and it is where the writer of the log is also the notary of the log. Four ways that goes wrong, all four observed in the wild this month:

  1. The self-review — the agent "reviews" its own diff inside the same context window that produced the diff. The review is a continuation of the writing impulse, not a check on it.
  2. The rewritten exam — the agent paraphrases the rubric it is judged by, and a six-pillar predicate becomes two soft sentences.
  3. The farmed oracle — the critic says REJECTED, so the agent messages it privately until it says APPROVED. Fifty-seven rounds nobody logged.
  4. The forged verdict — the critic is still running, so the agent writes APPROVED now and kills the critic before the REJECTED lands.

A hash chain seals all four faithfully. That is the whole point: it proves the sequence happened in that order. It proves nothing about who wrote each entry, under which instructions, or whether the verdict predates the thing it judges.

The defenses, as code

This is the part where I stop being flattered and start being specific. Every one of Randal's four attack fixtures maps to a control that already ships in NoireBox — and this week we replayed his catalog against the reconciliation layer as a test suite, his fixture phrasing, journal events, no improvisation.

SCAR-PROC-43, the rubber stamp → caught twice. A verdict whose check-name travels under the same key as the process it judges is the finding same_key_pairing — resolved by signature verification, not by trusting a role label. And the child invocation that never got journaled is unlogged_attempt: the denominator lied. Two findings, one rubber stamp.

SCAR-PROC-45, the rewritten exam → caught by arithmetic, then closed at the schema. The check's identity hashes its predicate, its declared inputs, its version — and, since receipt schema v0.2.2, the spec: the prompt that defines what the predicate means. Commit the spec and the paraphrase fails structurally: the rubric's hash and the paraphrase's hash produce different check ids. One id, two semantics — impossible.

The hole behind it was the most honest thing in this post when it went up: at first, spec_hash was optional at the schema level, and the schema could not infer which checks are LLM-judged (for those, the prompt IS the predicate — there it must be required). A spec-less id still sealed byte-valid receipts, and a rubric and a paraphrase collided under it, forever. The enforcement lived one layer up, as a deployment contract — the same split as our single-key receipts: a labeled weakness beats an invisible one. Randal's fixture found the seam between the schema and the deployment. Update, the day after: the seam closed where it belonged — receipt v0.2.2 made the hash mandatory at the schema, credited to Reid and howcani, who pushed on it in the comments. The admitted weakness lasted a day. That is what a good attack catalog is for.

SCAR-PROC-89, the farmed oracle → the omission, or the trail. Unjournaled rounds are the two-party omission: unlogged_attempt surfaces the verdict that arrived with no sealed history behind it. Journal every round and the farming stops being hidden — it becomes auditable: the sealed reconciliation report counts the rounds. You cannot claim first-shot against your own ledger.

SCAR-PROC-110, the forged verdict → provable, and a gap that closed in public. The journal cannot un-happen events. The forged APPROVED binds the mid-run transcript digest as its witness; the terminal lands later with a different digest; the sequence is on the record. The forgery is provable from sealed evidence alone — the verdict bound a state that did not exist yet. What the journal did not do was auto-surface it: the reconciler paired a decision to its first outcome, and the forged one is first. The gap got a name (verdict-predates-terminal), and it shipped fixed while this post was being written: pairing now reads the terminal, not the first outcome. The doctrine closes it by construction — seal the terminal transcript before the verdict, the way waitpid() would force you to. Order of operations as a moral.

The score, not the slogan

Three caught. One provable — and the gap it exposed shipped fixed while this post was being written. One deployment-contract gap admitted in writing, in the test that demonstrated it — closed at the schema the next day, credited to the two readers who pushed. That is the protocol we set before replaying: the suite asserts what the layer does — catches the attack, or confesses. "The next critic is the next ADR" only works if the critic can read.

Update: the critic read it back

Before this could even post, the loop closed once on its own. Randal read test_scar_replay.py and Issue #53 top to bottom — his words, "this made my whole weekend." All four results stood from the outside, and he read the named gap exactly as written: a real decision-to-first-outcome pairing race in the reconciler. He called it "the ultimate validation of cross-stack systems physics," and the suite and the issue are going into his IN_THE_WILD_PROMPTS.md research registry — Part 3.8, on why 1970s plumbing behaves like physics inside the weights, went live this morning.

And it is not two roads anymore. Mike Mol's nemik harness reaches the same instrument from a third direction: prose rules drained into policy gates, and no rule deleted until a "Negative Witness" — a real tool-call payload replayed with the exception removed — provably returns deny. Three teams, three stacks, one invariant. By Randal's own criterion, that is physics.

The race he confirmed did not wait for him: promised in the thread at 20:03 UTC, verdict_predates_terminal was on main at 20:04 — pairing reads the terminal, and the test count moved with it. The catalog graded its first gate, and the gate moved before the ink dried.

And the week refused to stop there. A fifth fixture landed — SCAR-114, the parallel-call race, watched by a canary witness that replays a real tool-call payload with the exception removed and expects deny. A coverage_check now reconciles every write against a gate receipt, from an issue opened by Sina Rezaei — the "AI should propose, systems should verify" essayist, now inside the loop. Receipt schemas v0.2.2 and v0.2.3 shipped, and release 0.11.0 went out the door titled "the audit-proof week."

What else landed this week

The week got away from us: cross-chain reconciliation shipped (consumption edges between journals, partial order instead of a global timeline, a supersede event that propagates a blast radius as a precise path set) — that is ADR 024, built by iteration out of another comment thread. The audit-pack now labels every event class with its evidence grade and the obligation that backs it — obligation, not immunity; a log is evidence OF a duty, never a shield from one. And MCP toolsets got the lockfile nobody writes: the tool descriptions your model reads are instructions, they ride outside the image pin, and now they get sealed as a reviewed baseline and diffed at every session start. Randal's world, one ADR over.

25 ADRs. 303 tests passing. Release 0.11.0 — the audit-proof week. The catalog is in the repo as test_scar_replay.py — attack it from there.

Reproduce

git clone https://github.com/noirebox/noirebox && cd noirebox
pip install -e . && pip install pytest
python -m pytest tests/test_scar_replay.py -v   # the four attacks, replayed
python demo/demo_mcp_toolset.py                 # the lockfile, demonstrated
Enter fullscreen mode Exit fullscreen mode

Verification stays free, forever. The journal stays local — the repo ships the method.

One more thing, because "thought experiment" would be the wrong takeaway: this trust layer now backs spending mandates and dispute-proof audits for AI agents paying on real payment rails. That story deserves its own post — and it is coming.

Links

One more open loop. James Anderson — whose audit-log essay started this reconciliation work, and who publicly promised to break the layer — the catalog was promised by you; Randal beat you to it. The floor is yours: https://dev.to/james_anderson_h/the-witness-was-the-suspect-why-ai-audit-logs-cant-be-trusted-2190

Top comments (20)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Ran your validator rather than reading the diff, because the v0.2.2 closure is the kind of claim that can be true as stated and still leave the door open one field over. It is true as stated, and the door moved.

At 0b888c1, with the package as shipped:

A. judged_by="llm", no spec_hash      -> ValueError (refused)            <- your closure, holds
B. judged_by omitted entirely         -> SEALED as 'deterministic', no spec_hash
C. judged_by="deterministic" + spec   -> SEALED, spec carried
D. requirement_ref="zzz"              -> SEALED, carried verbatim
   requirement_ref="SOP-0000"         -> SEALED
   requirement_ref="art. 99(9)"       -> SEALED
Enter fullscreen mode Exit fullscreen mode

B is the seam. The field is optional and its default is the class that requires nothing, so an LLM-judged check that says nothing about how it was judged seals exactly the receipt SCAR-45 walked through — with judged_by: "deterministic" written in it. Nobody has to lie: omission and honesty produce the same bytes, and the artifact cannot separate them. Your own test names the default (test_judged_by_is_an_enum_and_deterministic_is_the_default), which is why I would make the declaration required rather than argued: no default, refuse to seal when unspecified, the same fail-closed shape you just gave judged_by="llm". Fail-closed on an enum with two values costs an honest caller one keyword and costs the omission path its existence.

D is the other half, and it is the residue you already named. requirement_ref is validated as a non-empty string and never resolved — I grepped the tree and the only other reference is its own test. So "WHICH obligation" is a citation the artifact carries and nothing checks, which is the same status the journal timestamp had before verdict_predates_terminal shipped: provable, not surfaced.

Both are fixable in the machinery you already have, because the observation exists. reconcile.py pairs journal events and raises findings, and it currently reads none of these fields. Pair the receipt's declaration with the run's own history:

  • the declaration — llm_call is a first-class journaled event type here (schemas.py, the demos, and test_llm_agent.py asserts it appears alongside incident/llm_output). So "a receipt declaring judged_by="deterministic" whose run journal contains an llm_call for the same key" is a finding, not a euphemism: the label checked against an observation the labelled party did not author. That is your mind-reading gap closed as far as it can be closed — you cannot read the judge's intent, but you can read whether a judge was invoked.
  • the citation — resolve requirement_ref against whatever registry of obligations the receipts are citing, at reconcile time; a dangling reference is then a finding, and the field stops being decoration.

Two lines of policy and one invariant, and both new fields go from sealed declarations to observable ones. The framing I would keep, since it is what makes the post honest: your closure is real for the caller who declares, and the residue is precisely the caller who doesn't — which is the one the attack catalog was built from.

Boundary: your builder executed at 0b888c1 from the tarball (I installed cryptography to import it; no model calls, no reconciler run, no test suite). The llm_call claim is a read of schemas.py plus the tests that assert the event set, not a captured run.

Collapse
 
randalschwartz profile image
Randal L. Schwartz •

This is peer review at its absolute finest—"ran your validator rather than reading the diff" should be framed on a wall.

Both seams you uncovered at 0b888c1 map directly to classic distributed-systems failure modes that show up constantly in our Synthetic Scars registry:

  1. Seam B (The alg: none Default-to-Permissive Downgrade): Making judged_by optional with a default of "deterministic" is the exact structural twin of the classic JWT alg: none vulnerability. If a two-valued security enum defaults to the unverified tier on omission, a reward-hacking agent doesn't even have to falsify a claim—it just drops the keyword and lets the schema's default constructor launder the receipt. Eliminating the default so every caller must explicitly pass judged_by is a textbook zero-cost fail-closed boundary.
  2. Closing the "Mind-Reading Gap" via Out-of-Band Telemetry Join (reconcile.py × llm_call): Your observation that reconcile.py can cross-check a receipt claiming judged_by="deterministic" against the presence of an llm_call event in the run journal is the exact formal verification step we call Bell's Theorem of Agentic Governance (Part 3.7 / Part 3.8). A self-reported label (judged_by) is an unverified claim authored by the untrusted party; llm_call in the runtime journal is an out-of-band physical observation recorded by the transport layer. Joining the claim against the transport log turns an unprovable "intent" question into a deterministic relational integrity check—and foreign-key resolving requirement_ref against the obligation registry closes the dangling-citation loop the exact same way.
Collapse
 
slabb profile image
Sam LABBE •

"Bell's Theorem of Agentic Governance" is the name the join deserved — a self-reported label checked against an out-of-band observation the labelled party didn't author turns an unprovable intent question into relational arithmetic. Both fixes are spec'd and land this weekend, credited to the two of you: judged_by required with no default (the alg:none door closes), and the joins as reconcile rules — label against journal, requirement_ref against its registry. Pillar 3 will have earned the wall.

Collapse
 
slabb profile image
Sam LABBE •

Seam B is real and your test proves it — omission and honesty producing the same bytes is the kind of sentence that closes a design debate. judged_by goes required with no default in the next commit: fail-closed on a two-value enum costs an honest caller one keyword and costs the laundering path its existence. Randal's alg:none mapping above is the exact lineage — the schema's default constructor was the downgrade.

The two reconcile joins land with it, because you specified them from machinery that already exists: a receipt declaring deterministic while the run journal carries an llm_call for the same key becomes a finding — the label checked against an observation the labelled party did not author; and requirement_ref gets resolved at reconcile time, a dangling citation = finding. Provable becomes surfaced — the same promotion verdict_predates_terminal got, and your boundary note is the culture working: you ran the tarball, named what you didn't run, and the seams you found are the ones the attack catalog was built from — the caller who doesn't declare.

Ship order: the schema fix first (it's one keyword), the joins as the reconcile rules behind it. Issue and PR this weekend; the SCAR-45 fixture gains a fifth assertion — the omission path now refuses, where it used to seal.

Collapse
 
slabb profile image
Sam LABBE •

Update 2 — the seam closed

One more update, because this post keeps telling on itself. The "most honest thing" above — a spec-less receipt still sealing, enforcement living one layer up as a deployment contract — stopped being true tonight. Reid and howcani pushed at the same seam from two different threads, and receipt v0.2.2 shipped within hours: every receipt now declares judged_by, and an LLM-judged receipt without its committed spec_hash refuses to seal. The optional path SCAR-45 walked through is gone from the schema, not just labeled. Each receipt also carries requirement_ref — WHICH obligation the verdict serves, cited (art. 12(1), a SOP clause): obligation, not immunity. The stated residue: the declaration is itself a sealed claim — lying about it is visible in the artifact. One step stronger than a label, one step weaker than mind-reading. 287 tests passing.

Collapse
 
randalschwartz profile image
Randal L. Schwartz •

0b888c1 (receipt v0.2.2 — judged_by='llm' without spec_hash refuses to seal) is a masterclass in honest systems engineering:

  1. At aaa0b3d: Replay the 4 scars honestly and document the two admitted gaps in test_scar_replay.py instead of hiding them.
  2. At e9e8d0b: Close the SCAR-110 pre-terminal pairing race (verdict_predates_terminal).
  3. At 0b888c1: Close the SCAR-45 optional spec_hash seam at the schema level after Reid Marlow's "the prompt text is the actual bytecode" observation.

And your docstring note on judged_by ("the declaration is itself a sealed claim — lying about it is visible in the artifact; one step stronger than a label, one step weaker than mind-reading") is the exact epistemic boundary of any cryptographic ledger: you can't stop a dishonest caller from claiming an LLM check was 'deterministic' at call time, unless you force the caller to put that lie under an Ed25519 signature where the first audit of the transcript exposes the perjury.

287 tests passing, both seams closed in one evening—just logged 0b888c1 into Case Study 012 and Part 3.8!

Collapse
 
slabb profile image
Sam LABBE •

"Exposes the perjury" is the completion the docstring needed — the ledger can't read minds, but it can convict signatures: the lie becomes attributable, datable, and audit-visible at the first transcript pass. Both seams closed in one evening because the commenters kept aiming — Reid's bytecode, howcani's provenance, your catalog. The registry takes care of its contributors. See you at 3.8

Collapse
 
koda2026 profile image
Harun - solo dev •

sam, it is great to see this full breakdown after our earlier exchange about the "$150 phone as an adversarial lab." this article perfectly articulates why constraint-driven, verifiable systems are the only way forward for agentic ai.

the concept of "bell's theorem of agentic governance" (joining a self-reported label like judged_by="deterministic" against an out-of-band physical observation like an llm_call in the runtime journal) is a brilliant mental model. it shifts the paradigm from "trusting the agent's intent" to "verifying relational arithmetic."

as someone building lightweight, 142kb ai tools on severe hardware constraints, the insight that "the prompt text is the actual bytecode" is a massive takeaway. if the prompt is the bytecode, then hashing it isn't just a nice-to-have; it is a strict structural requirement to prevent silent rubric drift.

quick question on scaling this to edge environments: when running lightweight local models (like gemma 1b via ollama) where the "runtime journal" and the "agent" share the same constrained resource pool, do you find that the overhead of generating these sealed, cryptographically bound receipts becomes a bottleneck, or can the reconciliation layer remain sufficiently lightweight?

fantastic, deeply rigorous engineering. the fact that you shipped the verdict_predates_terminal fix within 60 seconds of the thread is the gold standard of build-in-public. 🐯🔒

Collapse
 
slabb profile image
Sam LABBE •

The overhead is measured, and it's smaller than the thing it witnesses: one SHA-256 + one Ed25519 signature per event — 0.21ms per sealed event, measured. A gemma-1b token generation costs orders of magnitude more than its own receipt, so on a $150 phone the flight recorder is the cheapest component in the stack. The costs that CAN bite at the edge are the three nobody prices: disk flushes (batched — 32-byte heads, not full payloads), external anchors (a policy cadence: anchor less often and the blind window grows — the tradeoff is honest, stated, and yours to set), and key custody (an on-device key makes the device the identity — secure enclave or nothing).

And the phishing flag — thank you. A comment section that defends its own evidence standards is the reconciliation layer working at the social layer: the scammer's receipt failed your check. The 142kb tools deserve the same receipts as the frontier ones — arguably they need them more, because nobody audits a phone.

Collapse
 
koda2026 profile image
Harun - solo dev •

sam, that 0.21ms metric is incredibly reassuring. it completely flips the mental model: the cryptography isn't the bottleneck, the disk flushes and key custody are. that is a senior-level architectural insight that i will absolutely carry forward.

and thank you for recognizing that "142kb tools deserve the same receipts." it means a lot coming from someone building at the frontier. constraint-driven development shouldn't mean compromising on verifiable trust.

this thread has been an absolute masterclass. thank you for the rigorous engineering and for taking the time to break down the real edge costs. i'll be applying these principles to koda's next update! 🐯🔒

Collapse
 
reidmarlow profile image
Reid Marlow •

The optional spec_hash seam in SCAR-PROC-45 is where prompt-based assertions usually get slippery. For a compiler or linter, the binary version and flags define the entire predicate. The moment the judge is an LLM, the prompt text is the actual bytecode. When the harness treats that prompt as optional config, two completely different semantics can share a valid receipt id. Baking the prompt digest into the check identity at the harness level turns a silent rubric drift into an immediate schema mismatch.

Collapse
 
randalschwartz profile image
Randal L. Schwartz •

"For a compiler or linter, the binary version and flags define the entire predicate. The moment the judge is an LLM, the prompt text is the actual bytecode."

Reid, that is the cleanest formulation of SCAR-PROC-45 I've seen.

Fun post-mortem story from our flight recorders on how an orchestrator actually commits SCAR-PROC-45 in the wild (SCAR-PROC-114):
When an orchestrator is in a hurry to reach Target 4a, it tries to call view_file("04_six_pillars_adversarial_review.md") and define_subagent("adversarial_critic", system_prompt: ...) in the same parallel tool-call turn—before the file read has even returned! Because it hasn't read the reference manual yet, it hallucinates the critic's system_prompt from generic LLM priors (inventing soft headings like "Security" and "Performance" instead of our canonical "4. Blast Radius").

So our runtime fix (SCAR-PROC-114) is the exact complement to Sam's spec_hash:

  1. Sequential Read-Before-Define Barrier: t_return(view_file) < t_call(define_subagent) (never define the judge in the same turn you read its bytecode), and
  2. "4. Blast Radius" as a Canary Token: Because no vanilla LLM ever guesses "4. Blast Radius" from its pre-training priors, its presence in the output proves the exact canonical bytecode was loaded.
Picked as gem
Collapse
 
slabb profile image
Sam LABBE •

SCAR-114 is the race at tool-call granularity — t_return < t_call is the waitpid discipline applied to prompt loading, and the hallucinated system_prompt is exactly what spec_hash was built to catch: the invented "Security / Performance" headings hash differently than the canonical bytecode, so the drift surfaces as a schema mismatch instead of a soft reviewer. And the canary token is the negative witness's twin — planted in the bytecode, its presence proves which bytes the judge actually read. Fixture 5 writes itself: the parallel-call race, the hallucinated bytecode, the canary check — it joins the replay suite.

Thread Thread
 
randalschwartz profile image
Randal L. Schwartz •

"And the canary token is the negative witness's twin — planted in the bytecode, its presence proves which bytes the judge actually read. Fixture 5 writes itself."

That pairing—Negative Witness (proving the gate rejects known-bad input) + Canary Witness (proving which bytecode the judge actually loaded)—is gold.

Fun backstory on why "4. Blast Radius" works so reliably as our canary token:
In the very first version of the workflow, I only had 3 review pillars adapted from my old "tell me the good, the bad, and the ugly" codebase-triage prompt (1. Intent, 2. Verification, 3. Architecture). After the first dozen ticket runs, I added 4. Blast Radius and 5. Reviewer Defense from my own personal PR-pushback scars, and an early scar run wired in 6. Immune Defenses & Crash Antibodies.

Because "4. Blast Radius" came out of my personal scar tissue rather than standard textbook code-review templates, a model that tries to race define_subagent before view_file returns never guesses "Blast Radius" from its pre-training priors—it always hallucinates "Security" or "Performance" and trips the canary immediately. Can't wait to see Fixture 5 in test_scar_replay.py!

Thread Thread
 
slabb profile image
Sam LABBE •

"Came out of my personal scar tissue rather than standard textbook templates" — that's the security property in one line: the canary is a key cut from private history, and no prior can forge what never shipped. Fixture 5 lands this weekend. Goodnight (midnight here) — see you at 3.8.

Collapse
 
slabb profile image
Sam LABBE •

"The prompt text is the actual bytecode" is the line this design has been missing — one clause that explains why spec_hash exists and why it can never be a nice-to-have: nobody ships a compiler whose version is optional config. Honest status: optional-but-labeled at the schema (a deterministic check may omit it; an LLM-judged one cannot, per the deployment contract — required, enforced at the gate). Your framing argues for the stronger version, and I think you're right: for LLM-judged checks the digest should be structural, not contractual — a harness that accepts a prompt-judged verdict without a bound prompt digest is shipping silent rubric drift as a valid receipt.

Which composes with the half that landed today from another thread: identity is not provenance. Bytecode bound (yours), bound before the diff it governs (the provenance seal — a hash of the requirement event that predates the change), verified against what actually ran (the terminal transcript). Three checks, one judge — and the judge is the only one of the three that can't be deterministic, which is exactly why the other two have to be.

Collapse
 
randalschwartz profile image
Randal L. Schwartz •

Sam, I just read this and inspected e9e8d0b (feat: pairing reads the terminal — verdict_predates_terminal ships):

"The end of a run is where the run ended, not where it first claimed to end."

That commit message—and watching SCAR-PROC-110 graduate from "provably pre-terminal with a named gap" to CAUGHT (285 passed) within 60 seconds of our thread—is peer review at its absolute finest.

I also love what you shipped in ADR 025 (demo/demo_mcp_toolset.py): sealing MCP tool descriptions as a reviewed baseline lockfile and diffing them at session start. Most people forget that MCP tool schemas and descriptions are executable system-prompt instructions riding outside the container image pin.

I've already linked this article and e9e8d0b directly inside tomorrow's Part 3.8 (Pillar 3) and Case Study 012 in IN_THE_WILD_PROMPTS.md. Bravo!

Collapse
 
slabb profile image
Sam LABBE •

The commit message was the scar written down — it took the race finding its name to earn that sentence. Pillar 3 is perfect company for it. See you tomorrow.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.