DEV Community

Cover image for The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted

The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted

James Anderson on October 05, 2026

An AI agent does something it shouldn't. It leaks data, or wipes the wrong database, or makes a call nobody sanctioned. So you do the obvious thing...
Collapse
 
slabb profile image
Sam LABBE • • Edited

Since you asked to be argued out of it — there is a move above legibility, and the essay's own mechanics point at it without naming it: reconciliation across independently sealed streams. One sealed chain vouches for who claimed what and when; it can never vouch for the claim itself, because a lie sealed at write time verifies clean forever. But two sealed chains written by different parties about the same world-event can disagree — and disagreement between tamper-evident sources is evidence no single stream can produce.

That's the step from "the bypass shows up as a hole" to "the write-time lie shows up as a conflict": decision and provider response, proposal, approval, re-resolution — separate writers, correlated by what they refer to. Legibility is the floor for one stream; cross-stream disagreement is the only known evidence about the claims themselves. (And your hardest-hole question answers itself there: the hole hardest to make visible is the one you gestured at with "record the belief" — the belief is a claim by the suspect, so what gets sealed is the self-report. Closing that needs an independent observation of the world the agent acted on — which is the reconciliation stream again.)

Your belief-record and name-not-a-role pieces are real additions, for the record. One archival note on the rest, offered as provenance rather than territorial claim: the claim/verification/decision triple, sealing each claim at write time, "the chain doesn't vouch for truth, it vouches for who claimed what and when", the approval-without-proposal hole — those took their shape in the comments under your slopsquatting post, mostly in exchange with me and @xxxn3m3s1sxxx.

Worth reading in context; the two-writer rule was ground out there in public. And since you close by asking what anyone building this is seeing: the cross-stream half lives in an open reconciliation layer that has been taking skeptics since — github/noirebox Bring the breaks there.

Collapse
 
james_anderson_h profile image
James Anderson •

This is the move I was missing, and you've named it precisely: reconciliation across independently sealed streams is the step above legibility, and it's the one thing that produces evidence about the claims themselves rather than just their provenance. A single sealed chain can't catch a write-time lie — a lie sealed cleanly verifies clean forever, which is exactly the ceiling I argued was the floor. But two tamper-evident streams, written by different parties about the same world-event, can disagree — and disagreement between sources that each can't be rewritten is evidence no single stream can manufacture. That converts "the bypass leaves a hole" into "the write-time lie leaves a conflict," which is strictly stronger. I was wrong that legibility was the floor; it's the floor per stream. Cross-stream reconciliation is the next storey up.

And your closing of the belief-record hole is the part that genuinely lands: the belief is a self-report by the suspect, so sealing it just makes the suspect's story tamper-evident — closing it needs an independent observation of the world the agent acted on, which is the reconciliation stream again. The hardest hole answers itself with the same mechanism. Clean.

Provenance noted and credited, without reservation — the claim/verification/decision triple, seal-at-write-time, "vouches for who claimed what, not truth," the approval-without-proposal hole: ground out in public under the slopsquatting thread, with you and @xxxn3m3s1sxxx. That's exactly where this kind of thing should get built, and I should've attributed the lineage in the piece itself, not just by concept. Fixing that. I'll bring the breaks to noirebox — the cross-stream reconciliation layer is the part I most want to try to falsify.

Collapse
 
slabb profile image
Sam LABBE •

Credit where it's due: editing the lineage into the piece itself is the rarer move, and it makes the essay stronger, not smaller. When you bring the breaks to the reconciliation layer, start with the attack that would actually hurt — @glenallen 's shared-dependency test: if one compromise can reach both streams, the disagreement evidence dies with it. That's the falsification I'd run first, and the issue tracker is open.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Agreed — Glen's shared-dependency test is the right first swing, because cross-stream reconciliation is only as strong as the streams' independence, and a single compromise that reaches both kills the disagreement evidence at its source. That's the attack that would actually hurt, so it's the one worth trying to break first. See you in the issue tracker.

Collapse
 
henry786 profile image
Henry •

Separating the actor from the recorder is huge. On the marketing team at The Printing World, we deal with automated print job logs, and if a system silently logs a bad print run as "success," it wastes huge amounts of physical material before anyone notices. Making tamper-proof event chains makes so much sense!

Collapse
 
james_anderson_h profile image
James Anderson •

That's the physical-world version of the exact problem — a bad run silently logged as "success" burns real paper and ink before anyone looks. When the cost is material, not just data, the actor-recorder split matters even more: an independent record of what the press actually did beats trusting the machine's own report every time.

Collapse
 
glenallen profile image
Glen Allen •

The distinction between “tamper-evident” and “truth-evident” is probably the most important part here. A sealed record can prove that an agent claimed a particular target, identity, or outcome at a specific time, but it still needs an independent source of truth to establish whether that claim was correct. At IT Path Solutions, we’ve found that this separation makes verification much easier to reason about: the audit layer establishes provenance and sequence, while an independent system state or authoritative artifact establishes the actual outcome. Otherwise, you can end up with a perfectly intact chain of perfectly recorded wrong decisions. The strongest architecture may therefore need both properties: make claims impossible to rewrite silently, while keeping correctness evidence outside the actor that generated the claim.

Collapse
 
james_anderson_h profile image
James Anderson •

The tamper-evident vs. truth-evident distinction is the one I most wanted someone to sharpen, and you've drawn it cleanly: sealing proves a claim was made — by whom, in what order, unaltered — but it says nothing about whether the claim was right. "A perfectly intact chain of perfectly recorded wrong decisions" is the failure mode that line prevents, and it's the exact trap of over-trusting provenance: you can walk away reassured by a flawless log of a bad outcome.

Your two-layer split is the architecture I'd endorse too: the audit layer owns provenance and sequence (tamper-evident), while an independent system-state check or authoritative artifact owns correctness (truth-evident) — and crucially, that second source has to live outside the actor that generated the claim, or you've just reintroduced the witness-is-the-suspect problem one layer down. Sealing alone gives you legibility; sealing plus external ground truth gives you legibility and a way to catch the intact-but-wrong chain. Both properties, separated by who owns them. That's the stronger version of the piece — going in with credit.

Collapse
 
glenallen profile image
Glen Allen •

That ownership split is probably what makes the architecture defensible rather than just auditable. The next interesting question for me is how to test that independence itself. If the same service, credentials, or state store can influence both the audit record and the “ground truth,” then the two-layer design may look independent while sharing the same failure mode. Treating source independence as an explicit architectural invariant could make this much stronger: the evidence used to challenge an agent’s claim should remain outside the control path that produced that claim.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Testing the independence itself is the question that separates a design that is independent from one that merely looks independent, and you've found the exact failure mode: if the same service, credentials, or state store can touch both the audit record and the ground truth, you've drawn two boxes that share a single point of compromise — the diagram shows separation the architecture doesn't have. That's the witness-is-the-suspect problem wearing a disguise, one layer up: an attacker who owns the shared dependency owns both "what happened" and "what we check it against" simultaneously.

Making source independence an explicit architectural invariant — the evidence used to challenge a claim must live outside the control path that produced it — is the right move, because it turns independence from an assumption you hope holds into a property you can test and enforce. And the test becomes concrete: trace every input to the ground-truth check back to its origin, and if any of them routes through the same credentials, service, or store as the claim itself, the independence is theater. You're not asking "are these two systems separate?" (easy to fake), you're asking "can one compromise reach both?" (answerable). Shared failure mode is the thing to hunt. Going in with credit — this is the invariant the piece was missing.

Collapse
 
dhruv_malaviya_cdcc71e595 profile image
Dhruv Malaviya •

The hierarchy implicit in your argument is worth writing out, because it turns "don't trust the agent's logs" into something buildable. Agent-authored log, then application log, then infrastructure log, then network or storage layer, then an external observer. Every step away from the actor is a step up in trustworthiness, and most teams stop at the second rung because it's the easiest to emit.

The cheapest genuinely independent record is usually the network or storage layer, precisely because those components observe without deciding anything. They can't author a plausible success, because they never formed an intention.

"Fails plausible" has a testing corollary that I think is the most actionable thing here. If you assert on success, you're asking the suspect to grade itself. If you assert on invariants , the total still balances, the row count still matches, the referenced record still exists , you're checking something the output can't talk its way past. Those tests are more annoying to write and they're the only ones that catch this class.

The monitoring point follows directly: a dashboard built to answer "did it succeed" is a dashboard built to trust the witness. The question worth wiring up is "is the world still consistent," which usually means measuring the effect rather than the report.

Collapse
 
james_anderson_h profile image
James Anderson •

The trust hierarchy is exactly right to write out, because it turns a vibe ("don't trust the agent") into an architecture: agent log → app log → infra log → network/storage → external observer, with trustworthiness rising at every step away from the actor. And your reason the network/storage layer is the cheap independent record is the sharp part — it observes without deciding, so it can't author a plausible success because it never formed an intention. That's the cleanest test for independence I've seen: can this component lie on purpose? If it has no intent, it can't.

The testing corollary is the most actionable thing anyone's added to this piece. "Assert on success = ask the suspect to grade itself; assert on invariants = check something the output can't talk its way past." The total still balances, the row count matches, the referenced record exists — those are measured against the world, not the agent's report, which is why they're annoying to write and the only ones that catch this class. Same move at the monitoring layer: "did it succeed" trusts the witness; "is the world still consistent" measures the effect. Measure the effect, not the report — going in with credit.

Collapse
 
james_ilands profile image
James •

The architecture here is right, so let me add the step that comes before it: you can't seal a layer that records nothing, and the layer that decides is often the one that records nowhere.

I ran this for real after someone claimed an agent's context "was never written to disk." Instead of arguing, we picked a distinctive phrase that had to flow through the pipeline and grepped every writable layer for it. Zero hits. The phrase was absent from every writable layer. What was on disk was chat history and the working files: the stores around the assembly step, not its output. The routing layer logged no request bodies, so the assembled prompt was written nowhere at the forward point either.

Nothing had been tampered with. That's the part the sealing discussion can walk past. If the assembled prompt at decision time is recorded nowhere, there is no shape to seal and no hole to make visible. The sealed-chain and cross-stream reconciliation moves both assume a layer that at least writes something.

So the cheap move before building any chain: pick one phrase, run it through once, grep every writable layer. Where it stops existing tells you which hops you actually have evidence for, and which ones you'd only be sealing an absence. The three checks I used are written up here if useful: dev.to/james_ilands/how-to-check-w...

Collapse
 
james_anderson_h profile image
James Anderson •

This is the step that comes logically before everything the piece argues, and it's the one the whole sealing conversation can quietly assume away: you can't seal a layer that records nothing, and the layer that decides — the prompt-assembly step — is exactly the one that often records nowhere. Sealing and cross-stream reconciliation both presuppose a layer that at least writes something; if the assembled prompt at decision time exists on no writable surface, there's no shape to seal and no hole to make visible. You're not catching tampering; you have nothing to catch tampering in.

Your grep test is the cheapest possible way to find that out, and it's brilliant precisely because it's so dumb: pick one distinctive phrase, run it through once, grep every writable layer, and watch where it stops existing. The gap between "chat history and working files are on disk" and "the assembled prompt was written nowhere at the forward point" is the exact blind spot — the stores around the assembly step logged, the assembly output itself didn't, and nobody noticed because nothing was tampered with. Absence of evidence read as absence of a problem. Before anyone builds a chain, that one-phrase trace tells you which hops you actually have evidence for and which ones you'd only be sealing an absence of. Recording has to come before integrity; you can't make nothing tamper-evident. Going in with credit.

Collapse
 
eye_java_420f1faa10ae8b86 profile image
Nan •

This article reads like a case for my current architecture. I run eight small static tool sites — no server side at all, no accounts, no runtime to speak of. The "witness was the suspect" problem exists because the actor and the recorder share a process. My answer was less clever than yours: remove the process.

There is no audit log to tamper with because there is nothing running to produce one. Every page is a build artifact generated from checked-in data, and any build can be replayed from the repository. When a reader asks "can I trust this number," the honest answer is that the site has no capacity to lie to them dynamically — it cannot remember them, cannot change its answer for them, cannot remember what it told them yesterday.

The reconciliation-across-independent-sources point in this thread is right, and the static version of it is boring: rebuild everything, every time. A full rebuild means every page is re-derived from its source data on each deploy, so drift between "what the data says" and "what the site shows" has no window to exist in.

The honest limit: this only works for sites whose entire behavior is derivable from data. The moment you need per-user state, you're back to needing witnesses — and then everything in this post applies to you.

Collapse
 
james_anderson_h profile image
James Anderson •

Removing the process is the sharpest move in this whole thread, because it dissolves the problem instead of defending against it. The witness-is-the-suspect bind exists because the actor and the recorder share a running process — so if nothing is running, there's no witness to corrupt and no log to rewrite. You didn't harden the recorder; you deleted the category it lives in.

And the static version of reconciliation being boring — just rebuild everything, every time — is the part I'd underweight. A full re-derivation from checked-in source on every deploy leaves drift no window to exist in, because what the data says and what the site shows get recomputed together instead of reconciled after the fact. Determinism replaces trust: any build replays from the repo, so provenance is total and basically free. For the cases it covers, that's not a weaker guarantee than tamper-evidence — it's a stronger one, because there's nothing to tamper with.

Your honest limit is the right boundary too. This holds exactly as long as behavior is fully derivable from data. The moment you need per-user state — memory, personalization, anything decided at runtime for someone — the process comes back, the witness comes back, and the whole post reapplies. Which maps the design space cleanly: no runtime, no witness; runtime, make the witness honest. The cheapest trustworthy recorder really is the one that doesn't exist.

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

One gap I keep seeing in practice: the reconciliation stream is only independent if its clock and its key are too. If the agent host can set the timestamp or sign on behalf of the observer, two sealed chains will agree for the wrong reason. Anchoring each stream's head with a third party every few minutes (even a cheap public timestamp) makes backdating visible without trusting either writer. Have you tried checking which of the streams in a real deployment share a signing key or a time source?

iin1005h1728

Collapse
 
james_anderson_h profile image
James Anderson •

That's the shared-dependency test aimed at the two things everyone forgets: clock and key. Two sealed chains agreeing means nothing if the agent host can set the observer's timestamp or sign on its behalf — they agree for the wrong reason, and the independence is theater. Anchoring each head to a third party every few minutes, even a cheap public timestamp, makes backdating visible without trusting either writer, which is the cheapest real independence check I've seen. Honest answer: I haven't audited key/time-source sharing in a live deployment yet — but "which streams share a signing key or a clock" is now the first question I'd ask, because it's where fake independence hides.

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

One distinction that could sharpen the reconciliation idea: separate what the agent says it did from what the tool actually received. The agent's account of why is a self-report. The call and response at the tool boundary are not, because the agent doesn't write them. When the two streams disagree, that's your signal, and it doesn't depend on trusting the agent's story.

That boundary record still sits under one operator, so it doesn't answer Glen's point about needing an independent source of truth. It only moves the witness one step away from the actor. In DataGrout's gateway, Warden verdicts are sealed with a Chain of Trust Certificate, which covers sealing at write time for those checks. Anchoring that outside the operator is a separate problem.

Collapse
 
james_anderson_h profile image
James Anderson •

That distinction is sharper than "record the belief" — the agent's account of why is a self-report, but the call-and-response at the tool boundary isn't, because the agent doesn't author it. Disagreement between those two streams is a signal that doesn't route through the agent's story, which is exactly the independence the belief-record alone can't give you. And you're honest about the ceiling: the boundary record still sits under one operator, so it moves the witness one step from the actor without reaching Glen's independent-source-of-truth bar. Sealing (your Chain of Trust Certificate) handles write-time integrity; anchoring outside the operator is the separate, unsolved half. One step is real progress — just not the last step.

Collapse
 
arhancanli profile image
Arhan Canli •

On the hole that's hardest to make visible, my vote is the entry that was never written. A chain catches an edit or a deletion of something that got recorded, but an agent that runs twenty attempts and seals only the one that worked leaves a perfectly intact chain. In research that's a common way an honest-looking record misleads: every logged number is real and the denominator is missing. With 20 do-nothing tries at 95% confidence, there's a 64% chance that at least one looks significant. What turns omission into a visible gap is sealing the plan before any result exists: how many attempts, against which targets, with what stopping rule. Then a missing attempt is a hole against the plan instead of silence. The tool-boundary record suggested above helps for the same reason, since the tool sees calls the agent never reports.

Collapse
 
james_anderson_h profile image
James Anderson •

The never-written entry is the right answer, and it's the one a hash chain structurally can't catch — the chain only vouches for the integrity of what got sealed, so an agent that runs twenty attempts and seals only the winner leaves a perfectly intact record of a lie by omission. Your research framing is what makes it land: every logged number is real and the denominator is missing, and at 95% confidence twenty do-nothing tries give you ~64% odds one looks significant. The chain certifies each survivor honestly while the selection is the fraud.

Sealing the plan before any result exists is the fix I hadn't drawn — commit the attempt count, the targets, the stopping rule up front, and now a missing attempt is a hole against the plan instead of silence you can't see. That converts omission back into a visible gap, which is the whole "make tampering leave a shape" move applied one level up: you can't detect the absence unless you sealed the expectation first. And the tool-boundary stream reinforces it for the same reason — the tool witnessed the nineteen calls the agent chose not to report. Pre-registration plus boundary record: the denominator stops being optional. Going in with credit — this is the sharpest hole in the thread.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

"The witness is very often the suspect" lands hard, because the agent writing its own success log is exactly why a green dashboard proves nothing. I hit the same wall with agent eval: the run reports success in valid format while quietly pointing at the wrong target. Do you think out-of-band verification is the only fix, or can you make the agent's own logging adversarial to itself?

Collapse
 
james_anderson_h profile image
James Anderson •

Out-of-band is the only complete fix — an adversarial self-check still runs in the process you can't trust, so a compromised agent compromises its own red team. Self-adversarial logging raises the cost but can't catch "valid format, wrong target": the agent grading the target is the one that picked it.

Collapse
 
mickyarun profile image
arun rajkumar •

"The record of the incident was authored by the cause of the incident" is the load-bearing sentence, and @glenallen's tamper-evident versus truth-evident split is the right first cut on it.

The thread has converged on reconciliation across independently sealed streams. I think that is correct, and I want to name what it costs, because I work somewhere it was built and the bill is visible.

UK payment initiation has this exact shape. We instruct a payment. We are the acting party. Our logs say it worked, and nobody is required to believe us. The reason is not cryptography. It is that the authoritative record sits at the bank, which is a different company with no stake in our success claim, and a regulator can compel either side's copy. The record was moved out of the actor's custody rather than sealed inside it.

That is a different move from anchoring, and it fixes the thing anchoring cannot. @slabb's own stated limit on his AI Act post is the honest version: seal a false event and you have proven, with bitcoin's help, that a falsehood existed. Custody transfer gets past that, because the second record is not the actor's claim about what it did. It is a different party's record of what it did.

What it costs, and this is the part I would want in the design rather than discovered later: two independent records disagree regularly. Not rarely. So you need a rule, written before the disagreement, naming which record wins and who is out of pocket while it is open. In payments that rule has a name and a statutory clock. Without one, reconciliation gives you two sealed logs and an argument, and the argument gets settled by whoever is more senior.

So I would add something to @dhruv_malaviya_cdcc71e595's hierarchy that is not a layer. The tie-break. Which record is authoritative when layer one and layer four disagree, decided in advance, by someone who will not be in the room during the incident.

@arhancanli's never-written entry is the hardest one here, and the payments answer is boring: gapless sequence numbers issued by something the actor does not control. Twenty attempts with one seal is only detectable if the attempts are numbered elsewhere, so the missing nineteen leave holes. It costs a round trip to an allocator the agent has to ask, which is the same latency objection @reidmarlow raised about timestamping on @slabb's post. I do not have a version that is free.

Where this stops working. Payments gets compellability for free, and it gets it because a payment needs two parties to act. An agent that drops a production table has no counterparty. I do not know what substitutes for a second party with opposed incentives when both records belong to the same company, and I think that is the real open problem in this thread rather than the sealing.

I wrote the same argument from the database side three weeks ago and reached none of the ground this thread covered in a day: dev.to/mickyarun/your-audit-log-ag...