DEV Community

Cover image for The AI Act went live two months ago. Your agent's logs are still editable.
Sam LABBE
Sam LABBE

Posted on

The AI Act went live two months ago. Your agent's logs are still editable.

In April, an AI coding agent took out a company's production database. One GraphQL mutation against the Railway API, nine seconds, and the production volume was gone. The backups attached to it went too. The founder told the story himself; the agent, reports said, then wrote a confession.

The confession made the rounds as a punchline. The detail that should have kept you up at night was quieter: the most complete witness to what happened was the agent itself, narrating its own actions in a transcript nobody outside the company could verify. And nobody doubts the story. But nobody doubting it isn't a property of the evidence. It's a property of the narrator's mood that day.

Two dev.to posts framed this summer's version of the debate. Prompt Injection Is the New SQL Injection is the one everybody quoted, and it ends where these conversations always end: defenses that shrink the blast radius from "catastrophic" to "survivable", plus a call to audit what your agent reads against what it's allowed to do. Its companion, Who's Accountable When the AI Was Just Following Instructions?, doesn't name a culprit. It asks for one thing instead: "a record of who authorized what — an unforgeable one, not logs the acting system can quietly rewrite."

That sentence is the whole problem. Neither thread shows how to build the thing it asks for. And both conversations live in the first gap. The unforgeable record is second-gap territory, and the law landed there two months ago.

Quick map, because I'll lean on it. Any AI system has two gaps. What it reads versus what it's allowed to do: the exploit gap. What it did versus what its logs say: the cover-up gap. This post lives in the second one.

The two gaps — Article 12 lives in the second one

Two months ago, the clock started

And while everyone argued about blame, the law moved. On August 2 the EU AI Act passed its general application date: it now applies everywhere its text doesn't schedule otherwise. The big "otherwise" is the high-risk block of Annex III, which lands on December 2, 2027. Article 12 sits in that block, and it says high-risk systems must technically allow for the automatic recording of events (logs) over their lifetime.

Fourteen months is nothing when the thing being retrofitted is logging that has to survive an auditor. Most agent stacks will check that box with a database table the operator owns. A log you control is a log you can edit, and on the day an authority asks to see it, "we have logs" and "we can prove these logs were never touched" are two very different sentences.

What the text actually asks for

Article 12 stops being abstract once you read it against an agent stack. The system must record events automatically over its lifetime, with enough detail to trace how it functioned, including the situations that pose risk, and to flag substantial modifications. The deployer keeps those logs for at least six months. On request, they go to the competent authority.

In agent terms:

  • every tool call, with arguments, results, and which model decision produced it
  • every LLM call. Including the transcript the model was fed — yes, that transcript
  • config changes and substantial modifications, versioned
  • timestamps that don't depend on the honesty of the machine writing them
  • and retention you can prove, not retention you describe in a policy nobody reads

Nothing in that list scares anyone. Every product in this space logs it today. The verb that hurts is the last one in the paragraph: provide. Whatever you hand over has to survive someone whose job is to distrust you.

"Automatic" is doing the heavy lifting

Automatic recording is satisfied by the agent writing its own log. Nobody types it, so it's automatic. But a log written by the system it describes, on infrastructure owned by the operator, is a confession written by the suspect.

Append-only flags stop the sloppy mistakes and nobody else. Hash chains stop outsiders. Neither stops the key holder: they can regenerate the entire history, fabricated events, recomputed hashes, re-signed seals, and end up with an internally consistent journal that verifies against its own key with flying colors. Which proves exactly one thing — that somebody with the key built it. Not when. And when is the whole game. A seal created after the fact tells you nothing about what existed before it.

A witness you can build from parts you already have

The fix is a timestamp from someone outside your trust boundary. The pieces are standard, and the core fits on one screen. The chain first:

import hashlib, json
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey

key = Ed25519PrivateKey.generate()

def seal(prev_head: bytes, event: dict) -> dict:
    payload = json.dumps(event, sort_keys=True).encode()
    head = hashlib.sha256(prev_head + payload).digest()
    return {
        "event": event,
        "head": head.hex(),
        "prev": prev_head.hex(),
        "sig": key.sign(head).hex(),
    }

GENESIS = bytes(32)  # head of an empty journal
journal = [seal(GENESIS, {"kind": "tool_call", "tool": "db.delete", "ok": True})]
Enter fullscreen mode Exit fullscreen mode

Change one event and every head after it changes. Forge the whole history and every head changes. At this point the chain only verifies against its own key, which proves nothing about when. So you take the head, all 32 bytes of it, and hand it to witnesses that don't work for you:

pip install opentimestamps-client        # the only install this post asks for
echo "<last head, hex>" | xxd -r -p > head.bin

ots stamp head.bin                       # lands in a bitcoin block header
ots upgrade head.bin.ots && ots verify head.bin.ots

openssl ts -query -data head.bin -sha256 -cert -out head.tsq
curl -s -H "Content-Type: application/timestamp-query" \
     --data-binary @head.tsq https://freetsa.org/tsr > head.tsr
openssl ts -verify -data head.bin -in head.tsr -CAfile tsa.crt
Enter fullscreen mode Exit fullscreen mode

Two different failure modes for a forger. OpenTimestamps buries your 32-byte head in a bitcoin Merkle tree, and redoing a network's proof of work is not a plan. The RFC 3161 receipt is the legal tier: a qualified timestamp authority carries a presumption of correctness under eIDAS, which is the difference between a technical measure and a measure a lawyer accepts. Notice what neither witness ever sees: an event. Only an opaque digest. The GDPR conversation stays short.

What the auditor gets is almost boring. Three files: the journal export, the anchoring receipts, and a verifier of about sixty lines that recomputes the chain from genesis, checks every signature against the public key you published, and checks each receipt against its external root. Exit 0 or the history is rejected, no matter when the tampering happened — any rewrite produces a different head than the one the witnesses recorded. Next to those files goes the logging description Annexe IV §2(f) asks for, written once, against a journal that can't quietly drift.

What the auditor actually receives: export.json, the receipts, exit 0

Two honest limits

Because a log is the easiest place to oversell compliance.

First: tamper-evidence is not truth-at-write. Anchoring proves the journal wasn't altered after the fact. If the agent records a false event and it gets sealed, you have now proven, with bitcoin's help, that a falsehood existed. Garbage in, sealed forever.

Second, this documents. It doesn't prevent. Article 12 lives in the cover-up gap and leaves the exploit gap exactly where it found it. An agent that shouldn't touch production will still touch production. It'll just do it on the record. That's worth something, but it's not the same thing, and I'd rather not blur them.

If you've been through an AI Act readiness review this summer, yours or a client's, I'm curious what they actually asked about your logs. Comments are open.

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

The operational snag with external TSAs on agent runs is query latency. Running an RFC 3161 request on every individual tool call adds 200 to 400ms of network overhead to tight loops. Batching to turn boundaries or session completion keeps the runtime fast, but an unexpected kill signal leaves the tail unanchored until the recovery pass seals the dangling head.

Collapse
 
slabb profile image
Sam LABBE •

Right instinct on the latency math — a design that pays a 300ms round-trip per tool call deserves to die. But the witness never sees a tool call. It sees a 32-byte head. The per-event cost is one SHA-256 and one Ed25519 signature, offline — 0.21ms per sealed event, measured. The 200-400ms belongs to the anchor step, and an anchor commits every event since the previous one, so it amortizes over whatever batch it covers. The cadence is a policy knob, not an architectural tax.

The dangling head is the sharper half of your comment, and it splits into two guarantees with different owners. Integrity never lapses: the chain and signatures on disk verify from genesis, and recovery chains from the last flushed head — nothing gets repaired or rewritten. What's missing is only the independent when. And a kill signal doesn't hand the key to anyone new — the adversary that tier exists for is the key holder, and SIGKILL grants them no fresh powers.

The honest limit stands, though: inside the anchor window, the key holder can rewrite the unanchored tail and no external witness would notice. Two operational habits bound that window to something you can defend in an audit. Anchor on a wall-clock timer, not just at turn boundaries — a killed session's blind spot is then minutes, not "whenever recovery runs." And anchor before consequential actions: the agent requests a timestamp, then runs the destructive tool. The tail before anything dangerous then starts at an anchored head, and the cost lands exactly where the stakes are. Tiering helps too — a co-located TSA for fast cadence, the qualified external one for the legal receipts; every receipt is itself a journaled event, so the tiers chain.

What cadence does your stack actually run — turn boundaries, or something tighter?