DEV Community

Anusha Mukka
Anusha Mukka

Posted on

Your Agent Forgot What It Did: Build an Audit Trail It Cannot Edit

An agent that writes its own logs can rewrite its own history. Put the log on the other side of the tool boundary, hash the chain, and keep a witness copy the agent cannot reach.

Your agent ran for forty minutes. It touched eleven files, called four APIs, and sent one email you did not expect to send. When you asked what happened, it gave you a tidy summary in perfect English. The summary said it read three files and drafted a reply. It did not mention the email.

I have seen that summary shape before, and it worries me more than the email. A human who forgets is careless. An agent that forgets on demand is a log problem.

Tell me about the last time you tried to debug an agent run from the agent's own messages. The transcript tells you what the model said it did. It does not tell you what the tool actually did when the arguments passed the boundary and hit your filesystem.

That gap is why audit trails keep coming up in reader mail. Not dashboards. A plain record, written somewhere the agent cannot edit, that answers three questions the morning after: what was attempted, what was allowed, and what actually left the process.

Two Bad Defaults and a Third Option That Ships

Most teams I talk to run one of two defaults, and both have the same failure shape.

The first default is no audit beyond chat history. The run lives in the conversation. The conversation scrolls away. If the agent called a tool, you find out when the customer does. This wins for speed. You pay for it the first time someone asks for a receipt.

The second default is a log the agent writes itself. A file called audit.log, a JSON blob the agent is prompted to append. That feels better because there is a file to point at. A process that can append can also overwrite. A process that can write can be prompted to delete the last line and write a nicer one.

I understand why both survive. The first avoids scope creep. The second avoids new infrastructure. On day one they are reasonable. On day thirty, with production access and customers who ask questions with lawyers nearby, they stop being reasonable.

The third option is the one that holds up under the only question that matters: can the thing that acted also change the record of the act? If the answer is yes, you have a story, not a trail.

The fix is not complicated, but the placement matters more than the format. Put the writer on the other side of the tool boundary. The agent proposes an action. A small gate decides yes or no. The gate writes the record before the tool runs, and it writes to a place the agent process does not control. Add a hash chain so a missing line is visible. Keep a second copy somewhere boring, like an append-only file with its own permissions. You now have a trail that survives a forgetful summary.

The model decides what to attempt. The log decides what you can prove. Keep the second decision out of the model's hands.

Here is the gap most explainers skip. They talk about logging what the model said. The useful log captures what the boundary saw: the tool name, the arguments after normalization, the policy decision, the credential that was presented, and the outbound effect. The transcript is not the evidence. The boundary crossing is.

What Belongs in a Record

Let me give you the fields I keep coming back to, because the schema matters more than the storage engine.

One record per tool attempt, not per conversation turn. Attempts matter. A denied attempt tells you what the data tried to talk the agent into doing.

  • run_id: one id for the whole task, so you can pull the sequence later.
  • seq: a monotonically increasing number inside the run. Gaps mean something left.
  • ts: when the gate saw the attempt, from the gate's clock, not the model's timestamp string.
  • principal: who the agent claimed to be, for example support-agent or triage-bot. Not a username, a role you can reason about.
  • tool: the tool name exactly as registered, like mail.send or fs.write.
  • args_digest: a hash of the normalized arguments, plus a redacted preview small enough to debug without leaking secrets.
  • decision: allow, deny, or ask. The policy result, not the model's intention.
  • reason: the rule that fired, for example scope_missing:mail.send or boundary_mismatch:account_8841.
  • effect: what happened after the decision: sent, written, refused, or error. If the tool timed out, the effect says so.
  • prev_hash and hash: a simple chain. Each record hashes the previous record's hash plus its own fields. Delete or edit a line and the chain breaks at the next check.

A few things are worth noting about this list. args_digest does two jobs at once. It lets you prove two calls were identical without storing a customer's email body in plaintext forever. And it makes replay detection trivial: same digest, same run, twice in ten seconds, worth a second look.

principal and decision together are the pair auditors actually ask for. Not what the agent thought about doing. What identity it presented, and what the system decided at the boundary. If you log only successes, you will miss the pattern where the agent probed three denied tools before it found an allowed one that did the same thing sideways.

The Gate in Plain Python

Here is a build that fits in an afternoon. Standard library only. The gate owns the log file. The agent gets a function that looks like a tool call. Behind that function, the gate checks a policy, writes the record, and only then runs the effect.

This is the minimum viable version. The point is to see the placement before you add queues and cloud storage.

import hashlib
import json
import time
from pathlib import Path

LOG_PATH = Path("audit.jsonl")
SECRET_PREVIEW_CHARS = 120

def canonical(obj):
    return json.dumps(obj, sort_keys=True, separators=(",", ":")).encode()

def args_digest(args):
    return hashlib.sha256(canonical(args)).hexdigest()[:16]

def preview(args):
    text = json.dumps(args, sort_keys=True)
    # Redact anything that smells like a secret or a long body
    for key in list(args.keys()):
        low = key.lower()
        if any(s in low for s in ("key", "token", "secret", "password", "body")):
            text = text.replace(str(args[key])[:40], "[redacted]")
    return text[:SECRET_PREVIEW_CHARS]

class AuditTrail:
    def __init__(self, path=LOG_PATH):
        self.path = Path(path)
        self.last_hash = "GENESIS"
        self.seq = 0
        if self.path.exists():
            # Recover chain head so restarts do not fork the story
            for line in self.path.read_text().splitlines():
                if line.strip():
                    rec = json.loads(line)
                    self.last_hash = rec.get("hash", self.last_hash)
                    self.seq = rec.get("seq", self.seq)

    def append(self, run_id, principal, tool, args, decision, reason, effect):
        self.seq += 1
        rec = {
            "run_id": run_id,
            "seq": self.seq,
            "ts": int(time.time()),
            "principal": principal,
            "tool": tool,
            "args_digest": args_digest(args),
            "args_preview": preview(args),
            "decision": decision,
            "reason": reason,
            "effect": effect,
            "prev_hash": self.last_hash,
        }
        rec["hash"] = hashlib.sha256(canonical(rec)).hexdigest()
        self.last_hash = rec["hash"]
        # The gate writes. The agent process never opens this file directly.
        with self.path.open("a") as f:
            f.write(json.dumps(rec, sort_keys=True) + "\n")
        return rec

def verify_chain(path=LOG_PATH):
    prev = "GENESIS"
    ok = True
    for i, line in enumerate(Path(path).read_text().splitlines(), start=1):
        if not line.strip():
            continue
        rec = json.loads(line)
        if rec.get("prev_hash") != prev:
            print(f"line {i}: prev_hash mismatch, expected {prev[:12]} got {rec.get('prev_hash','')[:12]}")
            ok = False
            break
        # Recompute hash without the hash field itself
        check = dict(rec)
        expected = check.pop("hash", None)
        actual = hashlib.sha256(canonical(check)).hexdigest()
        if actual != expected:
            print(f"line {i}: hash mismatch, record was edited")
            ok = False
            break
        prev = expected
    if ok:
        print("chain ok")
    return ok
Enter fullscreen mode Exit fullscreen mode

A few things are worth noting about this example. First, the log writer is the gate, not the agent. The agent calls attempt(). The gate decides and writes in one place. Second, the hash chain is cheap. One SHA-256 per record is nothing next to a model call. Third, preview() is deliberately lossy. You want enough to debug, not a second copy of the secret.

Now wire it to a real decision. Keep the policy embarrassingly simple at first: a map from principal to allowed tools, plus a boundary check.

POLICY = {
    "support-agent": {"mail.read", "mail.draft", "calendar.read"},
    "triage-bot": {"mail.read", "ticket.create"},
}

def attempt(trail, run_id, principal, tool, args, effect_fn):
    allowed = POLICY.get(principal, set())
    if tool not in allowed:
        rec = trail.append(run_id, principal, tool, args, "deny", f"scope_missing:{tool}", "refused")
        return {"ok": False, "reason": rec["reason"]}
    try:
        result = effect_fn(args)
        trail.append(run_id, principal, tool, args, "allow", "policy_ok", "done")
        return {"ok": True, "result": result}
    except Exception as exc:
        trail.append(run_id, principal, tool, args, "allow", "policy_ok", f"error:{type(exc).__name__}")
        raise
Enter fullscreen mode Exit fullscreen mode

Walk through a run:

python3 - << 'PY'
from audit_gate import AuditTrail, attempt, verify_chain

trail = AuditTrail("audit.jsonl")
run = "run_2026_10_10_001"

print(attempt(trail, run, "support-agent", "mail.read", {"account": "8841", "id": "msg_19"}, lambda a: "subject: invoice question"))
print(attempt(trail, run, "support-agent", "mail.send", {"account": "8841", "to": "customer@example.com", "body": "thanks"}, lambda a: "sent"))
print(attempt(trail, run, "triage-bot", "mail.send", {"account": "8841", "to": "x@example.com"}, lambda a: "sent"))

verify_chain("audit.jsonl")
PY
Enter fullscreen mode Exit fullscreen mode

You will see the first call allowed, the second denied because support-agent does not have mail.send in this toy policy, and the third denied for a different principal. The file now has three lines. Delete the middle line and run verify_chain again. It fails at the next record, because prev_hash no longer matches. That failure is the point. A log you can quietly edit is a log you cannot trust when it matters.

Here is the shape on disk:

run_id  seq  tool         decision  reason                 effect
------  ---  -----------  --------  ---------------------  --------
run_001   1  mail.read    allow     policy_ok              done
run_001   2  mail.send    deny      scope_missing:mail     refused
run_001   3  mail.send    deny      scope_missing:mail     refused
Enter fullscreen mode Exit fullscreen mode

The diagram is the same idea, drawn once:

                +------------------+
  user task --> |      Agent       |  can be persuaded, forgets politely
                +--------+---------+
                         | attempt(tool, args, key)
                         v
                +------------------+
                |       Gate       |  checks principal, scope, boundary
                +--------+---------+
                  allow  |  deny
                +--------+---------+
                |                  |
                v                  v
           tool runs          refused
                |                  |
                +--------+---------+
                         |
                         v
                +------------------+
                | Append-only log  |  gate writes, agent cannot open
                | hash chained     |
                +------------------+
                         |
                         v
                +------------------+
                | Witness copy     |  second path, different permissions
                +------------------+
Enter fullscreen mode Exit fullscreen mode

A few things are worth noting about the diagram. The agent never touches the log file. The witness copy is not a backup for disasters. It is there so the morning someone asks what happened, you have two places to compare if one copy looks too clean.

Where This Breaks

I will keep this section unhedged, because audit trails invite false confidence.

  • If the gate and the agent run in the same process with the same OS user, the agent can still find the file path and overwrite it. The placement only helps if the permissions differ. Run the gate as a separate process, a separate container, or at least a separate user, and make the log directory append-only for that user.
  • Hash chains detect edits after the fact. They do not prevent a denied tool from running through a second path that skips the gate. If your agent can import the mail library directly, you do not have a boundary. You have a suggestion.
  • Previews leak if you are careless. Logging the full tool arguments feels helpful during the first incident. It feels less helpful when the arguments include a customer email body, an API key pasted into a prompt, or a file the agent should never have read. Digest plus a redacted preview is the compromise that ages well.
  • Clocks lie. If you use the agent's timestamp, you will argue about order later. Use the gate's clock, and add a monotonic seq so order survives clock skew between machines.
  • Volume is real. A chatty agent that calls tools in a loop can write a million tiny records before lunch. Sample the heartbeats, keep the boundary crossings. Tool attempts are signal. Internal reasoning traces are a separate store with a shorter retention.
  • Retention is a policy, not a side effect. Decide how long you keep the raw log, how long you keep digests after the raw is gone, and who can query it. An audit trail nobody can read is just expensive storage.

None of these are reasons to skip the trail. They are reasons to place it carefully.

Build It If and Skip It If

Build it if your agent can take an action a human would have to explain later. Sending, writing, deleting, paying, and changing permissions all count. If a customer, a manager, or a regulator can ask what happened last Tuesday, you want the answer to come from a file the agent could not rewrite.

Build the minimal version if you can answer yes to two questions. Can you name the principal for each agent run, distinct from the human who started it? Can you list the tools that actually change state outside the process? If yes, the gate above fits in an afternoon.

Skip it, for now, if your agent only reads and drafts in a sandbox with no outbound effect. A read-only summarizer with no send, write, or delete does not need a hash chain. It needs a good transcript. Add the trail the sprint you give it its first write tool.

Minimal viable version, afternoon scope:

  1. One AuditTrail class as above, one JSONL file, one policy map.
  2. Gate every state-changing tool. Reads can log too, but start with writes.
  3. A verify_chain.py you run in CI or before you share the log.
  4. A second copy shipped off the box every few minutes: scp, object storage, or a log collector with different credentials. Even a cron that copies the file to a read-only bucket counts.

That version will not satisfy a formal audit on its own. It will satisfy the question that actually wakes people up: what did the agent do, and can we show the record was not edited after the fact?

Close

Pick one agent you already run. List its state-changing tools on a sticky note. Put the gate in front of just those tools this week, with the hash chain and the witness copy. Run it for one real task. Then try to edit the log quietly and run the verifier. Watch it catch you. That small failure is the confidence you want before the real incident asks for the log.

What is the one agent action you would least want to explain from memory alone, the one where you most wish you had a record the agent could not touch?

Resources

  1. W3C PROV Concepts The vocabulary for provenance that outlives any one vendor: entities, activities, and agents, and who was responsible for what.
  2. NIST SP 800-92 Guide to Computer Security Log Management The boring, durable guidance on what to log, how long to keep it, and how to protect the log itself.
  3. OWASP Logging Cheat Sheet Practical rules for what to include, what to redact, and how to keep logs useful without turning them into a second breach.

Top comments (0)