DEV Community

Umang Kumar
Umang Kumar

Posted on

Audit Logs for AI Coding Agents: What to Record, Where to Collect It, and What It Proves

When a coding agent does something surprising — deletes a directory, pushes to the wrong branch, reads a file it shouldn't have — the first question is always the same: what exactly happened, and what allowed it? Shell history and chat transcripts answer that badly. An audit log for AI coding agents answers it well, if you decide up front what to record.

This post covers the fields worth capturing, the difference between logging decisions and logging outcomes, where to collect events in common agents, and what a tamper-evident log does and does not prove. (For a from-scratch hash-chain implementation, see my earlier 40-line post; this one is about what to log.)

The questions an agent audit log must answer

Design backwards from the questions you will ask during an incident or review:

  1. Who acted? Which agent, in which session, on behalf of which person?
  2. What did it try to do? The tool, the normalised action, the exact target.
  3. What decided? Allowed, denied, or held — and by which rule or which human.
  4. Why? The reason attached to that decision, and the risk level it was assigned.
  5. What happened next? Did the action actually run, and did it succeed?
  6. What else happened in that run? The surrounding calls, in order.

If your log can't answer question 3 without re-reading a config file that may have changed since, it is a transcript, not an audit log.

The fields that matter

A practical record per tool call:

Field Why
decision_id, run_id Link one call to the run it belongs to
agent, human principal Separate "which bot" from "whose authority"
tool and normalised action read_file and Read are the same action; normalise so you can query across agents
canonical resource Absolute, traversal-collapsed path or normalised URL, so ./x/../.env and .env are one thing
verdict and rule What decided, by name
reason The explanation the agent (and you) saw
risk A stable classification you can filter on
approval id and approver For anything a person released
considered Which rules were evaluated and which matched
timestamp, policy version So the decision can be re-checked later against the same rules

Two things to leave out: raw secret values, and full arguments when they may contain credentials. Store a hash of arguments when you need to bind a record to an exact call without disclosing its contents.

Decision is not execution

This is the most common design mistake. A record that says permit means the call was authorised, not that it ran, and certainly not that it succeeded. The tool may have crashed; the network call may have timed out; the approval may have been granted and never used.

Keep two event types, or two fields: the authorisation decision (before) and the outcome (after). In incident review, "permitted but failed" and "permitted and succeeded" lead to very different conclusions.

Where to collect events

  • Claude Code: hooks run your program around tool use. A PreToolUse hook sees the tool name and input before execution and can return allow or deny; pairing it with a post-execution hook gives you the outcome side.
  • Codex CLI: supports opt-in OpenTelemetry export. Its event catalogue includes codex.tool_decision (approved or denied, and whether the source was configuration or the user) and codex.tool_result (duration, success, output snippet). Codex recommends keeping log_user_prompt = false unless policy explicitly permits storing prompt text.
[otel]
environment = "staging"
exporter = "none"        # or otlp-http / otlp-grpc to your own collector
log_user_prompt = false
Enter fullscreen mode Exit fullscreen mode
  • MCP traffic: a gateway between the client and its servers sees every routed tools/call and can record a decision for each, regardless of which client sent it.

Whatever the source, normalise into one schema. An auditor does not care which agent's naming convention a call used.

Tamper evidence, and its limits

Agents run with your user's permissions, so a plain log file is editable by the thing it is auditing. Hash chaining — each record includes the hash of the previous one — makes in-place edits detectable.

Be precise about what that proves. A chain verified on its own detects internal inconsistency. It does not detect deletion of the tail, or a whole history rewritten and re-hashed, unless you compare against a trusted external checkpoint: the head hash and record count copied somewhere the agent cannot write (a CI artefact, a separate log store, a ticket). An empty or missing log can verify as a valid empty chain, so check the record count you expect, not only the exit code.

Retention and access

  • Ship logs off the machine on a schedule; local-only logs die with the laptop.
  • Apply the same retention and access controls you use for CI logs.
  • Redact at the collector if arguments might contain sensitive data.

A concrete example

Here is what a governed decision looks like in Cirvix AgentControl, from the README's Node quickstart with an audit sink attached (@cirvix_ai/agent-control 0.3.0). The tool is an in-memory fixture; the second call is a denied .env.production read. A trimmed record from .cirvix/audit.jsonl:

{
  "decision_id": "dec_mv3wca7f2",
  "agent": "pr-triage",
  "tool": "read_file",
  "action": "fs.read",
  "resource": "/tmp/auditdemo/.env.production",
  "verdict": "deny",
  "rule": "deny-dotenv-read",
  "reason": "Reading .env files is denied outside an approved secrets flow. …",
  "risk": "critical",
  "risk_signals": ["credential-access", "read-only-tool"],
  "considered": [{ "rule": "deny-dotenv-read", "effect": "forbid", "matched": true }]
}
Enter fullscreen mode Exit fullscreen mode

Querying it:

cirvix logs --last 5
#   ✓ ALLOW   LOW       read_file   /tmp/auditdemo/src/index.mjs       allow-workspace-read
#   ✕ DENY    CRITICAL  read_file   /tmp/auditdemo/.env.production     Policy: deny-dotenv-read

cirvix why dec_mv3wca7f2          # explain one decision from local history
cirvix audit verify --file .cirvix/audit.jsonl --json
# { "ok": true, "records": 2, "head": "sha256:e934…" }
Enter fullscreen mode Exit fullscreen mode

Copy that head and records value somewhere the agent can't write, and you have the external checkpoint described above.

Where Cirvix fits

Cirvix AgentControl is an open-source authorization layer that evaluates governed tool calls (via its MCP gateway, Claude Code hook, or SDK wrappers) and, when an audit sink is configured, writes each decision to a local hash-chained JSONL file. Be clear on scope: it records authorisation, not successful execution; the Node SDK persists nothing unless you supply an audit sink (Python uses an on_decision callback); and calls that bypass the governed path are not in the log. A local chain is not independent compliance evidence on its own.

Checklist

  • [ ] One normalised schema across agents
  • [ ] Decision and outcome recorded separately
  • [ ] Rule, reason, risk and approver on every decision
  • [ ] No raw secrets; hash arguments when needed
  • [ ] Hash chain plus an external checkpoint of head and count
  • [ ] Logs shipped off-machine with retention and access controls

Repo: https://github.com/CIRVIX/agent-control
Site: https://cirvix.com

Top comments (0)