<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ventse</title>
    <description>The latest articles on DEV Community by Ventse (@traceseal).</description>
    <link>https://dev.to/traceseal</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3882611%2F9fea844c-d8a3-423f-9fd5-c12bd3e6960d.png</url>
      <title>DEV Community: Ventse</title>
      <link>https://dev.to/traceseal</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/traceseal"/>
    <language>en</language>
    <item>
      <title>OWASP Top 10 for Agentic Applications: what it asks you to record, and who should sign it</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:47:50 +0000</pubDate>
      <link>https://dev.to/traceseal/owasp-top-10-for-agentic-applications-what-it-asks-you-to-record-and-who-should-sign-it-k0g</link>
      <guid>https://dev.to/traceseal/owasp-top-10-for-agentic-applications-what-it-asks-you-to-record-and-who-should-sign-it-k0g</guid>
      <description>&lt;p&gt;In December 2025 the OWASP GenAI Security Project's Agentic Security Initiative published the &lt;a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/" rel="noopener noreferrer"&gt;Top 10 for Agentic Applications&lt;/a&gt;. Most write-ups walk through the ten risks, and this one does too, briefly. It then reads the prevention sections for a narrower question: what the list asks you to record about your agents, and what properties it expects those records to have. The answer is more demanding than "keep logs", and the most useful sentence on the subject comes in the last entry.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ten risks in one pass
&lt;/h2&gt;

&lt;p&gt;The entries are numbered ASI01 to ASI10. Condensed from the document's own descriptions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ASI01 Agent Goal Hijack.&lt;/strong&gt; An attacker redirects the agent's objectives or decision pathways, because agents "cannot reliably distinguish instructions from related content." The mechanics are in &lt;a href="https://traceseal.io/blog/indirect-prompt-injection/" rel="noopener noreferrer"&gt;our piece on indirect prompt injection&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI02 Tool Misuse and Exploitation.&lt;/strong&gt; Legitimate tools used harmfully through injection, misalignment, unsafe delegation or ambiguous instructions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI03 Identity and Privilege Abuse.&lt;/strong&gt; Escalation through delegation chains, role inheritance and cached credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI04 Agentic Supply Chain Vulnerabilities.&lt;/strong&gt; Third-party models, tools, plug-ins, MCP and A2A components or registries that are malicious or tampered with.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI05 Unexpected Code Execution.&lt;/strong&gt; Code the agent generates, turned into remote code execution or local misuse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI06 Memory &amp;amp; Context Poisoning.&lt;/strong&gt; Persistent corruption of stored memory, summaries, embeddings and retrieval stores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI07 Insecure Inter-Agent Communication.&lt;/strong&gt; Weak authentication, integrity or authorisation between agents that coordinate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI08 Cascading Failures.&lt;/strong&gt; A single fault propagating across agents into system-wide harm.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI09 Human-Agent Trust Exploitation.&lt;/strong&gt; An agent exploiting a person's trust so that they approve something they should not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI10 Rogue Agents.&lt;/strong&gt; Agents that deviate from their intended function or authorised scope, where each action "may individually appear legitimate".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The introductory letter frames all ten with a principle it calls Least-Agency: avoid autonomy you do not need, because "deploying agentic behavior where it is not needed expands the attack surface without adding value." In the same passage it says that "strong observability becomes non-negotiable". This piece picks up from that word, observability.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the logging requirement escalates
&lt;/h2&gt;

&lt;p&gt;Read the prevention sections in order and the logging advice does not stay at one level. It tightens.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ASI01:&lt;/strong&gt; on an unexpected goal shift, "surface the deviation for review, and record it for audit." A record, of unspecified quality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI02:&lt;/strong&gt; "Maintain immutable logs of all tool invocations and parameter changes." Now the record must not change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI08:&lt;/strong&gt; "Record all inter-agent messages, policy decisions, and execution outcomes in tamper-evident, time-stamped logs bound to cryptographic agent identities." Now any change must be detectable, and each entry is tied to a key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI09:&lt;/strong&gt; "Keep tamper-proof records of user queries and agent actions for audit and forensics."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI10:&lt;/strong&gt; "Maintain comprehensive, immutable and signed audit logs of all agent actions, tool calls, and inter-agent communication".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ASI08 gives the reason. It files the requirement under "logging and non-repudiation" and points back to threat T8, Repudiation and Untraceability, in OWASP's Agentic AI Threats and Mitigations guide, which it summarises as "the ability to trace, attribute, and audit cascading behaviors through resilient logging and non-repudiation mechanisms that prevent silent propagation."&lt;/p&gt;

&lt;p&gt;Non-repudiation is the word doing the work. A log supports it when the party that produced an entry cannot credibly deny producing it later, and nobody else can pass an entry off as theirs. An ordinary application log has neither property. Anyone with write access to the store, including a compromised process, can change it without leaving a trace. We set out the fields an agent log needs in &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;the audit trail piece&lt;/a&gt;. OWASP's list is concerned with whether anyone other than the author can believe those fields.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "immutable" is not enough
&lt;/h2&gt;

&lt;p&gt;The list says less about who writes the record, and ASI09 shows why that matters. Among its examples is "Explainability Fabrications: The agent fabricates plausible audit rationales to justify a risky configuration change." The same entry describes an agent that "acts as an untraceable 'bad influence,' manipulating the human into performing the final, audited action, making the agent's role in the compromise invisible to forensics."&lt;/p&gt;

&lt;p&gt;Take those two together. In the first, the agent writes a false entry into the audit trail. In the second, the audit trail is accurate: it records the human's action and misses the agent's part in it. Neither is fixed by immutability. Making a fabricated rationale immutable only preserves the fabrication. What helps is a record the agent did not write. It should describe what actually ran and what it produced, and be written at a boundary the agent cannot reach. That is the same objection to agent-written transcripts we raised in &lt;a href="https://traceseal.io/blog/soc-2-ai-agents-evidence/" rel="noopener noreferrer"&gt;the SOC 2 piece&lt;/a&gt;, and here OWASP makes it for us.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The distinction the list implies but does not state:&lt;/strong&gt; immutability protects a record after it is written. It says nothing about whether the writer was honest at the time. A log counts as evidence only if both hold.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The key belongs outside the agent
&lt;/h2&gt;

&lt;p&gt;ASI10 contains the sentence that says how to get there. After recommending signed behavioural manifests and per-run ephemeral credentials, it adds:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All signing and attestation mechanisms assume hardened cryptographic key management (e.g., HSM/KMS-backed keys, least-privilege access, rotation and revocation). Keys must never be directly available to agents; instead, orchestrators should mediate signing operations so that a compromised agent cannot simply exfiltrate or misuse long-lived keys.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This constraint is what gives "signed logs" any value. If the agent process can reach the signing key, a hijacked agent (ASI01) or a rogue one (ASI10) can sign whatever it likes, and the signature proves only that the key was used. The signer has to sit on the far side of a boundary the agent cannot cross. In practice that means the orchestrator or host, with the key in a separate process, a keystore or an HSM.&lt;/p&gt;

&lt;p&gt;ASI04 applies the same logic to what the agent runs. Its advice is to "Sign and attest manifests, prompts, and tool definitions", to "verify provenance before install or activation" and to "auto-reject unsigned or unverified." ASI01 goes further and suggests evaluating an "intent capsule", which it calls "an emerging pattern to bind the declared goal, constraints, and context to each execution cycle in a signed envelope". The direction is the same across the list: bind what was meant to run, what did run and what came out, then sign it with a key the agent cannot use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a signed execution receipt fits
&lt;/h2&gt;

&lt;p&gt;That is the problem a Traceseal execution receipt is built for, so here is the mapping, limits included. The &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;receipt specification&lt;/a&gt; defines a signed JSON document in three parts. The execution section holds hashes of the signed skill manifest, the sandbox configuration, the inputs and the outputs, plus the exit code, the timing, and a hash of the matching entry in the operator's hash-chained audit log. The provenance section holds the publisher's key, per-file content hashes and a transparency log reference. The attestation is an ed25519 signature by the operator over both.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ASI02 and ASI05:&lt;/strong&gt; the input, output and sandbox profile hashes record what a call received, what it returned and what configuration it ran under.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI04:&lt;/strong&gt; the manifest hash and publisher fingerprint tie the execution to one signed artefact. The transparency log reference lets a third party check that the manifest was published rather than swapped in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ASI08 and ASI10:&lt;/strong&gt; the operator signature gives each entry the non-repudiation property the list asks for. In our runner the operator key never enters the sandbox, which is the arrangement ASI10 describes. A third party can check the signature with &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;one command&lt;/a&gt; and nothing but the receipt file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three limits, stated plainly. First, the specification says outright that a receipt "does NOT prove the operator is honest"; it is "an &lt;em&gt;attestation&lt;/em&gt;, not a &lt;em&gt;proof of execution&lt;/em&gt;". A receipt moves trust from the agent to the operator. That is the point, but the trust has moved rather than disappeared. Publishing to a &lt;a href="https://traceseal.io/blog/transparency-log-witnesses/" rel="noopener noreferrer"&gt;witnessed transparency log&lt;/a&gt; narrows what the operator can rewrite afterwards. Second, receipts detect nothing. A hijacked goal (ASI01) or a manipulated human (ASI09) still produces well-formed receipts. What they give you is a reliable account of what happened when you reconstruct it afterwards. Third, a receipt covers only what passes through the receipting boundary. An agent with a second route to a tool that bypasses it leaves no receipt for that route, as &lt;a href="https://traceseal.io/blog/sandbox-an-ai-coding-agent/" rel="noopener noreferrer"&gt;our sandbox write-up&lt;/a&gt; discusses.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the list leaves open
&lt;/h2&gt;

&lt;p&gt;The document is guidance, not a standard. It does not define a log format, say which component should hold the keys beyond "orchestrators", or explain how a third party should check a log that someone else kept. It uses "immutable", "tamper-evident" and "tamper-proof" in different entries without separating them, although they are different properties. The first prevents change. The second makes change detectable. The third is something no software log fully achieves. A team that cites the list in a security review will have to decide which one it means.&lt;/p&gt;

&lt;p&gt;The useful reading of the list is not that agentic systems need more logging, since they already produce plenty. It is that the record has to be one the agent could not have written differently, signed with a key the agent could not reach. Every other prevention section assumes that record exists when something goes wrong.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://traceseal.io/blog/owasp-top-10-agentic-applications/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;. Traceseal issues signed execution receipts for AI agents: an open spec and an open verifier, one command to check.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>agents</category>
      <category>owasp</category>
    </item>
    <item>
      <title>MCP security: what to log when an AI agent calls tools, and what the log can prove</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 22 Sep 2026 08:48:11 +0000</pubDate>
      <link>https://dev.to/traceseal/mcp-security-what-to-log-when-an-ai-agent-calls-tools-and-what-the-log-can-prove-4l5e</link>
      <guid>https://dev.to/traceseal/mcp-security-what-to-log-when-an-ai-agent-calls-tools-and-what-the-log-can-prove-4l5e</guid>
      <description>&lt;p&gt;The Model Context Protocol has become the default way an AI agent reaches a tool. That is convenient for the people wiring agents up and awkward for the people who later have to answer for what those agents did, because the protocol standardises the &lt;em&gt;call&lt;/em&gt; and says almost nothing about the &lt;em&gt;record&lt;/em&gt;. This piece goes through what the current spec actually asks you to log, where a tool call gets written down today, and what each of those records can and cannot establish when someone who does not trust you asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consent is not a record
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://modelcontextprotocol.io/specification/2025-11-25/index" rel="noopener noreferrer"&gt;2025-11-25 revision of the specification&lt;/a&gt; is blunt about what a tool is. Under its key principles: "Tools represent arbitrary code execution and must be treated with appropriate caution," and "Hosts must obtain explicit user consent before invoking any tool." The &lt;a href="https://modelcontextprotocol.io/specification/2025-11-25/server/tools" rel="noopener noreferrer"&gt;tools section&lt;/a&gt; adds that clients should prompt for confirmation on sensitive operations and show tool inputs to the user before sending them to the server.&lt;/p&gt;

&lt;p&gt;All of that is about permission. None of it is about evidence. A consent dialog establishes that a person was shown a tool name and an argument block and clicked approve. It does not establish what the server did with those arguments, whether the result it returned reflected what changed, or whether the dialog the person saw was the one the model intended. Most hosts keep the consent decision in the same transcript store as everything else the model said, which puts it in the same position as any other line of that transcript: written by the runtime under review, editable by whoever administers the store. We went through why that fails an auditor's test in &lt;a href="https://traceseal.io/blog/soc-2-ai-agents-evidence/" rel="noopener noreferrer"&gt;the SOC 2 piece&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the spec does mention logging
&lt;/h2&gt;

&lt;p&gt;The protocol documents proper contain no logging requirement for tool calls. What exists lives in the &lt;a href="https://modelcontextprotocol.io/docs/2025-11-25/tutorials/security/security_best_practices" rel="noopener noreferrer"&gt;security best practices&lt;/a&gt; document, which sits alongside the authorisation spec rather than inside it, and it appears in three places.&lt;/p&gt;

&lt;p&gt;The first is the section on token passthrough, the anti-pattern where an MCP server forwards a client's upstream token to a downstream API without checking it was issued to the server. The document lists the consequences under a heading it calls "Accountability and Audit Trail Issues": the server cannot tell which client is calling, and "the downstream Resource Server's logs may show requests that appear to come from a different source with a different identity, rather than the MCP server that is actually forwarding the tokens." The spec's own words for the result are that "incident investigation, controls, and auditing" become more difficult. It is one of the clearer statements anywhere in the MCP corpus that a log written by the wrong party is worse than useful.&lt;/p&gt;

&lt;p&gt;The second is scope minimisation, where broad omnibus scopes are said to obscure audit trails because "a single omnibus scope masks user intent per operation." The mitigation includes a concrete instruction to servers: log elevation events, meaning the scope requested and the subset granted, with correlation IDs.&lt;/p&gt;

&lt;p&gt;The third is in the proxy guidance, where proxy services that spawn local servers over &lt;code&gt;stdio&lt;/code&gt; are told to "log all &lt;code&gt;stdio&lt;/code&gt; transport usage for security monitoring." That is a should, not a must, and it applies only to proxy architectures.&lt;/p&gt;

&lt;p&gt;So the sum of the spec's position is: log identity so the downstream system does not attribute your agent's actions to someone else, log scope changes so intent per operation is recoverable, and log process spawns if you proxy. Sensible, and thin. Nothing here says what a record of a tool call should contain, who should write it, or how anyone would check it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  An annotation is a claim, not a fact
&lt;/h2&gt;

&lt;p&gt;There is a detail in the tools spec that matters more for logging than it first appears. Tools can carry annotations, of which four are behavioural: &lt;code&gt;readOnlyHint&lt;/code&gt;, &lt;code&gt;destructiveHint&lt;/code&gt;, &lt;code&gt;idempotentHint&lt;/code&gt; and &lt;code&gt;openWorldHint&lt;/code&gt;. A host that wants to decide which calls need a human in the loop, or which calls to record in detail, will naturally reach for these. The &lt;a href="https://modelcontextprotocol.io/specification/2025-11-25/schema#toolannotations" rel="noopener noreferrer"&gt;schema&lt;/a&gt; says, in a note that is easy to skim past:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All properties in ToolAnnotations are hints. They are not guaranteed to provide a faithful description of tool behavior (including descriptive properties like title). Clients should never make tool use decisions based on ToolAnnotations received from untrusted servers.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The tools section restates it as a requirement: clients must consider annotations untrusted unless they come from trusted servers. Which means a log line reading &lt;em&gt;read-only tool invoked, no confirmation required&lt;/em&gt; is not recording what the tool did. It is recording what the tool's author said the tool does, and the author is a party whose code the agent is about to execute. If the server was swapped, updated or &lt;a href="https://traceseal.io/blog/indirect-prompt-injection/" rel="noopener noreferrer"&gt;compromised through the content it serves&lt;/a&gt;, the annotation is exactly the field an attacker would set to &lt;code&gt;true&lt;/code&gt;. A record that inherits its risk classification from the server being classified has the same structural flaw as a transcript written by the agent being audited.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four places a tool call gets written down
&lt;/h2&gt;

&lt;p&gt;Follow one call through a typical deployment and it leaves traces in four places, each written by a different party with a different interest.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The host's transcript.&lt;/strong&gt; The model's request, the arguments it chose, the consent decision, and whatever the server returned. Written by the agent runtime. Complete in the sense of containing everything the model saw, and worthless as independent evidence for the same reason.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The MCP server's own log.&lt;/strong&gt; Whatever the server operator chose to record. If you run the server, it is under your control and therefore not independent. If a vendor runs it, it is under theirs, and you will get it in an incident only if they choose to give it to you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The downstream system.&lt;/strong&gt; The repository, the database, the ticketing API. This is the closest thing to ground truth about effect, and the token passthrough section above explains the catch: unless the server authenticated correctly, this log may attribute the action to the wrong identity, or to a shared service account that a dozen agents use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The user's approval.&lt;/strong&gt; A click, or a standing allowlist. Usually stored in the host, so it collapses into the first bucket.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of the four is signed by anyone. None is bound to the code that ran: the host records a tool &lt;em&gt;name&lt;/em&gt;, and tool names are only required to be unique within a server, so the same name can point at different code tomorrow. And none of them can be handed to a third party without also handing over the trust you place in whoever wrote it. That is the gap between having logs and having evidence, and it is the gap that &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;the receipts primer&lt;/a&gt; is about.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The test worth applying:&lt;/strong&gt; for each record of a tool call, ask who wrote it and whether they had the ability, and any reason, to write something else. If the answer to the first is "the agent" or "the tool", the record is testimony, not evidence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What a record would need
&lt;/h2&gt;

&lt;p&gt;If the aim is a record that survives that test, the shape is not exotic. Per call, capture the server's identity as a hash of its code or manifest rather than its display name, the tool name and a hash of the exact arguments, a hash of the result, the identity that authorised the call and how (interactive consent, allowlist, or scope elevation with its correlation ID), and the constraints the call ran under, meaning the sandbox profile if there was one. Then sign the whole thing with a key the agent process cannot reach, and put the signature somewhere append-only that a third party can check.&lt;/p&gt;

&lt;p&gt;That is deliberately the shape of a Traceseal execution receipt. The &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;specification&lt;/a&gt; binds a skill manifest hash, a sandbox profile hash, input and output hashes, exit code and timing under an ed25519 signature from an operator key that never enters the sandbox, and the verifier is &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;a single command&lt;/a&gt; that needs nothing but the receipt. The &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;walkthrough&lt;/a&gt; shows what an &lt;code&gt;[OK]&lt;/code&gt; does and does not establish.&lt;/p&gt;

&lt;p&gt;Two limits, stated plainly. A receipt covers what passes through the receipting boundary. A local MCP server launched over &lt;code&gt;stdio&lt;/code&gt; runs with the client's privileges, and the security document is explicit that without sandboxing it can read anything the client can; if the agent can reach a tool through a path that produces no receipt, the receipt for the path it did use proves nothing about the other one. And a receipt records that a hashed input produced a hashed output under a hashed policy. It does not say the output was correct or the policy adequate. &lt;a href="https://traceseal.io/blog/sandbox-an-ai-coding-agent/" rel="noopener noreferrer"&gt;Our sandbox write-up&lt;/a&gt; covers where a containment boundary stops being demonstrable, and the same line applies here.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still open
&lt;/h2&gt;

&lt;p&gt;The spec does not define a log format for tool calls, does not require servers to identify their code by hash, and leaves annotation trust to each client's judgement of which servers are "trusted". Every host currently records tool use in its own shape, which means two organisations that both "log all MCP calls" have records that cannot be compared, let alone verified by the same tool. Whether that changes depends on the protocol's maintainers and the vendors shipping hosts, not on anything a single deployment can do.&lt;/p&gt;

&lt;p&gt;What a single deployment can do is stop treating the transcript as the record. The spec has already conceded that tools are arbitrary code execution and that a tool's own description of itself cannot be trusted. The consistent next step is to conclude that the tool's, and the agent's, description of what happened cannot be trusted either, and to keep a record they did not write.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://traceseal.io/blog/mcp-security-tool-call-logging/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;. Traceseal issues signed execution receipts for AI agents: an open spec and an open verifier, one command to check.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>mcp</category>
      <category>agents</category>
    </item>
    <item>
      <title>SOC 2 for AI agents: what evidence an auditor can actually accept</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:47:10 +0000</pubDate>
      <link>https://dev.to/traceseal/soc-2-for-ai-agents-what-evidence-an-auditor-can-actually-accept-16pl</link>
      <guid>https://dev.to/traceseal/soc-2-for-ai-agents-what-evidence-an-auditor-can-actually-accept-16pl</guid>
      <description>&lt;p&gt;At some point in the next twelve months, an auditor is going to ask what your AI agents did. Not what they were allowed to do — what they did. If you sell software to companies that care about SOC 2, the question arrives through your own audit; if you buy it, the question arrives through a vendor questionnaire. Either way, the answer "we have the chat logs" is about to stop being good enough, and it is worth understanding why before the fieldwork starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where agents land in the criteria
&lt;/h2&gt;

&lt;p&gt;SOC 2 does not have an AI section. A report is issued against the AICPA's &lt;a href="https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022" rel="noopener noreferrer"&gt;2017 Trust Services Criteria, with the points of focus revised in 2022&lt;/a&gt;, and nothing in that document mentions agents, models or prompts. That is not a loophole. The criteria are written around &lt;em&gt;systems&lt;/em&gt; and &lt;em&gt;system components&lt;/em&gt;, and an agent that reads a ticket, edits a repository and opens a pull request is a system component in every sense the criteria care about. It holds credentials, it changes production, and it does so without a human approving each step.&lt;/p&gt;

&lt;p&gt;Three criteria do most of the work. &lt;strong&gt;CC6&lt;/strong&gt; covers logical access: what the agent can reach, with which credentials, and who granted them. &lt;strong&gt;CC8&lt;/strong&gt; covers change management: whether changes the agent makes to infrastructure, data or code go through the same authorisation, testing and approval controls as changes a person makes. And &lt;strong&gt;CC7&lt;/strong&gt; covers system operations. The wording of &lt;a href="https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022" rel="noopener noreferrer"&gt;CC7.2&lt;/a&gt; is the one to read twice:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The entity monitors system components and the operation of those components for anomalies that are indicative of malicious acts, natural disasters, and errors affecting the entity's ability to meet its objectives; anomalies are analyzed to determine whether they represent security events.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To monitor the operation of a component for anomalies, you need a record of its operation that you can trust more than the component itself. For a database or a load balancer that is routine. For an agent, whose "operation" is a sequence of tool calls chosen at runtime by a model, it is the whole problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  ISO/IEC 42001 says the same thing more directly
&lt;/h2&gt;

&lt;p&gt;If you are also looking at &lt;a href="https://www.iso.org/standard/81230.html" rel="noopener noreferrer"&gt;ISO/IEC 42001:2023&lt;/a&gt;, the AI management system standard, the requirement is explicit rather than implied. Annex A control A.6.2.8, &lt;em&gt;AI system recording of event logs&lt;/em&gt;, requires the organisation to determine at which phases of the AI system's life cycle event logs are recorded, and sets a floor: at minimum, while the system is in use. The standard's text is paywalled, so we will not quote further, but the shape is clear. Logging the agent while it is operating is not an optional hardening step under 42001; it is a control you either have or do not.&lt;/p&gt;

&lt;p&gt;Both frameworks converge on one demand. You need a record of what the agent did, produced in a way that the auditor can rely on. Which raises the question of what "rely on" means.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a transcript is not evidence
&lt;/h2&gt;

&lt;p&gt;Most teams' first answer is the conversation log: the prompts, the model's responses, the tool calls and their results, stored by whatever agent framework they run. It feels complete, and for debugging it often is. For an auditor it has three problems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is written by the thing being audited.&lt;/strong&gt; The agent runtime produces the transcript, and the agent runtime is the component whose behaviour is in question. A model that has been &lt;a href="https://traceseal.io/blog/indirect-prompt-injection/" rel="noopener noreferrer"&gt;hijacked by the content it read&lt;/a&gt; will still produce a tidy transcript; it will just be a transcript of the wrong actions. An auditor testing CC7.2 wants a monitoring record that would catch the anomaly, not one that the anomaly gets to write.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It describes intent, not effect.&lt;/strong&gt; A tool-call record says the agent asked to run &lt;code&gt;git push&lt;/code&gt;. It does not say the push succeeded, what commit went where, or whether the working tree it pushed matched the diff shown to the reviewer. The distinction between what an agent reported and what happened on disk is the entire premise of &lt;a href="https://traceseal.io/blog/alibi-grade-agents-on-ground-truth/" rel="noopener noreferrer"&gt;grading agents on ground truth&lt;/a&gt;, and it is also the distinction an auditor is trained to draw.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It can be edited after the fact by anyone with write access to the store.&lt;/strong&gt; A Type 2 report covers a period, commonly six to twelve months, and the auditor samples from it. If the transcript store is a database your platform team administers, the auditor has to test the controls over &lt;em&gt;that&lt;/em&gt; before the transcripts are worth anything. We covered why a hash chain alone does not close this gap in &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;the receipts primer&lt;/a&gt;: a chain proves order and continuity, but whoever holds the chain can regenerate it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The test an auditor actually applies:&lt;/strong&gt; can I obtain this evidence, or something that corroborates it, from a source other than the system under review? A transcript fails that test by construction. It is the system under review describing itself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What passes
&lt;/h2&gt;

&lt;p&gt;Audit evidence has to be &lt;strong&gt;reperformable&lt;/strong&gt;. The auditor picks a sample — twenty agent runs from March, say — and needs to establish, for each one, what ran, under what constraints, and with what result, in a way that does not depend on taking your word for it. That maps onto four properties a record of agent activity needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bound to the code that ran.&lt;/strong&gt; A hash of the exact skill or tool code, so the auditor can confirm the sampled run used the approved version and not something patched in between. This is what makes the record useful under CC8 as well as CC7.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound to the constraints it ran under.&lt;/strong&gt; A hash of the sandbox profile or execution policy, so "the agent could not reach the network" is a checkable claim rather than an assertion. What a sandbox can and cannot prove on its own is in &lt;a href="https://traceseal.io/blog/sandbox-an-ai-coding-agent/" rel="noopener noreferrer"&gt;our sandbox write-up&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound to inputs and outputs.&lt;/strong&gt; Hashes rather than contents, so the auditor can confirm a specific input produced a specific output without the record itself leaking customer data. This matters for the confidentiality criteria: an evidence trail that stores raw prompts and outputs creates a second copy of every sensitive thing the agent touched.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Signed by a key the agent cannot use.&lt;/strong&gt; The attestation has to come from outside the agent's reach, otherwise you are back to the transcript problem. And the auditor has to be able to check the signature with a public tool, on their own machine, without asking you to run anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is, deliberately, the structure of a Traceseal execution receipt. The &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;specification&lt;/a&gt; is short enough to hand to an auditor: an &lt;code&gt;execution&lt;/code&gt; block with the skill manifest hash, sandbox profile hash, input and output hashes, exit code and timing; a &lt;code&gt;provenance&lt;/code&gt; block that names who signed the code and where it sits in the transparency log; and an &lt;code&gt;attestation&lt;/code&gt; block with an ed25519 signature over both from an operator key that never enters the sandbox. The verifier is &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;on PyPI&lt;/a&gt;, needs only the receipt file, and prints what the signature does and does not establish. The walkthrough is in &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;How to verify what an AI agent actually did&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Two honest limits. A receipt proves the operator attested that &lt;em&gt;this code&lt;/em&gt; ran under &lt;em&gt;this policy&lt;/em&gt; and produced &lt;em&gt;this output hash&lt;/em&gt;; it does not prove the output was correct or the policy was sensible. Those remain judgement calls, which is what the auditor is for. And a receipt only covers what was executed through the receipting boundary. If the agent can also act through a path that does not produce receipts, the auditor will find it, and should.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do before fieldwork
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Put agents on the system description.&lt;/strong&gt; If an agent holds credentials or changes production, it belongs in the scoped system, named as a component, with its access documented under CC6. Auditors find undisclosed components far less forgivable than disclosed ones with gaps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route every consequential action through one execution boundary&lt;/strong&gt; that produces a signed record, and make it the only path. One receipted path and one unreceipted path is worse than either alone, because it invites the question of which one was used.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the records for the whole audit period, and then some.&lt;/strong&gt; A Type 2 sample can reach back to the first day of the period. The EU AI Act's six-month log retention floor for high-risk systems, covered in &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;our audit-trail piece&lt;/a&gt;, is a reasonable minimum even where the Act does not apply.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test your own evidence before the auditor does.&lt;/strong&gt; Pick twenty runs at random, verify each receipt with the public tool, and confirm the manifest hash matches an approved release. If you cannot do that in an afternoon, the auditor cannot do it in a week.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;SOC 2 has never asked whether your systems are clever. It asks whether you can show, to someone who has no reason to believe you, that they did what you say they did. Agents make that harder, because for the first time the component in question decides its own actions. The answer is not a better transcript. It is a record the agent did not write.&lt;/p&gt;




&lt;p&gt;The receipt format, the verifier and the transparency log are open: the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;spec&lt;/a&gt;, the &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt; and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public log&lt;/a&gt;. Originally published at &lt;a href="https://traceseal.io/blog/soc-2-ai-agents-evidence/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>compliance</category>
      <category>devops</category>
    </item>
    <item>
      <title>The EU AI Liability Directive is dead: what actually governs AI agent harm now</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 01 Sep 2026 08:49:12 +0000</pubDate>
      <link>https://dev.to/traceseal/the-eu-ai-liability-directive-is-dead-what-actually-governs-ai-agent-harm-now-34f4</link>
      <guid>https://dev.to/traceseal/the-eu-ai-liability-directive-is-dead-what-actually-governs-ai-agent-harm-now-34f4</guid>
      <description>&lt;h1&gt;
  
  
  The EU AI Liability Directive is dead: what actually governs AI agent harm now
&lt;/h1&gt;

&lt;p&gt;On &lt;strong&gt;6 October 2025&lt;/strong&gt;, a short notice appeared in the EU's Official Journal confirming what the European Commission had already signalled eight months earlier: the &lt;a href="https://www.europarl.europa.eu/legislative-train/theme-a-europe-fit-for-the-digital-age/file-ai-liability-directive" rel="noopener noreferrer"&gt;AI Liability Directive&lt;/a&gt; was not coming. The withdrawal followed a Commission decision at its 2533rd meeting on 16 July 2025, and closed a proposal that had been on the table since 2022 — the one piece of EU law written specifically to make it easier to prove an AI system caused you harm.&lt;/p&gt;

&lt;p&gt;If you run agents that touch EU users, or sell into the EU, the practical question is not "what did we lose." It is: &lt;strong&gt;what governs an AI agent causing harm now, and does it cover the harm your agent could actually cause?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the directive was meant to do
&lt;/h2&gt;

&lt;p&gt;The proposed AI Liability Directive would have complemented the AI Act rather than duplicated it, easing a claimant's path to compensation for AI-caused harm along two specific mechanisms: a court's power to order &lt;strong&gt;disclosure of evidence&lt;/strong&gt; from a provider or deployer of a high-risk AI system, and a &lt;strong&gt;rebuttable presumption of a causal link&lt;/strong&gt; between an AI system's non-compliance and the resulting harm, where the claimant could show fault was likely but proving causation was disproportionately hard given how the system worked. It was designed for exactly the case that makes agents hard to litigate: the claimant cannot see inside the system, and the operator holds all the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it was dropped
&lt;/h2&gt;

&lt;p&gt;The Commission's &lt;a href="https://iapp.org/news/a/european-commission-withdraws-ai-liability-directive-from-consideration" rel="noopener noreferrer"&gt;2025 work programme&lt;/a&gt;, adopted 11 February 2025, listed the directive for withdrawal with a one-line justification: "no foreseeable agreement." Member states and industry groups had spent two years unable to converge on the scope and mechanics, and the withdrawal landed alongside a broader simplification push that also dropped the ePrivacy Regulation. The Commission's own text left the door open rather than closing it: it "will assess whether another proposal should be tabled or another type of approach should be chosen." As of this piece, no replacement has been tabled.&lt;/p&gt;

&lt;h2&gt;
  
  
  What already applies instead
&lt;/h2&gt;

&lt;p&gt;The gap is not empty. The EU's overhauled &lt;strong&gt;Product Liability Directive&lt;/strong&gt;, &lt;a href="https://eur-lex.europa.eu/eli/dir/2024/2853/oj" rel="noopener noreferrer"&gt;Directive (EU) 2024/2853&lt;/a&gt;, was adopted separately and was never at risk — it entered into force in the weeks after its publication on 18 November 2024, and member states have until &lt;strong&gt;9 December 2026&lt;/strong&gt; to transpose it. Unlike the withdrawn directive, it is not AI-specific, but it was substantially rewritten with software in mind: the definition of a covered "product" now explicitly extends to software, firmware and AI systems, closing a long-standing ambiguity about whether a defective model was a "product" at all.&lt;/p&gt;

&lt;p&gt;The mechanics carry over some of what the withdrawn directive was trying to do. A court can compel a defendant to &lt;strong&gt;disclose relevant evidence&lt;/strong&gt; once a claimant has made a sufficiently plausible case. And where the claimant can show a defect is likely but proving it is excessively difficult because of the product's technical complexity — or where the defendant fails to comply with a disclosure order — the product is &lt;strong&gt;presumed defective&lt;/strong&gt;. Practically, that is close to what the AI Liability Directive would have delivered for high-risk AI: the burden moves toward whoever holds the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap it doesn't close
&lt;/h2&gt;

&lt;p&gt;The catch is scope. The Product Liability Directive exists to protect natural persons, and the damage it compensates is &lt;strong&gt;death, personal injury, damage to property kept for private use, and destruction or corruption of data not used for professional purposes&lt;/strong&gt;. It expressly excludes &lt;strong&gt;pure economic loss&lt;/strong&gt; and property "used exclusively for professional purposes." A consumer whose smart device is bricked by a faulty AI update is covered. A business whose AI agent misfires a transaction, corrupts a business dataset, or ships a bad deployment is not — that is a commercial, economic loss, and it stays exactly where the withdrawn directive would have reached: general national contract and tort law, unharmonised across 27 member states.&lt;/p&gt;

&lt;p&gt;That is the case most companies running agents actually have. Agents deployed in a business context, causing business harm, to another business. The one EU proposal aimed squarely at that gap is the one that got withdrawn.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you run an agent
&lt;/h2&gt;

&lt;p&gt;Two different regimes now apply depending on who got hurt, and neither hands a B2B agent operator's counterparty an easy route:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Consumer-facing harm&lt;/strong&gt; — falls under the revised Product Liability Directive from December 2026 for newly placed products. Courts there can already order disclosure and will presume a defect against whoever cannot produce a credible answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business-to-business, economic harm&lt;/strong&gt; — the case a coding agent, a data pipeline agent, or an internal automation agent is most likely to produce — stays on national contract and tort law, with no EU-wide presumption to lean on either way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In both cases, the practical position is the same: nobody is going to take a log file's word for what an agent did. Where the Product Liability Directive applies, the mechanism rewards whoever can actually satisfy a disclosure order with something a court finds credible, not just present. Where it doesn't apply, and the dispute runs on ordinary national rules of evidence, the party who can independently prove what happened is arguing from a position the other side cannot easily contest. A &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;signed execution receipt&lt;/a&gt; — a canonical record of inputs, outputs, timing and sandbox policy, signed with a key the agent never touches — is built to be exactly that kind of artefact: something you hand over under a disclosure order, or into evidence, rather than something a claimant successfully argues you should be presumed to be hiding. What such a receipt does and does not establish, and the fields an audit trail needs regardless of which regime ends up applying, are covered in &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;How to verify what an AI agent actually did&lt;/a&gt; and &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;AI agent audit trails&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The uncomfortable read:&lt;/strong&gt; the one EU proposal written specifically for AI harm was withdrawn for lack of agreement. What replaced it either doesn't reach the B2B case most agents create, or rewards whoever can actually produce evidence once a court asks for it. Either way, "trust our logs" was never the safer bet.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What to do before December 2026
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Map your exposure.&lt;/strong&gt; Separate agent deployments that touch consumers and personal data from ones that only produce business-to-business, economic outcomes — they now sit under genuinely different legal regimes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't wait for the transposition deadline.&lt;/strong&gt; The Product Liability Directive only covers products placed on the market after 9 December 2026; anything shipped earlier, and anything B2B, is already governed by whatever your national law and your contracts say today.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build the evidence trail regardless of which regime applies.&lt;/strong&gt; Disclosure orders and presumptions of defect both cut against whoever cannot produce a credible record. That is true whether a court is applying the new EU rules or your counterparty's home jurisdiction's ordinary contract law.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The receipt format, the verifier and the transparency log are open: the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;spec&lt;/a&gt;, the &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt; and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public log&lt;/a&gt;. Regulation in this area is still being written and rewritten around you. The evidence you can produce when someone finally asks is not something you can backdate.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://traceseal.io/blog/eu-ai-liability-directive-withdrawn/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>legal</category>
      <category>compliance</category>
      <category>security</category>
    </item>
    <item>
      <title>Indirect prompt injection: how an AI agent gets hijacked by the data it reads</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 25 Aug 2026 08:50:13 +0000</pubDate>
      <link>https://dev.to/traceseal/indirect-prompt-injection-how-an-ai-agent-gets-hijacked-by-the-data-it-reads-5ceo</link>
      <guid>https://dev.to/traceseal/indirect-prompt-injection-how-an-ai-agent-gets-hijacked-by-the-data-it-reads-5ceo</guid>
      <description>&lt;p&gt;An agent that reads a web page, an issue tracker, an inbox or a pull-request description is reading text that somebody else wrote. &lt;strong&gt;Indirect prompt injection&lt;/strong&gt; is what happens when that text contains instructions addressed to the model rather than to the reader, and the model carries them out using the agent's tools and the agent's permissions.&lt;/p&gt;

&lt;p&gt;The uncomfortable part is not that the attack exists. It is that the people who study it hardest do not claim to have solved it. This piece covers the mechanism, what the standard mitigations genuinely buy you, and what a team can do about the residue: making what the agent actually did provable after the fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Direct and indirect
&lt;/h2&gt;

&lt;p&gt;OWASP's &lt;a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/" rel="noopener noreferrer"&gt;LLM01:2025 Prompt Injection&lt;/a&gt; entry splits the class in two. A &lt;strong&gt;direct&lt;/strong&gt; injection is one where a user's prompt input "directly alters the behavior of the model in unintended or unexpected ways". An &lt;strong&gt;indirect&lt;/strong&gt; injection occurs "when an LLM accepts input from external sources, such as websites or files" whose content alters the model's behaviour.&lt;/p&gt;

&lt;p&gt;The threat model changes completely between the two. A direct injection requires the attacker to be talking to your system. An indirect one does not. The attacker writes something once, somewhere your agent will eventually read, and waits for the agent to come to them.&lt;/p&gt;

&lt;p&gt;OWASP also records that these inputs "do not need to be human-visible/readable, as long as the content is parsed by the model". An HTML comment, pale text on a pale background, a document's metadata or wording rendered inside an image all qualify. A human reviewing the same source may see nothing at all.&lt;/p&gt;

&lt;p&gt;The technique was named by Simon Willison in &lt;a href="https://simonwillison.net/2022/Sep/12/prompt-injection/" rel="noopener noreferrer"&gt;September 2022&lt;/a&gt;. Its indirect form was set out systematically in &lt;a href="https://arxiv.org/abs/2302.12173" rel="noopener noreferrer"&gt;Not what you've signed up for&lt;/a&gt; (Greshake et al., February 2023), which argued that LLM-integrated applications "blur the line between data and instructions" and demonstrated working attacks against production systems, including Bing's GPT-4-powered chat and code-completion engines. The authors showed that "processing retrieved prompts can act as arbitrary code execution".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why there is no boundary to enforce
&lt;/h2&gt;

&lt;p&gt;Classical injection bugs have a structural fix. SQL injection ends when queries are parameterised, because the database is then told, in a channel the attacker cannot reach, which bytes are code and which are data. Prompt injection has no equivalent. Instructions and content arrive at the model as one undifferentiated token sequence. As Willison puts it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;LLMs are unable to reliably distinguish the importance of instructions based on where they came from. Everything eventually gets glued together into a sequence of tokens and fed to the model.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Delimiters, tagging of untrusted regions and system prompts that say "ignore any instructions in the content below" all raise the cost of an attack, sometimes considerably. What none of them provide is a rule the model is architecturally obliged to obey. They are strong suggestions to a system that treats every token as a suggestion.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lethal trifecta
&lt;/h2&gt;

&lt;p&gt;Willison's &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" rel="noopener noreferrer"&gt;lethal trifecta&lt;/a&gt; framing (June 2025) is the clearest way to work out whether a given agent is exposed. Three capabilities together create the risk:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Access to private data&lt;/strong&gt; — which is usually the entire point of giving the agent tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exposure to untrusted content&lt;/strong&gt; — any route by which attacker-controlled text or images reach the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The ability to communicate externally&lt;/strong&gt; in a way that could carry data back out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"If your agent combines these three features," Willison writes, "an attacker can easily trick it into accessing your private data and sending it to that attacker." The awkward observation is that most agents worth deploying hold all three by design. A coding agent that reads a private repository, browses documentation and opens a pull request has the full set. Dropping one leg is a real mitigation and frequently an unacceptable product decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the mitigations actually buy you
&lt;/h2&gt;

&lt;p&gt;OWASP lists seven countermeasures: constraining model behaviour, defining and validating output formats, input and output filtering, least-privilege access, human approval for high-risk actions, segregating external content, and adversarial testing. Each is worth doing. The entry prefaces all of them with a sentence that deserves quoting exactly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;NIST arrives somewhere similar from another direction. &lt;a href="https://csrc.nist.gov/pubs/ai/100/2/e2025/final" rel="noopener noreferrer"&gt;AI 100-2 E2025&lt;/a&gt;, &lt;em&gt;Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations&lt;/em&gt; (March 2025), describes its own purpose as identifying "current challenges in the life cycle of AI systems" and describing "corresponding methods for mitigating and managing the consequences of those attacks". Managing consequences is the language of an attack that landed.&lt;/p&gt;

&lt;p&gt;So the design question is not only how to stop injection. It is what your system can still tell you once one gets through.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes when the target is an agent
&lt;/h2&gt;

&lt;p&gt;Against a chatbot, a successful injection produces wrong text and a bad afternoon. Against an agent it produces actions: a file written, a secret read, an API called, a branch pushed, a message sent on your behalf. OWASP's impact list includes "executing arbitrary commands in connected systems" and "providing unauthorized access to functions available to the LLM".&lt;/p&gt;

&lt;p&gt;Two things follow. First, blast radius is set by the permissions you granted, not by the model's judgement, because the model's judgement is the component under attack. Second, and less widely appreciated: &lt;strong&gt;the agent's own account of events becomes attacker-influenced output&lt;/strong&gt;. If the injected text says "do X, then report that you did Y", the transcript says Y. Reasoning traces, tool-call summaries and self-reports are all generated downstream of the compromise. They are the last artefacts that should be treated as evidence, and often the only ones a team has kept.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the regulation already assumes
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://artificialintelligenceact.eu/article/15/" rel="noopener noreferrer"&gt;Article 15&lt;/a&gt; of the &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj" rel="noopener noreferrer"&gt;EU AI Act&lt;/a&gt; requires that high-risk AI systems "shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities". The technical measures it calls for must, where appropriate, "prevent, detect, respond to, resolve and control for" AI-specific attacks, expressly including "inputs designed to cause the AI model to make a mistake".&lt;/p&gt;

&lt;p&gt;Article 15 binds high-risk systems, so it will not reach every agent deployment. The drafting is instructive regardless: prevention is one verb out of five. Detection, response and resolution each presuppose an attack that succeeded, and each requires a record that stays trustworthy after the system it describes has been compromised.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two controls that survive a failed prevention
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Containment.&lt;/strong&gt; A kernel-namespace sandbox bounds what an injected instruction can reach, whatever the model decides to do. We published a working bubblewrap profile in &lt;a href="https://traceseal.io/blog/sandbox-an-ai-coding-agent/" rel="noopener noreferrer"&gt;How to sandbox an AI coding agent&lt;/a&gt;, along with the DNS leak we only found by running it. A sandbox limits consequences. It will not tell you that an injection arrived.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A record made outside the agent.&lt;/strong&gt; A &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;signed execution receipt&lt;/a&gt; is a canonical record of inputs, outputs, timing and sandbox policy, stored as SHA-256 hashes and signed with an ed25519 key the sandboxed process never sees. The agent does not write it and cannot amend it. A third party checks it offline, with no access to your systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;traceseal-verify
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;traceseal-verify receipt.json
&lt;span class="go"&gt;[OK] receipt.json — operator signature verified
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of this prevents injection. What it does is make the aftermath answerable. Given a suspected compromise last Tuesday, what did the agent read, what did it touch, and can you demonstrate that to somebody who has no reason to take your word for it? The walkthrough of what such a check does and does not establish is in &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;How to verify what an AI agent actually did&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to record so an injection stays provable
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The exact content the agent read&lt;/strong&gt;, hashed at fetch time. If the poisoned page is quietly edited afterwards, the hash still pins what was actually served to you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every tool call and its real arguments&lt;/strong&gt;, captured at the tool boundary rather than from the model's description of what it intended.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The sandbox policy in force&lt;/strong&gt; during the run, so containment becomes something a reader can check rather than something you assert.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observed side effects&lt;/strong&gt; — files written, network destinations reached, commits created — recorded from outside the agent process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A signature over all of it&lt;/strong&gt;, made with a key the agent cannot reach, so the record's integrity does not depend on the agent having behaved.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The field-level detail, including the EU AI Act's log-retention floor, is in &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;AI agent audit trails&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The question to sit with:&lt;/strong&gt; if one of your agents read a poisoned document last week, would you learn about it from your own records, or from the person its data ended up with?&lt;/p&gt;

&lt;p&gt;The receipt format, the verifier and the transparency log are open: the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;spec&lt;/a&gt;, the &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt; and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public log&lt;/a&gt;. Prevention research will keep moving, and it should. In the meantime, the attacks that get through are the ones your evidence has to account for.&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
    </item>
    <item>
      <title>Transparency log witnesses: how cosigned checkpoints make a log tamper-evident</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 18 Aug 2026 08:49:30 +0000</pubDate>
      <link>https://dev.to/traceseal/transparency-log-witnesses-how-cosigned-checkpoints-make-a-log-tamper-evident-2c4n</link>
      <guid>https://dev.to/traceseal/transparency-log-witnesses-how-cosigned-checkpoints-make-a-log-tamper-evident-2c4n</guid>
      <description>&lt;p&gt;On 17 August 2026, at 14:17:40 UTC, an independent witness operated by &lt;a href="https://witness.markovianprotocol.com/checkpoints" rel="noopener noreferrer"&gt;Markovian Protocol&lt;/a&gt; cosigned the checkpoint of our public transparency log. Nothing about the log changed that afternoon. What changed is what we can prove about it: from that moment, if we ever rewrote the log's history, a third party outside our infrastructure would be holding cryptographic evidence of the version we abandoned.&lt;/p&gt;

&lt;p&gt;This post explains what a transparency log witness is, what a cosigned checkpoint actually proves, and the obligations you take on when one starts watching your log. Most of it we learned by going through the process rather than by reading about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap witnesses close: a log can lie to you specifically
&lt;/h2&gt;

&lt;p&gt;An append-only log is the standard answer to "how do I show my records existed before the dispute did". You publish entries into a Merkle tree, the tree has a root hash, and anyone can check that a given entry is included. Certificate Transparency (&lt;a href="https://datatracker.ietf.org/doc/html/rfc6962" rel="noopener noreferrer"&gt;RFC 6962&lt;/a&gt;) runs the web's certificate issuance through this structure; the &lt;a href="https://go.dev/ref/mod#checksum-database" rel="noopener noreferrer"&gt;Go checksum database&lt;/a&gt; does it for module hashes; &lt;a href="https://docs.sigstore.dev/logging/overview/" rel="noopener noreferrer"&gt;Sigstore's Rekor&lt;/a&gt; does it for software signatures.&lt;/p&gt;

&lt;p&gt;But a Merkle tree on its own has a hole, and RFC 6962 was upfront about it: detecting a misbehaving log requires people to compare notes on what the log showed them. A log operator who controls the serving infrastructure can present one version of the tree to the auditor and a different version to everyone else. Each view is internally consistent, each proof checks out, and the two views never meet. This is the &lt;strong&gt;split-view attack&lt;/strong&gt;, and no amount of hashing inside the log prevents it, because the operator is the one doing the hashing.&lt;/p&gt;

&lt;p&gt;For our purposes the operator's honesty is exactly what is in question. Traceseal exists because &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;logs written by the party under scrutiny are not evidence&lt;/a&gt;. Our transparency log anchors signed agent receipts, and until last week it had the same weakness as every self-hosted log: you had to trust us not to serve you a private version of history.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a checkpoint is
&lt;/h2&gt;

&lt;p&gt;A checkpoint is a small signed statement of the log's current state, in the format specified by &lt;a href="https://c2sp.org/tlog-checkpoint" rel="noopener noreferrer"&gt;c2sp.org/tlog-checkpoint&lt;/a&gt;: the log's name (its origin), the number of entries, and the root hash of the tree at that size, wrapped in a &lt;a href="https://c2sp.org/signed-note" rel="noopener noreferrer"&gt;signed note&lt;/a&gt;. Ours currently reads:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;log.traceseal.io
19
OhQzzSdQHnpofrO/F2reZkiIDDPjUEM+wuGNc4mlD84=

— log.traceseal.io vLzsyMryP2doyfOtISb5...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nineteen entries. Certificate Transparency logs hold billions, and the format is the same, which is rather the point: the protocol does not care whether it is guarding the web PKI or a small receipts log. A checkpoint commits the operator to one specific history up to one specific size. What it cannot do alone is stop the operator issuing two different commitments to two different audiences.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a witness actually does
&lt;/h2&gt;

&lt;p&gt;A witness is a service, run by someone other than the log operator, that keeps the latest checkpoint it has seen for each log it follows. The protocol, &lt;a href="https://c2sp.org/tlog-witness" rel="noopener noreferrer"&gt;c2sp.org/tlog-witness&lt;/a&gt; (v1.0.0, March 2026), is short enough to read over coffee. The core of it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Witnesses verify that the checkpoint is consistent with their previously recorded state of the log (if any), and return a timestamped cosignature.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"Consistent" carries the whole load. When our log grows and asks the witness to cosign the new checkpoint, it must supply a consistency proof: a Merkle proof that the tree at the new size contains the tree the witness already recorded as a prefix. If the proof fails — if any old entry was altered or removed — the witness refuses. The &lt;a href="https://c2sp.org/tlog-cosignature" rel="noopener noreferrer"&gt;cosignature&lt;/a&gt; it returns on success embeds its own timestamp, so the statement a verifier ends up with is precise: &lt;strong&gt;at this time, this witness saw this exact view of the log, and that view extended every view the witness had seen before.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A split view now requires deceiving the witness as well, and the witness's records are not ours to edit. That is the entire trick. No new cryptography, just a signature from someone whose incentives are not our incentives.&lt;/p&gt;

&lt;p&gt;One implementation detail that cost us an afternoon: the log's own &lt;code&gt;/checkpoint&lt;/code&gt; endpoint serves only the log's signature. The cosignature lives on the witness's copy, at &lt;code&gt;/&amp;lt;sha256(origin)&amp;gt;/checkpoint&lt;/code&gt; on the witness's host. If you fetch your own checkpoint and wonder where the witness signature went, it was never going to be there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a cosignature does not prove
&lt;/h2&gt;

&lt;p&gt;Witnessing is content-blind, and it is worth being blunt about the edges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It says nothing about what the entries mean.&lt;/strong&gt; A witness checks append-only consistency, not truth. A log of fabricated receipts, faithfully appended and never rewritten, will be cosigned without complaint. Content verification is a separate job — ours is done by &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;signature checks on each receipt&lt;/a&gt;, and the honest-limits caveats in that post still apply.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One witness is one extra party to corrupt.&lt;/strong&gt; If the log operator and the witness collude, the split view returns. The mitigation is a quorum of independent witnesses, which is why a shared &lt;a href="https://github.com/transparency-dev/witness-network" rel="noopener noreferrer"&gt;witness network&lt;/a&gt; is being built around the C2SP protocols rather than each log recruiting its own friendly observer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A cosigned checkpoint covers the log up to that size.&lt;/strong&gt; Entries appended since the last cosignature are committed by the log's own signature only, until the next witnessing round picks them up.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The part nobody advertises: witnessing binds you
&lt;/h2&gt;

&lt;p&gt;Before the pin, our append-only claim was policy. We could have quietly regenerated the log, re-signed it, and nobody outside would have been the wiser. After the pin it is a commitment with teeth, and the teeth point at us.&lt;/p&gt;

&lt;p&gt;The log must now stay strictly append-only forever, because the witness's copy of every cosigned checkpoint is permanent evidence against any rewrite. And the log's signing key can no longer rotate casually: the witness pins the key it verified, and a checkpoint signed by an unknown key gets refused. We put both obligations in writing to the witness's operator — no history rewrites, and no key rotation without coordinating first — and they were confirmed back to us as the terms of witnessing.&lt;/p&gt;

&lt;p&gt;Worth sitting with: a tamper-evident log is not a feature you add. It is a constraint you accept. If you are not prepared to be caught by your own infrastructure, external witnessing is not for you — which is precisely what makes it persuasive when you do accept it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check our log yourself
&lt;/h2&gt;

&lt;p&gt;None of this asks for trust in the telling. Both views are public:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl https://log.traceseal.io/checkpoint
&lt;span class="nv"&gt;$ &lt;/span&gt;curl https://witness.markovianprotocol.com/&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'log.traceseal.io'&lt;/span&gt; | &lt;span class="nb"&gt;sha256sum&lt;/span&gt; | &lt;span class="nb"&gt;cut&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="s1"&gt;' '&lt;/span&gt; &lt;span class="nt"&gt;-f1&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/checkpoint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bodies should match byte for byte, with the witness's cosignature as an extra line on the second. We have published the verifier we run daily — it checks the log's Ed25519 signature, the witness cosignature against the witness's public key, and that live and pinned bodies agree — as &lt;code&gt;lib/tlog_witness.py&lt;/code&gt; in the &lt;a href="https://github.com/Traceseal/alibi" rel="noopener noreferrer"&gt;Alibi repository&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this fits in an agent accountability stack
&lt;/h2&gt;

&lt;p&gt;A signed receipt proves an execution record has not changed since it was sealed. &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;An audit trail&lt;/a&gt; tells you what to put in that record. The transparency log proves the receipt existed before anyone had a reason to dispute it. The witness closes the last gap in that chain: it stops the log itself being quietly replaced. Each layer covers a failure the previous one cannot.&lt;/p&gt;

&lt;p&gt;One witness is not the end state; a quorum of independent witnesses is. What ours bought is the step from "trust us" to "trust us, or catch us out through a party we do not control". How many witnesses a receipts log needs before a court or an auditor treats its timestamps as settled is a question nobody has an answer to yet. We would rather be early to finding out.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://traceseal.io/blog/transparency-log-witnesses/" rel="noopener noreferrer"&gt;Traceseal blog&lt;/a&gt;. The receipt format, verifier and transparency log are open: &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;spec&lt;/a&gt;, &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt;, &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public log&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
    <item>
      <title>AI agent audit trails: what to record, and what your records can prove</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 11 Aug 2026 08:51:37 +0000</pubDate>
      <link>https://dev.to/traceseal/ai-agent-audit-trails-what-to-record-and-what-your-records-can-prove-4e0e</link>
      <guid>https://dev.to/traceseal/ai-agent-audit-trails-what-to-record-and-what-your-records-can-prove-4e0e</guid>
      <description>&lt;p&gt;An AI agent does not just answer questions. It calls tools, edits files, opens pull requests, moves money, sends messages. Sooner or later one of those actions gets questioned, by a customer, an auditor, a regulator, or your own incident review, and what you can say at that point depends entirely on what you recorded at the time.&lt;/p&gt;

&lt;p&gt;Building a useful audit trail for an agent is really two problems, and teams usually only notice the first one. &lt;strong&gt;Completeness&lt;/strong&gt;: did you capture the right events? &lt;strong&gt;Credibility&lt;/strong&gt;: can anyone other than you believe the record? This post covers both, including the record-keeping duties the EU AI Act now attaches to high-risk systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI agent audit trail should record
&lt;/h2&gt;

&lt;p&gt;Chat transcripts are not an audit trail. The transcript records what the model said; the questions that matter later are about what the system &lt;em&gt;did&lt;/em&gt;. A trail worth keeping records, for every run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Which code and model ran.&lt;/strong&gt; Agent and skill versions, the model identifier, and content hashes of the code that executed, not just a version string someone can retag.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The inputs.&lt;/strong&gt; The task or prompt, the configuration, and the policy in force. Store sensitive payloads as SHA-256 digests rather than plaintext: a hash proves the input was what you say it was without retaining the data itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every tool call.&lt;/strong&gt; Name, arguments, result and exit status. Tool calls are where an agent touches the world; a trail that omits them records the commentary and drops the actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What changed.&lt;/strong&gt; Files written, diffs applied, messages sent, records created, again as digests, so the artefact can be checked against the record later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Timing.&lt;/strong&gt; Start and end of the run and of each step. This is not gold-plating: for biometric systems, &lt;a href="https://artificialintelligenceact.eu/article/12/" rel="noopener noreferrer"&gt;Article 12(3) of the EU AI Act&lt;/a&gt; requires logging "the period of each use of the system (start date and time and end date and time)" explicitly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human involvement.&lt;/strong&gt; Which steps ran under an approval, and who gave it. When something goes wrong, the first question after "what happened" is "who signed it off".&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The regulatory floor: EU AI Act record-keeping
&lt;/h2&gt;

&lt;p&gt;If your agent falls in the Act's high-risk category, record-keeping stops being a best practice and becomes a duty with numbers attached. Three articles of &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj" rel="noopener noreferrer"&gt;Regulation (EU) 2024/1689&lt;/a&gt; do the work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://artificialintelligenceact.eu/article/12/" rel="noopener noreferrer"&gt;Article 12&lt;/a&gt;&lt;/strong&gt; requires high-risk AI systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system", covering events relevant to identifying risk, post-market monitoring, and monitoring of operation by deployers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://artificialintelligenceact.eu/article/19/" rel="noopener noreferrer"&gt;Article 19&lt;/a&gt;&lt;/strong&gt; obliges providers to keep those automatically generated logs, to the extent they are under their control, for a period appropriate to the system's purpose and &lt;strong&gt;at least six months&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://artificialintelligenceact.eu/article/26/" rel="noopener noreferrer"&gt;Article 26(6)&lt;/a&gt;&lt;/strong&gt; places the mirror-image duty on deployers: keep the logs your high-risk system generates, again for at least six months, longer where other EU or national law (data protection in particular) requires it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On timing: under &lt;a href="https://artificialintelligenceact.eu/article/113/" rel="noopener noreferrer"&gt;Article 113&lt;/a&gt; the Act applies generally from 2 August 2026, which brings the obligations for high-risk systems listed in Annex III with it; systems that are high-risk because they are safety components of regulated products (Article 6(1)) follow from 2 August 2027.&lt;/p&gt;

&lt;p&gt;Most coding and operations agents are not high-risk systems under the Act. But the transparency duties in &lt;a href="https://traceseal.io/blog/eu-ai-act-article-50-ai-agents/" rel="noopener noreferrer"&gt;Article 50 already apply&lt;/a&gt; to systems that interact with people or generate content, and if you are going to keep records at all, the Act's six-month retention floor is the obvious number to copy rather than invent your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  A complete log can still prove nothing
&lt;/h2&gt;

&lt;p&gt;Suppose you record all of the above faithfully. You now have a trail that is excellent for debugging and incident response — and still worthless the moment someone outside your organisation needs convincing.&lt;/p&gt;

&lt;p&gt;The problem is authorship. Application logs are written by the same system they describe, stored on infrastructure the operator controls, and editable by anyone with write access to the store. When the record's author is the party whose conduct is in question, the record cannot clear them. Security practice has long treated missing or unprotected logs as a top-tier weakness (it is &lt;a href="https://owasp.org/Top10/A09_2021-Security_Logging_and_Monitoring_Failures/" rel="noopener noreferrer"&gt;A09 in the OWASP Top 10&lt;/a&gt;), but integrity is a separate property from existence. A log that exists, is complete, and could have been rewritten yesterday answers "what happened?" only for people who already trust you.&lt;/p&gt;

&lt;p&gt;We have written up the full argument in &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;why logs are not evidence&lt;/a&gt;; the short version is that WORM storage, centralised log collectors and hash chains each move the trust boundary without removing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the trail verifiable
&lt;/h2&gt;

&lt;p&gt;The fix is mechanical, and none of it is exotic cryptography:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Hash every artefact&lt;/strong&gt; the run touched, so the record commits to specific bytes rather than descriptions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Serialise the record as canonical JSON&lt;/strong&gt;, so there is exactly one byte sequence a signature can cover.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign it at execution time&lt;/strong&gt; with an ed25519 key held outside the agent's reach, so neither the agent nor a later intruder can rewrite history without breaking the seal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anchor it in a transparency log&lt;/strong&gt;, so the record provably existed before the dispute did.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That turns your audit trail into a receipt anyone can check offline, with no access to your systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;traceseal-verify
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;traceseal-verify receipt.json
&lt;span class="go"&gt;[OK] receipt.json — operator signature verified
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We walk through &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;what that verification actually proves&lt;/a&gt; in an earlier post, and the honest limits matter here too. A signed receipt proves the record has not changed since it was sealed and who sealed it. It does not prove the record is complete: an action taken outside the instrumented path leaves no receipt at all. That is why receipts pair naturally with &lt;a href="https://traceseal.io/blog/sandbox-an-ai-coding-agent/" rel="noopener noreferrer"&gt;sandboxing the agent&lt;/a&gt; — the sandbox narrows what can happen off the record, and the receipt seals what happened on it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The test worth applying:&lt;/strong&gt; pick one agent action from last month and try to show a third party what happened, without asking them to trust your database. If the exercise ends at "here is a row in our logging system", you have an audit trail. You do not yet have evidence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A checklist you can apply this week
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Record tool calls and artefacts, not just conversation.&lt;/li&gt;
&lt;li&gt;Hash sensitive inputs and outputs instead of hoarding them.&lt;/li&gt;
&lt;li&gt;Retain records for at least six months, the floor set by &lt;a href="https://artificialintelligenceact.eu/article/19/" rel="noopener noreferrer"&gt;Article 19&lt;/a&gt; and &lt;a href="https://artificialintelligenceact.eu/article/26/" rel="noopener noreferrer"&gt;Article 26(6)&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Sign records when they are made. A trail you start keeping after the dispute begins says nothing about what came before.&lt;/li&gt;
&lt;li&gt;Rehearse retrieval: if you have never pulled a specific run's record under time pressure, you do not know whether you can.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The receipt format, verifier and transparency log are open — the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;spec&lt;/a&gt;, the &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt;, and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public log&lt;/a&gt;. You can adopt the format without adopting us.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://traceseal.io/blog/ai-agent-audit-trail/" rel="noopener noreferrer"&gt;Traceseal blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>compliance</category>
      <category>devops</category>
    </item>
    <item>
      <title>How to sandbox an AI coding agent (and what a sandbox cannot prove)</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 04 Aug 2026 08:50:37 +0000</pubDate>
      <link>https://dev.to/traceseal/how-to-sandbox-an-ai-coding-agent-and-what-a-sandbox-cannot-prove-h9b</link>
      <guid>https://dev.to/traceseal/how-to-sandbox-an-ai-coding-agent-and-what-a-sandbox-cannot-prove-h9b</guid>
      <description>&lt;p&gt;The question usually arrives after the run rather than before it. An agent has spent forty minutes in a repository, the diff is larger than anyone expected, and somebody asks whether it touched anything outside the working directory. The instinct is to search the transcript. The trouble with that instinct is the one this blog keeps returning to: the transcript is written by the process you are asking about, so a clean transcript is exactly what a misbehaving run produces.&lt;/p&gt;

&lt;p&gt;The durable answer is to make the action impossible rather than to watch for it. This post is a working profile for doing that on Linux, an honest account of where it leaked when we tested it, and the point at which a sandbox stops helping.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a sandbox is, underneath the word
&lt;/h2&gt;

&lt;p&gt;"Sandbox" is used loosely enough to mean anything from a Docker container to a permissions prompt. On Linux it has a specific meaning, assembled from three kernel facilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Namespaces&lt;/strong&gt; give a process a private view of a global resource: its own mount table, process tree, or network stack (&lt;a href="https://man7.org/linux/man-pages/man7/namespaces.7.html" rel="noopener noreferrer"&gt;namespaces(7)&lt;/a&gt;). &lt;a href="https://man7.org/linux/man-pages/man7/user_namespaces.7.html" rel="noopener noreferrer"&gt;User namespaces&lt;/a&gt; are the piece that matters for developer tooling, because they let an ordinary account create the others without root.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Seccomp&lt;/strong&gt; filters system calls, so a process can be denied whole classes of kernel interface rather than denied a file (&lt;a href="https://man7.org/linux/man-pages/man2/seccomp.2.html" rel="noopener noreferrer"&gt;seccomp(2)&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Landlock&lt;/strong&gt; lets an unprivileged process restrict its own filesystem access and never get it back, which is the right shape for a wrapper that launches something it does not trust (&lt;a href="https://docs.kernel.org/userspace-api/landlock.html" rel="noopener noreferrer"&gt;kernel documentation&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://github.com/containers/bubblewrap" rel="noopener noreferrer"&gt;Bubblewrap&lt;/a&gt; is a small setuid-or-userns tool that composes the first two into a single command. It is what Flatpak uses underneath, it needs no daemon, and it starts in milliseconds, which matters when the thing you are wrapping is invoked hundreds of times a day. Everything below was run against bubblewrap 0.11.0 on Debian.&lt;/p&gt;

&lt;h2&gt;
  
  
  A profile that actually runs
&lt;/h2&gt;

&lt;p&gt;The shape you want for a coding agent is narrow: it may read the system, it may write to one directory, and it may not reach the network.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;bwrap &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--ro-bind&lt;/span&gt; / / &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--bind&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--dev&lt;/span&gt; /dev &lt;span class="nt"&gt;--proc&lt;/span&gt; /proc &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--unshare-net&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--unshare-pid&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--die-with-parent&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--chdir&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--&lt;/span&gt; your-agent-cli run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Line by line: &lt;code&gt;--ro-bind / /&lt;/code&gt; mounts the entire host filesystem read-only, so the agent keeps the toolchain, libraries and language runtimes it needs and can modify none of them. &lt;code&gt;--bind "$PWD" "$PWD"&lt;/code&gt; punches one writable hole at the working directory. &lt;code&gt;--unshare-net&lt;/code&gt; puts the process in a fresh network namespace with no interface but loopback. &lt;code&gt;--unshare-pid&lt;/code&gt; stops it seeing or signalling other processes on the box, and &lt;code&gt;--die-with-parent&lt;/code&gt; means an orphaned agent is killed rather than left running.&lt;/p&gt;

&lt;p&gt;Two classes of action stop being things you monitor and become things that cannot happen. A write outside the working directory fails at the kernel:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;touch&lt;/span&gt; /home/tim/SHOULD_NOT_EXIST
&lt;span class="go"&gt;touch: cannot touch '/home/tim/SHOULD_NOT_EXIST': Read-only file system
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And there is no route off the machine, so a connection to a raw address fails to establish. That is a better guarantee than any amount of transcript review, because it does not depend on the agent's cooperation, its honesty, or your diligence in reading the output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it leaked
&lt;/h2&gt;

&lt;p&gt;Testing the profile above produced a result we did not expect. Egress was properly gone: &lt;code&gt;ip route&lt;/code&gt; was empty inside the namespace and &lt;code&gt;curl&lt;/code&gt; to a bare IP address failed to connect. But name resolution still worked.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;getent hosts example.com
&lt;span class="go"&gt;2606:4700:10::6814:179a  example.com
2606:4700:10::ac42:93f3  example.com
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The explanation is in &lt;code&gt;/etc/nsswitch.conf&lt;/code&gt;. Debian's hosts line includes the &lt;code&gt;resolve&lt;/code&gt; module, which does not send a DNS packet at all: it talks to systemd-resolved over a Unix socket at &lt;code&gt;/run/systemd/resolve/io.systemd.Resolve&lt;/code&gt;. That socket arrived inside the sandbox as an ordinary file, carried in by the read-only bind of &lt;code&gt;/&lt;/code&gt;. Masking the directory confirms it, and the resolution stops:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;bwrap ... &lt;span class="nt"&gt;--unshare-net&lt;/span&gt; &lt;span class="nt"&gt;--tmpfs&lt;/span&gt; /run/systemd/resolve &lt;span class="nt"&gt;--&lt;/span&gt; getent hosts example.com
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="c"&gt;# no output&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The general lesson is worth more than the specific fix. A network namespace governs interfaces, routes and sockets in the network stack. It has nothing to say about a Unix socket to a host daemon that still has full network access, and a filesystem bind will hand you one of those without mentioning it. Anyone auditing an agent sandbox should assume the same is true of the container runtime socket, the SSH agent socket, and the D-Bus session bus.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Worth stating plainly: we wrote the confident version of the paragraph above first, then ran it, and the run disagreed. A sandbox profile that has not been driven end to end is a hypothesis about isolation, not a control.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three further limits are structural rather than fixable. The working directory is writable by design, and for most repositories that is where the interesting secrets live: &lt;code&gt;.env&lt;/code&gt; files, and &lt;code&gt;.git/config&lt;/code&gt; remotes with an access token embedded in the URL. Environment variables cross the boundary untouched unless you add &lt;code&gt;--clearenv&lt;/code&gt;, so an API key exported in the parent shell is available inside. And the kernel is shared, which is the standing caveat on all container-style isolation and the reason &lt;a href="https://csrc.nist.gov/pubs/sp/800/190/final" rel="noopener noreferrer"&gt;NIST SP 800-190&lt;/a&gt; treats it as weaker than a virtual machine boundary. Docker's &lt;a href="https://docs.docker.com/engine/network/drivers/none/" rel="noopener noreferrer"&gt;&lt;code&gt;none&lt;/code&gt; network driver&lt;/a&gt; is the same idea with the same caveats, if a container is already in your stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody budgets for
&lt;/h2&gt;

&lt;p&gt;Isolation changes what the agent does, not only what it is able to do, and this is a genuine operational cost rather than a footnote. An agent that cannot reach the network will still try to install a package. When that fails, it improvises: it writes a stub, or vendors something from a cache, or quietly adjusts a test so the missing dependency stops mattering. The run does not stop; it goes sideways.&lt;/p&gt;

&lt;p&gt;The second effect is on diagnosis. A run that failed because the sandbox blocked it and a run that failed because the task was wrong produce output that looks much the same, and the agent's own explanation of which one happened is not evidence either. Teams adopting sandboxing tend to lose a week to this before they start logging the policy denials separately from the task failures. It is worth doing on day one. This is also why &lt;a href="https://traceseal.io/blog/alibi-grade-agents-on-ground-truth/" rel="noopener noreferrer"&gt;grading agents against ground truth on disk&lt;/a&gt; beats grading them on what they said they did.&lt;/p&gt;

&lt;h2&gt;
  
  
  From control to evidence
&lt;/h2&gt;

&lt;p&gt;Here is the limit of everything above. A sandbox constrains; it does not communicate. You now know the agent could not write outside the repository. A customer, an auditor or a regulator asking the same question next quarter still has nothing but your description of a profile you say you applied.&lt;/p&gt;

&lt;p&gt;Closing that gap is what an execution receipt is for. The &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;receipt specification&lt;/a&gt; carries a &lt;code&gt;sandbox_profile_hash&lt;/code&gt; field: a SHA-256 hash of the sandbox configuration, meaning the bubblewrap argument list plus the environment. A verifier holding a known-good profile can hash it and check whether the run used that profile or a different one, offline, with &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;one command&lt;/a&gt; and no access to your systems. Combined with the signature over the whole record, an altered profile cannot be presented as the original. For what the signature covers and what an &lt;code&gt;[OK]&lt;/code&gt; means, see &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;the walkthrough&lt;/a&gt; and &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;the primer on why logs are not evidence&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;What that does not do is more interesting, and the specification says so itself rather than leaving a reader to discover it. Among the things a valid receipt does &lt;em&gt;not&lt;/em&gt; prove:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That the sandbox actually enforced the declared profile (the receipt records what the operator &lt;em&gt;claims&lt;/em&gt; the sandbox was; a compromised operator could lie).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Section 6.1 of the same document puts it more bluntly still: the receipt is "an &lt;em&gt;attestation&lt;/em&gt;, not a &lt;em&gt;proof of execution&lt;/em&gt; in the zero-knowledge sense". An operator who never ran bubblewrap at all can sign a receipt naming a profile hash they took from the documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still open
&lt;/h2&gt;

&lt;p&gt;That last gap does not close by thinking harder about receipt design, and we are not going to pretend a signature solves it. It closes only by moving the seal outside the party being trusted, and each route out has a real cost. A second party can run or observe the sandbox and cosign the record, which works and requires someone willing to be that party. Hardware-rooted attestation can bind the measurement to a chip rather than an assertion, at the price of a much heavier deployment. A third-party execution environment removes the operator's discretion entirely, along with the operator's control of their own infrastructure.&lt;/p&gt;

&lt;p&gt;Which of those becomes normal is going to be settled by procurement departments and insurers rather than by cryptographers, and probably not this year. In the meantime the honest position is that a signed profile hash moves an unverifiable claim to a checkable one and stops short of proof, which is still a considerable distance from a log file nobody can check at all.&lt;/p&gt;

&lt;p&gt;The regulatory direction is at least consistent with the effort. &lt;a href="https://artificialintelligenceact.eu/article/15/" rel="noopener noreferrer"&gt;Article 15&lt;/a&gt; of &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj" rel="noopener noreferrer"&gt;Regulation (EU) 2024/1689&lt;/a&gt; requires high-risk systems to be designed for an appropriate level of accuracy, robustness and cybersecurity across their lifecycle. A demonstrable containment boundary is easier to argue than a policy document describing one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions to ask of your own setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Has the profile been run, or only written?&lt;/strong&gt; Try to write outside the workdir and try to reach the network from inside it. Ours resolved DNS through a socket we had not thought about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What sockets came in with the filesystem?&lt;/strong&gt; Enumerate what is under &lt;code&gt;/run&lt;/code&gt; and &lt;code&gt;/var/run&lt;/code&gt; in the mount. Each one is a channel your network namespace does not cover.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What is in the environment?&lt;/strong&gt; If you have not used &lt;code&gt;--clearenv&lt;/code&gt;, every secret in the parent shell is inside the sandbox.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Are policy denials logged separately from task failures?&lt;/strong&gt; If not, you cannot tell a blocked agent from a confused one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can anyone outside your team check which profile ran?&lt;/strong&gt; If the answer is that you would tell them, that is a description, not a record.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The receipt format, the verifier and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;transparency log&lt;/a&gt; are open. You can adopt the format without adopting us.&lt;/p&gt;

</description>
      <category>security</category>
      <category>linux</category>
      <category>ai</category>
      <category>devops</category>
    </item>
    <item>
      <title>Alibi: grade coding agents on what actually happened on disk</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Thu, 30 Jul 2026 08:36:27 +0000</pubDate>
      <link>https://dev.to/traceseal/alibi-grade-coding-agents-on-what-actually-happened-on-disk-4lan</link>
      <guid>https://dev.to/traceseal/alibi-grade-coding-agents-on-what-actually-happened-on-disk-4lan</guid>
      <description>&lt;p&gt;Ask a coding agent whether it finished the job and it will usually say yes. The closing summary may be fluent, itemised and confident. None of that proves what happened.&lt;/p&gt;

&lt;p&gt;We have written before about why &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;logs are not evidence&lt;/a&gt;: a record produced by the system it describes inherits that system's honesty. An agent's account of its own run has the same defect one layer up.&lt;/p&gt;

&lt;p&gt;So we built &lt;a href="https://github.com/Traceseal/alibi" rel="noopener noreferrer"&gt;Alibi&lt;/a&gt;, an MIT-licensed harness that ignores the account entirely. Alibi runs the same missions through agentic CLIs and grades each run against ground truth on disk: the files that exist and the changes that were made. It never grades what the agent says it did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The grading rule
&lt;/h2&gt;

&lt;p&gt;Each configured CLI, currently Claude Code and OpenAI Codex, receives an identical mission brief inside its own &lt;a href="https://github.com/containers/bubblewrap" rel="noopener noreferrer"&gt;bubblewrap sandbox&lt;/a&gt;. Each lane gets a private working tree, an isolated home and an isolated &lt;code&gt;/tmp&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When the process exits, Alibi inspects the tree:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the requested file exist?&lt;/li&gt;
&lt;li&gt;Does it contain what the mission required?&lt;/li&gt;
&lt;li&gt;Did the specified tests pass?&lt;/li&gt;
&lt;li&gt;Did anything change outside the permitted area?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent's exit message is not consulted. A mission passes if the disk says it passed.&lt;/p&gt;

&lt;p&gt;The sandbox also makes containment measurable. A write outside the allowed tree is a recorded violation, not a judgement about style. Alibi retains the working directory, sandbox home, temporary files and logs after the run, so someone can derive the grade again and dispute it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Alibi refuses to claim
&lt;/h2&gt;

&lt;p&gt;Alibi is not a leaderboard. The mission set is small and the runs happened on our machines. A handful of missions is a probe, not a census, so we are not publishing rankings from it.&lt;/p&gt;

&lt;p&gt;It also says nothing about consent behaviour. Lanes run with approvals granted by design because the object of study is execution, not permission seeking. How a CLI behaves when it has to ask first is a separate question.&lt;/p&gt;

&lt;p&gt;The grades remain Alibi's judgement. Every assertion is ordinary code in the mission definition. You can read a check, run it again and dispute it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The grader cannot be trusted either
&lt;/h2&gt;

&lt;p&gt;A grading harness has the same structural problem as the agents it grades: it produces the record of its own run. If our scores mattered to anyone, we could quietly edit them.&lt;/p&gt;

&lt;p&gt;Alibi now supports signed Traceseal receipts. Point one environment variable at an Ed25519 key and each CLI and mission result is sealed into a receipt containing hashes of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the mission definition&lt;/li&gt;
&lt;li&gt;the result record&lt;/li&gt;
&lt;li&gt;the retained evidence tree&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A run-level receipt chains those results into one verifiable unit. Signing happens on the host after the contained processes exit, and the key is never mounted inside the sandbox.&lt;/p&gt;

&lt;p&gt;A second party can act as a witness. The witness tool derives every hash again from the retained evidence, checks each signature and the chain, and cosigns only if everything matches. It refuses to cosign with the original signer's key.&lt;/p&gt;

&lt;p&gt;There is an important limit. Two keys on one machine prove key separation, not operator separation. Full independence requires the witness to verify on hardware the original signer cannot write to.&lt;/p&gt;

&lt;p&gt;The signature means "this run occurred as described." The scores remain Alibi's judgement. Signing makes them tamper-evident, not true.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to ask of any agent benchmark
&lt;/h2&gt;

&lt;p&gt;First, what record was it graded against? If the answer is the agent's final message, the benchmark measured prose.&lt;/p&gt;

&lt;p&gt;Second, could the operator rewrite the results afterwards? A results table in a repository the grader controls is still a log.&lt;/p&gt;

&lt;p&gt;Third, can you run it yourself? The missions, assertions and harness should be readable and executable without the author's cooperation.&lt;/p&gt;

&lt;p&gt;Finally, what was the consent model? Auto-approved runs measure capability. They do not measure conduct, and a result that blurs the two flatters everyone.&lt;/p&gt;

&lt;p&gt;Alibi is our attempt to answer those questions cleanly for our own testing. Whether its missions generalise beyond the behaviours we care about remains open. The repository is small enough to read in an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Traceseal/alibi" rel="noopener noreferrer"&gt;Read the missions, dispute the assertions and grade your own agent on GitHub.&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
      <category>security</category>
    </item>
    <item>
      <title>Signed execution receipts: a primer on why logs are not evidence</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 28 Jul 2026 08:47:31 +0000</pubDate>
      <link>https://dev.to/traceseal/signed-execution-receipts-a-primer-on-why-logs-are-not-evidence-2dfg</link>
      <guid>https://dev.to/traceseal/signed-execution-receipts-a-primer-on-why-logs-are-not-evidence-2dfg</guid>
      <description>&lt;p&gt;Every team running AI agents keeps logs. Almost none of them can answer the question that actually gets asked in a dispute: &lt;strong&gt;why should anyone believe them?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A log is a record your system keeps about itself. It is useful for debugging, indispensable for operations, and — on its own — close to worthless as proof. This is a primer on the difference, and on what a signed execution receipt adds that a log cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What evidence actually demands
&lt;/h2&gt;

&lt;p&gt;Strip away the legal vocabulary and a record has to satisfy four things before a sceptical outsider should rely on it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Integrity.&lt;/strong&gt; It has not been altered since it was created — and if it has, that fact is detectable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attribution.&lt;/strong&gt; Someone specific is on the hook for it. A record nobody vouched for commits nobody.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completeness.&lt;/strong&gt; You are seeing the whole set, not a selection made after the fact by an interested party.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independence.&lt;/strong&gt; The check can be performed by the doubter, without the cooperation of the party being doubted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Note that &lt;em&gt;admissibility&lt;/em&gt; is a much lower bar than any of this. English civil procedure will receive a business record in evidence without requiring a witness to speak to it (&lt;a href="https://www.legislation.gov.uk/ukpga/1995/38/section/9" rel="noopener noreferrer"&gt;Civil Evidence Act 1995, s.9&lt;/a&gt;). Getting a log in front of a tribunal is easy. What the tribunal then decides it is &lt;em&gt;worth&lt;/em&gt; is the entire question, and that is where ordinary logging collapses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where logs fail, property by property
&lt;/h2&gt;

&lt;p&gt;An application log fails all four, and it fails them structurally rather than through bad implementation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Integrity:&lt;/strong&gt; the log is a mutable file on infrastructure the operator controls. Anyone with disk or database access can rewrite history and leave no trace in the artefact itself. This is not a novel observation — computer-security guidance has said for two decades that log data requires explicit protection precisely because it is alterable by whoever holds the system (&lt;a href="https://csrc.nist.gov/pubs/sp/800/92/final" rel="noopener noreferrer"&gt;NIST SP 800-92&lt;/a&gt;), and the federal control catalogue carries a dedicated control, AU-9 &lt;em&gt;Protection of Audit Information&lt;/em&gt;, for exactly this problem (&lt;a href="https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final" rel="noopener noreferrer"&gt;NIST SP 800-53 Rev. 5&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attribution:&lt;/strong&gt; the agent writes the log describing the agent. The witness and the accused are the same process. A misbehaving or compromised agent writes a clean log by construction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Completeness:&lt;/strong&gt; what reaches an auditor is a filtered export — a date range, a grep, a dashboard screenshot. The absence of an entry is indistinguishable from an entry that was never emitted, dropped under load, or removed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independence:&lt;/strong&gt; verification means asking the operator for the file and believing what arrives. That is not verification; it is a courtesy.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;The structural problem: a log asks the reader to trust the writer. No amount of log quality fixes that, because the defect is in the direction of trust, not in the contents.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The partial fixes, and where each one stops
&lt;/h2&gt;

&lt;p&gt;The industry has good, real answers to pieces of this. It is worth being precise about what each one buys, because teams routinely believe they have solved evidence when they have solved storage.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Centralised logging / SIEM.&lt;/strong&gt; Moves the log off the machine that produced it, so a single compromised host cannot quietly rewrite it. Real improvement. But the collector is still operator infrastructure, so integrity now rests on the operator's word about the collector instead of about the host.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Append-only and WORM storage.&lt;/strong&gt; Prevents edits at the storage layer. Also a real improvement — and enforced by a policy the operator configures, can reconfigure, and is asking you to take on trust. It constrains the operator; it does not let an outsider check anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hash chaining.&lt;/strong&gt; Each entry commits to the hash of the previous one, so removing or editing an entry breaks the chain — the idea behind digital timestamping since &lt;a href="https://link.springer.com/article/10.1007/BF00196791" rel="noopener noreferrer"&gt;Haber and Stornetta (1991)&lt;/a&gt;. This is genuinely strong, and it is why Traceseal keeps a hash-chained audit log underneath the receipts. Its limit: a chain proves &lt;em&gt;internal consistency&lt;/em&gt;. Whoever holds the chain can rebuild it end to end and present a perfectly consistent alternative history, unless the chain is anchored to something outside their control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trusted timestamps and transparency logs.&lt;/strong&gt; The anchoring step. An &lt;a href="https://www.rfc-editor.org/rfc/rfc3161" rel="noopener noreferrer"&gt;RFC 3161&lt;/a&gt; timestamp binds data to a time via a third party; a Merkle-tree transparency log in the &lt;a href="https://www.rfc-editor.org/rfc/rfc9162" rel="noopener noreferrer"&gt;RFC 9162&lt;/a&gt; style makes it cryptographically awkward to show two different histories to two different observers. This is the piece that closes the "rebuild the chain" gap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read that list in order and a pattern emerges: each step moves a claim from "trust the operator's process" towards "check the maths yourself". A signed receipt is simply the end of that road, packaged so a single file carries the whole argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the signature adds
&lt;/h2&gt;

&lt;p&gt;An execution receipt is a self-contained JSON document recording one execution: which signed code ran, what it consumed and produced (as SHA-256 hashes, per &lt;a href="https://csrc.nist.gov/pubs/fips/180-4/upd1/final" rel="noopener noreferrer"&gt;FIPS 180-4&lt;/a&gt;, so the record proves integrity without exposing the data), the sandbox policy it ran under, and an Ed25519 signature (&lt;a href="https://www.rfc-editor.org/rfc/rfc8032" rel="noopener noreferrer"&gt;RFC 8032&lt;/a&gt;) over the lot.&lt;/p&gt;

&lt;p&gt;Three design choices do the actual work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The signature covers a canonical encoding.&lt;/strong&gt; Keys sorted at every level, no whitespace, no booleans or nulls — so a given receipt has exactly one valid byte sequence. There is no room for a semantically identical but differently-encoded document to pass. Any edit anywhere breaks verification.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The public key travels inside the receipt.&lt;/strong&gt; Verification is offline and self-contained: no key server, no API call, no account, no request to the operator. The doubter is not dependent on the doubted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hashes, not contents.&lt;/strong&gt; Because inputs and outputs appear only as hashes, a receipt can be published without leaking the data it describes — and anyone holding the original data can recompute the hash and confirm the match.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Set against the four properties: integrity comes from the canonical signature, attribution from the operator key that signed it, independence from offline verification, and completeness from anchoring the record in a &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public transparency log&lt;/a&gt;. That last one is worth stating plainly, because it is the property most evidence schemes quietly skip.&lt;/p&gt;

&lt;p&gt;The format is fully specified — the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;receipt specification&lt;/a&gt; is written so that a developer can implement a verifier in any language without touching our code, and the reference verifier is &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;on PyPI&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;traceseal-verify
&lt;span class="nv"&gt;$ &lt;/span&gt;traceseal-verify receipt.json
&lt;span class="o"&gt;[&lt;/span&gt;OK] receipt.json — operator signature verified
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a field-by-field walkthrough of a real receipt and what an &lt;code&gt;[OK]&lt;/code&gt; does and does not prove, see &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;how to verify what an AI agent actually did&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is becoming a compliance question
&lt;/h2&gt;

&lt;p&gt;The EU AI Act already distinguishes keeping records from being able to stand behind them. &lt;a href="https://artificialintelligenceact.eu/article/12/" rel="noopener noreferrer"&gt;Article 12&lt;/a&gt; of &lt;a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj" rel="noopener noreferrer"&gt;Regulation (EU) 2024/1689&lt;/a&gt; requires high-risk systems to allow the automatic recording of events over their lifetime, to a standard appropriate for traceability. &lt;a href="https://artificialintelligenceact.eu/article/19/" rel="noopener noreferrer"&gt;Article 19&lt;/a&gt; requires providers to keep those automatically generated logs where they are under their control. The obligation is to produce a record that supports traceability — and a record that the producing party can silently rewrite supports it only as far as that party's credibility extends.&lt;/p&gt;

&lt;p&gt;Nothing in the regulation mandates cryptographic receipts. But when the moment comes to demonstrate traceability to a regulator, an enterprise customer or an insurer, the difference between a log export and a signed receipt is the difference between asking to be believed and inviting a check.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to ask of your own stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If an insider altered our records, how would we know?&lt;/strong&gt; If the answer relies on that insider's access being correctly restricted, you have a control, not evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can an outsider verify without our help?&lt;/strong&gt; If verification requires you to hand over a file and be believed, the record proves nothing about your good faith to someone who doubts it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the record anchored outside our control?&lt;/strong&gt; Internal consistency is not the same as an unrewritable history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Was it sealed at the time?&lt;/strong&gt; Receipts only cover executions that were instrumented when they ran. Evidence collection that begins when the dispute begins proves nothing about what came before.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://traceseal.io/blog/signed-execution-receipts-primer/" rel="noopener noreferrer"&gt;traceseal.io/blog&lt;/a&gt;. The receipt spec, verifier and transparency log are open — you can adopt the format without adopting us.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>devops</category>
      <category>compliance</category>
    </item>
    <item>
      <title>How to verify what an AI agent actually did</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Tue, 21 Jul 2026 08:46:24 +0000</pubDate>
      <link>https://dev.to/traceseal/how-to-verify-what-an-ai-agent-actually-did-251f</link>
      <guid>https://dev.to/traceseal/how-to-verify-what-an-ai-agent-actually-did-251f</guid>
      <description>&lt;p&gt;Your agent says it finished the job. Its log file agrees. Neither of those is evidence: the log was written by the same process it describes, and it can be rewritten afterwards by anyone with disk access. If you need to show a customer, an auditor or a court what an agent did, you need a record that &lt;strong&gt;fails loudly when altered&lt;/strong&gt; and that a stranger can check without trusting you. This is a practical walkthrough of doing exactly that with an open verifier and a signed execution receipt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: install the verifier
&lt;/h2&gt;

&lt;p&gt;The verifier is a small open-source Python package with a published &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;receipt specification&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;traceseal-verify
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;traceseal-verify receipt.json
&lt;span class="go"&gt;[OK] receipt.json — operator signature verified
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The check runs &lt;strong&gt;offline&lt;/strong&gt;. It needs no account, no API call and no access to the operator's infrastructure — which is the point. Verification that requires the operator's cooperation is not independent verification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: read the receipt — the anatomy
&lt;/h2&gt;

&lt;p&gt;A receipt is a single JSON document with three blocks. Here is a real one, abridged:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"attestation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"attested_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-04-15T05:29:04Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"operator_fingerprint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ed25519:d8d13a6f..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"operator_public_key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"3ba02728..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"signature"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"4678a52c..."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"execution"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"skill_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"yoast-seo-audit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"skill_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.1.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"inputs_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="s2"&gt;"sha256:6ac78ba8..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"outputs_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sha256:66858bb9..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"sandbox_profile_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sha256:e7a3e6b8..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"skill_manifest_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="s2"&gt;"sha256:79c14974..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"exit_code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ok"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"true"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"wall_time_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;47&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provenance"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"manifest_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sha256:79c14974..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"publisher_fingerprint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ed25519:d8d13a6f..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"transparency_log_seq"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"receipt_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1.0"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each block answers a different question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;provenance&lt;/code&gt; — which code ran.&lt;/strong&gt; The publisher signs a content-addressed manifest: every artefact (the skill definition, its capability declaration, the code itself) is listed by its SHA-256 hash (&lt;a href="https://csrc.nist.gov/pubs/fips/180-4/upd1/final" rel="noopener noreferrer"&gt;FIPS 180-4&lt;/a&gt;). Change one byte of the code and the manifest hash no longer matches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;execution&lt;/code&gt; — what happened.&lt;/strong&gt; Inputs, outputs and the sandbox policy are recorded as hashes of their canonical JSON form. That proves integrity without exposing the data — you can demonstrate the output is unchanged without publishing the output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;attestation&lt;/code&gt; — who vouches for it.&lt;/strong&gt; The operator signs the whole record with an Ed25519 key (&lt;a href="https://www.rfc-editor.org/rfc/rfc8032" rel="noopener noreferrer"&gt;RFC 8032&lt;/a&gt;). The signature covers a canonical JSON encoding — sorted keys, fixed separators — so there is exactly one valid byte sequence for a given receipt. Any edit anywhere breaks the seal.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 3: understand what the check proves
&lt;/h2&gt;

&lt;p&gt;When &lt;code&gt;traceseal-verify&lt;/code&gt; prints &lt;code&gt;[OK]&lt;/code&gt;, four things have been established:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The receipt is byte-for-byte what the operator signed — nothing was edited afterwards, by the agent, the operator's tooling, or anyone who handled the file since.&lt;/li&gt;
&lt;li&gt;The code that ran is exactly the code the publisher signed, down to the hash of each file.&lt;/li&gt;
&lt;li&gt;The recorded inputs, outputs and sandbox policy are the ones present at execution time — anything you are later shown can be checked against the hashes.&lt;/li&gt;
&lt;li&gt;The signing key is identified by fingerprint, so repeated receipts from the same operator are linkable, and the receipt can be anchored in a &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;public transparency log&lt;/a&gt; (the &lt;code&gt;transparency_log_seq&lt;/code&gt; field), which prevents the operator quietly maintaining two versions of history.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What a receipt does not prove
&lt;/h2&gt;

&lt;p&gt;Honest tools state their limits, and this matters if you ever rely on a receipt in a dispute:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It does not prove the work was good.&lt;/strong&gt; A receipt proves what ran and what it produced — not that the output was correct. Checking outcomes against ground truth is a separate, complementary verification step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It does not prove events it never covered.&lt;/strong&gt; A receipt seals one execution. Actions taken outside sealed executions are simply absent — which is why instrumentation has to start before the incident, not after.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It binds a key, not a person.&lt;/strong&gt; The signature proves the holder of the operator key vouched for the record. Tying that key to a legal identity is a key-management question, the same as with any signing scheme.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The trust inversion:&lt;/strong&gt; with logs, the burden is on the reader to trust the writer. With signed receipts, the burden is on the record to survive verification. That is the difference between "our system says it behaved" and "check it yourself".&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why this beats screenshots and log exports
&lt;/h2&gt;

&lt;p&gt;The evidence teams typically produce today — screenshots, log exports, dashboard PDFs — shares one flaw: it is all produced by, or under the control of, the party whose behaviour is in question. Computer-security guidance on log management has warned for years that logs require protection precisely because they are alterable by whoever controls the system (&lt;a href="https://csrc.nist.gov/pubs/sp/800/92/final" rel="noopener noreferrer"&gt;NIST SP 800-92&lt;/a&gt;). A detached signature over a canonical record removes the alterability, and an open verifier removes the need to take anyone's word for it.&lt;/p&gt;

&lt;p&gt;Everything shown here is open: the &lt;a href="https://github.com/traceseal/traceseal-verify/blob/main/RECEIPT-SPEC.md" rel="noopener noreferrer"&gt;receipt spec&lt;/a&gt;, the &lt;a href="https://pypi.org/project/traceseal-verify/" rel="noopener noreferrer"&gt;verifier on PyPI&lt;/a&gt;, and the &lt;a href="https://log.traceseal.io" rel="noopener noreferrer"&gt;transparency log&lt;/a&gt;. You can implement the format yourself; the spec even includes a reference canonical-JSON encoder.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://traceseal.io/blog/verify-what-an-ai-agent-actually-did/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>agents</category>
      <category>devops</category>
    </item>
    <item>
      <title>EU AI Act Article 50: what it means for teams running AI agents</title>
      <dc:creator>Ventse</dc:creator>
      <pubDate>Sun, 19 Jul 2026 10:14:59 +0000</pubDate>
      <link>https://dev.to/traceseal/eu-ai-act-article-50-what-it-means-for-teams-running-ai-agents-p28</link>
      <guid>https://dev.to/traceseal/eu-ai-act-article-50-what-it-means-for-teams-running-ai-agents-p28</guid>
      <description>&lt;p&gt;On 2 August 2026 the transparency obligations in Article 50 of the EU AI Act (Regulation (EU) 2024/1689) became applicable. I run agents daily and I build tooling for them, so I have spent more time inside this article of the Act than is probably healthy. The short version: if your organisation deploys AI systems that interact with people or produce content, this now applies to you. Agents do both. And it applies wherever the system's output is used in the EU, so "we're not an EU company" is not the exit it sounds like.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Article 50 actually requires
&lt;/h2&gt;

&lt;p&gt;Four duties, in practical terms:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;People must know they are interacting with an AI system, unless it is obvious from context.&lt;/li&gt;
&lt;li&gt;Systems that generate text, audio, image or video must mark outputs as artificially generated, in a machine-readable form "where technically feasible". That quoted phrase carries a lot of weight, and I have not seen a settled answer on what it means for plain text.&lt;/li&gt;
&lt;li&gt;Organisations deploying emotion recognition, biometric categorisation or deepfake-style generated content must inform the people affected.&lt;/li&gt;
&lt;li&gt;AI-generated text published to inform the public on matters of public interest must be disclosed as such, unless a human took editorial responsibility for it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Penalties for non-compliance run up to €15 million or 3% of worldwide annual turnover, whichever is higher (Article 99(4)).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents are the hard case
&lt;/h2&gt;

&lt;p&gt;For a chat window, disclosure is a banner. Done.&lt;/p&gt;

&lt;p&gt;An agent plans, calls tools, edits files, sends messages and commits code. The transparency question stops being "did we tell the user it's AI?" and becomes "can we show what the system actually did?" Those are very different engineering problems.&lt;/p&gt;

&lt;p&gt;The uncomfortable part is that the standard answer, application logs, proves nothing to anyone outside your organisation. Logs are written by the same system they describe. They can be edited after the fact. When an agent's action is challenged (a wrong transaction, a leaked file, a published article), "our logs say it behaved" is an assertion, not evidence. A regulator, an auditor or a journalist has no reason to take your word for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the record verifiable instead of trusted
&lt;/h2&gt;

&lt;p&gt;The approach we took with Traceseal is to make the execution record tamper-evident and checkable by a third party:&lt;/p&gt;

&lt;p&gt;The publisher signs the agent skill with an ed25519 key over a content-addressed manifest, so it is provable which code ran. The operator runs it in a kernel-namespace sandbox and signs a record of inputs, outputs, timing and sandbox policy. The record stores SHA-256 hashes rather than the data itself, so it proves integrity without exposing anything sensitive.&lt;/p&gt;

&lt;p&gt;Anyone can then check the receipt offline, with no access to the operator's systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;traceseal-verify
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;traceseal-verify receipt.json
&lt;span class="go"&gt;[OK] receipt.json — operator signature verified
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The receipt is canonical JSON under a signature, so any tampering breaks the seal, whether it comes from the agent, the operator or someone downstream.&lt;/p&gt;

&lt;p&gt;An honest limitation: a receipt only covers executions that were sealed at the time. If you start keeping records after a dispute begins, you have proof of nothing that came before. That cuts both ways. It is an argument for instrumenting now, and it is also a real gap if you are hoping to retrofit compliance onto last quarter's agent runs. You cannot.&lt;/p&gt;

&lt;p&gt;I also do not know yet how regulators will weigh a signed receipt against an ordinary log file when the first Article 50 disputes actually land. The Act sets out the duties. Enforcement practice will be written by cases that have not happened. My bet is that "check it yourself" evidence beats "trust me" evidence in front of any tribunal, but it is a bet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before your next agent ships
&lt;/h2&gt;

&lt;p&gt;Inventory your agents: anything that talks to people or generates content is in scope for at least one of the four duties. Then decide what your evidence standard is. Screenshots and logs are trust-me evidence. Signed receipts are check-it-yourself evidence.&lt;/p&gt;

&lt;p&gt;The receipt format, the verifier and the transparency log are open: the spec, the verifier on PyPI, and the public log are all at &lt;a href="https://traceseal.io/" rel="noopener noreferrer"&gt;traceseal.io&lt;/a&gt;. You can adopt the format without adopting us.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>compliance</category>
      <category>agents</category>
      <category>eu</category>
    </item>
  </channel>
</rss>
