<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Quentin Natoly</title>
    <description>The latest articles on DEV Community by Quentin Natoly (@quentin_natoly).</description>
    <link>https://dev.to/quentin_natoly</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4130405%2F472e6ed9-dfa1-47fb-a634-e2bd70edfc16.png</url>
      <title>DEV Community: Quentin Natoly</title>
      <link>https://dev.to/quentin_natoly</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/quentin_natoly"/>
    <language>en</language>
    <item>
      <title>Your agent issued a €1,180 payment. Your traces can't prove it.</title>
      <dc:creator>Quentin Natoly</dc:creator>
      <pubDate>Fri, 18 Sep 2026 10:56:42 +0000</pubDate>
      <link>https://dev.to/quentin_natoly/your-agent-issued-a-eu1180-payment-your-traces-cant-prove-it-2958</link>
      <guid>https://dev.to/quentin_natoly/your-agent-issued-a-eu1180-payment-your-traces-cant-prove-it-2958</guid>
      <description>&lt;p&gt;A claims-processing agent read a file, checked a policy, and issued a €1,180 transfer. The function ran. The record exists. The money moved.&lt;/p&gt;

&lt;p&gt;I had the official OpenTelemetry GenAI auto-instrumentation switched on the whole time, with content capture enabled and the latest semantic conventions opted in. Then I went looking for the span that attested to the payment.&lt;/p&gt;

&lt;p&gt;There isn't one. Not on Anthropic, not on OpenAI. Four spans per run, none of them &lt;code&gt;execute_tool&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This isn't a bug report. Everything behaved exactly as specified. The problem is that what the specification produces and what an auditor needs are two different things, and nothing in the docs tells you where the gap is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observability and forensics are not the same problem
&lt;/h2&gt;

&lt;p&gt;An observability trace tells you what happened &lt;em&gt;if you trust it&lt;/em&gt;. A forensic record has to hold up &lt;em&gt;against someone who disputes it&lt;/em&gt; — a regulator, a customer, an insurer, an internal audit committee.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Observability&lt;/th&gt;
&lt;th&gt;Forensics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sampled&lt;/td&gt;
&lt;td&gt;Exhaustive on consequential actions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mutable, deletable&lt;/td&gt;
&lt;td&gt;Tamper-evident&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Short retention&lt;/td&gt;
&lt;td&gt;Retention matched to liability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User identity&lt;/td&gt;
&lt;td&gt;Distinct agent identity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most teams running agents in production have the first and believe they have the second. I wanted to know exactly how wide that gap was, so I built a scenario small enough to reason about completely and measured it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the traces actually contain
&lt;/h2&gt;

&lt;p&gt;The tool call &lt;em&gt;is&lt;/em&gt; in the trace. It just isn't where the spec normalises it. Inside &lt;code&gt;gen_ai.output.messages&lt;/code&gt;, each assistant turn carries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1180&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"beneficiary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"client_44190"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"issue_payment"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"toolu_018RmF8kpxzy8vambkdmyMzz"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tool_call"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both providers emit the same shape. So the information isn't lost — but read what it is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provable&lt;/strong&gt;: what the model decided to do, and when it decided it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not provable&lt;/strong&gt;: that it happened, when, or what the system returned.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"The model requested a transfer" and "a transfer was made" are different claims. Only the first one is in your traces. Every dispute you will ever have lives in that gap.&lt;/p&gt;

&lt;p&gt;There's a second-order problem that I find worse. An investigator following the specification looks for &lt;code&gt;gen_ai.tool.call.arguments&lt;/code&gt;, the normalised location, and finds nothing. The data is recoverable — but only if you already know to open a nested JSON blob that the conventions never point you to. A trace that hides its own evidence from someone reading the spec correctly is not much of a record.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is not what the spec asks for
&lt;/h2&gt;

&lt;p&gt;I expected to find that the conventions simply delegated tool execution to application developers. That's not what they say. On the &lt;code&gt;execute_tool&lt;/code&gt; span:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GenAI instrumentations that can instrument tool execution calls SHOULD do so&lt;/strong&gt;, unless another instrumentation can reliably cover all supported tool types.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Neither package does, on either provider. The same note adds that application developers are &lt;em&gt;encouraged&lt;/em&gt; to instrument tool calls that automatic instrumentation doesn't cover — and today that second sentence is carrying all the weight. The burden the spec places on instrumentations has landed entirely on you.&lt;/p&gt;

&lt;p&gt;Worth noting too: &lt;code&gt;gen_ai.tool.call.arguments&lt;/code&gt; and &lt;code&gt;gen_ai.tool.call.result&lt;/code&gt; are &lt;code&gt;opt_in&lt;/code&gt;. Without explicit activation you know an action occurred, not which one. Not the amount, not the recipient.&lt;/p&gt;

&lt;h2&gt;
  
  
  One provider loses the agent's mandate
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;gen_ai.tool.definitions&lt;/code&gt; — the list of tools exposed to the model — appears on OpenAI spans and not on Anthropic ones. Same tool specification sent to both.&lt;/p&gt;

&lt;p&gt;This one isn't cosmetic. The first question after an incident is whether the action was within the agent's mandate. Without the tool definitions captured at call time, you cannot reconstruct what the agent was &lt;em&gt;allowed&lt;/em&gt; to do at the moment it acted. If &lt;code&gt;issue_payment&lt;/code&gt; was added to the agent the day before the incident, the trace will never say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix is structural, and it's yours to write
&lt;/h2&gt;

&lt;p&gt;Adding &lt;code&gt;execute_tool&lt;/code&gt; spans with a propagated &lt;code&gt;tool_call_id&lt;/code&gt;, under an enclosing &lt;code&gt;invoke_agent&lt;/code&gt; span, takes executions proven from 0 to 2 and normalised forensic coverage from 4/11 to 8/11. Roughly two dozen lines.&lt;/p&gt;

&lt;p&gt;But the attributes aren't what makes it work — the &lt;strong&gt;tree&lt;/strong&gt; is. The &lt;code&gt;tool_call_id&lt;/code&gt; on the execution span matches the id the model emitted, so the chain from decision to action holds. Provider and model resolve by walking up to the parent span.&lt;/p&gt;

&lt;p&gt;Which produces an operational consequence I didn't anticipate and now think is the most practically important finding here: &lt;strong&gt;sampling is unsafe on consequential actions.&lt;/strong&gt; The attributes are spread across the tree. Drop a parent span and the child action becomes uninterpretable even though its own span is perfectly intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I got it wrong
&lt;/h2&gt;

&lt;p&gt;While verifying every claim against the conventions YAML, I found that my own benchmark was wrong.&lt;/p&gt;

&lt;p&gt;I had listed &lt;code&gt;gen_ai.agent.version&lt;/code&gt; as something the conventions don't provide, and injected a custom attribute to fill the gap. It does provide it — &lt;code&gt;conditionally_required&lt;/code&gt; on &lt;code&gt;invoke_agent&lt;/code&gt;. My scenario had the value sitting in a variable and simply never set it in the right place.&lt;/p&gt;

&lt;p&gt;That single fix moved coverage by a full point. A gap I had attributed to the specification was a gap in my own instrumentation.&lt;/p&gt;

&lt;p&gt;I'm including this because it &lt;em&gt;is&lt;/em&gt; the lesson. Classifying why an attribute is missing — not sent, sent but not emitted, or genuinely absent from the spec — is the entire discipline, and it's easy to get wrong in the direction that flatters your argument. Read &lt;code&gt;model/*.yaml&lt;/code&gt; at a pinned commit, not the generated docs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What no amount of instrumentation fixes
&lt;/h2&gt;

&lt;p&gt;Three attributes stay absent even after correct instrumentation. Two are real gaps:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;gen_ai.agent.id&lt;/code&gt;&lt;/strong&gt; exists, but the spec scopes it to &lt;em&gt;hosted&lt;/em&gt; agent resources — a Bedrock ARN, a GCP Agent Registry identifier — and explicitly discourages recording in-memory instance ids. For a self-hosted agent, which covers most enterprise deployments, there's no appropriate attribute. The agent's actions stay indistinguishable from those of the human whose credentials it runs under.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;gen_ai.conversation.id&lt;/code&gt;&lt;/strong&gt; exists, but the spec says instrumentations should not invent one when no natural identifier is available. No UUID, no trace id, no hash. So it's frequently absent, and nothing ties a sequence of actions to a business object.&lt;/p&gt;

&lt;p&gt;The third, &lt;code&gt;gen_ai.system_instructions&lt;/code&gt;, is absent because my scenario sends no separate system prompt — an artefact of the setup, not a gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;Two providers, one scenario, one agent shape. Nothing here generalises to multi-agent systems, MCP servers, streaming, or frameworks like LangChain.&lt;/p&gt;

&lt;p&gt;Everything in &lt;code&gt;gen_ai.*&lt;/code&gt; is still marked Development in the conventions. These numbers are pinned to specific versions and will drift — that's expected, and tracking the drift is part of the problem.&lt;/p&gt;

&lt;p&gt;And none of this addresses integrity. The spans are freely mutable and deletable. Hash chaining and signing are a separate problem that I have not solved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Everything is in the repository, including the four span dumps the numbers come from. Both tools are pure functions of those dumps, so you can re-derive every figure offline with no API key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python compare.py anthropic.json openai.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Quentin-NA/agent-trace-forensics" rel="noopener noreferrer"&gt;https://github.com/Quentin-NA/agent-trace-forensics&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Pinned to semantic-conventions-genai &lt;code&gt;b5d8440&lt;/code&gt; (2026-09-08), instrumentation packages &lt;code&gt;1.1b0&lt;/code&gt;/&lt;code&gt;1.1b1&lt;/code&gt;, models &lt;code&gt;claude-sonnet-4-6&lt;/code&gt; and &lt;code&gt;gpt-4o-mini&lt;/code&gt;, run of 2026-09-16.&lt;/p&gt;




&lt;p&gt;If you're running agents that touch money, records, or anything a regulator cares about, the question isn't whether you have traces. It's whether a span exists that proves the action happened — and whether it's still interpretable once the parent has been sampled away.&lt;/p&gt;

&lt;p&gt;I'd be glad to hear from anyone who has measured this differently, or on other providers.&lt;/p&gt;

</description>
      <category>opentelemetry</category>
      <category>observability</category>
      <category>ai</category>
      <category>python</category>
    </item>
  </channel>
</rss>
