<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Asuran</title>
    <description>The latest articles on DEV Community by Asuran (@zkasuran).</description>
    <link>https://dev.to/zkasuran</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4072471%2F23ca2796-c59f-4181-abb0-4dec9252b824.png</url>
      <title>DEV Community: Asuran</title>
      <link>https://dev.to/zkasuran</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zkasuran"/>
    <language>en</language>
    <item>
      <title>The LangGraph agent trace that quietly lost its input and output</title>
      <dc:creator>Asuran</dc:creator>
      <pubDate>Tue, 11 Aug 2026 16:44:40 +0000</pubDate>
      <link>https://dev.to/zkasuran/the-langgraph-agent-trace-that-quietly-lost-its-input-and-output-3cbd</link>
      <guid>https://dev.to/zkasuran/the-langgraph-agent-trace-that-quietly-lost-its-input-and-output-3cbd</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;Sentry's JavaScript SDK can auto instrument LangGraph so every agent run shows up in the AI Agents view with its input and its output attached to a &lt;code&gt;gen_ai.invoke_agent&lt;/code&gt; span. That is the whole reason you reach for it: when an agent misbehaves in production you open the trace and read what went in and what came out. Issue &lt;a href="https://github.com/getsentry/sentry-javascript/issues/19628" rel="noopener noreferrer"&gt;#19628&lt;/a&gt; is about the case where that view is quietly empty and the fix makes the agent trace carry its data no matter how the graph state is shaped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;LangGraph lets you define the graph state two ways. The common tutorials use &lt;code&gt;MessagesAnnotation&lt;/code&gt;, where the state is a list of chat messages under a &lt;code&gt;messages&lt;/code&gt; key. But you can also define a custom state annotation with your own keys, which is normal once an agent tracks more than a chat log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CustomState&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Annotation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Root&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;idea&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Annotation&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;expanded&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Annotation&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
  &lt;span class="na"&gt;validated&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Annotation&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;StateGraph&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;CustomState&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addNode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;expand&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expandNode&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEdge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;START&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;expand&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEdge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;expand&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;END&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;Sentry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;instrumentStateGraph&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;recordInputs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;recordOutputs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;compiled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;my_agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;compiled&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;idea&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;test idea&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The instrumentation read the input as &lt;code&gt;args[0].messages&lt;/code&gt; and the output as &lt;code&gt;result.messages&lt;/code&gt;. For a custom state there is no &lt;code&gt;messages&lt;/code&gt; key, so the input attribute came out as an empty array and the output attribute was never set. No error and no warning. The span looked healthy in the UI with nothing inside it, which is the worst kind of observability bug because you only notice when you already need the data and it is not there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The defect is in &lt;code&gt;packages/server-utils/src/ai/langgraph&lt;/code&gt;. The input read collapsed a missing key into an empty array:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inputMessages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="nx"&gt;LangChainMessage&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;)?.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and the output helper bailed out the moment there was no &lt;code&gt;messages&lt;/code&gt; array:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outputMessages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;resultObj&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;outputMessages&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isArray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outputMessages&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// records nothing for a custom state&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fix keeps the &lt;code&gt;MessagesAnnotation&lt;/code&gt; path exactly as it was. When there is no &lt;code&gt;messages&lt;/code&gt; array, it serializes the whole state and records it, wrapped as a single role and content message so the attribute stays a valid chat array:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// input, in instrumentCompiledGraphInvoke&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputState&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;inputState&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;object&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAttributes&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;GEN_AI_INPUT_MESSAGES&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;([{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputState&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}]),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// output, in setResponseAttributes&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;object&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAttribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;GEN_AI_RESPONSE_TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;([{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;assistant&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}]));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wrapping the state into a role and content message is not a random choice. The same package already wraps raw string prompts that way in &lt;code&gt;extractLLMRequestAttributes&lt;/code&gt; "to align with the chat schema used elsewhere", so this follows an existing convention rather than inventing a new shape. A bare object in &lt;code&gt;gen_ai.input.messages&lt;/code&gt; would risk the UI failing to render it, which would trade one silent empty for another. Serialization uses the SDK's own circular safe &lt;code&gt;stringify&lt;/code&gt; so an unusual state object cannot throw inside the span callback.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;I added three tests in the package's own vitest suite. One drives a graph on a custom state annotation and asserts the input and output state land on the span. One is a regression test that a &lt;code&gt;MessagesAnnotation&lt;/code&gt; graph still records its real messages and its response delta. One asserts a null input (the resume case) records nothing rather than a misleading empty array. To prove the tests actually catch the bug, I ran them against the unfixed source and they fail with &lt;code&gt;expected [] to have a length of 1 but got +0&lt;/code&gt;, which is the empty input the issue describes. With the fix they pass.&lt;/p&gt;

&lt;p&gt;Everything is green on the repo's own gates. The full &lt;code&gt;@sentry/server-utils&lt;/code&gt; suite goes from 335 passing to 338 passing with the three added tests and zero failures. &lt;code&gt;oxlint --type-aware&lt;/code&gt; is clean, &lt;code&gt;oxfmt --check&lt;/code&gt; is clean and &lt;code&gt;tsc&lt;/code&gt; is clean on the changed source and the changed test. The change is about 14 lines of real source across two files.&lt;/p&gt;

&lt;p&gt;I kept the diff focused and matched the maintainer's stated direction. On the issue, a Sentry maintainer wrote that they would "have a go at serializing the entire state here", so serializing the whole state is exactly the approach taken.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;This is a fix to Sentry's own JavaScript SDK, in the LangGraph auto instrumentation that powers the AI Agents product. The value is in what the trace now shows. Before the fix a custom state agent produced an &lt;code&gt;invoke_agent&lt;/code&gt; span whose input and output were empty, so the AI Agents view told you an agent ran and nothing about what it did. After the fix the same span carries the input state and the full output state, so the trace is useful for the exact debugging session it exists for. It also stops emitting a misleading empty array on a resume, which reads as real but blank data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Google AI
&lt;/h2&gt;

&lt;p&gt;I reproduced the bug end to end with real Google Gemini, because a unit test with a mock LLM proves the code path but not the product story. The reproduction builds a LangGraph &lt;code&gt;StateGraph&lt;/code&gt; on a custom state annotation &lt;code&gt;Annotation.Root({ topic, poem })&lt;/code&gt; where one node calls Gemini through &lt;code&gt;@langchain/google-genai&lt;/code&gt; &lt;code&gt;ChatGoogleGenerativeAI&lt;/code&gt; (model &lt;code&gt;gemini-flash-latest&lt;/code&gt;) to write a short poem into the state. The graph is instrumented with the local build of the SDK and a real Sentry client with tracing on, then the &lt;code&gt;invoke_agent&lt;/code&gt; span is read straight off the client.&lt;/p&gt;

&lt;p&gt;Same real agent run, two SDK builds. Against the unfixed build Gemini returned a real poem and the graph state held it, yet the &lt;code&gt;invoke_agent&lt;/code&gt; span reported &lt;code&gt;gen_ai.input.messages&lt;/code&gt; as &lt;code&gt;[]&lt;/code&gt; and had no &lt;code&gt;gen_ai.response.text&lt;/code&gt; at all. Against the fixed build the input state showed up as &lt;code&gt;[{"role":"user","content":"{\"topic\":\"the sea\"}"}]&lt;/code&gt; and the response text carried the whole output state including the Gemini generated poem. Google Gemini produced the content and Sentry is where the missing content reappears once the SDK reads the state correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  A note on how this was built
&lt;/h2&gt;

&lt;p&gt;AI assistance (Claude, Anthropic) was used in developing this change. The design, review and verification were done by the author. Verified locally before submitting: the full &lt;code&gt;@sentry/server-utils&lt;/code&gt; vitest suite (338 passed, up from 335, zero failures), the new tests failing on the unfixed source to prove they catch the bug, &lt;code&gt;oxlint --type-aware&lt;/code&gt; clean, &lt;code&gt;oxfmt --check&lt;/code&gt; clean, &lt;code&gt;tsc&lt;/code&gt; clean on the changed files, plus a real Google Gemini LangGraph run showing the span input and output empty before the fix and populated after.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>javascript</category>
      <category>ai</category>
    </item>
    <item>
      <title>Nine green tests and a parser that never worked once</title>
      <dc:creator>Asuran</dc:creator>
      <pubDate>Tue, 11 Aug 2026 11:51:29 +0000</pubDate>
      <link>https://dev.to/zkasuran/nine-green-tests-and-a-parser-that-never-worked-once-4m3i</link>
      <guid>https://dev.to/zkasuran/nine-green-tests-and-a-parser-that-never-worked-once-4m3i</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two bugs, one story. A parser that failed on every real call, plus the test that certified it as working for two review rounds. The second one is the one worth reading about.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the code does
&lt;/h2&gt;

&lt;p&gt;I was writing adapters for a protocol indexer. Each adapter has a Receipt parser: it receives the on-chain state changes from one transaction and produces a structured outcome, "this account swapped 100 of token A for 98 of token B". Downstream consumers trust that outcome. If the parser picks the wrong movement out of the transaction, the system reports a swap that never happened, with no error anywhere.&lt;/p&gt;

&lt;p&gt;The tests ran against live mainnet, which is where this starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug one: the test that could not fail
&lt;/h2&gt;

&lt;p&gt;The end-to-end test for one adapter looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;halted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;// the call reverted on chain, fine&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;    &lt;span class="c1"&gt;// the call succeeded, check the payload&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reasoning was sensible. This is a live chain, the contract call might legitimately revert depending on pool state, so the test tolerates a revert instead of failing the build on someone else's liquidity.&lt;/p&gt;

&lt;p&gt;Read it again. &lt;strong&gt;There is no input that makes this test fail.&lt;/strong&gt; Success gets checked. Failure gets accepted. The assertion has no opinion.&lt;/p&gt;

&lt;p&gt;Nine cases against Monad mainnet, all green, through two rounds of review. What they were actually reporting was my own parser throwing on every single call. The halt the test tolerated was never the chain rejecting anything. It was my code.&lt;/p&gt;

&lt;p&gt;The reviewer did not catch it. I did not catch it. The suite was green, which is what we both looked at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug two: the parser underneath
&lt;/h2&gt;

&lt;p&gt;With the tolerance removed the real bug surfaced immediately, a one-liner of the worst kind:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;leg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;transfers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="cm"&gt;/* right endpoints, right amount, not our own token */&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.find()&lt;/code&gt; returns the first match. The predicate described a shape that any ERC-20 transfer of the right size could satisfy. So in a transaction carrying two structurally identical movements, the parser reported whichever came first in the log.&lt;/p&gt;

&lt;p&gt;Two ways that goes wrong. I reproduced both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Put a decoy transfer with the same shape ahead of the real one and the decoy is reported as the outcome asset.&lt;/li&gt;
&lt;li&gt;Duplicate the real movement and two genuine transfers collapse into one outcome, while coverage still looks complete.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither raises anything. You get a plausible, wrong answer.&lt;/p&gt;

&lt;p&gt;The fix is a rule rather than a patch: &lt;code&gt;filter&lt;/code&gt; the candidates, then require &lt;strong&gt;exactly one&lt;/strong&gt;. Zero keeps the existing missing-evidence error. Two or more throws and names the count. Ambiguity becomes a loud failure instead of a coin flip, which is the only correct behavior when the data genuinely cannot distinguish two movements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I care more about the test than the parser
&lt;/h2&gt;

&lt;p&gt;The parser bug is ordinary. &lt;code&gt;.find()&lt;/code&gt; where you meant "the only one" is a mistake every codebase has. Once you see it you fix it in a minute.&lt;/p&gt;

&lt;p&gt;The test bug is the interesting one, because &lt;strong&gt;it consumed the evidence I was using to judge the parser.&lt;/strong&gt; A green suite is what I hand a reviewer. A branch in the assertion means the suite cannot tell the happy path from the exact failure the code has, so "green before submit" becomes a sentence nobody checked.&lt;/p&gt;

&lt;p&gt;Three things I now do because of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Any &lt;code&gt;if (failure) {...} else {...}&lt;/code&gt; in a test is unverified until probed.&lt;/strong&gt; So is any comment excusing a failure mode. When a test says a failure is acceptable because external state might reject the call, go find out which one actually happened. Dump the raw trace once, then assert the single real behavior and let it fail loudly if the world changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grep for first-match selection.&lt;/strong&gt; &lt;code&gt;.find(&lt;/code&gt;, &lt;code&gt;[0]&lt;/code&gt;, &lt;code&gt;.at(0)&lt;/code&gt;, &lt;code&gt;.shift()&lt;/code&gt;. For each one, ask what stops a second, equally valid item from matching. If the predicate cannot name something fixed, it is guessing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prove the new assertion fails without the fix. Be careful how you undo it.&lt;/strong&gt; This one bit me a second time. To show an assertion catches the bug you revert the source and run the suite. I used &lt;code&gt;git checkout -- &amp;lt;file&amp;gt;&lt;/code&gt; to undo the experiment, but that file also held the round's real work, so my new assertions vanished and the next run went green for an entirely different reason. Caught by diffing the working tree, not by the suite. So: toggle only the file the experiment needs, never one carrying the change, then capture the failing output to a file the moment it fails, because you cannot reproduce it once the state is restored.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of both bugs
&lt;/h2&gt;

&lt;p&gt;Neither of these was a crash. Nothing threw, nothing logged, no monitor fired. One produced a wrong answer that looked like a right answer. The other produced a green test that looked like a passing test.&lt;/p&gt;

&lt;p&gt;That is the class of bug I have started looking for first, because it is the class that survives review. A crash gets fixed the day it happens. A confident wrong answer ships, then it gets built on.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written with AI assistance (Claude, Anthropic). The bugs, the reproductions and the review history are from my own work on a real open source adapter. Both defects were reproduced from both directions, the decoy and the duplicate, before the fix was submitted. The corrected assertions were confirmed failing against the unpatched parser.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>testing</category>
      <category>typescript</category>
    </item>
    <item>
      <title>The sparring partner was reading my config file</title>
      <dc:creator>Asuran</dc:creator>
      <pubDate>Tue, 11 Aug 2026 11:51:27 +0000</pubDate>
      <link>https://dev.to/zkasuran/the-sparring-partner-was-reading-my-config-file-ljm</link>
      <guid>https://dev.to/zkasuran/the-sparring-partner-was-reading-my-config-file-ljm</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I was tuning an agent for a simulation competition. Thirty in-game days, two players, one shared commodity market. Plant crops, hire hands, buy livestock, sell into a market whose prices move with supply. Whoever ends with more coins wins.&lt;/p&gt;

&lt;p&gt;The agent has about forty knobs. How much cash to keep before buying an animal, how many hands to hire on each day, when to buy more land. Tuning it means picking a knob, moving it, then finding out whether the agent got better.&lt;/p&gt;

&lt;p&gt;Which needs an opponent. So I generated one from my own agent: same file, frozen configuration, saved as &lt;code&gt;bench/rivalsim.py&lt;/code&gt;. Sixteen seeds, sixteen games, count the wins. Standard practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result that should have scared me
&lt;/h2&gt;

&lt;p&gt;I lowered &lt;code&gt;animal_reserve&lt;/code&gt; from 400 to 150, so the agent would keep buying livestock instead of hoarding cash, bought eight cows instead of five and pushed the crew to twelve hands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;16 out of 16. Plus 40,000 coins of margin.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every previous experiment had moved the number by two or three thousand. This one moved it by forty. I had been fiddling with the same parameters for days and suddenly one change swept the board.&lt;/p&gt;

&lt;p&gt;I shipped it to the live ladder. Then I shipped a second variation on the same idea.&lt;/p&gt;

&lt;p&gt;Both lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The knob was wired to both players
&lt;/h2&gt;

&lt;p&gt;My agent reads its overrides from an environment variable. One line near the top:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;_ov&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;KAGG_P&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# tuning overrides, parsed into the config
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The sparring partner was generated by copying the agent source. It copied that line too.&lt;/p&gt;

&lt;p&gt;So when I set &lt;code&gt;animal_reserve=150&lt;/code&gt; and ran the match, the environment variable was set for the whole process. &lt;strong&gt;Both agents read it.&lt;/strong&gt; My opponent also stopped keeping cash, also bought livestock it could not feed, also starved its herd.&lt;/p&gt;

&lt;p&gt;I was not measuring a better strategy. I was measuring an opponent I had quietly sabotaged, using the same lever I thought I was pulling on myself.&lt;/p&gt;

&lt;p&gt;Here is the same configuration re-measured with the override hook stripped out of the partner:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;config&lt;/th&gt;
&lt;th&gt;wins&lt;/th&gt;
&lt;th&gt;margin&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;11/16&lt;/td&gt;
&lt;td&gt;+1,953&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the "breakthrough"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/16&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;-15,314&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the variation I shipped after it&lt;/td&gt;
&lt;td&gt;1/16&lt;/td&gt;
&lt;td&gt;-13,044&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A 16-0 win became 0-16. The margin swung by 55,000 coins. The change was one of the worst regressions I had produced and it presented as the best result of the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it read as a triumph
&lt;/h2&gt;

&lt;p&gt;This is the part worth generalising, because the bug is not specific to farming simulations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sabotaging your opponent and improving yourself move the metric in the same direction.&lt;/strong&gt; Win rate cannot tell them apart. A benchmark exists to hold everything constant except the thing under test. A shared configuration channel breaks that invariant without leaving a trace. No exception, no warning, no failing assertion. Just a number that got better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The magnitude was the tell and I read it as the reward.&lt;/strong&gt; Forty thousand against a usual two or three thousand should have been the moment I stopped and asked what else changed. Instead the size of the jump is what convinced me it was real. That is exactly backwards. I would not have said so out loud before it happened to me.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The generator strips the hook.&lt;/strong&gt; The script that produces a sparring partner now cuts the override block out of the copy, so the partner is a fixed strategy rather than a mirror of whatever I am currently testing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A partner has to prove it is inert.&lt;/strong&gt; Run it against an unrelated third agent, once with the environment variable set to something deliberately destructive and once without. Byte-identical scores or the partner is not frozen. Mine now returns 155,875 either way. That check takes about a minute and it is the cheapest thing in the pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One opponent is not a benchmark.&lt;/strong&gt; Even with an inert partner, a single opponent rewards whatever exploits that opponent. A later build went 16 out of 16 against one partner and then sat around 27th of the ~30 opponents it met on the real ladder. So the gate became a seven-agent round robin with the seating permuted, 144 games each. Two identical agents score about 42% there with 16% ties, which gives me a floor: anything not clearly above that is noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat a great result as a measurement bug until proven otherwise.&lt;/strong&gt; Not as false, just as unverified in a specific direction. The question is no longer "did this help" but "what else did this touch".&lt;/p&gt;

&lt;h2&gt;
  
  
  The line I keep coming back to
&lt;/h2&gt;

&lt;p&gt;A sparring partner that reads the same configuration as the candidate is not a sparring partner. It is a mirror. A mirror always tells you that you are winning.&lt;/p&gt;

&lt;p&gt;If you benchmark anything against a copy of itself (common in agent work: self-play, A/B harnesses, model comparisons) go and check what your two sides share. Environment variables, config files, a singleton, a cache, a random seed. Anything the candidate can reach that the opponent can also reach is a channel where your experiment leaks into your control.&lt;/p&gt;

&lt;p&gt;Two bad submissions went out before I found mine.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written with AI assistance (Claude, Anthropic). The bug, the diagnosis and the numbers above are from my own work and were re-measured before publishing: every figure in the tables comes from a captured 16 seed run, plus the inert-partner check is reproducible with one command.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>testing</category>
      <category>ai</category>
    </item>
    <item>
      <title>A documented prop that silently stacked every component at the origin</title>
      <dc:creator>Asuran</dc:creator>
      <pubDate>Tue, 11 Aug 2026 11:43:35 +0000</pubDate>
      <link>https://dev.to/zkasuran/a-documented-prop-that-silently-stacked-every-component-at-the-origin-5h0f</link>
      <guid>https://dev.to/zkasuran/a-documented-prop-that-silently-stacked-every-component-at-the-origin-5h0f</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/tscircuit/core" rel="noopener noreferrer"&gt;tscircuit&lt;/a&gt; lets you describe a printed circuit board in React. You write &lt;code&gt;&amp;lt;resistor name="R1" footprint="0805" /&amp;gt;&lt;/code&gt; and the renderer produces real manufacturing output: component placement, copper, silkscreen, a netlist, Gerbers. &lt;code&gt;tscircuit/core&lt;/code&gt; is the renderer at the centre of that, so a placement bug there ships into every board anyone exports.&lt;/p&gt;

&lt;p&gt;I found this one by building a board, not by reading the source. That distinction matters. I come back to it at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;I was laying out an LED matrix and reached for &lt;code&gt;pcbLayout={{ matchAdapt: true }}&lt;/code&gt;, a documented prop. Every component landed on top of every other component at the group origin.&lt;/p&gt;

&lt;p&gt;Not an error. Not a warning. Four components, one coordinate, a board that looks plausible until you notice all the footprints are in the same place.&lt;/p&gt;

&lt;p&gt;The cause is four lines of dispatch in &lt;code&gt;Group.ts&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pcbLayoutMode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;grid&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_doInitialPcbLayoutGrid&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pcbLayoutMode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pack&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_doInitialPcbLayoutPack&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pcbLayoutMode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;flex&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_doInitialPcbLayoutFlex&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// "match-adapt" falls off the end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Group._getPcbLayoutMode()&lt;/code&gt; happily returns &lt;code&gt;"match-adapt"&lt;/code&gt;. The dispatch handles grid, pack and flex. So the mode is valid on the way in, no branch claims it, no layout runs, so every component keeps the default position it was born with. Schematic match-adapt works fine, which is what makes it convincing: the same prop name does something real on the other view.&lt;/p&gt;

&lt;p&gt;That is the shape of bug I have learned to trust. Not a crash you can grep for in a log, but a documented input that quietly produces geometry a fab house would happily build wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;PR:&lt;/strong&gt; &lt;a href="https://github.com/tscircuit/core/pull/3137" rel="noopener noreferrer"&gt;tscircuit/core#3137&lt;/a&gt;, which fixes &lt;a href="https://github.com/tscircuit/core/issues/3136" rel="noopener noreferrer"&gt;issue #3136&lt;/a&gt;. +76 lines, no deletions.&lt;/p&gt;

&lt;p&gt;Six lines of source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pcbLayoutMode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;match-adapt&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// There is no dedicated PCB match-adapt layout. Without this branch the&lt;/span&gt;
  &lt;span class="c1"&gt;// mode falls through and every component stays on the group origin,&lt;/span&gt;
  &lt;span class="c1"&gt;// producing overlapping footprints. Fall back to packing so the board is&lt;/span&gt;
  &lt;span class="c1"&gt;// still laid out. Schematic match-adapt is handled separately.&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_doInitialPcbLayoutPack&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus a regression test and a PCB snapshot.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The test asserts the bug, not the feature.&lt;/strong&gt; It would be easy to write a test that renders the board and checks it looks right. That passes for the wrong reasons. Instead it asserts the two things that were actually false before:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Components must be laid out, not collapsed onto a single origin point.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;distinctCenters&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;pcbComponents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;center&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toFixed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;center&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toFixed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;distinctCenters&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pcbComponents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;// Packing keeps footprints apart, so there should be no overlap DRC errors.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;overlapErrors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;circuit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getCircuitJson&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;[]).&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pcb_courtyard_overlap_error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="nx"&gt;el&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pcb_footprint_overlap_error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;overlapErrors&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toHaveLength&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four distinct centres and zero overlap errors. Both fail on the old code, both for the real reason. The PCB snapshot comes along so a reviewer can see the board rather than trust the assertions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I said what the fix is not.&lt;/strong&gt; Match-adapt on the PCB side does not exist in this repo. My branch routes it to the packing layout, which lays the board out correctly. I put that in the PR body as a limitation rather than presenting it as a match-adapt implementation. A maintainer might want the real algorithm instead. Either way the current behavior, silent overlap on a documented prop, is not defensible. The fallback is at least honest about being a fallback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Green before submit, on their gates not mine.&lt;/strong&gt; CI runs format-check, a smoke test and ten test shards. All green on the PR.&lt;/p&gt;

&lt;h2&gt;
  
  
  What building a board taught me that reading code did not
&lt;/h2&gt;

&lt;p&gt;This is the part I would actually tell someone.&lt;/p&gt;

&lt;p&gt;I had spent a stretch scanning this codebase for bugs by reading it. I filed several. Two maintainers pushed back on the batch. The substance of it was: these are nitpicks in the wrong places, they are not user issues, go build a board.&lt;/p&gt;

&lt;p&gt;That was fair and it was useful. So I built boards. An LED matrix, a Bluetooth speaker front end. Within one afternoon of actually using the library the way a user does, I hit a documented prop that silently destroys the layout, worth more than everything the reading pass produced. Reading finds code that looks wrong. Using finds behavior that is wrong.&lt;/p&gt;

&lt;p&gt;The bug was in a file I had already read. I did not see it, because a dispatch handling three of four cases reads like a dispatch. It only becomes a bug when you type the fourth case into a real board and watch the components pile up.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written with AI assistance (Claude, Anthropic). The board that surfaced the bug, the diagnosis, the fix and the tests are mine, verified before submitting: the repo's own test suite plus format checks green locally and in CI across ten shards, the new assertions confirmed failing against unpatched code, plus the issue checked live for an existing fix PR before opening mine.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Sentry's trace said gemini-flash-latest. Google served gemini-3.6-flash.</title>
      <dc:creator>Asuran</dc:creator>
      <pubDate>Tue, 11 Aug 2026 11:43:00 +0000</pubDate>
      <link>https://dev.to/zkasuran/sentrys-trace-said-gemini-flash-latest-google-served-gemini-36-flash-2j7a</link>
      <guid>https://dev.to/zkasuran/sentrys-trace-said-gemini-flash-latest-google-served-gemini-36-flash-2j7a</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;The project is Sentry's own Python SDK, &lt;code&gt;getsentry/sentry-python&lt;/code&gt;. Its &lt;code&gt;google_genai&lt;/code&gt; integration is what turns a Gemini call into a &lt;code&gt;gen_ai.chat&lt;/code&gt; span, so your AI monitoring dashboard has token counts, finish reasons, cost and the model that answered. It ships in the same SDK as the error monitoring, which means for a lot of teams it is the only record of what their agent actually did in production.&lt;/p&gt;

&lt;p&gt;I went in to fix a five month old issue about two missing span attributes. What I found was worse than missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/getsentry/sentry-python/issues/5812" rel="noopener noreferrer"&gt;Issue #5812&lt;/a&gt; says streaming responses never capture &lt;code&gt;response_id&lt;/code&gt; or &lt;code&gt;model_version&lt;/code&gt;. It is exact and it is correct. &lt;code&gt;accumulate_streaming_response&lt;/code&gt; in &lt;code&gt;google_genai/streaming.py&lt;/code&gt; declares both locals, reads text, finish reasons, tool calls and five token counts out of the chunks, then hands the two declared locals to the span without ever reading them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;usage_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;span class="n"&gt;response_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;          &lt;span class="c1"&gt;# declared
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;                &lt;span class="c1"&gt;# declared
&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="p"&gt;...&lt;/span&gt;                     &lt;span class="c1"&gt;# everything else: read from the chunk
&lt;/span&gt;                            &lt;span class="c1"&gt;# response_id and model: never touched
&lt;/span&gt;
&lt;span class="n"&gt;accumulated_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AccumulatedResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;response_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# always None
&lt;/span&gt;    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;             &lt;span class="c1"&gt;# always None
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The non-streaming path in the same integration reads both off the response, so the two halves of one integration disagree.&lt;/p&gt;

&lt;p&gt;Before writing a line of fix I wanted to know the true blast radius, so I built a differential probe. It sends identical content through the real integration twice, once as one whole &lt;code&gt;GenerateContentResponse&lt;/code&gt; and once as a chunk sequence, then diffs the &lt;code&gt;gen_ai.*&lt;/code&gt; attributes that land on the chat span. Six response shapes: plain text, reasoning plus cached tokens, a tool call, a &lt;code&gt;MAX_TOKENS&lt;/code&gt; finish, metadata only on the final chunk, then no usage metadata at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SUMMARY: attributes broken on the streaming path
  gen_ai.response.id                 in 6/6 cases
  gen_ai.response.model              in 6/6 cases
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing else diverged. Text, finish reasons, tool calls and all five token counts match the non-streaming path in every shape. &lt;code&gt;AsyncModels.generate_content_stream&lt;/code&gt; is affected identically. So this is one narrow defect, not a broken integration, which is exactly what you want to know before touching anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then Sentry showed me the real bug
&lt;/h3&gt;

&lt;p&gt;I wired my Sentry DSN in, ran real Gemini traffic through the unfixed code, then opened the span. &lt;code&gt;gen_ai.response.model&lt;/code&gt; was &lt;strong&gt;there&lt;/strong&gt;. My own bug report appeared to be wrong.&lt;/p&gt;

&lt;p&gt;It was not. I dumped the eleven attributes the SDK actually put on the wire and &lt;code&gt;gen_ai.response.model&lt;/code&gt; was not among them. Sentry was filling the row in from &lt;code&gt;gen_ai.request.model&lt;/code&gt;, the model I asked for.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;gemini-2.5-flash&lt;/code&gt; you cannot see the difference, because ask and served are the same string. So I ran it again against &lt;code&gt;gemini-flash-latest&lt;/code&gt;, an alias, with the source reverted to the parent commit. Same prompt, same commit, one streaming and one not:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Span&lt;/th&gt;
&lt;th&gt;&lt;code&gt;gen_ai.response.id&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;gen_ai.response.model&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;streaming (the bug)&lt;/td&gt;
&lt;td&gt;no row at all&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-flash-latest&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;non-streaming (control)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1_l6armtIJOKqfkPnai9yQM&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.6-flash&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Google served &lt;code&gt;gemini-3.6-flash&lt;/code&gt;. The streaming trace says &lt;code&gt;gemini-flash-latest&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fun1nonj55faelyof4tx6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fun1nonj55faelyof4tx6.png" alt="Streaming vs non-streaming gen_ai.chat attributes in Sentry, same prompt and same commit" width="800" height="863"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Same prompt, same commit, side by side in Sentry. Left, the streaming span reports &lt;code&gt;gemini-flash-latest&lt;/code&gt;, the alias I asked for. Right, the non-streaming control reports &lt;code&gt;gemini-3.6-flash&lt;/code&gt;, the model Google actually served. The streaming span also has no &lt;code&gt;response.id&lt;/code&gt; row at all.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That reframes the whole issue. A blank field is honest, you know you have nothing. A field silently backfilled from the request is a confident wrong answer. &lt;code&gt;gen_ai.response.model&lt;/code&gt; exists precisely to record the case where those two differ. Every production streaming span, which is to say every user-facing chat, attributes the generation to whatever alias the caller happened to type.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;gen_ai.response.id&lt;/code&gt; has no fallback, so it is simply gone. It is the only handle that ties a Sentry span to a specific Gemini response, which is what you need when you take a bad generation to Google.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; &lt;a href="https://github.com/zkasuran/sentry-python/tree/fix/google-genai-streaming-response-metadata" rel="noopener noreferrer"&gt;&lt;code&gt;getsentry/sentry-python&lt;/code&gt;, branch &lt;code&gt;fix/google-genai-streaming-response-metadata&lt;/code&gt;&lt;/a&gt;, commit &lt;code&gt;5b4c291&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Six lines of source, inside the loop that was already walking every chunk:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;        &lt;span class="c1"&gt;# Gemini repeats the response id and the served model version on every
&lt;/span&gt;        &lt;span class="c1"&gt;# chunk, but keep the last non-None value so a chunk that omits either
&lt;/span&gt;        &lt;span class="c1"&gt;# one does not discard it.
&lt;/span&gt;        &lt;span class="n"&gt;response_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;response_id&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model_version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus 85 lines of test. Two assertions added to the existing streaming test, plus one new async streaming test where only the first chunk carries the metadata.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlup7cqp705wox85pheh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlup7cqp705wox85pheh.png" alt="The streaming gen_ai.chat span before and after the fix" width="800" height="863"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The streaming span before and after the fix. &lt;code&gt;gen_ai.response.id&lt;/code&gt; goes from absent to recorded, &lt;code&gt;response.model&lt;/code&gt; now carries the served value instead of leaning on the UI fallback.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Per Sentry's &lt;code&gt;CONTRIBUTING.md&lt;/code&gt; a contribution needs maintainer agreement on the issue before a PR exists. A bot enforces that. Two earlier PRs on this exact issue were auto-closed for skipping that step, neither reviewed on merit. So I posted the verified finding on #5812 and asked to pick it up. The PR follows once a maintainer answers. The branch above is the full change either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Last non-&lt;code&gt;None&lt;/code&gt;, not first.&lt;/strong&gt; Real traffic decided this. I measured live Gemini streams and every chunk carries both fields: 4 of 4 chunks on &lt;code&gt;gemini-2.5-flash&lt;/code&gt;, 8 of 8 on &lt;code&gt;gemini-flash-latest&lt;/code&gt;. But this repo's own streaming fixture puts them on the first chunk only. So first-wins passes the tests while last-wins-without-a-guard would let a metadata-free chunk clear a good value. Last non-&lt;code&gt;None&lt;/code&gt; satisfies both, matching the convention the &lt;code&gt;anthropic&lt;/code&gt; integration in the same SDK already uses (&lt;code&gt;model = event.message.model or model&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prove the test catches it.&lt;/strong&gt; A new assertion that passes alongside a fix proves nothing. I reverted the source, kept the tests, then ran again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5 failed, 68 passed
FAILED test_streaming_generate_content[True-True]  - KeyError: 'gen_ai.response.id'
FAILED test_streaming_generate_content[True-False] - KeyError: 'gen_ai.response.id'
FAILED test_streaming_generate_content[False-True] - KeyError: 'gen_ai.response.id'
FAILED test_streaming_generate_content[False-False]- KeyError: 'gen_ai.response.id'
FAILED test_async_streaming_response_id_and_model  - KeyError: 'gen_ai.response.id'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why 510 passing tests missed it, which I think is the interesting part.&lt;/strong&gt; Two reasons, neither of them carelessness.&lt;/p&gt;

&lt;p&gt;The fixture at &lt;code&gt;test_google_genai.py:743&lt;/code&gt; already carries &lt;code&gt;"responseId": "response-id-stream-123"&lt;/code&gt; and &lt;code&gt;"modelVersion": "gemini-1.5-flash"&lt;/code&gt;. The data was on the way in. No streaming test asserted either attribute reached the span. The only assertions on those two constants in the whole 4,760 line file are in the non-streaming test. The suite tested the door, not the delivery.&lt;/p&gt;

&lt;p&gt;Then a hardening commit made the dead code look deliberate. &lt;code&gt;a66281a&lt;/code&gt;, "Guard against &lt;code&gt;None&lt;/code&gt; response ID and response model", rewrote the consumer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;-&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;accumulated_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt;        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SPANDATA&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GEN_AI_RESPONSE_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;accumulated_response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;    &lt;span class="n"&gt;response_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;accumulated_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response_id&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;        &lt;span class="nf"&gt;set_on_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SPANDATA&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GEN_AI_RESPONSE_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a careful &lt;code&gt;None&lt;/code&gt; guard on a value that is always &lt;code&gt;None&lt;/code&gt;. It reads as intentional, so the producer never got a second look. A guard is not a test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I deliberately did not report.&lt;/strong&gt; &lt;code&gt;element_wise_usage_max&lt;/code&gt; takes an independent per-field maximum across chunks, so a stream whose candidate count dips on the final chunk reports &lt;code&gt;output_tokens=9, total_tokens=19&lt;/code&gt; where the final chunk said &lt;code&gt;5&lt;/code&gt; and &lt;code&gt;15&lt;/code&gt;. I measured it and it is real. But the code says what it is doing and why: "Use last possible chunk, in case of interruption, and gracefully handle missing intermediate tokens by taking maximum". That is a stated design choice about interrupted streams, not an oversight. Filing it as a bug would have been me not reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gates, all captured:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pytest tests/integrations/google_genai/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;511 passed, 1 skipped (510 before, one new test)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;new assertions, source reverted&lt;/td&gt;
&lt;td&gt;5 failed, all &lt;code&gt;KeyError: 'gen_ai.response.id'&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ruff check&lt;/code&gt; and &lt;code&gt;ruff format --check&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;clean&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uv run --group typing mypy sentry_sdk&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;28 errors before, 28 after, diff identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the differential probe&lt;/td&gt;
&lt;td&gt;6/6 shapes now report parity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;real Gemini through a real DSN&lt;/td&gt;
&lt;td&gt;both attributes present on the streaming span&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 28 mypy errors are pre-existing on &lt;code&gt;master&lt;/code&gt;. I diffed the output rather than claiming zero, because "mypy passes" would not have been true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;Sentry is not decoration on this one. It is where the second, worse bug was visible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI agent monitoring is the surface.&lt;/strong&gt; The integration emits &lt;code&gt;gen_ai.chat&lt;/code&gt; spans that Sentry renders as first-class AI spans with the model, a token breakdown, cost and context utilisation, plus an Agent Timeline tab beside the waterfall. That panel is what showed me &lt;code&gt;gen_ai.response.model&lt;/code&gt; populated on a span where the SDK had sent no such attribute. A local test could not have told me that, because the fallback lives on Sentry's side, not in the SDK.&lt;/p&gt;

&lt;p&gt;To be exact about what I verified: I know the eleven attributes the SDK sent and I know what the UI displayed. That the fallback is &lt;code&gt;gen_ai.request.model&lt;/code&gt; is an inference from those two facts plus the alias run, not from reading Sentry's server code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distributed tracing gave me the control.&lt;/strong&gt; Each run is one transaction with the &lt;code&gt;gen_ai.chat&lt;/code&gt; span nested under &lt;code&gt;http.server&lt;/code&gt; and the outbound &lt;code&gt;POST generativelanguage.googleapis.com&lt;/code&gt; under that. Running streaming and non-streaming in separate transactions in the same environment made the non-streaming span a control I could put beside the broken one. Same prompt, same commit, one attribute panel each. The bug is not an argument at that point, it is a diff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Environments carried before and after.&lt;/strong&gt; &lt;code&gt;bugsmash-before&lt;/code&gt;, &lt;code&gt;bugsmash-after&lt;/code&gt; and &lt;code&gt;bugsmash-unfixed-alias&lt;/code&gt;, one release string per run, so the two states coexist in one project and neither screenshot depends on my memory of which build was live.&lt;/p&gt;

&lt;p&gt;Everything is real traffic. Real Gemini API, real DSN, EU region ingest. The evidence script sends both runs and reads the attributes off the transport, which is worth a note for anyone else instrumenting this: &lt;code&gt;stream_gen_ai_spans&lt;/code&gt; defaults to &lt;code&gt;True&lt;/code&gt;, so &lt;code&gt;gen_ai.*&lt;/code&gt; spans leave as their own &lt;code&gt;span&lt;/code&gt; envelope items rather than inside the transaction payload. My first capture read &lt;code&gt;before_send_transaction&lt;/code&gt; and found zero gen_ai spans. If your AI spans look missing, that is why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Google AI
&lt;/h2&gt;

&lt;p&gt;Google AI is where the ground truth came from. It changed the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemini's own traffic settled the accumulation rule.&lt;/strong&gt; The repo's fixture says the metadata is on the first chunk. That is a guess. I ran real streams and counted: &lt;code&gt;response_id&lt;/code&gt; and &lt;code&gt;model_version&lt;/code&gt; on 4 of 4 chunks for &lt;code&gt;gemini-2.5-flash&lt;/code&gt;, 8 of 8 for &lt;code&gt;gemini-flash-latest&lt;/code&gt;, one distinct value each, stable across the stream. Fixtures tell you what someone assumed. The API tells you what happens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Gemini alias produced the smoking gun.&lt;/strong&gt; The wrong-model bug is invisible on a pinned model id. It only appears when the served model differs from the requested one, which is what &lt;code&gt;gemini-flash-latest&lt;/code&gt; does: I asked for the alias and &lt;code&gt;model_version&lt;/code&gt; came back &lt;code&gt;gemini-3.6-flash&lt;/code&gt;. That is not a contrived case, it is the documented purpose of a &lt;code&gt;-latest&lt;/code&gt; alias. It is exactly the case &lt;code&gt;gen_ai.response.model&lt;/code&gt; exists to record. Without a real Google-native key there is no way to discover that, because a fixture would have whatever value I typed into it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Longer generations made the stream multi-chunk.&lt;/strong&gt; A short prompt returns a single chunk, which hides any accumulation bug. Asking Gemini for a 120 word paragraph gave 4 to 8 chunks per run, so the last-non-&lt;code&gt;None&lt;/code&gt; rule was tested against real chunk boundaries rather than a two-element list I invented.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell the next person
&lt;/h2&gt;

&lt;p&gt;Read what your instrumentation sends, not what the dashboard shows you. The dashboard is allowed to be helpful. Helpful is not the same as accurate. A row that is populated from a plausible neighbour looks exactly like a row that is right, so it passes every eyeball review you throw at it.&lt;/p&gt;

&lt;p&gt;And when a five month old issue has a careful &lt;code&gt;None&lt;/code&gt; guard sitting on top of it, check whether the value can ever be anything else.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written with AI assistance (Claude, Anthropic). The bug hunt, the probe design, the fix and every number above are mine and were verified before publishing: the differential probe over six response shapes, the google_genai suite at 511 passed and 1 skipped, the same suite at 5 failed with the source reverted, &lt;code&gt;ruff&lt;/code&gt; clean, &lt;code&gt;mypy&lt;/code&gt; diffed against &lt;code&gt;master&lt;/code&gt; rather than claimed clean, plus both Sentry runs captured from real Gemini traffic against a real DSN. Every screenshot is an unedited capture of my own project.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>python</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
