<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Raj Murugan</title>
    <description>The latest articles on DEV Community by Raj Murugan (@rajmurugan).</description>
    <link>https://dev.to/rajmurugan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1906575%2F92e08690-ea8e-4b95-93ce-525ed9f2668c.png</url>
      <title>DEV Community: Raj Murugan</title>
      <link>https://dev.to/rajmurugan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rajmurugan"/>
    <language>en</language>
    <item>
      <title>The harness is one integer column</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Mon, 28 Sep 2026 20:03:54 +0000</pubDate>
      <link>https://dev.to/rajmurugan/the-harness-is-one-integer-column-1mlg</link>
      <guid>https://dev.to/rajmurugan/the-harness-is-one-integer-column-1mlg</guid>
      <description>&lt;p&gt;A background job was re-billing an AI model every sixty seconds, forever. Not a rewrite. Not a redesign. The fix was one INT column, capped at five. And the fix itself shipped with a gap: it forgot to log why a row gave up, so the very rows it saved became impossible to explain. This is what I now do differently, after both halves of that day.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftwqzxsqv8m0e2hexobeb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftwqzxsqv8m0e2hexobeb.png" alt="Diagram of a bounded retry: a poison-pill row polled forever re-bills a model on every lap until an atomic database counter, incremented before each risky attempt and capped at five, flips the row to abandoned and logs why it gave up." width="799" height="658"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure, in one breath
&lt;/h2&gt;

&lt;p&gt;A row that can never reach its terminal state, sitting behind a poller that checks it every tick, is an infinite retry loop wearing a queue's clothes. Nothing was broken in the traditional sense: no exception, no crash, no alarm. A background close job kept picking up the same "in progress" row, calling the model again, failing to close it again, and handing it back to the next poll. Every lap re-billed the model. Nobody was watching for a row that could not finish, because "could not finish" is not an error, it is just still running.&lt;/p&gt;

&lt;p&gt;This is a level-triggered trigger against a condition that never clears. An edge-triggered retry fires once on the transition into failure. A level-triggered poller fires every time it observes the bad state, and if nothing ever changes that state, it fires forever. The row was the poison pill. The poll was the mouth that kept swallowing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the counter belongs in the database, not memory or metrics
&lt;/h2&gt;

&lt;p&gt;The instinctive fix is a counter: retry, but only a few times. Where you put that counter decides whether it works.&lt;/p&gt;

&lt;p&gt;Not in Lambda memory. A retry counter held in the function's own memory resets on every cold start, and a poison-pill row cold-starts its handler just as often as anything else. The counter never reaches five, because it is reintroduced to zero constantly.&lt;/p&gt;

&lt;p&gt;Not in a CloudWatch metric. Metrics observe. They tell you a number happened. They do not gate the next attempt, and a dashboard with a suspicious spike on it does not stop the next poll from firing. Watching is not the same instrument as stopping.&lt;/p&gt;

&lt;p&gt;What actually holds a ceiling under concurrency is a durable, atomic transition in the database itself, and it has to do two jobs, not one: stop the count from being lost, and stop two pollers from both grabbing the same row on the same tick.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Claim the row and count the attempt in one statement. Zero rows back means&lt;/span&gt;
&lt;span class="c1"&gt;-- someone else already has it, or it is already at the cap. Either way, skip&lt;/span&gt;
&lt;span class="c1"&gt;-- it this tick.&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;jobs&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="k"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'claimed'&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'pending'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="n"&gt;RETURNING&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;state = 'pending'&lt;/code&gt; clause is what stops two poller instances from both picking up the same row and both calling the model, not the increment on its own: an increment alone only protects the count, it does not claim the row. The &lt;code&gt;retry_count &amp;lt; 5&lt;/code&gt; clause is what makes the cap a fact the database enforces, rather than something application code has to remember to check after reading the count back. (This is the behaviour under the default Read Committed isolation Postgres and Aurora ship with; under a stricter isolation level a second concurrent writer gets a serialization error instead of quietly waiting its turn, and needs its own retry to match.)&lt;/p&gt;

&lt;p&gt;If the risky work then fails, a second atomic statement decides what happens to the row next:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;jobs&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="k"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;CASE&lt;/span&gt; &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'abandoned'&lt;/span&gt; &lt;span class="k"&gt;ELSE&lt;/span&gt; &lt;span class="s1"&gt;'pending'&lt;/span&gt; &lt;span class="k"&gt;END&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;last_close_error&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="n"&gt;RETURNING&lt;/span&gt; &lt;span class="k"&gt;state&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under the cap, the row goes back to &lt;code&gt;pending&lt;/code&gt; for the next poll. At the cap, the same statement flips it straight to &lt;code&gt;abandoned&lt;/code&gt;, and the poller's own &lt;code&gt;state = 'pending'&lt;/code&gt; clause means it will never claim that row again.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bounded-retry pattern
&lt;/h2&gt;

&lt;p&gt;Three decisions make this actually work, and each one is a place the naive version gets it wrong.&lt;/p&gt;

&lt;p&gt;Claim and increment in the same statement, before the risky work, not after. If you increment on success or on a clean failure, a crash mid-attempt, which is exactly what a poison pill causes, never gets counted, and the cap never bites. Count attempts, not successes.&lt;/p&gt;

&lt;p&gt;Force a terminal state at the cap, enforced by the database, not read back and checked by application code. At five, the same &lt;code&gt;WHERE retry_count &amp;lt; 5&lt;/code&gt; that gates every claim stops matching, and the follow-up statement flips the row to &lt;code&gt;abandoned&lt;/code&gt;. The retry loop has an exit now, one that does not depend on some later code path remembering to check.&lt;/p&gt;

&lt;p&gt;Do the arithmetic on the worst case, honestly. Five attempts at the model's per-call cost is a bounded, small number, call it a dollar. Uncapped, the same bug run for a year at one call a minute is the same unbounded shape as any retry loop nobody put a ceiling on: a small per-call cost, multiplied by nothing ever stopping it. The whole value of the cap is that it turns an open-ended bill into a number you can write down in advance, whatever that number turns out to be for your own per-call cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix needed its own observability
&lt;/h2&gt;

&lt;p&gt;Here is the part I would tell you to do differently if I were doing it again. Version one had the claim, the cap, and the terminal state. It did not have the &lt;code&gt;last_close_error&lt;/code&gt; column shown above: it flipped rows to &lt;code&gt;abandoned&lt;/code&gt; at five and wrote nothing about why.&lt;/p&gt;

&lt;p&gt;So the abandoned rows carried a &lt;code&gt;null&lt;/code&gt; reason. They were safe, in the sense that they had stopped costing money. They were also undiagnosable: nobody could look at a batch of abandoned rows and tell you whether they were all the same bug, five different bugs, or a symptom of something upstream getting worse. Was recovering them safe? Nobody could say, because nobody knew what had actually gone wrong on attempt one through five. &lt;code&gt;last_close_error&lt;/code&gt; was a follow-up migration, added after that gap got noticed, not part of the design from day one.&lt;/p&gt;

&lt;p&gt;A guardrail with no telemetry is a black box you will have to re-debug from scratch the next time it fires, using none of the information the guardrail itself was sitting on the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same shape, one layer up
&lt;/h2&gt;

&lt;p&gt;None of this is specific to a background job. Swap poller for agent loop and the mapping is exact: bound the iterations, force a terminal state at the bound, and log why it stopped. &lt;a href="https://rajmurugan.com/blog/agents-need-a-harness" rel="noopener noreferrer"&gt;I have written before about what happens when an agent loop skips all three&lt;/a&gt;: a runaway that ran for days because nothing capped it, and would have been just as undiagnosable as this one if it had been capped without the third piece.&lt;/p&gt;

&lt;p&gt;The stopping half is the part everyone remembers to build. The observability half is the part that gets cut when the deadline is close, because a capped, silent failure still looks like the incident is over. It is over in the sense that the bill stopped. It is not over in the sense that you can tell anyone what actually happened, and the next stuck row, or the next runaway agent, gets the exact same treatment: caught, capped, and still a mystery.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I now do
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4cqqry1yhyk1a8rd047i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4cqqry1yhyk1a8rd047i.png" alt="Four rows: the counter lives in the row itself, not Lambda memory or a CloudWatch metric; increment before the risky work so the cap counts attempts; force a terminal state at the cap so the loop has an exit; log the reason on every attempt, not just the last one." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A retry ceiling lives in the row it protects, incremented atomically, not in a process's memory and not only in a metric.&lt;/p&gt;

&lt;p&gt;Claim and count in the same statement, gated by the cap, so the database enforces the ceiling instead of application code checking it after the fact.&lt;/p&gt;

&lt;p&gt;A cap without a logged reason is half a guardrail. Write down why it gave up on every attempt, not just the last one.&lt;/p&gt;

&lt;p&gt;Do the worst-case arithmetic before you ship the cap, honestly, so you know the number you are actually bounding.&lt;/p&gt;

&lt;p&gt;If a bounded retry in your system does not write down why it stopped, what would it cost you to find out, the next time it fires?&lt;/p&gt;

&lt;p&gt;This pairs with &lt;a href="https://rajmurugan.com/blog/agents-need-a-harness" rel="noopener noreferrer"&gt;Every dashboard was green while the agent burned six figures a year&lt;/a&gt;, the incident that made me go looking for this pattern everywhere else. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; or via &lt;a href="https://rajmurugan.com" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>observability</category>
      <category>reliability</category>
      <category>agents</category>
    </item>
    <item>
      <title>Not a Python tutorial: the patterns that bite in production agents</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:33:01 +0000</pubDate>
      <link>https://dev.to/rajmurugan/not-a-python-tutorial-the-patterns-that-bite-in-production-agents-3g2n</link>
      <guid>https://dev.to/rajmurugan/not-a-python-tutorial-the-patterns-that-bite-in-production-agents-3g2n</guid>
      <description>&lt;p&gt;This is not a Python tutorial. If you can already write a decorator, skip ahead. The agent loop does not care that you know the language. It cares whether your AWS SDK call just froze every other request your process was supposed to be handling while it waited on a model.&lt;/p&gt;

&lt;p&gt;Rung 1 of the &lt;a href="https://rajmurugan.com/roadmap" rel="noopener noreferrer"&gt;AI Architect Roadmap&lt;/a&gt; is Foundations, and it has sat marked as a gap since the page went up: the Python and tooling that matter once code is running an agent loop in production, not what an intro course teaches. This closes it. Five patterns, all of them things I've either hit myself or watched a production Bedrock workload hit, in roughly the order they bite.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi34ayxpijy1nlr20vm8d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi34ayxpijy1nlr20vm8d.png" alt="Diagram comparing a synchronous boto3 Bedrock Converse call, which blocks every other coroutine in the agent loop for the life of the call, against the same call through aioboto3, awaited, which yields control and lets other coroutines keep running." width="800" height="683"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  boto3 is synchronous, and that is expensive inside an event loop
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;boto3&lt;/code&gt; has no native async support. &lt;code&gt;bedrock-runtime.converse_stream()&lt;/code&gt; returns a plain synchronous &lt;code&gt;EventStream&lt;/code&gt;, and the call itself is a blocking network round trip. Call it directly from inside an &lt;code&gt;asyncio&lt;/code&gt; agent loop and you have not made an async mistake in the abstract, you have stopped every other coroutine in that process from making progress for the entire duration of the model's response, streaming or not.&lt;/p&gt;

&lt;p&gt;The fix is &lt;a href="https://github.com/terricain/aioboto3" rel="noopener noreferrer"&gt;&lt;code&gt;aioboto3&lt;/code&gt;&lt;/a&gt;, which wraps &lt;code&gt;aiobotocore&lt;/code&gt; to give you an async Bedrock client with the same &lt;code&gt;converse_stream()&lt;/code&gt; shape, this time awaited. The gotcha that catches people even after they switch: since aioboto3 8.0, &lt;code&gt;.client()&lt;/code&gt; is an async context manager, and opening a fresh one inside every request re-authenticates and re-establishes the connection each time, the same tax you were trying to avoid, just hidden one layer down. The client, not just the &lt;code&gt;Session&lt;/code&gt;, has to live for the process's lifetime, held open with an &lt;code&gt;AsyncExitStack&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;aioboto3&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;contextlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AsyncExitStack&lt;/span&gt;

&lt;span class="n"&gt;session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;aioboto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Session&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;exit_stack&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AsyncExitStack&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;  &lt;span class="c1"&gt;# populated at startup, reused by every request
&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;startup&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;global&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;exit_stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enter_async_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-runtime&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;shutdown&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;exit_stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;aclose&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle_turn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;converse_stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;modelId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu.anthropic.claude-sonnet-4-6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}]}],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="p"&gt;...&lt;/span&gt;  &lt;span class="c1"&gt;# yields control back to the loop between chunks
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Call &lt;code&gt;startup()&lt;/code&gt; once when the process boots and &lt;code&gt;shutdown()&lt;/code&gt; once when it stops (a framework's lifespan hook, if it has one). Every request in between reuses the same client and the same underlying connection pool.&lt;/p&gt;

&lt;p&gt;If you cannot touch the call site (a library that only exposes the sync client), the fallback is &lt;code&gt;loop.run_in_executor(None, sync_call)&lt;/code&gt; to push the blocking call onto a thread pool instead of the event loop thread. It costs a thread, but it stops the freeze.&lt;/p&gt;

&lt;h2&gt;
  
  
  Typed tool signatures are not optional decoration
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ToolSpecification.html" rel="noopener noreferrer"&gt;Bedrock Converse API's &lt;code&gt;toolSpec&lt;/code&gt;&lt;/a&gt; takes an &lt;code&gt;inputSchema&lt;/code&gt; that must be real JSON Schema, with &lt;code&gt;type: "object"&lt;/code&gt; at the top level and a &lt;code&gt;required&lt;/code&gt; array naming which properties are mandatory. A bare Python type hint on your tool function does not become this automatically if you are calling Bedrock directly with &lt;code&gt;boto3&lt;/code&gt;. Nothing in &lt;code&gt;boto3&lt;/code&gt; converts &lt;code&gt;def get_weather(city: str, unit: str = "celsius")&lt;/code&gt; into a schema on its own, you build one by hand, via a Pydantic model's &lt;code&gt;model_json_schema()&lt;/code&gt;, or by using an agent framework that does the conversion for you (Strands Agents' &lt;code&gt;@tool&lt;/code&gt; decorator parses a function's type hints and docstring into the schema automatically, which is exactly why reaching for &lt;code&gt;boto3&lt;/code&gt; directly instead of a framework is a decision worth making on purpose, not the path of least resistance).&lt;/p&gt;

&lt;p&gt;Get the &lt;code&gt;required&lt;/code&gt; array wrong, drop a type, or leave a field's description empty, and the failure is not a clean exception. The model either omits an argument it needed, or invents a plausible-looking value for a field whose constraints it was never told, and the tool call fails validation two layers downstream from where the real bug is. &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/tool-use-inference-call.html" rel="noopener noreferrer"&gt;AWS's own guidance is explicit that the description is what the model uses to decide when a tool applies&lt;/a&gt;, not just what it does, so a thin schema does not just risk a malformed call, it risks the model never reaching for the tool at all. A correct schema is not a security boundary either: it constrains shape, not content, so a syntactically valid string can still be a path-traversal or injection payload the schema never sees. Validate and sanitise inside the tool handler regardless of how tight the schema looks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming to a browser is a second translation, not a continuation
&lt;/h2&gt;

&lt;p&gt;Bedrock's stream is a boto3 &lt;code&gt;EventStream&lt;/code&gt; of typed chunk events, not a wire format a browser understands. If you are proxying that stream to a frontend over Server-Sent Events, you are writing a translator, not a pass-through: every chunk has to become its own &lt;code&gt;data: ...\n\n&lt;/code&gt; frame, and the frame boundary matters as much as the content.&lt;/p&gt;

&lt;p&gt;The failure mode looks identical to a hung connection: your backend is emitting frames correctly, but something between it and the browser is buffering, an ASGI server, a reverse proxy, a &lt;code&gt;gzip&lt;/code&gt; middleware that waits for enough bytes before flushing. &lt;code&gt;curl --no-buffer&lt;/code&gt; against your own endpoint is the fastest way to tell the difference between "the server is not sending anything yet" and "the server sent it, something in the middle is holding it."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the GIL actually still matters
&lt;/h2&gt;

&lt;p&gt;Python 3.13 shipped a free-threaded build experimentally in October 2024. Python 3.14, a year later, &lt;a href="https://docs.python.org/3/howto/free-threading-python.html" rel="noopener noreferrer"&gt;promoted it to officially supported under PEP 779&lt;/a&gt;. It is still opt-in: the default build keeps the GIL on, you have to explicitly install or build the free-threaded variant, and as of this year package compatibility across the ecosystem sits at roughly half, not all, of what people actually depend on.&lt;/p&gt;

&lt;p&gt;For most agent workloads that barely matters, because the loop is I/O-bound: it spends nearly all of its time waiting on a Bedrock response, not burning CPU, and &lt;code&gt;asyncio&lt;/code&gt; was already the right tool for that, GIL or no GIL. Where it still bites is CPU-bound work sharing the same process as the loop, local tokenisation, embedding a document, post-processing a response with a regex pass over a few hundred KB of text. That work still serialises behind the GIL on a stock interpreter no matter how many coroutines you write, and I have watched a "slow Bedrock call" get blamed on the model when the actual bottleneck was a synchronous embedding step competing with the agent loop for the same core.&lt;/p&gt;

&lt;h2&gt;
  
  
  The order these actually bite in
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A synchronous &lt;code&gt;boto3&lt;/code&gt; call inside the agent loop, freezing every other request.&lt;/li&gt;
&lt;li&gt;A fresh &lt;code&gt;aioboto3&lt;/code&gt; client per call, after the async fix, quietly re-paying the connection cost.&lt;/li&gt;
&lt;li&gt;A tool schema that drifted from the Python type hints it was meant to describe.&lt;/li&gt;
&lt;li&gt;An assumption that a boto3 stream is already browser-ready SSE.&lt;/li&gt;
&lt;li&gt;CPU-bound work blamed on "the model" when it is the GIL serialising a step that shares the process with the loop.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Rung 2, the ML you actually need to operate an LLM rather than train one, is still open. If there is a Python pattern that has bitten you in a production agent that is not on this list, I'd genuinely like to hear which one.&lt;/p&gt;

&lt;p&gt;This post is paired with &lt;a href="https://rajmurugan.com/blog/introducing-the-ai-architect-roadmap" rel="noopener noreferrer"&gt;Introducing the AI Architect Roadmap&lt;/a&gt;, and Evals for Production AI, the series currently climbing rung 7, continues. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; or via &lt;a href="https://rajmurugan.com" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>python</category>
      <category>agents</category>
    </item>
    <item>
      <title>Introducing the AI Architect Roadmap</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:33:01 +0000</pubDate>
      <link>https://dev.to/rajmurugan/introducing-the-ai-architect-roadmap-58e9</link>
      <guid>https://dev.to/rajmurugan/introducing-the-ai-architect-roadmap-58e9</guid>
      <description>&lt;p&gt;"AI Architect" is one of the vaguest titles in this industry right now. It can mean someone who wired up a chatbot demo in an afternoon, or someone who ships and operates a production Agentic AI system that survives real traffic, a security review, and a finance team asking where the AWS bill went. I got tired of pretending those are the same skill set, so I wrote down what I actually think closes the gap between them, as eight rungs, and put it live at &lt;a href="https://rajmurugan.com/roadmap" rel="noopener noreferrer"&gt;rajmurugan.com/roadmap&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4y6ejuq483dhgum5aadh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4y6ejuq483dhgum5aadh.png" alt="The AI Architect Roadmap drawn as a vertical timeline of eight rungs, Foundations through Enterprise Solutions. Rung 7, Observability and Cost, is the flagship in gold. Rungs 3, 4 and 6 are covered in blue, rungs 5 and 8 are strong in green, and rungs 1 and 2 are marked not yet in grey." width="800" height="1360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a map instead of a course
&lt;/h2&gt;

&lt;p&gt;Most "AI architect learning path" content is a listicle wearing a ladder's clothes: a dozen tutorial links in a row, sorted roughly by vibes. It reads like progress and teaches like a syllabus nobody finishes. I wanted something closer to an audit than a course: eight rungs, foundations through enterprise delivery, each one marked against what I have actually shipped and written about in production, not what a junior engineer might eventually get around to in theory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Foundations&lt;/strong&gt;: the Python and tooling that matter once code is running an agent loop in production, not an intro course.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learn AI &amp;amp; ML&lt;/strong&gt;: the machine learning an operator needs, not a trainer, tokens, context windows, the knobs that move cost and behaviour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Master GenAI&lt;/strong&gt;: prompting, RAG, tools, and the SDKs that wire an agent together on Bedrock.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design AI Systems&lt;/strong&gt;: agents that survive contact with production, intent versus state, a harness that bounds a runaway loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build AI Infrastructure&lt;/strong&gt;: CDK, containers, CI/CD with OIDC, cost controls baked in from day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security &amp;amp; Governance&lt;/strong&gt;: deterministic authz in code the model never touches, guardrails as a backstop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability &amp;amp; Cost&lt;/strong&gt;: where the depth is. More on that below.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise Solutions&lt;/strong&gt;: production builds a business can actually adopt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Six of the eight were covered by writing here before today, two of those (5 and 8, with 6 close behind) with enough posts behind them that "covered" undersells it. Two were still marked gap, and I mean that literally, not "coming soon" dressed up as done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the depth actually is
&lt;/h2&gt;

&lt;p&gt;Rung 7, Observability and Cost, is the flagship, and it is not close. It is the rung where I have put the most measured, specific writing: an agent that trended toward a six-figure annual run rate while every dashboard stayed green, prompt caching that ships off by default and, in the workload I measured it against, was worth 55 to 78 percent off the recurring system-prefix bill once switched on, an AgentCore Memory write that returned success and, in my own testing, read back empty for up to thirty seconds on the wrong read path (a rough ballpark from a handful of runs, not a published guarantee). None of that is exotic. It is the ordinary way a probabilistic system hides its failures behind metrics built for a deterministic one, and it is the rung I would tell anyone shipping agentic AI on AWS to take most seriously, because it is the one uptime and error-rate dashboards are structurally blind to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two rungs marked gap, on purpose
&lt;/h2&gt;

&lt;p&gt;Foundations and ML fundamentals were the two I had not written yet, and I thought about quietly filling them with generic content just so the page would read as finished. Every other "complete path" I have seen does exactly that: pads the early rungs with tutorial-grade material because it is the easiest to write, and it is exactly the material a working engineer already knows and skips past. I would rather the page tell the truth. A gap marked gap is more useful than a gap dressed up as coverage.&lt;/p&gt;

&lt;p&gt;That changed today for one of the two. &lt;a href="https://rajmurugan.com/blog/python-patterns-that-bite-in-production-agents" rel="noopener noreferrer"&gt;Not a Python tutorial: the patterns that bite in production agents&lt;/a&gt; closes rung 1: the senior lens on the Python patterns that actually bite in a production agent loop, not another "hello world" first chapter.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Rung 2, the operator's mental model for the machine learning underneath an LLM, is still open. So is finishing Evals for Production AI, the series currently climbing rung 7 (four parts in, no finale yet). Both are queued, not promised for a date, because that is the same honesty this roadmap is trying to hold itself to.&lt;/p&gt;

&lt;p&gt;Which rung would you actually want filled next?&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>genai</category>
      <category>agents</category>
    </item>
    <item>
      <title>You probably don't need to fine-tune</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:29:05 +0000</pubDate>
      <link>https://dev.to/rajmurugan/you-probably-dont-need-to-fine-tune-1e</link>
      <guid>https://dev.to/rajmurugan/you-probably-dont-need-to-fine-tune-1e</guid>
      <description>&lt;p&gt;A post crossed my feed a fortnight ago: "30 fine-tuning interview questions, with answers." The framing was sharp. "Should we fine-tune?" is a trap question in 2026 interviews, the author argued, because the expected answer is usually "not yet", and how you get there is the whole test. Fair call. Two lines from it are worth keeping: every fine-tune has to beat the best prompt on the base model, not the lazy one, and you fine-tune the interface while you retrieve the content, because knowledge changes daily and behaviour changes quarterly.&lt;/p&gt;

&lt;p&gt;I agree with all of it, and none of it told me what happens the day after a team says yes and someone opens the Bedrock console. The questions underneath were model-agnostic: LoRA, QLoRA, DPO, the method zoo. That's the gap I want to close. Not another "should you fine-tune" explainer. What the ladder looks like before you touch weights, what "fine-tune" actually decomposes into once you're inside Bedrock, and what each path costs to run, priced today, not at launch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0cp8wywni8dlt2xkmu4d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0cp8wywni8dlt2xkmu4d.png" alt="Dark scorecard: Provisioned Throughput price per model unit for Amazon Nova Micro and Amazon Nova Pro, both $60.50 an hour with no commitment, despite Micro costing a fraction of Pro on demand." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The ladder, before you touch weights
&lt;/h2&gt;

&lt;p&gt;Prompt, then RAG, then caching, then fine-tune, in that order, and only past whichever step actually stops solving your problem. Most teams that reach for fine-tuning first are trying to fix one of three things: the model doesn't know something (that's retrieval), the model is slow or expensive per call (that's caching, or a smaller model), or the model's tone drifts (that's a better system prompt, most of the time). None of those need a training job.&lt;/p&gt;

&lt;p&gt;The three reasons that genuinely do:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Structured output you can't prompt your way to.&lt;/strong&gt; A rigid schema, a house format, a tool-call convention the base model won't hold consistently across thousands of calls even with a strong system prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost or latency distillation.&lt;/strong&gt; You've proven a big model does the job well and you want a small, cheap, fast model to do the same job for a fraction of the price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Narrow domain vocabulary or tone the prompt can't carry.&lt;/strong&gt; Not "be more formal", but a genuine specialist register: clinical shorthand, legal drafting conventions, a product's own internal taxonomy, the kind of thing a few-shot prompt approximates and a fine-tune nails.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The industry has mostly converged on that ladder. What's missing from every version of it I've read is the other side of the "yes": once you actually commit, the decision forks twice, once into different job types and again into different ways to serve the result, each with its own bill, and almost everything written about it treats "fine-tuning" as one thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Fine-tune" is four different jobs on Bedrock
&lt;/h2&gt;

&lt;p&gt;Bedrock's &lt;code&gt;CreateModelCustomizationJob&lt;/code&gt; API takes a &lt;code&gt;customizationType&lt;/code&gt; parameter, and the values are not flavours of the same thing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;FINE_TUNING&lt;/code&gt;&lt;/strong&gt;: supervised fine-tuning on your own labelled examples. The classic case: input/output pairs, a training job, a customised model at the end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;CONTINUED_PRE_TRAINING&lt;/code&gt;&lt;/strong&gt;: unlabelled domain data, for building broader domain adaptation rather than a specific task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;DISTILLATION&lt;/code&gt;&lt;/strong&gt;: you don't write training pairs at all. You point Bedrock at a stronger "teacher" model and a weaker "student" model, hand it your prompts or invocation logs, and Bedrock generates the synthetic training set and runs the job for you. &lt;a href="https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available/" rel="noopener noreferrer"&gt;General availability landed in May 2025&lt;/a&gt;, and AWS's own published figures put distilled models at up to 500% faster and 75% cheaper than the teacher, with under 2% accuracy loss on retrieval-style tasks. If your actual reason to fine-tune is reason 2 above, cost and latency distillation, this is very often the more direct route to it than hand-rolled supervised fine-tuning, because you skip building the labelled dataset entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;REINFORCEMENT_FINE_TUNING&lt;/code&gt;&lt;/strong&gt;: the newest of the four, Nova-only today. You supply a reward function, either a Bedrock model acting as judge or your own Lambda grading logic, and the job trains against that signal instead of a fixed answer key. It's the most direct fit for reason 1 above, structured output and tool-calling reliability, because you can grade "did the tool call parse" programmatically rather than needing labelled pairs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Serving the result forks again
&lt;/h2&gt;

&lt;p&gt;Bedrock's own documentation still says, plainly, "if you customized a model, you must purchase Provisioned Throughput to be able to use it." Read literally, and read against an older doc page in isolation, that sounds like the whole story. It isn't, and the gap between the two is exactly the kind of thing that's easy to get wrong reading one page instead of the console.&lt;/p&gt;

&lt;p&gt;Since &lt;strong&gt;16 July 2025&lt;/strong&gt;, a second path exists: &lt;code&gt;CreateCustomModelDeployment&lt;/code&gt;, which AWS explicitly describes as complementing Provisioned Throughput, not replacing it. Deploy a natively customised model this way and you pay the base model's own on-demand token rate, no reservation, no hourly floor. The catch is eligibility, and it's narrow: only &lt;strong&gt;Amazon Nova Lite, Nova 2 Lite, Nova Micro, Nova Pro&lt;/strong&gt; (all &lt;code&gt;us-east-1&lt;/code&gt;) and &lt;strong&gt;Meta Llama 3.3 70B Instruct&lt;/strong&gt; (&lt;code&gt;us-west-2&lt;/code&gt;), and only if the underlying model was customised &lt;strong&gt;on or after 16 July 2025&lt;/strong&gt; for Nova, or 15 September 2025 for Llama. Everything else, every Titan customisation, every model customised before those dates, every region outside those two, still has exactly one way to serve it: Provisioned Throughput.&lt;/p&gt;

&lt;p&gt;So the real shape is three serving paths, not two:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Provisioned Throughput&lt;/strong&gt;: universal. Works for any customised model, any region Bedrock customisation supports. Reserved capacity, billed hourly, whether a request arrives or not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom Model Deployment&lt;/strong&gt;: on-demand, same per-token price as the base model, but only for the specific Nova and Llama models above, customised after their respective GA dates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom Model Import&lt;/strong&gt;: for models you trained entirely outside Bedrock (SageMaker, your own PEFT or LoRA run) in a supported open-weight architecture, imported and served on demand, billed per Custom Model Unit-minute, scaling to zero when idle. More on this one below.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcurof890beuzj8ykw09e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcurof890beuzj8ykw09e.png" alt="Architecture diagram: a Bedrock base model forks into native customisation (FINE_TUNING, CONTINUED_PRE_TRAINING, DISTILLATION, REINFORCEMENT_FINE_TUNING), producing a customised model that can be served two ways: Provisioned Throughput, billed hourly like an EC2 reservation, for any customised model; or Custom Model Deployment, on demand at the base model's own token rate, restricted to specific Nova and Llama models customised after mid-2025. A separate path trains outside Bedrock, exports weights to S3, and imports them via Custom Model Import, which bills on demand per Custom Model Unit-minute and scales to zero." width="800" height="890"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Provisioned Throughput numbers, priced today
&lt;/h2&gt;

&lt;p&gt;If your model or region falls outside Custom Model Deployment's narrow eligibility, this table is what you're actually pricing. Pulled from the AWS Price List API for &lt;code&gt;us-east-1&lt;/code&gt;, effective 2026-08-01, in US dollars per model unit per hour:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;No commitment&lt;/th&gt;
&lt;th&gt;1-month commitment&lt;/th&gt;
&lt;th&gt;6-month commitment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Titan Text Lite&lt;/td&gt;
&lt;td&gt;$7.10&lt;/td&gt;
&lt;td&gt;$6.40&lt;/td&gt;
&lt;td&gt;$5.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Titan Text Express&lt;/td&gt;
&lt;td&gt;$20.50&lt;/td&gt;
&lt;td&gt;$18.40&lt;/td&gt;
&lt;td&gt;$14.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Titan Text Premier (custom)&lt;/td&gt;
&lt;td&gt;$32.25&lt;/td&gt;
&lt;td&gt;$27.95&lt;/td&gt;
&lt;td&gt;$16.38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nova Micro&lt;/td&gt;
&lt;td&gt;$60.50&lt;/td&gt;
&lt;td&gt;$55.00&lt;/td&gt;
&lt;td&gt;$30.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nova Lite&lt;/td&gt;
&lt;td&gt;$60.50&lt;/td&gt;
&lt;td&gt;$55.00&lt;/td&gt;
&lt;td&gt;$30.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nova Pro&lt;/td&gt;
&lt;td&gt;$60.50&lt;/td&gt;
&lt;td&gt;$55.00&lt;/td&gt;
&lt;td&gt;$30.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nova Canvas&lt;/td&gt;
&lt;td&gt;$60.50&lt;/td&gt;
&lt;td&gt;$55.00&lt;/td&gt;
&lt;td&gt;$30.25&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that Nova block again. On demand, Nova Micro costs a small fraction of Nova Pro per token, that's the entire point of the model family. Provisioned, they are the same price. If you're on Provisioned Throughput at all, whether because your model doesn't qualify for Custom Model Deployment or because you deliberately want reserved, latency-consistent capacity, the tier you picked to save money on inference is invisible once you reserve capacity for it: you're renting a model unit, not a token budget, and AWS prices the unit the same across the family regardless of which model is sitting on it. This is exactly why Custom Model Deployment matters when it's available: on Nova Lite, Micro or Pro, customised after 16 July 2025, in &lt;code&gt;us-east-1&lt;/code&gt;, you skip this table entirely and pay Nova's own on-demand token price instead.&lt;/p&gt;

&lt;p&gt;There's a third path that doesn't use &lt;code&gt;CreateModelCustomizationJob&lt;/code&gt; at all: &lt;strong&gt;Custom Model Import.&lt;/strong&gt; Train an open-weight model yourself, anywhere (SageMaker, your own PEFT or LoRA run, wherever), then import the resulting weights into Bedrock for a supported architecture (Llama and Mistral families among them). Bedrock serves it &lt;strong&gt;on demand&lt;/strong&gt;: billed in five-minute windows by Custom Model Unit, currently &lt;strong&gt;$0.05718 per CMU-minute&lt;/strong&gt; in &lt;code&gt;us-east-1&lt;/code&gt;, auto-scaling to zero when nothing has invoked it for five minutes. The trade-off is cold start: AWS and early adopters report anywhere from several seconds to a couple of minutes on the first request after the model has scaled to zero, depending on model size, so this isn't a free lunch against Provisioned Throughput for a genuinely latency-sensitive workload, it's a different point on the cost-versus-latency curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that actually costs, worked through
&lt;/h2&gt;

&lt;p&gt;Illustrative maths, not a measured production bill, built from the table above and the live Custom Model Import rate:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Steady, 24/7 traffic, one model unit, one month (730 hours):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Titan Text Lite, native fine-tune, no commitment: 730 × $7.10 = &lt;strong&gt;$5,183/month&lt;/strong&gt;. Six-month commitment: 730 × $5.10 = &lt;strong&gt;$3,723/month&lt;/strong&gt;, but that's a six-month lock-in, $22,338 total, whether traffic shows up or not.&lt;/li&gt;
&lt;li&gt;Custom Model Import, a small Llama-class model needing 2 CMUs, run flat out: 730 × 60 × 2 × $0.05718 = &lt;strong&gt;$5,009/month&lt;/strong&gt;. Roughly the same as the Titan no-commitment price, with zero commitment, plus whatever cold starts cost you in latency along the way.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Bursty traffic, same imported model, genuinely invoked 4 hours a day:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Custom Model Import: 4 × 30 × 60 × 2 × $0.05718 = &lt;strong&gt;$823/month&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Provisioned Throughput can't do this at all. You are renting capacity, not metering usage, so the bill is identical whether the model answers one request that day or ten thousand. Custom Model Deployment, where it's eligible, can: it's metered like the base model, not reserved like Provisioned Throughput.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the actual decision hiding inside "should we fine-tune": it is two decisions stacked as one. Do you need a customised model, and can you commit to buying dedicated, always-on capacity for it, or does your model and region qualify for one of the on-demand routes instead. Most of what's written addresses only the first.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1hjseoqbo9rnzggf7cua.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1hjseoqbo9rnzggf7cua.png" alt="Five numbered rows summarising the post: climb the ladder (prompt, RAG, caching, then fine-tune), the three real reasons to fine-tune, one API with four job types, three ways to serve the result, Provisioned Throughput as the universal but reserved default, and the two narrower on-demand routes, Custom Model Deployment and Custom Model Import." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How you'd know it worked
&lt;/h2&gt;

&lt;p&gt;Whichever job type and serving path you take, the trigger post's baseline rule is the right one, and I'd go further: don't trust your own eyeball on "it feels better," measure it. A fine-tune, a distillation, or an RFT run has to beat the best prompt on the base model on a real evaluation, the same way, every time you change it. That's not a throwaway line. I spent the last four posts on exactly how that measurement goes wrong: an LLM-as-judge score isn't a fixed property of the output you're grading, it shifts depending on what sits next to it in the batch, and a continuous quality signal doesn't come with a threshold until you derive one from your own traffic. If you're leaning on an LLM judge to decide whether the customised model earned whatever it costs to serve, the same failure modes apply to that decision as to any other regression gate.&lt;/p&gt;

&lt;p&gt;Before you commit to serving anything, you should be able to answer three questions with numbers, not confidence: what does the best prompt on the base model score, on the same eval, today; what does the customised model score; and what is the smallest sample size at which that gap is actually distinguishable from noise. If you can't answer the third one, you don't know if the first two differ at all.&lt;/p&gt;

&lt;p&gt;If you've actually shipped a fine-tune on Bedrock, which serving path did you end up on, and did Provisioned Throughput's bill, or the narrower eligibility for the on-demand routes, change the decision after the fact rather than before it?&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>finetuning</category>
      <category>genai</category>
    </item>
    <item>
      <title>Your quality alert needs 32 samples</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 08 Sep 2026 09:40:24 +0000</pubDate>
      <link>https://dev.to/rajmurugan/your-quality-alert-needs-32-samples-4l7e</link>
      <guid>https://dev.to/rajmurugan/your-quality-alert-needs-32-samples-4l7e</guid>
      <description>&lt;p&gt;I finished Part 3 with a signal I could compute on live traffic with no answer key: trigram grounding, the fraction of a summary's word-trigrams that appear in its source. It responds to a real regression. I wrote that it was a lead worth replicating, not a result, and I stand by that.&lt;/p&gt;

&lt;p&gt;Then I tried to alert on it, and found the part nobody writes down. &lt;strong&gt;A continuous signal does not come with a threshold.&lt;/strong&gt; You have to derive one, and deriving it tells you something uncomfortable about how much traffic you need before an alert means anything.&lt;/p&gt;

&lt;p&gt;For a regression that drops the signal by a third, a single output catches &lt;strong&gt;2%&lt;/strong&gt; of the time. You need a window of &lt;strong&gt;32&lt;/strong&gt; before you can page anyone.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqv4vgjgj3bj8rnw7nq3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqv4vgjgj3bj8rnw7nq3.png" alt="Dark scorecard, three tiles. One output: 2 percent of regressions caught at a one-page-a-month budget. Thirty-two outputs: 91 percent caught. Mean minus 1.5 sigma: a threshold of minus 0.0062, below the floor of the scale." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The reflex answer, and why it is wrong here
&lt;/h2&gt;

&lt;p&gt;Ask anyone where to put a threshold on a metric and you will get some version of "a couple of standard deviations below the mean". It is the right instinct and it does not survive contact with this signal.&lt;/p&gt;

&lt;p&gt;Here is the control arm from Part 3, 64 summaries generated with the guardrail intact, at the 2000-character tier, one of the two truncated tiers where the signal separates the arms:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;mean&lt;/td&gt;
&lt;td&gt;0.0830&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;median&lt;/td&gt;
&lt;td&gt;0.0746&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;standard deviation&lt;/td&gt;
&lt;td&gt;0.0594&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;minimum&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p05&lt;/td&gt;
&lt;td&gt;0.0167&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That distribution is right-skewed with a hard floor at zero, and one perfectly healthy output scores exactly zero. Now apply the reflex:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Pages on&lt;/th&gt;
&lt;th&gt;Catches&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;mean − 1.0σ&lt;/td&gt;
&lt;td&gt;+0.0235&lt;/td&gt;
&lt;td&gt;14% of normal outputs&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean − 1.5σ&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−0.0062&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean − 2.0σ&lt;/td&gt;
&lt;td&gt;−0.0359&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 1.5 sigma the threshold is a negative number, on a scale that cannot go below zero. It will never fire, on anything, ever. An alert that cannot fire looks exactly like an alert that is working.&lt;/p&gt;

&lt;p&gt;The 1.0 sigma line is worse in a way that is easier to miss: it pages you on &lt;strong&gt;14 out of every 100 normal outputs&lt;/strong&gt; to catch 20% of real ones. Healthy outputs are essentially all of your traffic, so that is a pager firing on ordinary work almost every time it fires at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deriving the line properly
&lt;/h2&gt;

&lt;p&gt;The fix is to stop assuming a shape and use the one you measured. Resample the observed control distribution, build the null distribution of the window mean at size n, and put the threshold at whatever quantile matches your false-page budget. Then resample the regression arm and see how often it lands below that line.&lt;/p&gt;

&lt;p&gt;Two budgets, because they are the two people actually argue about: &lt;strong&gt;one false page a month&lt;/strong&gt; (roughly 3% of daily checks) and &lt;strong&gt;one a quarter&lt;/strong&gt; (1%).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Window size&lt;/th&gt;
&lt;th&gt;Catches, 1 page/month&lt;/th&gt;
&lt;th&gt;Catches, 1 page/quarter&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 output&lt;/td&gt;
&lt;td&gt;2%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;td&gt;3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;29%&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;58%&lt;/td&gt;
&lt;td&gt;36%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;79%&lt;/td&gt;
&lt;td&gt;58%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;32&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;td&gt;94%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9lw0by2vtnxwhe87009q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9lw0by2vtnxwhe87009q.png" alt="Line chart, detection rate against alerting window size, two series. At a budget of one false page a month the curve runs 2 percent at one output, 29 at eight, 58 at sixteen, 91 at thirty-two, 99 at forty-eight. At one false page a quarter it runs 0, 13, 36, 78, 94. A dashed line marks the 80 percent power target. The monthly curve sits just under it at n equals 24 and clears it by n equals 32; the quarterly curve is still just under at n equals 32 and clears it by n equals 48." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On this grid the monthly budget clears 80% at &lt;strong&gt;n = 32&lt;/strong&gt;, and the quarterly one at &lt;strong&gt;n = 48&lt;/strong&gt;. A finer sweep puts the actual monthly crossing nearer &lt;strong&gt;n = 24&lt;/strong&gt;, which sits within simulation noise of the line itself. I would still build at 32: 24 is exactly on 80% and nobody designs an alert to sit on its own threshold.&lt;/p&gt;

&lt;p&gt;Read the first row again, because it is the one that changes what you build. A single output, thresholded at a budget you could actually live with, catches &lt;strong&gt;one regression in fifty&lt;/strong&gt;. Not a weak signal. Functionally no signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that means if you are wiring this up
&lt;/h2&gt;

&lt;p&gt;The instinct with a per-output score is to check it per output. Every eval harness I have seen encourages that: score the row, compare to threshold, flag the row. It is the wrong unit for this signal, and the arithmetic says so before you write any code.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Alert on a window, not a row.&lt;/strong&gt; The unit is an aggregate over n outputs, and n is a number you compute rather than pick.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Work out what n means in your traffic.&lt;/strong&gt; 32 outputs is a rounding error for a high-volume summariser and half a week for a low-volume internal tool. If it is half a week, you have not built a quality alert, you have built a weekly report, and you should call it that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose the false-page budget before the threshold.&lt;/strong&gt; The budget is a product decision about how much trust you can spend. The threshold is arithmetic that follows from it. Doing it in the other order is how you end up at 14% of normal outputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibrate the threshold on several hundred controls, not sixty-four.&lt;/strong&gt; The threshold is an
estimate too, and mine is built from 64 outputs. Simulate the real thing (draw m controls, set the
line, apply it to fresh windows) and a nominal 3% budget calibrated on 64 comes back at &lt;strong&gt;6.3% on
average and 16.8% at the unlucky end&lt;/strong&gt;. That is two to five times the pages you signed up for. At
m = 512 it settles to 3.4%. Budget the control data before you budget the pages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The budget is per check, and I have quoted it at one check a day.&lt;/strong&gt; 3% of daily checks is about
one page a month. If you evaluate every window as it closes, multiply by how often that happens: a
high-volume summariser running 300 windows a day at 3% is nine pages a day, not one a month. And
rolling windows make consecutive checks correlated, which breaks the conversion altogether.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use quantiles of the observed distribution, not mean and sigma.&lt;/strong&gt; Bounded, skewed metrics are the normal case in eval work, not the exception, and sigma-based lines on them are how you get a threshold below the floor of the scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1xybmw6bmq0kwdp6dqd8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1xybmw6bmq0kwdp6dqd8.png" alt="Alerting flow, top to bottom. In live traffic with no answer key, a Bedrock summariser in production feeds a Python trigram-overlap score computed per output from input and output alone. Those scores collect into a window of n outputs, n equals 32, computed rather than chosen. Separately, a false-page budget of one a month sets the threshold, which is a quantile of the null distribution. The window feeds the threshold as an aggregate, not a row." width="800" height="1132"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The threshold does not sit on the model output. It sits on an aggregate of them, and the size of that aggregate is the number this post is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest caveats
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The effect being detected here is large.&lt;/strong&gt; A 33.5% drop in the signal, from a deliberately deleted guardrail. Smaller regressions need bigger windows, and the table above is the optimistic end.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is the truncated-input regime, and that is load-bearing.&lt;/strong&gt; Everything above is the&lt;br&gt;
2000-character tier. At 12000 characters, which is normal full-context operation, the same deleted&lt;br&gt;
guardrail moves the signal only &lt;strong&gt;12.8%&lt;/strong&gt; instead of 33.5%, and the window arithmetic changes with&lt;br&gt;
it: n = 32 gives &lt;strong&gt;31%&lt;/strong&gt; power, n = 64 gives 49%, and 80% needs more than &lt;strong&gt;128&lt;/strong&gt;. So if your inputs&lt;br&gt;
are not already thin, the honest answer is that this signal needs a much larger window than the&lt;br&gt;
number in the title, or a different signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One signal, one corpus, one summariser family.&lt;/strong&gt; Sixteen posts written by one person, graded by one rubric. The n = 32 is mine. The method transfers; the number does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This assumes your window is stationary.&lt;/strong&gt; I resampled a control arm collected in one sitting, so it carries no daily or weekly cycle. Real traffic has both, and a window that spans a Monday and a Sunday has variance my numbers do not include. That is not something a larger n fixes: with a fixed threshold and a drifting baseline the false-page rate drifts too. It needs re-baselining on a trailing control window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And it is still built on a signal that did not survive correction.&lt;/strong&gt; Part 3 was clear that trigram grounding is a lead, not a result: raw p = 0.020, Bonferroni 0.120. This post derives a threshold for it anyway, which is the right thing to do with a lead you intend to test, and the wrong thing to treat as settled. If the signal does not replicate in your setup, the threshold arithmetic is still the transferable part.&lt;/p&gt;

&lt;p&gt;Every round, every script, and the raw JSON: &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;github.com/rajmurugan01/do-you-trust-it-evals&lt;/a&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;p&gt;This is Part 4 of &lt;strong&gt;Evals for Production AI&lt;/strong&gt;, on how you actually know an AI system is good once it is live. &lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; asked what it takes to trust an LLM judge. &lt;a href="https://rajmurugan.com/blog/regression-gate-needs-a-power-calculation" rel="noopener noreferrer"&gt;Part 2&lt;/a&gt; pointed that judge at a deploy gate and watched it go blind. &lt;a href="https://rajmurugan.com/blog/your-golden-dataset-is-too-easy" rel="noopener noreferrer"&gt;Part 3&lt;/a&gt; found the blind spot was in the dataset, and left a label-free signal worth replicating. This one works out where its line goes.&lt;/p&gt;

&lt;p&gt;If you are running a quality alert on an LLM output score, I would genuinely like to know what window size you landed on and how you chose it. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; or via &lt;a href="https://rajmurugan.com" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>evals</category>
      <category>observability</category>
    </item>
    <item>
      <title>Your golden dataset is too easy</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Mon, 31 Aug 2026 14:41:52 +0000</pubDate>
      <link>https://dev.to/rajmurugan/your-golden-dataset-is-too-easy-572f</link>
      <guid>https://dev.to/rajmurugan/your-golden-dataset-is-too-easy-572f</guid>
      <description>&lt;p&gt;I spent two posts trying to detect a regression I had planted myself, and failed three separate ways. An LLM judge over a golden dataset: p = 1.000. A judge-free deterministic assertion: p = 1.000. Six label-free signals computed on the same outputs: nothing below p = 0.17.&lt;/p&gt;

&lt;p&gt;Three instruments, one answer. At some point the honest move is to stop suspecting the instrument.&lt;/p&gt;

&lt;p&gt;The prompt change was always there, and one input tier away it produces an effect the same gate catches easily. On the corpus I had written down it produces none worth measuring. Hold every single thing constant, give the summariser less source to work with, and the gate that read p = 1.000 reads p = 0.0020.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fikq06ft7jqb4i6n7vcu1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fikq06ft7jqb4i6n7vcu1.png" alt="Dark scorecard headed 'Your golden dataset is too easy', subhead 'the regression was always there, the dataset gave it nothing to do'. Three tiles, one per source length given to the summariser. All three are clustered over the 16 posts. At 12000 characters the judge gate reads p equals 1.000, control 0 of 64 versus regression 1 of 64. At 2000 characters p equals 0.625, 3 of 64 versus 5 of 64. At 600 characters p equals 0.0020, 7 of 64 versus 23 of 64." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is following on from
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; established that the judge in this rig is trustworthy: 16/16 on clean inputs, 5/5 on single-variable corruptions, and a self-consistency check. &lt;a href="https://rajmurugan.com/blog/regression-gate-needs-a-power-calculation" rel="noopener noreferrer"&gt;Part 2&lt;/a&gt; pointed that trusted judge at a regression gate and watched it go blind, then blamed sample size.&lt;/p&gt;

&lt;p&gt;The setup has not changed. A Claude Haiku 4.5 summariser writes a two to three sentence summary of each of the 16 published posts on this site, at temperature 0.3, capped at 300 output tokens. A Claude Sonnet 4.5 judge scores each summary against its source for faithfulness at temperature 0, with a strict rubric. The regression is one thing: three guardrail sentences deleted from the summariser's system prompt, exactly what a prompt looks like after somebody tidies it up. Four repeats per arm, so 64 gradings per arm.&lt;/p&gt;

&lt;p&gt;Everything below ran against real Bedrock calls in my own account. Rounds 6 and 7 are new here, and all of it is in the &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;repo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question production actually asks
&lt;/h2&gt;

&lt;p&gt;A golden dataset gives you the one thing production never does: the right answer, written down in advance. On live traffic you have the input, you have the output, and that is the entire inventory.&lt;/p&gt;

&lt;p&gt;So before blaming the dataset I tried the other obvious thing. What can you compute from an (input, output) pair alone, with no labels anywhere?&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;novel_numbers&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;numbers in the output that are not in the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;grounding&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;fraction of the output's content words that appear in the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;trigram_grounding&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;fraction of the output's word-trigrams that appear in the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;novel_words&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;count of output content words absent from the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;novel_caps&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;capitalised entity-shaped tokens in the output, absent from the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;length&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;characters&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Round 6 applied all six to the 128 summaries rounds 4 and 5 had already produced. No new Bedrock calls: same outputs, same regression, different question asked of them.&lt;/p&gt;

&lt;p&gt;Nothing fired. Every content signal that moved at all pointed the same direction, the regression arm being consistently less grounded than the control, and not one reached significance. The best was trigram grounding at p = 0.168.&lt;/p&gt;

&lt;h2&gt;
  
  
  The detector that was detecting a markdown heading
&lt;/h2&gt;

&lt;p&gt;Before that null meant anything I had to check the instrument, and it is as well I did.&lt;/p&gt;

&lt;p&gt;The first version reported that 96 of 128 summaries contained a fabricated number or entity. Seventy-five percent, against a judge that had just failed exactly one of those same summaries. When your label-free signal and your judge disagree by that margin, the signal is wrong. It was:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;86 of the hits were the word &lt;code&gt;Summary&lt;/code&gt;.&lt;/strong&gt; The summariser likes to open with a &lt;code&gt;# Summary&lt;/code&gt; markdown heading. My entity regex saw a capitalised token absent from the source and called it a fabricated entity. I had built a hallucination detector that was mostly detecting a heading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sentence-initial words.&lt;/strong&gt; &lt;code&gt;Instead&lt;/code&gt;, &lt;code&gt;Rather&lt;/code&gt;, &lt;code&gt;Yes&lt;/code&gt;, &lt;code&gt;Key&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plurals.&lt;/strong&gt; &lt;code&gt;Macs&lt;/code&gt;, &lt;code&gt;LLMs&lt;/code&gt;, &lt;code&gt;ACLs&lt;/code&gt; flagged against sources that say Mac, LLM, ACL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Roundings.&lt;/strong&gt; One summary said &lt;code&gt;99%+&lt;/code&gt; cache hit ratios. The source says 99.8% and 99.9%. My matcher saw &lt;code&gt;99&lt;/code&gt; absent from the source and called it fabricated. It is a true statement, and conservative.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fixed: strip markdown before extraction, require an entity to actually look like one rather than merely start a sentence, stem plurals, and allow a number to be grounded if the source states something it is a faithful rounding of. The count went from 96 to 11.&lt;/p&gt;

&lt;p&gt;The broken version is committed as &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals/blob/main/scripts/unlabelled_signals_v1.py" rel="noopener noreferrer"&gt;&lt;code&gt;unlabelled_signals_v1.py&lt;/code&gt;&lt;/a&gt; so you can run it and get the 96 yourself. It was originally only a claim in a code comment, which is not good enough for a post whose whole argument is that you must read what your signal flagged.&lt;/p&gt;

&lt;p&gt;That is the part I would want someone to take from this post even if they skip the rest. A label-free signal is cheap to compute and cheap to get wrong, and there is no judge behind it to catch you. The 96 would have looked like a crisis on a dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Moving the input instead
&lt;/h2&gt;

&lt;p&gt;Three instruments, three nulls, one dataset.&lt;/p&gt;

&lt;p&gt;The 16 posts are dense, tightly-written technical writing that already contains every number a short summary would want. An anti-hallucination guardrail has nothing to suppress on an input that offers no temptation to invent. Delete it and the output barely moves, because the guardrail was not doing any work in the first place.&lt;/p&gt;

&lt;p&gt;That is testable. Hold the summariser, both system prompts, the judge, the rubric, the posts and the repeat count fixed, and move exactly one thing: how much of each source the summariser is given.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;12000 characters.&lt;/strong&gt; Rounds 4 to 6. Not the whole post: ten of the sixteen are longer than that, up to 29,171 characters, so the baseline tier is already a truncation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2000 characters.&lt;/strong&gt; Intro plus a section.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;600 characters.&lt;/strong&gt; The opening paragraph. The model is asked to summarise a post it has mostly not been shown, and must either hedge or invent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The judge sees the same truncated source the summariser saw, so nobody is scored for omitting text they were never given.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fedgl8i4fkel9rmmik9oq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fedgl8i4fkel9rmmik9oq.png" alt="Architecture of the round 7 experiment. 16 published posts, truncated to 12000, 2000 or 600 characters, feed two Claude Haiku 4.5 summariser arms: a control on the baseline prompt and a regression arm with the guardrail deleted. Their 128 summaries per source length go to two instruments in parallel. The labelled instrument adds the source answer and a Claude Sonnet 4.5 judge, and is blind until 600 characters at clustered p equals 1.000, then 0.625, then 0.0020. The label-free instrument uses input and output only, a trigram overlap with no model in it, and fires at 2000 characters on a raw p before correction. 512 Bedrock calls, published to the repo." width="800" height="346"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Only one of those two paths has a model in it. The label-free instrument is a set intersection over a log line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;

&lt;p&gt;512 new Bedrock calls, all of them for the two short tiers; the 12000 row is reused from rounds 4 to 6. Both instruments were built before this round and neither is tuned to it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source given&lt;/th&gt;
&lt;th&gt;Judge gate, counts&lt;/th&gt;
&lt;th&gt;Fisher&lt;/th&gt;
&lt;th&gt;Clustered&lt;/th&gt;
&lt;th&gt;Label-free trigram&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;12000 chars&lt;/td&gt;
&lt;td&gt;0/64 vs 1/64&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.168&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2000 chars&lt;/td&gt;
&lt;td&gt;3/64 vs 5/64&lt;/td&gt;
&lt;td&gt;0.718&lt;/td&gt;
&lt;td&gt;0.625&lt;/td&gt;
&lt;td&gt;0.020 raw&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;600 chars&lt;/td&gt;
&lt;td&gt;7/64 vs 23/64&lt;/td&gt;
&lt;td&gt;0.0015&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0020&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.019 raw&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two columns for the judge gate because they answer different questions. Fisher exact treats all 64 gradings as independent, which they are not. The clustered column is the exact sign-flip over the 16 posts. Part 2 reported Fisher and only reached for the 16 clusters in its power calculation, which was half the problem. This is the column I am quoting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The dataset was the problem.&lt;/strong&gt; The same gate, the same regression, the same judge that read p = 1.000 on full posts reads p = 0.0020 when the source is short. 23 failures out of 64 against a control of 7. It survives Holm correction across the 18 tests in that family at p = 0.033, and it survives dropping the outlier post at p = 0.0039. The prompt change does not manifest on the inputs I had chosen to write down.&lt;/p&gt;

&lt;p&gt;That is the finding. It is also the one I was least interested in when I started, which is worth noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number I wanted to headline, and why I am not
&lt;/h2&gt;

&lt;p&gt;The result I actually wanted was the second column. At 2000 characters the judge over ground truth cannot separate the arms, p = 0.625, and a signal computable on a production log line with no ground truth anywhere separates them at p = 0.020. Same outputs, same n. I had a whole post built around that sentence.&lt;/p&gt;

&lt;p&gt;It does not survive its own correction. The family is 18 tests: the five label-free hallucination signals at each of three tiers, plus the three judge gates. Length is in the analysis as a diagnostic, not as a candidate signal, so it sits outside the family. Holm over the 18 leaves the 600-character gate at p = 0.033 and takes trigram grounding at 2000 to p = 0.289. Putting length in as well, at 21 tests, gives 0.037 and 0.328, so nothing here turns on that choice. Bonferroni over just the six signals I searched still gives 0.120. A post that spent Part 2 lecturing about power calculations does not get to correct hard where it kills a result it likes and lightly where it saves one.&lt;/p&gt;

&lt;p&gt;Worse, and more instructive: &lt;strong&gt;the two signals that do survive Holm are the two I threw away.&lt;/strong&gt; &lt;code&gt;novel_words&lt;/code&gt; at 600 characters comes in at Holm-adjusted p = 0.004, and &lt;code&gt;novel_caps&lt;/code&gt; at p = 0.033. Both are confounded, and I will show why below, but the honest summary is that my statistically strongest signals are the ones I have mechanistic reasons to distrust, and my mechanistically cleanest signal does not clear correction.&lt;/p&gt;

&lt;p&gt;So the label-free result is a lead worth replicating, not a result. It is directionally right at both short tiers, 14 of 16 posts move the predicted way at 2000 characters, it is not a length artefact, and it is nowhere near significant once you account for how it was found. If you take one number from this post, take the clustered p = 0.0020 from the gate, not p = 0.020 from the signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why trigram grounding, and not the other five
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;novel_words&lt;/code&gt; is confounded by length.&lt;/strong&gt; It is a raw count, and at 600 characters the unguarded model writes 31% longer summaries than the control, 538 characters against 411, p = 0.0006. More words, more novel words. Normalised by content-word count it becomes &lt;code&gt;grounding&lt;/code&gt;, which reads p = 0.060 at 600 and p = 0.115 at 2000. But normalised per 1000 output characters instead, it reads p = 0.010 at 600. Two defensible normalisations, two different answers, which is itself a warning about how much freedom you have when you choose a rate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;novel_caps&lt;/code&gt; fires at p = 0.002 and is measuring at least two different things.&lt;/strong&gt; I read the unguarded arm's full hit list at 600 characters: 40 hits, and every one is an acronym or an inference. &lt;code&gt;AWS&lt;/code&gt; twelve times, &lt;code&gt;AI&lt;/code&gt; six, &lt;code&gt;LLM&lt;/code&gt;/&lt;code&gt;LLMs&lt;/code&gt; six, &lt;code&gt;API&lt;/code&gt;/&lt;code&gt;APIs&lt;/code&gt; six, &lt;code&gt;ARM&lt;/code&gt; four. On &lt;code&gt;llm-is-not-a-security-boundary&lt;/code&gt; the source says "language model" and the summary says "LLM": abbreviation, not fabrication. On &lt;code&gt;part-4-local-dev-docker&lt;/code&gt; the model added "ARM" to a post about the &lt;code&gt;--platform linux/amd64&lt;/code&gt; flag on a Mac, which is correct, useful, and genuinely not in the source. That second kind is interesting rather than wrong, because reaching for outside knowledge is exactly what the deleted guardrail existed to suppress. My favourite is &lt;code&gt;AM&lt;/code&gt;, which is the regex catching the clock in "cryptic CloudFormation errors at 2 AM". There is no outright fabrication anywhere in that list, and I am not going to claim one I cannot point at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;novel_numbers&lt;/code&gt; never separates the arms.&lt;/strong&gt; It is exactly zero at 12000 characters. At 2000 it flags four bare digits per arm, p = 1.000, and at 600 it flags eighteen across both arms in the wrong direction, control 11 against regression 7, p = 0.500. The most intuitive hallucination signal, the one everybody reaches for first, has nothing in it at any source length.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;trigram_grounding&lt;/code&gt; is the cleanest of them.&lt;/strong&gt; It is the fraction of the summary's own word-trigrams that appear in the source, so it is a rate rather than a count. A rate can still track length, so I checked: at 600 characters the two correlate at r = -0.15 pooled and +0.06 inside the control arm. At 2000 characters it moves from 0.083 to 0.055 and 14 of 16 posts move in the predicted direction, which is p = 0.004 on the sign test alone. Two posts reverse, &lt;code&gt;prompt-caching-bedrock-strands&lt;/code&gt; at +0.076 and &lt;code&gt;agentcore-memory-read-after-write&lt;/code&gt; at +0.002, and I have no account of the first one.&lt;/p&gt;

&lt;p&gt;The likely mechanism: a prompt that forbids adding content also, in practice, pushes the model to reuse the source's wording. The baseline prompt never mentions phrasing, so this is a side effect of the content constraint rather than the thing it asks for. Trigram overlap measures exactly that, and unigram overlap is too coarse to see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I now do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Treat the golden dataset as a hypothesis about where failure lives, and test it.&lt;/strong&gt; This is the finding that survived. If your eval corpus is the tidy end of your traffic, a real regression can sit at p = 1.000 in CI and p = 0.0020 one input-difficulty tier away. Truncating your own inputs is a crude but cheap way to find out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate a label-free signal before trusting it.&lt;/strong&gt; Print what it flagged and read the list. Mine would have reported a 75% hallucination rate that was mostly a markdown heading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Correct for the search.&lt;/strong&gt; If you try six signals, the one that fires needs the correction, and you have to be willing to publish it after the correction rather than before.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefer rates to counts, and then check the other rate.&lt;/strong&gt; Every count I tried was confounded by output length. Two reasonable normalisations of the same count disagreed at p = 0.060 and p = 0.010.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Label-free to compute is not label-free to calibrate.&lt;/strong&gt; I only know trigram grounding responds to this regression because I ran a controlled comparison against a control arm, which is exactly what you cannot do on live traffic. Deploying it still needs a baseline and a threshold.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The honest caveats
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Truncation is a proxy for "traffic you did not anticipate", not the thing itself.&lt;/strong&gt; It holds domain, topic, style and vocabulary constant and moves only the density of grounding material, which is the cleanest single variable available without leaving the corpus. It does not simulate a novel domain, an adversarial user, or a shifted register.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 300-token output cap is doing some of the work at 600 characters, and it is all one post.&lt;/strong&gt; Four unguarded outputs at that tier ran all the way into the cap, up to 1531 characters, two of them cut off mid-sentence, against a control maximum of 505. All four are the same post, &lt;code&gt;year-10-study-system-production-ai-failure-modes&lt;/code&gt;, all four of its repeats, and all four were judged FAIL. So part of that headline 23 is one post's instruction-following blowout rather than hallucination. Dropping that post entirely leaves 15 clusters and the gate still reads p = 0.0039, with trigram grounding at p = 0.036.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;At 600 characters the control is failing too, 7 times in 64.&lt;/strong&gt; The guarded system degrades at that tier as well, which is worth knowing before you adopt truncation as a diagnostic: a tier where your control also falls over tells you less about the regression than one where it holds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"14 of 16 posts" is a 2000-character fact and does not replicate at 600.&lt;/strong&gt; At 600 it is 9 posts in the predicted direction, 6 against, 1 exactly level, which is chance. The 600 p-value comes from the size of a few large moves, not from agreement across posts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;At the 12000 tier the judge saw more of the post than the summariser did.&lt;/strong&gt; &lt;code&gt;summarize()&lt;/code&gt; sends &lt;code&gt;body[:12000]&lt;/code&gt; and &lt;code&gt;judge()&lt;/code&gt; sends &lt;code&gt;source[:14000]&lt;/code&gt;, so on the eight posts longer than 14000 characters the judge held up to 2000 characters the summariser never got. Round 7's 2000 and 600 tiers pass the already-truncated body to both, so they are clean. The bias runs toward the top row's null, not against it, but the three rows are not protocol-identical and I would rather say so.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cross-tier comparison spans an inference-profile change.&lt;/strong&gt; Rounds 4 to 6 ran on &lt;code&gt;us.&lt;/code&gt; inference profiles and round 7 on &lt;code&gt;global.&lt;/code&gt; ones, same model version strings, different routing, because I switched for the cost saving between sittings. Both arms within any one tier always ran on the same profile, so the control-versus-regression comparisons that carry every finding are unaffected. AWS documents routing, monitoring and price differences between profiles and makes no output-equivalence guarantee in either direction, so I would not lean on the trend down the table as if the tiers were perfectly comparable. Worth flagging if you copy the cost saving: &lt;code&gt;global.&lt;/code&gt; routes outside the US geography, which matters in a regulated shop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One corpus, one summariser family, one judge, one rubric, one regression.&lt;/strong&gt; Sixteen posts written by one person. The mechanism is general enough to be worth checking in your setup. The numbers are mine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 64 are 4 repeats over 16 posts, not 64 independent observations.&lt;/strong&gt; Every label-free p-value above is clustered: an exact sign-flip test enumerating all 2^16 assignments over the posts. The judge gate is reported both ways in the results table, because Fisher exact on the counts is the number Part 2 used and the clustered one is the number I am standing behind. Clustering is not automatically the more conservative choice: across the 15 label-free tests it is stricter in 8, looser in 4 and identical in 3. It is stricter in every test that came near significance, which is the case that matters.&lt;/p&gt;

&lt;p&gt;All seven rounds, every script including the broken one, and the raw JSON are in the repo: &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;github.com/rajmurugan01/do-you-trust-it-evals&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhgnmf7wjjb15va8sji5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwhgnmf7wjjb15va8sji5.png" alt="Numbered recap. Row 01, three nulls: judge, deterministic assertion and six label-free signals, all on the same 16 posts, nothing below p = 0.168, marked blind. Row 02, the dataset: same gate, shorter source at 600 characters, 23 of 64 fail against 7 of 64, p = 1.000 becomes p = 0.0020, marked too easy. Row 03, the correction: six signals searched, Bonferroni takes the label-free result from 0.020 to 0.120, a lead not a result, marked replicate it." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;p&gt;This is Part 3 of &lt;strong&gt;Evals for Production AI&lt;/strong&gt;, a series on how you actually know an AI system is good once it is live, not according to the dashboard. &lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; asked what it takes to trust an LLM judge. &lt;a href="https://rajmurugan.com/blog/regression-gate-needs-a-power-calculation" rel="noopener noreferrer"&gt;Part 2&lt;/a&gt; pointed that judge at a deploy gate and watched it go blind. This one found the blind spot was in the dataset, and killed my preferred explanation on the way.&lt;/p&gt;

&lt;p&gt;If you are running an eval corpus you suspect is the tidy end of your traffic, happy to compare notes. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; or via &lt;a href="https://rajmurugan.com" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next in &lt;em&gt;Evals for Production AI&lt;/em&gt;:&lt;/strong&gt; a continuous signal with no labels behind it does not come with a threshold. Where you set the line, what a false page costs, and how to tell a real drop from a Tuesday.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>evals</category>
      <category>genai</category>
    </item>
    <item>
      <title>Your regression gate needs a power calculation</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 25 Aug 2026 01:48:13 +0000</pubDate>
      <link>https://dev.to/rajmurugan/your-regression-gate-needs-a-power-calculation-4j8h</link>
      <guid>https://dev.to/rajmurugan/your-regression-gate-needs-a-power-calculation-4j8h</guid>
      <description>&lt;p&gt;I deleted the anti-hallucination guardrail from my summariser's system prompt on purpose, to see whether my eval would notice.&lt;/p&gt;

&lt;p&gt;It did not. Then I removed the LLM judge from the eval entirely and replaced it with a deterministic string assertion, on the theory that the judge was the weak link. That did not notice either. Two measurement approaches, one with a model in the loop and one without, both returning the same answer: no detectable difference, Fisher exact p = 1.000.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypam9wlya2sncd37jghz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypam9wlya2sncd37jghz.png" alt="Dark scorecard, three tiles. With an LLM judge: p equals 1.000, control 0 of 64 versus regression 1 of 64. With no judge, a deterministic assertion: p equals 1.000, control 1 of 64 versus regression 1 of 64. Minimum detectable effect: 16 to 43 percent, exact, against a 1.6 percent baseline, needing a 10 times to 27 times rise." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The interesting part is not that the gate failed. It is that I had a confident, wrong explanation ready to publish, and the second experiment is the only reason I did not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is following on from
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; tested whether an LLM judge can be trusted at all. Five deliberately corrupted summaries, 5 out of 5 caught and correctly named. Four repeat gradings of an unmodified input, identical scores every time. The judge discriminates and it is stable.&lt;/p&gt;

&lt;p&gt;A regression gate is a different job. It runs in CI, nobody injected an error, nobody knows what changed, and its whole purpose is to notice that quality dropped before your users do. I assumed that a judge which clears Part 1's three rounds is fit for that job. This post is what happened when I checked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The regression
&lt;/h2&gt;

&lt;p&gt;Sixteen published posts from this blog, summarised by Claude Haiku 4.5 on Bedrock, graded by Claude Sonnet 4.5 against a strict faithfulness rubric at temperature 0. The judge never changes anywhere in this experiment. It is the instrument.&lt;/p&gt;

&lt;p&gt;The break is one variable: three sentences deleted from the summariser's system prompt, the task sentence untouched.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Baseline
&lt;/span&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You summarize technical blog posts in 2-3 sentences for a reader deciding whether to &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;click through. Base the summary ONLY on the text provided. Do not add claims, numbers, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;product names, or examples that are not explicitly present in the source text. If the &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;post does not state a specific number or outcome, do not invent one.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Regression: the guardrail is gone, the task is identical
&lt;/span&gt;&lt;span class="n"&gt;DEGRADED_SYSTEM_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You summarize technical blog posts in 2-3 sentences for a reader deciding whether to &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;click through.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is what a prompt looks like after somebody tidies it up, and it is the change a reviewer waves through.&lt;/p&gt;

&lt;p&gt;Alongside it, two controls. A &lt;strong&gt;control arm&lt;/strong&gt; runs the identical model and the identical prompt, byte for byte, changing nothing. The summariser runs at temperature 0.3, so the baseline is not deterministic and I need to know how much it moves on its own. And a &lt;strong&gt;positive control&lt;/strong&gt;: the same prompt with a much weaker summariser, Amazon Nova Micro, which AWS describes as its "fastest text-only model, optimized for speed and low cost in tasks like summarization, translation, and classification", at roughly a thirtieth of Haiku 4.5's input price. Worth noting AWS names summarisation first, so this is not a model being used outside its stated purpose, it is a model built for speed and cost on short work being asked for faithfulness on long documents. That arm exists to show the gate detects anything at all.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filaf0jijjvx5x2w83ek0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filaf0jijjvx5x2w83ek0.png" alt="Architecture: sixteen source posts fan into three summariser arms on Amazon Bedrock, a control and a prompt regression both on Claude Haiku 4.5 across four runs each, and a model regression on Nova Micro across one run, then all three into a single unchanged Claude Sonnet 4.5 judge at temperature 0. The control returns 0 fail of 64 and the prompt regression 1 fail of 64, marked indistinguishable, while the model regression returns 5 fail of 16, marked detected. All three feed a box asking whether this can gate a deploy." width="800" height="402"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt one: the judge over the golden dataset
&lt;/h2&gt;

&lt;p&gt;One run of each arm looked like a clean gradient.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Mean faithfulness&lt;/th&gt;
&lt;th&gt;Items moved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline (Part 1)&lt;/td&gt;
&lt;td&gt;16/16 PASS&lt;/td&gt;
&lt;td&gt;5.00&lt;/td&gt;
&lt;td&gt;reference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Control&lt;/td&gt;
&lt;td&gt;16/16 PASS&lt;/td&gt;
&lt;td&gt;5.00&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt regression&lt;/td&gt;
&lt;td&gt;15/16 PASS&lt;/td&gt;
&lt;td&gt;4.94&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model regression&lt;/td&gt;
&lt;td&gt;11/16 PASS&lt;/td&gt;
&lt;td&gt;4.19&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing changed, nothing moved. Subtle change, small movement. Blunt change, big movement. The gate works, just coarsely.&lt;/p&gt;

&lt;p&gt;I nearly wrote that post. Part 1 had already taught me why not: its most interesting first-run finding did not survive a controlled re-run and was cut before publishing. One item moving, once, out of sixteen is the same shape. So I ran both arms three more times each.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm&lt;/th&gt;
&lt;th&gt;Failures per run&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;th&gt;Mean faithfulness per run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Control&lt;/td&gt;
&lt;td&gt;0, 0, 0, 0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 / 64&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5.00, 5.00, 5.00, 4.94&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt regression&lt;/td&gt;
&lt;td&gt;1, 0, 0, 0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1 / 64&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4.94, 5.00, 5.00, 5.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Fisher exact on 0/64 against 1/64 gives &lt;strong&gt;p = 1.000&lt;/strong&gt;. The gradient did not survive repetition. There was one event.&lt;/p&gt;

&lt;p&gt;That one event is real, and worth looking at, because it is the only direct evidence in this whole post that the guardrail does anything. The judge failed a summary containing "caused real production issues". That phrase is not in the source, which says the author "burned an afternoon". The judge's reasoning named it exactly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The summary accurately captures the main structure and three gotchas, but 'caused real production issues' overstates the source, which mentions 'burned an afternoon' and a separate OIDC gotcha in the blog's own pipeline, not that these three gotchas caused production issues.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A textbook scope-broadening hallucination, correctly caught and correctly explained, of exactly the category the deleted sentences existed to prevent. Once, in 64 gradings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt two: delete the judge
&lt;/h2&gt;

&lt;p&gt;Here is the explanation I was ready to publish. Faithfulness was pinned at 5.00 on every baseline item, so the metric had no headroom. The verdict is a binary step function over a continuous quality change. The judge cannot see small drops.&lt;/p&gt;

&lt;p&gt;It is a tidy story and it is testable, so I tested it. If the judge is the bottleneck, an eval with no judge in it should do better.&lt;/p&gt;

&lt;p&gt;So the trap moves to the input, and the assertion becomes deterministic. Each source post already states things at low intensity ("burned an afternoon"). A guarded summariser is told not to escalate; an unguarded one is free to. I defined a vocabulary of escalations ("production outage", "caused an outage", "data loss", "in every case", "guaranteed", "catastrophic", and similar), then filtered it per post to the phrases &lt;strong&gt;verified absent from that post's own source text&lt;/strong&gt;, so a hit can only ever be something the model introduced. Then a substring check. No judge, no rubric, no scores, no ceiling.&lt;/p&gt;

&lt;p&gt;One caveat worth stating rather than letting a reader find it: this vocabulary is not independent of the first experiment. It contains "production issues", the phrase the judge flagged in attempt one. I built the second gate partly from what the first one saw, which makes this a weaker replication than two genuinely independent designs would be.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;th&gt;Prompt regression&lt;/th&gt;
&lt;th&gt;Fisher exact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LLM judge over golden dataset&lt;/td&gt;
&lt;td&gt;0 / 64&lt;/td&gt;
&lt;td&gt;1 / 64&lt;/td&gt;
&lt;td&gt;p = 1.000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deterministic assertion, no judge&lt;/td&gt;
&lt;td&gt;1 / 64&lt;/td&gt;
&lt;td&gt;1 / 64&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;p = 1.000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The control tripped on "critical failure". The regression arm tripped on "production incident". One each.&lt;/p&gt;

&lt;p&gt;Removing the judge changed nothing. My explanation did not survive it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually wrong
&lt;/h2&gt;

&lt;p&gt;Removing the judge changed nothing, so "the judge cannot see small drops" does not stand on its own as the explanation. That is weaker than saying the judge is exonerated, and the weaker version is what the data supports: both gates are blunt, and two blunt instruments returning the same null cannot tell you which one is blunt. What the next number shows is that the sample size alone is enough to explain both nulls, whatever the instruments were doing.&lt;/p&gt;

&lt;p&gt;Both designs were asked to separate two rates that are, at everything measured here, roughly 1.6% each. So the real question is what size difference this design could have found at all.&lt;/p&gt;

&lt;p&gt;With 64 gradings per arm at 80% power, the answer is &lt;strong&gt;16.1%&lt;/strong&gt;. The rate would have to rise more than &lt;strong&gt;tenfold&lt;/strong&gt; before either gate reliably noticed.&lt;/p&gt;

&lt;p&gt;That has to be computed exactly rather than with the usual two-proportion normal approximation, which wants around five expected events per arm. At a 1.6% baseline and n = 64 there is &lt;strong&gt;one&lt;/strong&gt; expected event. Out of range, not borderline. I ran the normal approximation anyway and first published 15.4%, and the giveaway I ignored is that the standard methods disagree with each other by nearly a factor of two at these rates. The number above enumerates every outcome pair under two binomials and applies the same two-sided Fisher test used everywhere else here. It is &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals/blob/main/scripts/power.py" rel="noopener noreferrer"&gt;twelve lines, in the repo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;It is worse than 64, too. Those 64 gradings are 4 repeats over the same 16 posts, so repeating them buys precision on each post's score rather than more posts. Counting each post once, the same calculation gives &lt;strong&gt;42.6%&lt;/strong&gt;, a 27-fold rise. So the honest answer is a range: between about 16% and about 43%, depending on how correlated repeat gradings of one post are. Evan Miller is explicit in &lt;a href="https://arxiv.org/abs/2411.00640" rel="noopener noreferrer"&gt;Adding Error Bars to Evals&lt;/a&gt; (arXiv 2411.00640) that this sits on a sliding scale, from "perfectly correlated (in which case each cluster acts as a single independent observation)" to perfectly uncorrelated. His second recommendation covers "[w]hen questions are drawn in related groups, computing clustered standard errors" and his third covers reducing variance by resampling answers, which is the one that actually bites here. His fifth is the one I skipped:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Using power analysis to determine whether an eval (or a random subsample) is capable of testing a hypothesis of interest&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the calculation I did not do. It runs in under a second and it answers the question both experiments spent real money failing to answer. An eval that cannot resolve the effect you care about does not return "no regression". It returns nothing, in a format that looks exactly like "no regression".&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gate does catch
&lt;/h2&gt;

&lt;p&gt;The positive control lands hard. Swapping to Nova Micro moved seven items and dropped mean faithfulness from 5.00 to 4.19. The distortions are the interesting part, because they are not invented facts. On the prompt-caching post the summary reported "up to 99.9% hit ratios, proving the cost benefits of enabling caching". The 99.9% is real: it is that post's own headline measurement. What the weaker model added was "proving the cost benefits", welding a cache hit ratio onto a billing claim the source measures separately at 55% and 78%. The judge caught the conflation, not the number.&lt;/p&gt;

&lt;p&gt;So these gates are not broken. They are &lt;strong&gt;coarse&lt;/strong&gt;, and the boundary is computable in advance rather than discoverable in production. Ten times the hallucination rate: caught. A change you would actually ship on a Tuesday: invisible.&lt;/p&gt;

&lt;h3&gt;
  
  
  The row that neither passed nor failed
&lt;/h3&gt;

&lt;p&gt;Sixteen items in that arm. Eleven passed, four failed. That leaves one.&lt;/p&gt;

&lt;p&gt;One grading came back as well-formed JSON with &lt;code&gt;faithfulness: 4&lt;/code&gt;, &lt;code&gt;completeness: 3&lt;/code&gt;, a populated &lt;code&gt;hallucinations&lt;/code&gt; list and a sentence of correct reasoning. It had no &lt;code&gt;verdict&lt;/code&gt; key. The judge answered every part of the question except the one the gate reads. By the rubric (PASS requires faithfulness 4 or higher &lt;strong&gt;and&lt;/strong&gt; an empty hallucination list) a populated list makes it a FAIL, so the arm's true count is five non-passing out of sixteen, which is what the charts show.&lt;/p&gt;

&lt;p&gt;It happened once, in the sixteen gradings that arm got, and not in the 128 Haiku gradings. That arm was never re-run, so this is not evidence that it is rare: conditional on the arm where it appeared, one in sixteen is all I can say. Now consider a run of 15 PASS, 0 FAIL, and one of these:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;                    &lt;span class="c1"&gt;# 15 != 16, red. Catches it.
&lt;/span&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;                      &lt;span class="c1"&gt;# 0 == 0, green. Misses it.
&lt;/span&gt;&lt;span class="n"&gt;pass_rate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;failed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# 15/15 = 100%, green. Misses it.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only the first survives. The other two treat a malformed response as if the item did not exist. The third is the version people write, because computing a rate from the two buckets you have feels more careful than counting, and it is the one that silently drops the row from the denominator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this sits against what is already written
&lt;/h2&gt;

&lt;p&gt;"Your eval gate might be noisy" is not my discovery. The variance problem is named in places, and Miller's paper is the rigorous treatment of it.&lt;/p&gt;

&lt;p&gt;What I could not find is anyone publishing the control arm. The standard practitioner write-up, for instance &lt;a href="https://www.traceloop.com/blog/automated-prompt-regression-testing-with-llm-as-a-judge-and-ci-cd" rel="noopener noreferrer"&gt;Traceloop's guide to prompt regression testing with LLM-as-a-judge in CI/CD&lt;/a&gt;, recommends running new versions against a curated dataset, scoring with a judge, and failing the build when quality declines. Reasonable on its face, and it says nothing about how far that score moves when nothing has changed, which is the number that decides whether any of it works.&lt;/p&gt;

&lt;p&gt;So the delta is not the idea. It is the measurement: a re-run-with-nothing-changed arm, the same regression measured 64 times by two unrelated methods, the raw JSON published, and a power calculation attached to the result rather than left implicit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The suite I was about to recommend, which would not have worked
&lt;/h2&gt;

&lt;p&gt;I had a fix section written. It said: keep a suite of known-bad inputs, each asserting a specific failure the judge must still catch, and if somebody deletes the guardrail the fabricated-number case stops failing.&lt;/p&gt;

&lt;p&gt;Then I read my own code. Part 1's corruption suite hard-codes its five corrupted summaries as string literals and only ever calls the &lt;strong&gt;judge&lt;/strong&gt;. The summariser is never invoked. Deleting the guardrail from &lt;code&gt;SYSTEM_PROMPT&lt;/code&gt; cannot change that suite's inputs, its outputs, or its 5 out of 5 result. It is a judge-regression test. It guards the instrument, not the thing being measured, and it has nothing to say about the regression this entire post is about.&lt;/p&gt;

&lt;p&gt;I would have shipped that as the takeaway. It was caught by an independent review pass with no knowledge of how the post was built, which is the only reason it is in this section instead of the conclusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I now do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compute the minimum detectable effect before running the eval, not after.&lt;/strong&gt; If you cannot state what size regression your gate would catch, you do not have a gate, you have a ritual. Mine could catch a tenfold increase and nothing smaller, and one calculation would have told me that on day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run the control arm.&lt;/strong&gt; Re-run with nothing changed and measure how far it moves on its own. Without that number, every difference you find is unattributed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count clusters, not rows.&lt;/strong&gt; Four repeats of sixteen items is sixteen independent units. Repeating a small dataset raises confidence in the mean, not the sample size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assert on the count, never on a rate derived from buckets.&lt;/strong&gt; &lt;code&gt;passed == total&lt;/code&gt; catches a malformed response. &lt;code&gt;failures == 0&lt;/code&gt; and &lt;code&gt;passed / (passed + failed)&lt;/code&gt; both quietly drop it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check that your test exercises the thing you changed.&lt;/strong&gt; A suite can be rigorous, pass cleanly, and be pointed at a different component entirely.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The honest caveats
&lt;/h2&gt;

&lt;p&gt;The load-bearing limit is the one this whole post is about: &lt;strong&gt;I cannot tell you how much worse the guardrail deletion makes the summariser.&lt;/strong&gt; I can show it produced at least one real hallucination, verified absent from its source. I cannot show it happens more often than with the guardrail in place, because neither design has the resolution to distinguish 1.6% from 1.6%. There is a reading of this data where the true post-deletion rate genuinely is about 1.6%, the gates are correctly reporting a tiny effect, and "underpowered" rather than "blind" is the right word. That reading is consistent with every number here, and the practical conclusion does not change: compute the effect you need to detect before you trust the answer.&lt;/p&gt;

&lt;p&gt;Sixteen posts, one judge model, one rubric, one summariser family, four repeats per arm. I am not claiming a general result about LLM-as-judge regression detection. The mechanism is general enough to be worth checking in your setup; the numbers are mine.&lt;/p&gt;

&lt;p&gt;The positive control is weaker evidence than the other two arms, deliberately. My intended swap was to an older model in the same family, which would have been a clean single variable. Nova Micro crosses model families, so tokenizer, training and instruction-following all move at once. It is there to show the gate detects something, not to attribute what.&lt;/p&gt;

&lt;p&gt;Part 1's runs were on 2026-08-18 and these on 2026-08-24: six days apart, same judge model, same rubric, same account, not the same sitting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model that went away mid-experiment
&lt;/h2&gt;

&lt;p&gt;The within-family swap was blocked, for a reason worth stating accurately, because my first instinct was to write it up as "the provider retired a model out from under me with no warning" and that is not what happened.&lt;/p&gt;

&lt;p&gt;Bedrock's &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html" rel="noopener noreferrer"&gt;model lifecycle&lt;/a&gt; has three states: Active, Legacy, End-of-Life. Claude 3 Haiku moved to Legacy on 10 March 2026 with a published EOL of 10 September 2026, sixteen days after this post goes up. Six months of notice on a public page. No ambush.&lt;/p&gt;

&lt;p&gt;What actually bit me is narrower. Access lapses on an &lt;strong&gt;inactivity timer&lt;/strong&gt;: new customers cannot use a Legacy model at all, and existing customers "may lose access" after a period of not calling it. The exact period is worth flagging because the sources disagree. The documentation says access may be lost "after 15 days of inactivity". The error Bedrock returned to me says 30.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ResourceNotFoundException: Access denied. This Model is marked by provider as
Legacy and you have not been actively using the model in the last 30 days.
Please upgrade to an active model on Amazon Bedrock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I am quoting both rather than picking one, because I do not know which governs and I did not test the boundary.&lt;/p&gt;

&lt;p&gt;The documented model-access flow is to re-accept the agreement with &lt;code&gt;aws bedrock list-foundation-model-agreement-offers&lt;/code&gt; and &lt;code&gt;aws bedrock create-foundation-model-agreement&lt;/code&gt;, or through model access in the console. AWS documents those as the model-access steps; it does not connect them to Legacy inactivity specifically, and I did not test whether they recover a lapsed Legacy model. To check where you stand before any of this matters, &lt;code&gt;aws bedrock get-foundation-model-availability --model-id &amp;lt;id&amp;gt;&lt;/code&gt; returns an &lt;code&gt;agreementAvailability&lt;/code&gt; field. There is also a phase worth knowing about: since 10 June 2026 Claude 3 Haiku has been in &lt;strong&gt;public extended access&lt;/strong&gt;, during which the docs say to expect higher, provider-set pricing. In this case that has not bitten. Bedrock's published extended-access price table currently lists only Claude 3.5 Sonnet and 3.5 Sonnet v2, not Claude 3 Haiku.&lt;/p&gt;

&lt;p&gt;So the operational lesson is not that providers retire models without warning. It is that a benchmark baseline you only run occasionally is exactly the workload an inactivity timer catches, that reactivating may cost more than it did, and that after the EOL date it stops being recoverable at all. Pin the recorded outputs rather than assuming you can regenerate them.&lt;/p&gt;

&lt;p&gt;All five rounds, every script, and the raw JSON for every run and every arm are in the repo: &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;github.com/rajmurugan01/do-you-trust-it-evals&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7wqlld8uhjfsn5zyjr1t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7wqlld8uhjfsn5zyjr1t.png" alt="Numbered recap under the banner: removing the judge changed nothing. Row 01, the control: nothing changed across 16 items and 4 runs, detection floor 0 of 64, marked baseline. Row 02, both gates: guardrail deleted, judge then no judge, both p equals 1.000, marked no difference. Row 03, the calculation: clusters not rows, 16 of them, power before you run, needs a 10 times to 27 times rise, marked compute it." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;p&gt;This is Part 2 of &lt;strong&gt;Evals for Production AI&lt;/strong&gt;, a series on how you actually know an AI system is good once it's live, not according to the dashboard. &lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; asked what it takes to trust an LLM judge's score. This one pointed that trusted judge at the job most teams want it for, and found the instrument was not what decided the answer.&lt;/p&gt;

&lt;p&gt;If you are wiring an eval into a deploy gate and want a second pair of eyes on where its detection floor actually sits, happy to compare notes. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; or via &lt;a href="https://rajmurugan.com" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next in &lt;em&gt;Evals for Production AI&lt;/em&gt;:&lt;/strong&gt; an offline eval can only ever test the inputs you thought to write down. What live production signals catch that a golden dataset structurally cannot, and how to spot a quality drop in traffic you never anticipated.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>evals</category>
      <category>genai</category>
    </item>
    <item>
      <title>What a Year 10 study system taught me about production AI failure modes</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 18 Aug 2026 02:26:58 +0000</pubDate>
      <link>https://dev.to/rajmurugan/what-a-year-10-study-system-taught-me-about-production-ai-failure-modes-1mfa</link>
      <guid>https://dev.to/rajmurugan/what-a-year-10-study-system-taught-me-about-production-ai-failure-modes-1mfa</guid>
      <description>&lt;p&gt;I spend my professional time shipping production AI agents on AWS for enterprise clients. Bedrock AgentCore, Strands, the integration layer between LLMs and the systems that actually matter. The work that does not show up in demo videos.&lt;/p&gt;

&lt;p&gt;A few weeks ago I built an AI system for my Year 10 son. Not on AWS. Anthropic's Claude projects and Cowork routines, because the data lives in Google Workspace and the user is a teenager who needs a frictionless experience. Different stack, same architectural disciplines.&lt;/p&gt;

&lt;p&gt;It took three iterations and one architecture pivot inside the third before it worked. What I did not expect was how cleanly the failure modes mapped to patterns I see in production Bedrock work. The five lessons below are the ones I will be taking straight back into client engagements.&lt;/p&gt;

&lt;h2&gt;
  
  
  The build, briefly
&lt;/h2&gt;

&lt;p&gt;Three subject-specific Claude projects with Socratic tutoring. A weekly Cowork routine that aggregates the week and produces three differentiated emails: a full dossier to me, a focused brief to each tutor, a forward-looking plan for him.&lt;/p&gt;

&lt;p&gt;Cowork is Anthropic's scheduled-routine service for Claude. Think of it as cron for Claude projects, with built-in connectors to Gmail, Drive, and Calendar.&lt;/p&gt;

&lt;p&gt;The architecture had to do two things: read what happened in the projects during the week, and produce structured output across multiple channels on a schedule. Sounds simple. Most agentic systems do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three iterations, one architecture pivot
&lt;/h2&gt;

&lt;p&gt;The version above is v3. There were two earlier shapes I had to rule out before I got there, and inside v3 there was a further pivot at the integration layer. The full diagram is on the canonical post — &lt;a href="https://rajmurugan.com/blog/year-10-study-system-production-ai-failure-modes" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt; — here's the prose walkthrough.&lt;/p&gt;

&lt;h2&gt;
  
  
  v1: the answer machine
&lt;/h2&gt;

&lt;p&gt;The first build was the obvious one. Ask Claude a question, get an answer. It worked on day one and failed the same week. He was not learning; he was copying. The pedagogical intent — that the system make him think — was nowhere in the design. I rebuilt.&lt;/p&gt;

&lt;h2&gt;
  
  
  v2: Socratic tutor, per subject
&lt;/h2&gt;

&lt;p&gt;v2 made the model refuse to give direct answers. One Claude project per subject, system prompts tuned to ask back rather than tell, hint at the next step rather than skip to the end. That worked. He had to actually do the maths. His tutors started commenting that the homework conversations were sharper.&lt;/p&gt;

&lt;p&gt;The constraint we hit next was visibility. Three subjects across a week is a lot of context for a parent and three different tutors to absorb. We needed a layer on top that aggregated the week and split it for different audiences. That extension is what became v3.&lt;/p&gt;

&lt;h2&gt;
  
  
  v3's first architecture failed at the integration layer
&lt;/h2&gt;

&lt;p&gt;My first design for the v3 aggregation layer had the Sunday routine reading project chats directly and updating a rolling Google Doc. First real run: neither side worked. Cowork has no API surface to read project chats on the same account. The Drive connector is read-only for content. Two load-bearing assumptions, both wrong.&lt;/p&gt;

&lt;p&gt;I have seen this exact pattern in Bedrock work. A team designs an agent that "queries the knowledge base, then updates the ticket." First deployment surfaces the truth: the knowledge base query returns chunks the agent cannot reason over, or the ticket system's API has a write surface that does not match the read surface. The system worked on paper because nobody tested the primitives in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture that works flips the data flow
&lt;/h2&gt;

&lt;p&gt;The redesign inverted both directions. Each subject project now drafts a session-summary email at end-of-session with a fixed subject prefix. I review the draft and click Send. The Sunday routine searches Gmail by subject prefix instead of reading project chats. The routine creates a new dated doc each week instead of modifying an existing one.&lt;/p&gt;

&lt;p&gt;Same intent. Different primitives. Push beats pull, and append beats mutate.&lt;/p&gt;

&lt;p&gt;This is the same shift I keep recommending in Bedrock production work (see &lt;a href="https://rajmurugan.com/blog/part-3-strands-agent-sdk" rel="noopener noreferrer"&gt;Part 3 of the AgentCore series&lt;/a&gt; where I walk through the same pattern with Strands tool calls). When an agent needs to "look inside" a system that cannot expose its state cleanly, the answer is almost always to make the producer emit, not to make the consumer introspect. Pull architectures have one failure mode for every integration. Push architectures have one failure mode total: the producer doesn't emit. That is debuggable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five lessons that map directly to production AWS AI
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Verify the primitives before you design around them.&lt;/strong&gt; v3's first design cost me an evening because I drew a diagram before I tested whether Cowork could read project chats. Most failed Bedrock POCs I see fail for the same reason at a bigger scale: the team designs around assumed Knowledge Base behaviour, assumed AgentCore session state, assumed Strands tool-call semantics, without the kind of isolated &lt;a href="https://rajmurugan.com/blog/part-2-cdk-infrastructure-bedrock-agentcore" rel="noopener noreferrer"&gt;primitive testing&lt;/a&gt; I'd insist on at the start of any CDK deployment. Save the diagram for after the smoke test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. When introspection isn't possible, push beats pull.&lt;/strong&gt; This is the single most useful pattern I have learned this year. Half the production Bedrock work I do involves rearranging data flow from pull to push because the producer can be modified and the consumer cannot. If your agent needs to read state from a system that doesn't expose it cleanly, the right move is almost never a better retrieval strategy. It's an emit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Human-in-the-loop is a feature, not a bug.&lt;/strong&gt; v3 has me clicking Send on two emails per cycle. Ten seconds of human attention, full audit trail, kill switch on every outbound message. I argue for this on every enterprise engagement and lose half the time, because someone wants the autonomy metric. The teams that ship the human-in-the-loop version tend to still be running their agents six months later. The teams that skip it learn the value of an audit trail the hard way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Calibration windows beat go-live confidence.&lt;/strong&gt; Three Sunday runs on real data before flipping to multi-recipient. Same discipline I apply to Bedrock agents before they touch a customer-facing channel. Skipping calibration is the single biggest predictor of an enterprise AI rollback I have seen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The "things we got wrong" doc is more useful than the spec.&lt;/strong&gt; I now write a decision log on every client engagement, with the versions-that-didn't-work explicitly preserved. Six months later, when someone joins the project and asks "why is it built this way," that doc is the answer. The polished spec is for review committees. The decision log is for the team.&lt;/p&gt;

&lt;p&gt;There is a sixth lesson — about the pedagogical move from v1 to v2, and what it taught me about scoping AI agents around &lt;em&gt;intent&lt;/em&gt; rather than &lt;em&gt;output&lt;/em&gt;. That one belongs in a separate post on when to graduate from Claude projects to AgentCore. It's in draft.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I'm posting this on an AWS-focused blog
&lt;/h2&gt;

&lt;p&gt;Because the lessons are stack-agnostic. AgentCore, Strands, Bedrock Knowledge Bases, Lambda-backed tools, plain old Claude projects with Cowork: the architectural disciplines are the same. The model layer is interchangeable. The integration layer is where production AI lives or dies. That's true on AWS, that's true on Anthropic's stack, and that's true on whatever comes next.&lt;/p&gt;

&lt;p&gt;Side builds like this are how I sharpen patterns I then apply at scale to enterprise AWS engagements. Enterprise engagements take longer to write up because they have to be anonymised. Side builds let me publish the pattern faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repo
&lt;/h2&gt;

&lt;p&gt;Full design history, the v1 and v2 specs, the v3 architecture that failed, the v3 architecture that works, the prompts, the build plan, and the calibration log are open-source: &lt;a href="https://github.com/rajmurugan01/study-loop" rel="noopener noreferrer"&gt;github.com/rajmurugan01/study-loop&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you ship production AI on AWS and any of these patterns ring true, I would like to compare notes. Find me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;, or comment below.&lt;/p&gt;

&lt;p&gt;More posts on production AWS AI: &lt;a href="https://rajmurugan.com/blog" rel="noopener noreferrer"&gt;browse the blog&lt;/a&gt; or &lt;a href="https://rajmurugan.com/rss.xml" rel="noopener noreferrer"&gt;subscribe by RSS&lt;/a&gt;. The next post in this thread, on when to graduate from Claude projects to AgentCore, is in draft.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>agentcore</category>
      <category>aiagents</category>
    </item>
    <item>
      <title>Field Notes: Turning prompt caching on for a production Bedrock workload</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 18 Aug 2026 02:26:54 +0000</pubDate>
      <link>https://dev.to/rajmurugan/field-notes-turning-prompt-caching-on-for-a-production-bedrock-workload-2l8b</link>
      <guid>https://dev.to/rajmurugan/field-notes-turning-prompt-caching-on-for-a-production-bedrock-workload-2l8b</guid>
      <description>&lt;p&gt;Two kwargs in Strands' &lt;code&gt;BedrockModel&lt;/code&gt; cut a Bedrock workload's system-prefix billing by 78%. The Strands tutorial doesn't mention them. Almost none of the production Strands code I've audited this year has them set.&lt;/p&gt;

&lt;p&gt;This is Part 2 of &lt;a href="https://rajmurugan.com/blog/three-things-bedrock-workload/" rel="noopener noreferrer"&gt;AI Operations Services&lt;/a&gt; — the deep dive on prompt caching that Part 1 promised. Short, specific, entirely evidence-led: two kwargs to enable, one per-model gotcha that took half a day to find, one measurement technique that gives you the answer in seconds without waiting for CloudWatch to aggregate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fubss1jrwuvmrr1uan96k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fubss1jrwuvmrr1uan96k.png" alt="Per-turn billing pattern measured on Amazon Nova Pro and Anthropic Sonnet 4.6 across 10 spaced turns against the real PENNY_SYSTEM_PROMPT system prefix (8,156 tokens on Nova, 8,788 on Sonnet). Both models reach a stable read-only steady state after the first call; Nova needs two writes before the cache propagates, Sonnet hits on the first read." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The default is None
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands.models.bedrock&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BedrockModel&lt;/span&gt;

&lt;span class="nc"&gt;BedrockModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu.amazon.nova-pro-v1:0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the example shape you find in most Strands tutorials and the agentic-AI content on Bedrock. It is also the shape that produces a Bedrock call with no &lt;code&gt;cachePoint&lt;/code&gt; block, no system-prefix caching, no tool-registry caching, and a full-input bill on every turn.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;BedrockModel&lt;/code&gt; exposes two kwargs that turn caching on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;BedrockModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu.amazon.nova-pro-v1:0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cache_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# caches the system prompt
&lt;/span&gt;    &lt;span class="n"&gt;cache_tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="c1"&gt;# caches the tool registry
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both default to &lt;code&gt;None&lt;/code&gt;. The Strands docs mention the kwargs in the API reference but not in the getting-started flow, so they are easy to miss on first build and easy to forget on the second. Every workload I have walked into this year had them unset. The fix is two kwargs and a &lt;code&gt;boto3&lt;/code&gt; version bump if your environment is more than a quarter behind.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the two kwargs actually do
&lt;/h2&gt;

&lt;p&gt;Under the hood, Bedrock's &lt;code&gt;converse&lt;/code&gt; and &lt;code&gt;converseStream&lt;/code&gt; accept a &lt;code&gt;cachePoint&lt;/code&gt; block at specific positions in the request body. The block tells Bedrock "cache the prefix up to this marker, and on a subsequent call with the same prefix, bill it as a cache read instead of a full input."&lt;/p&gt;

&lt;p&gt;&lt;code&gt;cache_prompt="default"&lt;/code&gt; inserts a &lt;code&gt;cachePoint&lt;/code&gt; at the end of the &lt;code&gt;system&lt;/code&gt; block, so the entire system prompt becomes a cacheable prefix. &lt;code&gt;cache_tools="default"&lt;/code&gt; inserts a &lt;code&gt;cachePoint&lt;/code&gt; inside &lt;code&gt;toolConfig.tools&lt;/code&gt;, so the tool registry becomes the next cacheable prefix after the system block. Both points compose: a call with both set caches &lt;code&gt;system + tools&lt;/code&gt; together, which is the right thing to want when both are large and stable.&lt;/p&gt;

&lt;p&gt;The TTL is five minutes from the most recent cache write or read on a given prefix. Inside the TTL, subsequent calls bill the prefix tokens at the cache-read rate. Outside the TTL, the next call pays a fresh cache write and the meter restarts.&lt;/p&gt;




&lt;h2&gt;
  
  
  The per-model gotcha
&lt;/h2&gt;

&lt;p&gt;This is the thing that cost me half a day, because it is not in the Bedrock docs and the error message points at the request shape rather than the underlying constraint:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Malformed input request: extraneous key [cachePoint] is not permitted&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The error appeared on every Nova call after I set both &lt;code&gt;cache_prompt&lt;/code&gt; and &lt;code&gt;cache_tools&lt;/code&gt;, but only when both were set. Sonnet 4.6 took the same config without complaint. The difference is per-model: Bedrock's server-side validator gates &lt;code&gt;cachePoint&lt;/code&gt; placement per model family, not per feature.&lt;/p&gt;

&lt;p&gt;What works on each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model family&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;cachePoint&lt;/code&gt; in &lt;code&gt;system&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;cachePoint&lt;/code&gt; in &lt;code&gt;toolConfig.tools&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova (Pro, Lite, Micro)&lt;/td&gt;
&lt;td&gt;accepted&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;rejected server-side&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic (Sonnet, Haiku, Opus)&lt;/td&gt;
&lt;td&gt;accepted&lt;/td&gt;
&lt;td&gt;accepted&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern that survives a fallback chain with both families in it is one config-of-config: pass &lt;code&gt;cache_prompt="default"&lt;/code&gt; to every model in the chain, and pass &lt;code&gt;cache_tools="default"&lt;/code&gt; only to the Anthropic-family models. If you do not split it, the Nova path fails on every call with the malformed-input error and the SDK retry loop swallows the failures into the fallback chain. From the dashboard, the symptom is "Nova is throttling at 100%" with no further clue. From the agent logs, the symptom is the actual error string above.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands.models.bedrock&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BedrockModel&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;make_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;BedrockModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;is_anthropic&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;anthropic.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;BedrockModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cache_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;cache_tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_anthropic&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the smallest piece of code that handles both families correctly. Half a day saved.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to verify, in seconds
&lt;/h2&gt;

&lt;p&gt;The default reflex for verifying a Bedrock change is to wait for CloudWatch metrics to aggregate, then look at &lt;code&gt;cacheReadInputTokenCount&lt;/code&gt; per &lt;code&gt;ModelId&lt;/code&gt; over a 15-minute window. That works, but it is the slow path. The fast path is the per-call &lt;code&gt;usage&lt;/code&gt; block returned inline by &lt;code&gt;bedrock-runtime.converse(...)&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-runtime&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;converse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;modelId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eu.amazon.nova-pro-v1:0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cachePoint&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}}],&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ping&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;response["usage"]&lt;/code&gt; carries four keys you want: &lt;code&gt;inputTokens&lt;/code&gt;, &lt;code&gt;outputTokens&lt;/code&gt;, &lt;code&gt;cacheReadInputTokens&lt;/code&gt;, &lt;code&gt;cacheWriteInputTokens&lt;/code&gt;. On call 1 against a freshly-seeded prefix, &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; is large and &lt;code&gt;cacheReadInputTokens&lt;/code&gt; is zero. On call 2 against the same prefix inside the TTL, &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; drops to zero and &lt;code&gt;cacheReadInputTokens&lt;/code&gt; is large.&lt;/p&gt;

&lt;p&gt;If you instead see &lt;code&gt;cacheWriteInputTokens: 0&lt;/code&gt; and &lt;code&gt;cacheReadInputTokens: 0&lt;/code&gt; on every call, your config did not take effect: the &lt;code&gt;cachePoint&lt;/code&gt; block is missing from the request, or the SDK version is too old to emit it, or the prefix is too short to be cacheable (Bedrock has a per-model minimum).&lt;/p&gt;

&lt;p&gt;The per-call &lt;code&gt;usage&lt;/code&gt; block is the right measurement primitive because it is exact, immediate, per-turn, and free. No CloudWatch lag, no metric aggregation, no dashboard to build. Three calls and you know.&lt;/p&gt;




&lt;h2&gt;
  
  
  The five-second propagation lag
&lt;/h2&gt;

&lt;p&gt;A subtlety the docs do not flag: Bedrock takes a few seconds to make a freshly-written cache entry available for reads. Fire two calls inside a second against the same prefix and the second one will pay a full cache write rather than a cheap read. The lag I measured on Nova and Sonnet in &lt;code&gt;eu-central-1&lt;/code&gt; was around five seconds; six seconds between calls is enough to clear it.&lt;/p&gt;

&lt;p&gt;This matters for two reasons. First, when you measure caching with a tight loop, you will conclude caching does not work, because turn 2 of your driver will still be a write. Use spaced calls or accept that your measurement run wastes the first call or two on writes. Second, in production, the lag means a burst of three calls in the first second of a user turn pays one write plus two writes, not one write plus two reads, on a fresh prefix. After the first burst, every subsequent call inside the TTL is a read.&lt;/p&gt;




&lt;h2&gt;
  
  
  The measured results
&lt;/h2&gt;

&lt;p&gt;Methodology: real production system prompt (8,156 tokens on the Nova tokeniser and 8,788 on the Anthropic tokeniser, varying by tokeniser not by content), 10-turn driver against the workload's staging Bedrock account, 6-second intra-call spacing, both Nova Pro and Sonnet 4.6 in the same run. &lt;code&gt;cache_prompt="default"&lt;/code&gt; set on every model, &lt;code&gt;cache_tools="default"&lt;/code&gt; set on Sonnet only.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Turns&lt;/th&gt;
&lt;th&gt;Hit ratio&lt;/th&gt;
&lt;th&gt;System-prefix billing reduction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova Pro&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic Sonnet 4.6&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99.8%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The asymmetry is pricing-driven. Anthropic publishes its cache-read multiplier directly: 10% of input price per the &lt;a href="https://platform.claude.com/docs/en/docs/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic pricing page&lt;/a&gt; (retrieved 2026-06-26). At 99.8% hit ratio on the system-prefix tokens, that is 78% off the full-input bill. Nova's cache-read multiplier is on the &lt;a href="https://aws.amazon.com/bedrock/pricing/" rel="noopener noreferrer"&gt;AWS Bedrock pricing page&lt;/a&gt;; plug your contracted rate against the measured 99.9% hit ratio to compute your own reduction.&lt;/p&gt;

&lt;p&gt;Two caveats. The hit ratios are measured on a 10-turn driver in a single run, not aggregated across days of production traffic. In production, calls drift in and out of the TTL window depending on user-burst patterns, so the steady-state hit ratio will be lower than 99.9%. The number to track post-deploy is the per-day ratio of &lt;code&gt;cacheReadInputTokens&lt;/code&gt; to &lt;code&gt;cacheReadInputTokens + cacheWriteInputTokens + non-cached inputTokens&lt;/code&gt; per &lt;code&gt;ModelId&lt;/code&gt;. Second: prompt caching only helps the prefix tokens, not the per-turn user message or output.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I now do on every Strands &lt;code&gt;BedrockModel&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Three things, in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Set &lt;code&gt;cache_prompt="default"&lt;/code&gt; on every model in the fallback chain.&lt;/li&gt;
&lt;li&gt;Set &lt;code&gt;cache_tools="default"&lt;/code&gt; on the Anthropic-family models only.&lt;/li&gt;
&lt;li&gt;After deploy, fire three calls against the prefix with five to six seconds between them and print &lt;code&gt;response["usage"]&lt;/code&gt;. Confirm call 1 has a non-zero &lt;code&gt;cacheWriteInputTokens&lt;/code&gt; and calls 2 and 3 have a non-zero &lt;code&gt;cacheReadInputTokens&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then add &lt;code&gt;CacheReadInputTokenCount&lt;/code&gt; per &lt;code&gt;ModelId&lt;/code&gt; to the workload's CloudWatch dashboard for the steady-state ratio. The dashboard is not how you verify the deploy, it is how you spot regressions: a hit ratio that drifts down over time is usually a sign that the system prompt is being mutated per request and the cache is being invalidated on every call.&lt;/p&gt;




&lt;h2&gt;
  
  
  Series
&lt;/h2&gt;

&lt;p&gt;Part 1: &lt;a href="https://rajmurugan.com/blog/three-things-bedrock-workload/" rel="noopener noreferrer"&gt;Three things I learned diagnosing a production Bedrock workload&lt;/a&gt; — load tests can lie, latency isn't always model speed, prompt caching is almost never on.&lt;/p&gt;

&lt;p&gt;Part 2: this post, the deep dive on the caching kwargs and the per-model gotcha.&lt;/p&gt;

&lt;p&gt;Part 3 (coming): attributing a mixed change — the latency improvement on the engagement above shipped a model and region swap together; how much was each.&lt;/p&gt;

&lt;p&gt;If you are running Strands with both Nova and Anthropic in your fallback chain, &lt;strong&gt;have you hit the Nova &lt;code&gt;toolConfig.tools&lt;/code&gt; rejection?&lt;/strong&gt; Curious whether anyone has solved it differently than splitting &lt;code&gt;cache_tools&lt;/code&gt; per model family. Drop a note in the comments or DM me on &lt;a href="https://www.linkedin.com/in/muruganraj/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>agentcore</category>
      <category>finops</category>
    </item>
    <item>
      <title>A clean pass rate is not calibration</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 18 Aug 2026 02:23:42 +0000</pubDate>
      <link>https://dev.to/rajmurugan/a-clean-pass-rate-is-not-calibration-5gh9</link>
      <guid>https://dev.to/rajmurugan/a-clean-pass-rate-is-not-calibration-5gh9</guid>
      <description>&lt;p&gt;Sixteen out of sixteen. Every summary my eval graded came back faithful, first try. That number should make you suspicious of the eval, not proud of the system, and it's the reason this post exists: a same-day Bedrock experiment on my own blog, built to find out what it actually takes to trust an LLM-as-judge score before you wire one into anything that matters.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj7dmv2sg9t8h9s4igaoq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj7dmv2sg9t8h9s4igaoq.png" alt="Dark scorecard: Round 1 baseline 16/16 summaries passed, Round 2 calibration 5/5 single-variable injected errors caught, Round 3 self-consistency 4/4 identical re-grades matched." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Sixteen posts, two Bedrock models in my own AWS account, one small repo: &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;do-you-trust-it-evals&lt;/a&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Haiku 4.5&lt;/strong&gt; (&lt;code&gt;us.anthropic.claude-haiku-4-5-20251001-v1:0&lt;/code&gt;) writes a 2-3 sentence summary of each of my 16 published posts, instructed to use only claims present in the source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Sonnet 4.5&lt;/strong&gt; (&lt;code&gt;us.anthropic.claude-sonnet-4-5-20250929-v1:0&lt;/code&gt;) grades each summary against its source: faithfulness 1-5, completeness 1-5, a list of unsupported claims, and a PASS/FAIL verdict. The rubric treats a conditional claim stated as a universal ("sometimes" becoming "always") as a hallucination, not just an invented fact.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Round 1: the baseline that should worry you
&lt;/h2&gt;

&lt;p&gt;Cold run, no tuning: 16 out of 16 summaries passed, faithfulness 5 across the board. I didn't take the judge's word for it. The summary for the CDK post claims "nine specific pitfalls", and the source has exactly nine, numbered &lt;code&gt;## Gotcha #1&lt;/code&gt; through &lt;code&gt;## Gotcha #9&lt;/code&gt;. The summaries were genuinely faithful, not a rubber stamp catching nothing because there was nothing to catch.&lt;/p&gt;

&lt;p&gt;A 100% pass rate proves the eval didn't break on the easy case. It proves nothing about whether the judge would catch a hard one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round 2: calibrate with a single variable
&lt;/h2&gt;

&lt;p&gt;Five corrupted summaries, each with exactly one injected error and nothing else touched, so a FAIL verdict can only be explained by that one change: a fabricated number, a conditional claim broadened to a universal, a real number misattributed to the wrong post, a fabricated named entity, an inflated count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5 out of 5 caught&lt;/strong&gt;, and in every case the judge's own hallucination list named the exact injected error. If you can't point to the one thing you changed in a corrupted test case, you don't have a calibration result, you have a guess with a percentage attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round 3: does the judge agree with itself
&lt;/h2&gt;

&lt;p&gt;Fiona Lau's &lt;a href="https://arxiv.org/abs/2603.04417" rel="noopener noreferrer"&gt;"Same Input, Different Scores"&lt;/a&gt; (2026) found substantial LLM-judge score variability even at temperature 0, with completeness scoring showing the largest fluctuations. So I re-ran two inputs, an easy clean one and a hard borderline one, four times each, tracking both faithfulness and completeness. Both held steady across all eight runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I now do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Treat a 100% pass rate as an instruction to calibrate, not a result to report.&lt;/li&gt;
&lt;li&gt;Corrupt one variable per test case, and diff it against the original to check.&lt;/li&gt;
&lt;li&gt;Cover more than one failure category: a fabricated fact, a scope-broadened claim, and a misattributed-but-real number all fail differently.&lt;/li&gt;
&lt;li&gt;Re-grade the same unmodified input more than once before trusting a single run.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full methodology, all three scripts, and the raw JSON for every round: &lt;a href="https://github.com/rajmurugan01/do-you-trust-it-evals" rel="noopener noreferrer"&gt;github.com/rajmurugan01/do-you-trust-it-evals&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is Part 1 of &lt;em&gt;Do You Trust It?&lt;/em&gt;, a series on how you actually know an AI system is good once it's live. Full write-up with the honest caveats (small n, benign-only corruptions, what this doesn't test) on &lt;a href="https://rajmurugan.com/blog/clean-pass-rate-is-not-calibration" rel="noopener noreferrer"&gt;rajmurugan.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>evals</category>
      <category>genai</category>
    </item>
    <item>
      <title>Your LLM security diagram defends the wrong layer</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Tue, 21 Jul 2026 09:42:01 +0000</pubDate>
      <link>https://dev.to/rajmurugan/your-llm-security-diagram-defends-the-wrong-layer-1346</link>
      <guid>https://dev.to/rajmurugan/your-llm-security-diagram-defends-the-wrong-layer-1346</guid>
      <description>&lt;p&gt;Search "LLM security architecture" and you will meet the same diagram again and again. A tidy left-to-right pipeline, user input to retrieval to the model to output to tools, and hanging underneath it four red boxes: prompt injection, retrieval poisoning, context poisoning, output injection. It is a good diagram. I have drawn versions of it myself. And if you build your defences the way it is laid out, you will spend a year patching the wrong layer.&lt;/p&gt;

&lt;p&gt;I argued in the last post that &lt;a href="https://rajmurugan.com/blog/llm-is-not-a-security-boundary" rel="noopener noreferrer"&gt;the LLM is not a security boundary&lt;/a&gt;, that the controls which actually hold are deterministic and live in code the model never touches. This is the same idea from the other side. The four-box diagram is not wrong about the threats. It is wrong about where the defence goes, and the error is baked so far into the layout that it is hard to see.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx84qcwsi46a7w86k7auc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx84qcwsi46a7w86k7auc.png" alt="Dark infographic titled Four attack names, one wrong reflex. Across the top, four red boxes: prompt injection, retrieval poisoning, context poisoning, output injection, each with a dashed amber probabilistic patch beneath it. Below them a single solid blue bar labelled the deterministic boundary the diagram omits, with the sensitive data safe underneath." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What the diagram gets right
&lt;/h3&gt;

&lt;p&gt;The four attacks are real, and the diagram names them well. Prompt injection: an attacker hides an instruction in text your model reads and steers it. Retrieval poisoning: a malicious document lands in the corpus and gets pulled into context. Context poisoning: retrieved data carries an embedded instruction. Output injection: an unvalidated model response gets executed as a command downstream. Every one of these has put a real system on an incident call. Naming them is useful.&lt;/p&gt;

&lt;p&gt;The diagram is also comprehensive-feeling, which is most of its appeal. It walks the pipeline stage by stage and hangs a threat under each stage, so it looks like a complete accounting. Nothing is missing. That completeness is exactly what makes the next step feel obvious, and the next step is the trap.&lt;/p&gt;

&lt;h3&gt;
  
  
  The trap is in the layout
&lt;/h3&gt;

&lt;p&gt;Read the diagram as a to-do list and it hands you one mitigation per box. Worse, the obvious mitigation for each box sits at the same stage as the threat, which is the stage the attacker controls.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt injection sits at the input, so you reach for input scanning and prompt hardening. The attacker writes the input.&lt;/li&gt;
&lt;li&gt;Retrieval poisoning sits at the corpus, so you reach for document scanning. The attacker writes the document.&lt;/li&gt;
&lt;li&gt;Context poisoning sits at the model's reading of the context, so you reach for a grounding or relevance score. You are now scoring meaning.&lt;/li&gt;
&lt;li&gt;Output injection sits at the output, so you reach for an output filter. That filter is a classifier.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every one of those is a probabilistic classifier you tune against an adversary. Thresholds, false negatives, a curve you push toward zero and never reach. Build all four and you have not built a boundary. You have built four leaky nets stacked on top of each other and called the result secure. A determined injection that scores just under every threshold walks the whole length of the pipeline untouched.&lt;/p&gt;

&lt;p&gt;And there is a tell that the list itself is the problem: it grows. Tool poisoning, memory injection, confused-deputy attacks across agents. Next quarter there is a fifth box, and a defence organised as one-patch-per-named-attack is permanently a step behind the naming. It is an open set sold to you as a closed one.&lt;/p&gt;

&lt;h3&gt;
  
  
  The layer the diagram leaves out
&lt;/h3&gt;

&lt;p&gt;Here is the fact the layout hides. All four attacks are harmless right up until model output crosses into a consequence: a tool call, a database query, a retrieval, or an answer leaving the building. That crossing is a single surface. It is finite, it is enumerable, and it is the one layer the attacker does not control, because it is your code.&lt;/p&gt;

&lt;p&gt;The reflex the diagram trains looks like this, and it is quietly futile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;injection_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;        &lt;span class="c1"&gt;# tuned threshold
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Blocked&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;malice_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# tuned threshold
&lt;/span&gt;    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;output_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;               &lt;span class="c1"&gt;# tuned threshold
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Blocked&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three gates, three thresholds, and an injection tuned to score 0.89 everywhere sails through all of them. Now the boundary version, at the layer the diagram skips:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Model output is a proposal, never an instruction. Every proposal is
# checked against the finite set of things allowed to happen.
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;act_on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TOOLS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;          &lt;span class="c1"&gt;# enumerated capability, not a classifier
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;NotAllowed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;check_authz&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;principal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;# token-derived principal, per call
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;TOOLS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="n"&gt;proposal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And retrieval filtered before the model, keyed off the verified principal, so a chunk the principal is not cleared to see is never loaded in the first place:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;principal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;entitlements&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Be precise about what this filter does and does not do. It is a confidentiality control, not an anti-poisoning one: it stops chunks outside the principal's entitlements, but a poisoned document that sits inside their authorised scope carries a valid ACL and passes. That residue is caught downstream, at the action boundary and by the grounding backstop, not here. What the filter does buy you, and it is the reason the fourth threat surface (an answer leaving the building) never needed its own allow-list, is that the model can only ever repeat data the principal was already cleared to retrieve. The answer is bounded on the way in, not policed on the way out.&lt;/p&gt;

&lt;p&gt;None of this asks what the attack was called. Prompt injection, context poisoning, some technique that does not have a name yet: they all arrive at the same door, and the door checks the action against a list, not the intent against a classifier. The model can be fooled into proposing anything. It cannot be fooled into a proposal that the door has no entry for.&lt;/p&gt;

&lt;h3&gt;
  
  
  Enumerate the actions, not the attacks
&lt;/h3&gt;

&lt;p&gt;That is the whole reframe. You cannot enumerate the ways a model can be fooled. That set is open, adversarial, and growing while you read this. You can enumerate the actions your system is allowed to take: the tools in the registry, the tables in the grant, the entitlements on the index, the destinations on the egress allow-list. That set is small, closed, and yours to write down and review.&lt;/p&gt;

&lt;p&gt;This is why the deterministic boundary is load-bearing and the classifiers are not. One defends a finite set you control. The other chases an infinite set the attacker controls. When people say defence in depth, they usually mean "more layers." The layer that matters is the one where an infinite problem becomes a finite one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where the four boxes still earn their place
&lt;/h3&gt;

&lt;p&gt;I am not telling you to throw the diagram out, and I want to be precise about the limits of my own argument, because the reframe oversells if you let it.&lt;/p&gt;

&lt;p&gt;The boundary stops security failures, not correctness ones. If a poisoned document convinces the model to give a wrong but fully authorised answer, no allow-list catches that. The action is within the user's rights. It is just wrong. Grounding checks, provenance, and citation-of-source are the right tools there, and those are exactly the probabilistic layer I just spent five paragraphs demoting. Demoted, not deleted. They move from load-bearing to backstop, which is where they belong.&lt;/p&gt;

&lt;p&gt;The boundary also constrains which actions and whose, not what rides inside them. An allowed tool called with attacker-shaped arguments is still an allowed tool: talk the model into calling a permitted &lt;code&gt;send_report&lt;/code&gt; with an exfiltrating recipient, and a registry check that only asks "is this tool allowed" waves it through. This is why the action boundary is not just the tool list. It is the tool list, plus the grant, plus what each action is allowed to carry and where it is allowed to send it. Enumerate the arguments and the destinations too, not only the verbs. The egress allow-list is doing as much work as the tool registry.&lt;/p&gt;

&lt;p&gt;Input and output scanning also earn a place. They raise the cost of the low-effort attacks and, more usefully, they give you signal to detect the attempt. Keep them. Just never let a tuned classifier be the only thing standing between the model and the data.&lt;/p&gt;

&lt;p&gt;So the correction is not "the diagram is wrong." The threats are real and the diagram names them cleanly. The correction is "stop reading a threat map as a defence architecture." Put the deterministic boundary in first. Then let the classifiers backstop it, labelled honestly as backstops.&lt;/p&gt;

&lt;p&gt;The companion post walks through where each of these lives in code. Here is the shape of the boundary itself, control by control:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5evcm031yv1hy5a4i8vi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5evcm031yv1hy5a4i8vi.png" alt="Dark infographic, six numbered rows, each a control in the deterministic boundary. 01 Propose never dispose: prompt to LLM to proposed action, not yet executed. 02 Tool allow-list: proposed call to allow-list check, named tool passes, anything else rejected. 03 Read-only SQL plus grant: query to read-only grant, SELECT only, writes rejected. 04 Per-call authz: each call checked individually, not cached once at session start. 05 ACL pre-filter: all rows to ACL filter, permitted rows only reach the model. 06 Probabilistic backstop: model output to content filter, response, boundary already held upstream." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What I now do
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Before drawing a single mitigation, list the actions the system can take: tools, tables, entitlements, egress destinations. That list, not the attack list, is the security surface.&lt;/li&gt;
&lt;li&gt;Put a deterministic check on each action, at the point of the action, keyed off a verified principal, never off anything the model produced.&lt;/li&gt;
&lt;li&gt;Treat every input scanner, grounding check, and output filter as a backstop, and give it a name that says so in the design doc. Budget for it and tune it. Do not let it hold the line alone.&lt;/li&gt;
&lt;li&gt;When a new attack name starts trending, ask one question before building anything: does my action boundary already stop it? Most of the time the answer is yes, and the correct amount of new work is none.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The four-box diagram will keep circulating, because it is a genuinely good map of where the water comes in. Just remember that a map of the leaks is not a plan for the wall. The wall goes lower than the diagram draws it, at the layer the attacker cannot reach, and it is made of boring deterministic code that does not care what the flood is called.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part 6, and the close, of **Production AI, Honestly&lt;/em&gt;&lt;em&gt;: the series on the AWS AI work between the demo and something you'd trust a customer behind. It runs from measuring where the probabilistic layer lies (cost, caching, memory) to building the boundary that holds (&lt;a href="https://rajmurugan.com/blog/llm-is-not-a-security-boundary" rel="noopener noreferrer"&gt;the LLM is not a security boundary&lt;/a&gt; walks those controls in detail) to this, the reframe of the diagram that had everyone patching the wrong layer. If you are building agents over sensitive data and want a second set of eyes on where your real boundary sits, &lt;a href="https://rajmurugan.com/contact" rel="noopener noreferrer"&gt;get in touch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Field Notes: The AgentCore Memory write that returns success and reads back empty</title>
      <dc:creator>Raj Murugan</dc:creator>
      <pubDate>Mon, 13 Jul 2026 04:58:59 +0000</pubDate>
      <link>https://dev.to/rajmurugan/field-notes-the-agentcore-memory-write-that-returns-success-and-stores-nothing-ng8</link>
      <guid>https://dev.to/rajmurugan/field-notes-the-agentcore-memory-write-that-returns-success-and-stores-nothing-ng8</guid>
      <description>&lt;p&gt;I wired long-term memory into an agent on Amazon Bedrock AgentCore, wrote a record, got a &lt;code&gt;201&lt;/code&gt;, and read back nothing. The record existed. The API said success. The read returned zero rows. It took me longer than I would like to admit to work out that all three of those were true at the same time.&lt;/p&gt;

&lt;p&gt;AgentCore went GA on 13 October 2025, after a July preview. Memory is one of its newer pieces, and the docs are good on the happy path and quiet on the parts that bite. This is the operational truth of the write side, the bit you only learn by running it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbr2kyyrv5y21azvjj2ch.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbr2kyyrv5y21azvjj2ch.png" alt="BatchCreateMemoryRecords returns 201 Created while a namespace read returns zero records for fifteen to thirty seconds, across three measured runs." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two things bit me. One was an API that did not exist. The other was a write that succeeds and is not yet readable. Neither is in the tutorial.&lt;/p&gt;

&lt;h2&gt;
  
  
  The API that never existed
&lt;/h2&gt;

&lt;p&gt;The codebase had a helper for persisting a fact to memory. It called &lt;code&gt;client.ingest_memory_records(...)&lt;/code&gt;. It read as correct. It had a docstring. It had a sensible name.&lt;/p&gt;

&lt;p&gt;It had never run. It was written, never called, and so had never thrown. When I finally wired it into a real path, I checked the client first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;
&lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-agentcore&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;hasattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ingest_memory_records&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# -&amp;gt; False
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no &lt;code&gt;ingest_memory_records&lt;/code&gt; operation on the AgentCore data plane. The helper would have raised &lt;code&gt;AttributeError&lt;/code&gt; the first time anyone called it. Dead code that mirrors a real-sounding API is worse than no code, because it passes the eye test. A method name is not a fact. It is a claim, and an unrun claim is a guess wearing a fact's clothes.&lt;/p&gt;

&lt;p&gt;The real operation is &lt;code&gt;BatchCreateMemoryRecords&lt;/code&gt;. Confirm it against the SDK you actually ship, not against what the name suggests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ops&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service_model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;operation_names&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ops&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Memory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Record&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="c1"&gt;# BatchCreateMemoryRecords, BatchDeleteMemoryRecords, BatchUpdateMemoryRecords,
# DeleteMemoryRecord, GetMemoryRecord, ListMemoryRecords, RetrieveMemoryRecords
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lesson one, before any of the memory detail: verify the operation exists in your pinned SDK version. The write API and the read API are not symmetric in name, and one of the plausible names is a trap.&lt;/p&gt;

&lt;p&gt;Better than checking one call by hand, make the build check every call. The service model is the ground truth for what exists, so a CI test can fail the build the moment source names an operation that does not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# CI: fail the build if the codebase calls a bedrock-agentcore method that
# does not exist in the pinned SDK. Kills the ingest_memory_records class,
# including the next hallucinated API someone commits.
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-agentcore&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;called&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;batch_create_memory_records&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;get_memory_record&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retrieve_memory_records&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ingest_memory_records&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;  &lt;span class="c1"&gt;# grep these from source
&lt;/span&gt;&lt;span class="n"&gt;missing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;called&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;hasattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no such AgentCore operation: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;   &lt;span class="c1"&gt;# -&amp;gt; ['ingest_memory_records']
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Two tiers, and which one you are writing to
&lt;/h2&gt;

&lt;p&gt;AgentCore Memory has two tiers, and the whole confusion comes from not knowing which one a given call touches.&lt;/p&gt;

&lt;p&gt;Short-term memory is raw events. You write them with &lt;code&gt;CreateEvent&lt;/code&gt;, one per turn or in batches, scoped to an actor and a session. This is conversation history. You do not semantically search it.&lt;/p&gt;

&lt;p&gt;Long-term memory is extracted records, organised into namespaces. Normally these are produced asynchronously: a memory strategy runs in the background, reads your short-term events, and extracts or consolidates records into a namespace. The AWS docs are explicit that this generation is an async background process. You read long-term records with &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt;, a semantic search scoped to a namespace.&lt;/p&gt;

&lt;p&gt;So if you want an agent to recall a durable fact on the next turn, it has to live in a long-term namespace that &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt; reads. Writing a &lt;code&gt;CreateEvent&lt;/code&gt; and hoping the strategy extracts the right fields is slow and non-deterministic. There is a better path.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;BatchCreateMemoryRecords&lt;/code&gt; writes directly into a long-term namespace. It is the bring-your-own-extraction door: you have already structured the fact, so you skip the strategy and put the record where the reader will look. The request is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;batch_create_memory_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;memoryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MEMORY_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;clientToken&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# batch idempotency; a retried identical batch dedupes
&lt;/span&gt;    &lt;span class="n"&gt;records&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;requestIdentifier&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seed-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# correlation key, NOT idempotency
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;namespaces&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/orgs/&amp;lt;tenant&amp;gt;/user/&amp;lt;user&amp;gt;/preferences/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Role: AE, mid-market SaaS. Prefers blunt feedback.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="c1"&gt;# memoryStrategyId is optional; see below
&lt;/span&gt;    &lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note there is no &lt;code&gt;actorId&lt;/code&gt; argument. The actor is encoded into the namespace string. Get the namespace wrong and the write goes somewhere the reader never queries, which is its own quiet failure.&lt;/p&gt;

&lt;p&gt;That has a security edge too. Because the actor is just part of a string, your tenant isolation is only as strong as the code that builds it, and a bug there is a cross-tenant read. The same reasoning behind &lt;a href="https://rajmurugan.com/blog/llm-is-not-a-security-boundary" rel="noopener noreferrer"&gt;the LLM is not a security boundary&lt;/a&gt; applies to your own string formatting: do not let it be the only thing standing between tenants. &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt; honours the &lt;code&gt;bedrock-agentcore:namespace&lt;/code&gt; (exact) and &lt;code&gt;bedrock-agentcore:namespacePath&lt;/code&gt; (subtree) IAM condition keys, so a policy can pin a principal to its own &lt;code&gt;/orgs/&amp;lt;tenant&amp;gt;/&lt;/code&gt; prefix and the service refuses an off-tenant namespace whatever the code passes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Success is not retrievability
&lt;/h2&gt;

&lt;p&gt;Here is the part that cost me the afternoon. The write returns &lt;code&gt;201&lt;/code&gt; with a &lt;code&gt;successfulRecords&lt;/code&gt; entry and a &lt;code&gt;memoryRecordId&lt;/code&gt;. Fetch that id directly and the record is there immediately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;batch_create_memory_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MEMORY_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;records&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;rid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;successfulRecords&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memoryRecordId&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_memory_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MEMORY_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;memoryRecordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;rid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# exists, right away
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now query the namespace the way an agent actually would, and at five seconds it is empty:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retrieve_memory_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;memoryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MEMORY_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;namespace&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/orgs/&amp;lt;tenant&amp;gt;/user/&amp;lt;user&amp;gt;/preferences/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;searchCriteria&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;searchQuery&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role preferences&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;topK&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# +5s -&amp;gt; 0 records
&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_memory_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MEMORY_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;namespace&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;NS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# +5s -&amp;gt; 0 records too
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both the semantic read and the plain namespace list returned nothing, while the record was fetchable by id the whole time. Poll for longer and the rows appear. So the record was created, addressable, and not yet indexed for namespace or semantic retrieval.&lt;/p&gt;

&lt;p&gt;I ran that write-then-poll loop three times, fresh namespace each time, checking every three seconds. The semantic read first returned the record at 16, 27 and 15 seconds. The namespace list was similar, at 16, 23 and 15 seconds. Same account, same region, one sitting, so treat it as a rough ballpark, not a benchmark. And treat the fresh-namespace part as a confound, not a control: some of that time may be the namespace itself warming up rather than the record indexing, so a warm namespace already holding thousands of records could behave differently. I have not measured that steady state yet, and it is the number production would actually care about. Three samples also cannot see a tail, and the tail is the whole operational question, so the right move is not to trust the ballpark at all.&lt;/p&gt;

&lt;p&gt;The docs tell you long-term &lt;em&gt;generation&lt;/em&gt; from events is asynchronous. They do not tell you that a &lt;em&gt;direct&lt;/em&gt; &lt;code&gt;BatchCreateMemoryRecords&lt;/code&gt; write also indexes asynchronously. I went looking, in the API reference and the memory guide, and could not find the read-after-write behaviour stated anywhere. You would reasonably assume the direct door skips the wait, because you did the extraction yourself. It does not skip the indexing.&lt;/p&gt;

&lt;p&gt;The mental model that would have saved me the afternoon: &lt;code&gt;201&lt;/code&gt; means accepted, not queryable. Read-after-write on a namespace is eventually consistent, on the order of tens of seconds. If you need certainty that a specific record landed, read it by id with &lt;code&gt;GetMemoryRecord&lt;/code&gt;, which is immediate. If you need it to appear in a namespace search, poll until it does rather than sleep on a fixed guess (more on that below).&lt;/p&gt;

&lt;p&gt;One consequence follows straight from that lag: because the write returns before it is searchable, a retry inside the window is easy, and &lt;code&gt;requestIdentifier&lt;/code&gt; will not save you. It is a correlation key, so two writes with the same one create two distinct records. The idempotency control is &lt;code&gt;clientToken&lt;/code&gt; on the batch call. Retry the identical batch with the same token and the service dedupes it. If your write path retries on timeout, set it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one that was simpler than the docs implied
&lt;/h2&gt;

&lt;p&gt;A smaller finding while I was in there. Long-term records carry an optional &lt;code&gt;memoryStrategyId&lt;/code&gt;. I assumed retrieval might require the record's strategy to match the strategy that owns the namespace. It does not. I wrote two records to the same namespace, one with a &lt;code&gt;memoryStrategyId&lt;/code&gt; and one without, and &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt; returned both. So an unfiltered retrieve does not gate on strategy id: give it a namespace and a &lt;code&gt;searchQuery&lt;/code&gt; and it returns matching records whether or not they carry a &lt;code&gt;memoryStrategyId&lt;/code&gt;. You can add that gate yourself, &lt;code&gt;searchCriteria&lt;/code&gt; also takes a &lt;code&gt;memoryStrategyId&lt;/code&gt; and &lt;code&gt;metadataFilters&lt;/code&gt;, but retrieval does not impose it by default. The strategy id is for associating a record with a strategy's consolidation, not a mandatory gate on reads. One less thing to get exactly right.&lt;/p&gt;

&lt;p&gt;That association is the part worth thinking about, and it outranks the read behaviour. Consolidation exists precisely to merge and dedupe the records a strategy owns, so a directly-written "durable fact" sitting under a built-in strategy is not something I would assume stays byte-for-byte as I wrote it. The documented pairing for bring-your-own extraction is a &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/memory-self-managed-strategies.html" rel="noopener noreferrer"&gt;self-managed strategy&lt;/a&gt;: direct &lt;code&gt;BatchCreateMemoryRecords&lt;/code&gt; writes bypass the extraction pipeline entirely, and a self-managed strategy leaves extraction and consolidation to code you control rather than a built-in pass you did not write. If you are seeding durable facts by hand, that is the strategy to put them under.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I now do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Treat &lt;code&gt;201&lt;/code&gt; from &lt;code&gt;BatchCreateMemoryRecords&lt;/code&gt; as accepted, not queryable. The record is addressable by id immediately and searchable by namespace tens of seconds later.&lt;/li&gt;
&lt;li&gt;If I need read-after-write certainty on a specific record, read it by id with &lt;code&gt;GetMemoryRecord&lt;/code&gt;, never by a namespace search.&lt;/li&gt;
&lt;li&gt;Keep my own durable map of record ids. &lt;code&gt;GetMemoryRecord&lt;/code&gt; is only immediate because I already hold the &lt;code&gt;memoryRecordId&lt;/code&gt;, which means I persisted it somewhere (an actor-to-ids table in DynamoDB, say). So AgentCore Memory is the semantic-recall layer, not the system of record: if a fact has to be readable the instant it is written, my store is the source of truth and AgentCore is the index that catches up.&lt;/li&gt;
&lt;li&gt;Do not gate a "saved, now ask me" experience on instant recall. If a user saves a profile and immediately asks the agent what it knows about them, the honest answer for a few tens of seconds is nothing. For an onboarding flow this is fine, because there is natural delay before the first real turn. For a health check that writes then reads a namespace, it is a flake generator.&lt;/li&gt;
&lt;li&gt;Gate on a readiness probe, never a timer. A fixed &lt;code&gt;sleep(30)&lt;/code&gt; is both slow and a p99 flake generator, and it hides the silent namespace-typo failure behind a wait that looks deliberate. Poll &lt;code&gt;ListMemoryRecords&lt;/code&gt;, or the &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt; the reader actually uses, until the record appears or a timeout fires, and alarm on the timeout. That converts "I guessed thirty seconds" into a measured, monitored wait, and turns a wrong namespace into a real error instead of a silently empty read.&lt;/li&gt;
&lt;li&gt;If the write path can retry, and a timeout inside the index-lag window is the obvious case, set &lt;code&gt;clientToken&lt;/code&gt; on the batch so a replay dedupes. &lt;code&gt;requestIdentifier&lt;/code&gt; will not, it is a correlation key and two identical ones make two records.&lt;/li&gt;
&lt;li&gt;Enforce tenant isolation on the namespace in IAM, not just in the code that builds the string. Pin the principal with a &lt;code&gt;bedrock-agentcore:namespacePath&lt;/code&gt; condition so an off-tenant read is refused by the service, not by a code review.&lt;/li&gt;
&lt;li&gt;Verify the operation exists in the pinned SDK before trusting a helper. A method name, a docstring, and a green diff are not evidence that a call is real.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a reason to avoid direct writes. They are the right call when you already have a structured fact, because the alternative is waiting on async extraction that may drop or reshape the fields you care about. It just means you design around two facts the docs bury: the name might not be a real operation, and the &lt;code&gt;201&lt;/code&gt; means the service took your record, not that a reader can find it. Both looked like success. Neither was, until I actually read it back.&lt;/p&gt;

&lt;p&gt;The hard part was never the memory model. It was the gap between what the API reports and what is true a second later.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Latency measured in a dev account, one region, three runs, polled at three-second granularity, so the real figure sits a little under each number. Retention and consolidation behaviour of directly-created long-term records I have not fully characterised yet. If your lag numbers differ, I would like to hear them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>bedrock</category>
      <category>agentcore</category>
      <category>aiagents</category>
    </item>
  </channel>
</rss>
