<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: vadim albarov</title>
    <description>The latest articles on DEV Community by vadim albarov (@vadim_albarov).</description>
    <link>https://dev.to/vadim_albarov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4058639%2Fae7afcec-637e-4f73-bea2-21b8cd0d9c4f.jpg</url>
      <title>DEV Community: vadim albarov</title>
      <link>https://dev.to/vadim_albarov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vadim_albarov"/>
    <language>en</language>
    <item>
      <title>Can a Local LLM Wash Out a Watermark Without Washing Out the Meaning? I Tested 300 Rewrites</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:49:07 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/can-a-local-llm-wash-out-a-watermark-without-washing-out-the-meaning-i-tested-300-rewrites-15jb</link>
      <guid>https://dev.to/vadim_albarov/can-a-local-llm-wash-out-a-watermark-without-washing-out-the-meaning-i-tested-300-rewrites-15jb</guid>
      <description>&lt;p&gt;AI text watermarks hide a signal in &lt;em&gt;which words&lt;/em&gt; a model picks. A common idea is that you can remove that signal by running the text through another model: "rewrite this in your own words."&lt;/p&gt;

&lt;p&gt;That raises a simple question. &lt;strong&gt;If you rewrite a text enough to break its word patterns, do you still keep its meaning?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I tested it on a single GPU with 16 GB of VRAM. In short: &lt;strong&gt;yes, mostly — but the model you pick matters a lot, and there is a real trade-off between "changed a lot" and "kept everything."&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Here is the pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Claude Opus 5.5 wrote 100 source texts.&lt;/strong&gt; There were 10 genres (news, technical explainers, science, opinion, personal stories, marketing, how-to guides, business memos, history, and short fiction), with 10 topics each. Lengths were from about 190 to 900 words. Every text was packed with checkable details: names, numbers, dates, steps, and plot events.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three local models rewrote every text&lt;/strong&gt; with &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; on a GPU with 16 GB VRAM:

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;qwen3:4b-instruct&lt;/code&gt; (small)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemma3:12b&lt;/code&gt; (mid)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;qwen3:14b&lt;/code&gt; (upper mid, thinking mode off)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Fable 5.1 judged all 300 rewrites blind.&lt;/strong&gt; Each packet held the original and three rewrites labeled A/B/C in random order. The judge listed every lost, changed, or added fact, and gave each rewrite a score from 1 to 5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I measured how much the wording changed&lt;/strong&gt; by counting how many of the original's word n-grams survive in the rewrite.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;All models used the same prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Rewrite the text below completely in your own words.
- Change the vocabulary and the sentence structure; do not copy phrases.
- Keep ALL of the meaning: every fact, name, number, date, step, argument, and plot event.
- Do not add new information, comments, or opinions.
- Keep the same genre, tone, and roughly the same length. Keep the title line, rewritten.
Output only the rewritten text, nothing else.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Settings: temperature 0.7, &lt;code&gt;num_ctx&lt;/code&gt; 8192, one run per text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why n-grams are a fair proxy for "watermark removed"
&lt;/h2&gt;

&lt;p&gt;Most text watermarks (green-list schemes like Kirchenbauer et al., and Google's SynthID-Text) work at the token level. While the model writes, they nudge it toward certain tokens based on the tokens just before. A detector then counts how often these "preferred" token pairs or sequences appear.&lt;/p&gt;

&lt;p&gt;So if most of the original word sequences are gone, most of the signal is gone too. I did &lt;strong&gt;not&lt;/strong&gt; run a watermark detector, because I have no access to one for these texts. Treat the n-gram numbers as a proxy, not a proof.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Meaning, by model
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Avg score (1–5)&lt;/th&gt;
&lt;th&gt;"Preserved"&lt;/th&gt;
&lt;th&gt;"Mostly preserved"&lt;/th&gt;
&lt;th&gt;"Not preserved"&lt;/th&gt;
&lt;th&gt;Same understanding for the reader&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:14b&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4.68&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;82%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;18%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:12b&lt;/td&gt;
&lt;td&gt;4.09&lt;/td&gt;
&lt;td&gt;40%&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;98%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b-instruct&lt;/td&gt;
&lt;td&gt;3.76&lt;/td&gt;
&lt;td&gt;21%&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;3%&lt;/td&gt;
&lt;td&gt;81%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;All 300&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4.18&lt;/td&gt;
&lt;td&gt;48%&lt;/td&gt;
&lt;td&gt;51%&lt;/td&gt;
&lt;td&gt;1%&lt;/td&gt;
&lt;td&gt;93%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Same understanding" was a yes/no question to the judge: &lt;em&gt;would a reader of the rewrite come away with the same understanding as a reader of the original?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When the judge ranked the three rewrites of each text (still blind):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Ranked best&lt;/th&gt;
&lt;th&gt;Ranked worst&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:14b&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:12b&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b-instruct&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Average errors per rewrite, as listed by the judge:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Facts lost&lt;/th&gt;
&lt;th&gt;Facts changed&lt;/th&gt;
&lt;th&gt;Things added&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:14b&lt;/td&gt;
&lt;td&gt;0.59&lt;/td&gt;
&lt;td&gt;0.96&lt;/td&gt;
&lt;td&gt;0.05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:12b&lt;/td&gt;
&lt;td&gt;1.42&lt;/td&gt;
&lt;td&gt;2.04&lt;/td&gt;
&lt;td&gt;0.30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b-instruct&lt;/td&gt;
&lt;td&gt;1.22&lt;/td&gt;
&lt;td&gt;2.94&lt;/td&gt;
&lt;td&gt;0.38&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Wording change, by model
&lt;/h3&gt;

&lt;p&gt;"Kept" means the share of the original's n-grams that still appear in the rewrite. Lower means a heavier rewrite.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Words kept (1-gram)&lt;/th&gt;
&lt;th&gt;2-grams kept&lt;/th&gt;
&lt;th&gt;3-grams kept&lt;/th&gt;
&lt;th&gt;5-grams kept&lt;/th&gt;
&lt;th&gt;Longest copied run (words)&lt;/th&gt;
&lt;th&gt;Length vs original&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:12b&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;34%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11.7&lt;/td&gt;
&lt;td&gt;1.12×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b-instruct&lt;/td&gt;
&lt;td&gt;64%&lt;/td&gt;
&lt;td&gt;37%&lt;/td&gt;
&lt;td&gt;23%&lt;/td&gt;
&lt;td&gt;9%&lt;/td&gt;
&lt;td&gt;12.4&lt;/td&gt;
&lt;td&gt;1.05×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:14b&lt;/td&gt;
&lt;td&gt;72%&lt;/td&gt;
&lt;td&gt;51%&lt;/td&gt;
&lt;td&gt;37%&lt;/td&gt;
&lt;td&gt;21%&lt;/td&gt;
&lt;td&gt;20.7&lt;/td&gt;
&lt;td&gt;1.07×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  The trade-off
&lt;/h3&gt;

&lt;p&gt;Put the two tables side by side:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;qwen3:14b&lt;/strong&gt; kept the meaning best by far, but it also kept the most of the original wording. About 1 in 5 five-word sequences survived, and on average it copied a run of about 21 words straight from the original.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;gemma3:12b&lt;/strong&gt; changed the wording the most (only 8% of 5-grams survived) and still scored 4.09. Almost all its rewrites (98%) gave the reader the same understanding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;qwen3:4b-instruct&lt;/strong&gt; changed the wording almost as much as Gemma, but it was the least accurate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Across all 300 rewrites, the correlation between "meaning score" and "5-grams kept" was &lt;strong&gt;0.36&lt;/strong&gt;. So yes, rewrites that copy more wording tend to keep more meaning. But inside each model the link was weak (0.03–0.22). Most of the difference comes from &lt;em&gt;which model&lt;/em&gt; you use, not from how hard a given model happens to rewrite a given text.&lt;/p&gt;

&lt;h3&gt;
  
  
  What kinds of mistakes?
&lt;/h3&gt;

&lt;p&gt;The type of error mattered more than the count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;qwen3:14b&lt;/strong&gt; mostly made tiny changes in nuance:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"two placebo pills work better than one" was hedged to "may be more effective"&lt;/li&gt;
&lt;li&gt;"ransomware" was widened to "malware attacks"&lt;/li&gt;
&lt;li&gt;an overall risk of "moderate" became "average"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;gemma3:12b&lt;/strong&gt; made the wording vaguer or swapped in a near-synonym:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Most houseplants" became "Many indoor plants"&lt;/li&gt;
&lt;li&gt;"hiring manager" became "recruiter"&lt;/li&gt;
&lt;li&gt;Kubernetes init containers that "run to completion, one after another" became a &lt;em&gt;single&lt;/em&gt; init container that "executes a sequence of tasks". That one is a real technical error.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is a typical example. Original:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;SafeNest Products announced on Monday that it is recalling about 86,000 infant car seats sold in the United States because the harness buckle may fail to latch fully.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Gemma:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;On Monday, SafeNest Products declared a recall of roughly 86,000 car seats intended for infants, distributed across the United States, due to a potential issue with the harness buckle's functionality.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every fact is still there, but "may fail to latch fully" became "a potential issue with ... functionality." It reads fine, but it says less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;qwen3:4b-instruct&lt;/strong&gt; made the dangerous mistakes: wrong names and wrong numbers.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It &lt;strong&gt;renamed products&lt;/strong&gt;: "Quietwave Q7" headphones became "SilentHush S7" all through the text, and the "Elevate Pro Dual-Motor" desk became "Rise Pro Dual-Drive."&lt;/li&gt;
&lt;li&gt;It garbled numbers: a desk height range "from 150 to 200 centimeters" became "from 150 to 20."&lt;/li&gt;
&lt;li&gt;It changed amounts: British tea imports of "millions of pounds" a year became "hundreds of thousands of pounds."&lt;/li&gt;
&lt;li&gt;It cut a farmer's name, "Marcos Ribeiro," to "Marcos Ribe."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A rough keyword search over the judge's notes found number-related changes in about &lt;strong&gt;31 of 100&lt;/strong&gt; qwen3:4b rewrites, compared with about 11 for Gemma and 3 for qwen3:14b. This count is approximate, but the gap is clear.&lt;/p&gt;

&lt;p&gt;All 3 rewrites judged "not preserved" in the whole test came from qwen3:4b. Two were the renamed products, and the third was a short story whose ending stopped making sense.&lt;/p&gt;

&lt;h3&gt;
  
  
  By genre and length
&lt;/h3&gt;

&lt;p&gt;The genre made less difference than I expected. The average score ranged from 4.03 (marketing copy) to 4.40 (science writing). Marketing copy had the most "not preserved" verdicts. Product names and spec numbers are exactly where the small model slips, and those details &lt;em&gt;are&lt;/em&gt; the meaning of marketing copy. Business memos and technical explainers were most often fully "preserved" (77–80%).&lt;/p&gt;

&lt;p&gt;Length did not hurt. The 750–900 word texts scored about the same as, or a little better than, the short ones (4.33 vs 4.19). With an 8K context, all three models handled 900 words easily.&lt;/p&gt;

&lt;h3&gt;
  
  
  Speed
&lt;/h3&gt;

&lt;p&gt;Total time for 100 rewrites on the 16 GB GPU:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;th&gt;Per text&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b-instruct&lt;/td&gt;
&lt;td&gt;11 min&lt;/td&gt;
&lt;td&gt;~7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:14b&lt;/td&gt;
&lt;td&gt;25 min&lt;/td&gt;
&lt;td&gt;~15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:12b&lt;/td&gt;
&lt;td&gt;25 min&lt;/td&gt;
&lt;td&gt;~15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  So — does the rewrite keep the meaning?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;With a good 12–14B model: yes.&lt;/strong&gt; Readers would get the same understanding from 98–100% of the rewrites, and none were judged "not preserved." The remaining issues are small shifts in nuance or precision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With a 4B model: usually, but you can't trust the details.&lt;/strong&gt; One in five rewrites changed what a reader would understand. Names and numbers are the weak spot, and they are often the facts that matter most.&lt;/p&gt;

&lt;p&gt;And there is a trade-off you can't ignore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you want &lt;strong&gt;maximum fidelity&lt;/strong&gt;, &lt;code&gt;qwen3:14b&lt;/code&gt; is the clear winner. But it keeps more of the original wording, so if your goal is to break token-level patterns, it does the weakest job.&lt;/li&gt;
&lt;li&gt;If you want &lt;strong&gt;maximum change in wording with good meaning&lt;/strong&gt;, &lt;code&gt;gemma3:12b&lt;/code&gt; is the sweet spot. Only 8% of 5-grams survive, and the reader gets the same understanding 98% of the time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No real watermark detector.&lt;/strong&gt; N-gram overlap is a proxy. Some watermark research claims robustness to paraphrase, and semantic-level watermarks would not be hurt by word changes at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One judge, and it is an LLM.&lt;/strong&gt; Fable 5.1 was strict and specific (it listed things like "price changed from X to Y"), but I did not check it against human ratings. The generator and the judge are both Claude models, which could add some bias toward Claude-like phrasing. That should affect all three local models equally, though.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One run per text&lt;/strong&gt; at temperature 0.7. Another seed would give somewhat different rewrites.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;English only&lt;/strong&gt;, synthetic texts, one prompt. A stricter prompt ("never change any name or number") would probably help the 4B model.&lt;/li&gt;
&lt;li&gt;The keyword counts of name and number errors are approximate.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Next time: SynthID?
&lt;/h2&gt;

&lt;p&gt;This test answered only half of the question. It showed that a good local model keeps the meaning, but it only &lt;em&gt;estimated&lt;/em&gt; watermark removal from n-gram overlap.&lt;/p&gt;

&lt;p&gt;A real check would use &lt;strong&gt;SynthID-Text&lt;/strong&gt;, Google's open-source text watermark. Would rewriting with a local model really remove a watermark that a detector can see?&lt;/p&gt;

&lt;p&gt;Should I run that test next? Tell me in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>ollama</category>
      <category>discuss</category>
    </item>
    <item>
      <title>How Much Does Your Lunch Break Cost in Claude Code Tokens?</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Thu, 01 Oct 2026 15:06:06 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/how-much-does-your-lunch-break-cost-in-claude-code-tokens-4k0o</link>
      <guid>https://dev.to/vadim_albarov/how-much-does-your-lunch-break-cost-in-claude-code-tokens-4k0o</guid>
      <description>&lt;p&gt;You step away for lunch with a big Claude Code session open. You come back, type one message, and continue. Nothing looks different. But that one message may have cost 20 times more than the one before it.&lt;/p&gt;

&lt;p&gt;I wanted a real number, so I scanned a month of my own Claude Code logs. The short answer: a typical break past the one-hour mark cost me about 125k–135k tokens of cold cache writes. Over the month, these restarts added about 15% to my input usage.&lt;/p&gt;

&lt;p&gt;This post explains why it happens, shows the script I used, and lists what I changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How prompt caching works in Claude Code
&lt;/h2&gt;

&lt;p&gt;Every turn in a Claude Code session sends the whole conversation to the model again: system prompt, tool definitions, project files you read, tool output, everything. A long session can carry 100k–400k tokens of context.&lt;/p&gt;

&lt;p&gt;Prompt caching makes this affordable. After the first request, the stable prefix of the prompt is stored in a cache. Later requests read it from the cache instead of processing it again. In API pricing terms, relative to the base input price:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Token type&lt;/th&gt;
&lt;th&gt;Relative price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Uncached input&lt;/td&gt;
&lt;td&gt;1.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache write, 1-hour TTL&lt;/td&gt;
&lt;td&gt;2.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache read&lt;/td&gt;
&lt;td&gt;0.1x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A warm turn reads the context at 0.1x. A cold turn writes the whole context again at 2.0x. That is where the &lt;strong&gt;roughly 20x&lt;/strong&gt; comes from.&lt;/p&gt;

&lt;p&gt;The cache has a time to live (TTL). In my Claude Code sessions it is &lt;strong&gt;one hour&lt;/strong&gt;, and each request resets the timer. Tool calls count as requests, so an agent working on its own keeps the cache warm. The risk is when &lt;em&gt;you&lt;/em&gt; are the slow part: a meeting, lunch, the end of the day.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: the 1-hour TTL is what my sessions used. The default API TTL is 5 minutes, and Claude Code can drop to 5 minutes in some cases, for example usage overage. The script below reports which TTL each write used, so you can check your own.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Where the data lives
&lt;/h2&gt;

&lt;p&gt;Claude Code stores each session as a JSONL file under &lt;code&gt;~/.claude/projects/&amp;lt;project&amp;gt;/&amp;lt;session-id&amp;gt;.jsonl&lt;/code&gt;. Every assistant turn includes a &lt;code&gt;usage&lt;/code&gt; block. Trimmed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-01T13:28:05.993Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"msg_..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"cache_creation_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;32912&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"cache_read_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;211&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"cache_creation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"ephemeral_1h_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;32912&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"ephemeral_5m_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three fields matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;cache_read_input_tokens&lt;/code&gt;: context served from the cache (cheap).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cache_creation_input_tokens&lt;/code&gt;: context written to the cache (expensive).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cache_creation.ephemeral_1h_input_tokens&lt;/code&gt; / &lt;code&gt;ephemeral_5m_input_tokens&lt;/code&gt;: which TTL the write used.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On a normal warm turn, reads are large and writes are small (just the new tool output). A cache miss looks the other way round: &lt;strong&gt;a large write, little or no read&lt;/strong&gt;. If the previous turn in the same file was more than an hour earlier, the idle gap is the cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  The script
&lt;/h2&gt;

&lt;p&gt;The script walks every session file, removes duplicate messages by ID, and flags turns where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;cache_creation_input_tokens&lt;/code&gt; is at least a threshold (default 50,000), and&lt;/li&gt;
&lt;li&gt;the write is larger than the read.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It records the gap since the previous turn, groups the misses by gap length, and estimates a weighted cost share using the prices above.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Click to see the full script (100 lines)
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timedelta&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;defaultdict&lt;/span&gt;

&lt;span class="n"&gt;ROOT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expanduser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;~/.claude/projects&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;MIN_MISS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;50_000&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromisoformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Z&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;+00:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;misses&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="n"&gt;totals&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;per_project&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ROOT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.jsonl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;project&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dirname&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)[:&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;prev_ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;seen_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;replace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSONDecodeError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
            &lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;mid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;mid&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seen_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;seen_ids&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse_ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;read&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_read_input_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="n"&gt;created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_creation_input_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="n"&gt;cc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_creation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
            &lt;span class="n"&gt;w1h&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral_1h_input_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="n"&gt;w5m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ephemeral_5m_input_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;turns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
            &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt;
            &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;
            &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;raw&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt;
            &lt;span class="n"&gt;per_project&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt;
            &lt;span class="n"&gt;per_project&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;
            &lt;span class="n"&gt;gap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;prev_ts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;prev_ts&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;MIN_MISS&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;misses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                                   &lt;span class="n"&gt;gap&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;gap&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                                   &lt;span class="n"&gt;w1h&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;w1h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;w5m&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;w5m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="n"&gt;prev_ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;

&lt;span class="n"&gt;misses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fmt_gap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;g&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;total_seconds&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;3600&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;6.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;total_seconds&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;m &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cache misses &amp;gt;= &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;MIN_MISS&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; created tokens, read &amp;lt; created&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;date&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gap&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ttl&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;project&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;28&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; sess&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;buckets&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defaultdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;misses&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;ttl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w1h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w5m&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5m&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;%Y-%m-%d %H&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;M&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;fmt_gap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gap&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;project&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;28&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;28&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;session&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;5m&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;minutes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5m-1h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hours&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1h-24h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;timedelta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hours&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;gt;24h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;buckets&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="n"&gt;buckets&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Misses by idle gap before the request:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;5m&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5m-1h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1h-24h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;gt;24h&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;buckets&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; misses  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens written&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;miss_tok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;misses&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;All assistant turns: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;turns&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  cache read tokens    : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  cache created tokens : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;miss_tok&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; in the misses listed above)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  uncached input tokens: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;raw&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# cost weights: read 0.1, 1h write 2.0, raw 1.0
&lt;/span&gt;&lt;span class="n"&gt;w_read&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;
&lt;span class="n"&gt;w_created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;
&lt;span class="n"&gt;w_raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;totals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;raw&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;
&lt;span class="n"&gt;w_miss&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;miss_tok&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;
&lt;span class="n"&gt;tot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;w_read&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;w_created&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;w_raw&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Weighted input cost share (read=0.1, write=2.0, raw=1.0):&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  cache reads : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;w_read&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;tot&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;%&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  cache writes: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;w_created&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;tot&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;%   of which listed misses: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;w_miss&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;tot&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;%&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  raw input   : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;w_raw&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;tot&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;5.1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;%&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Top projects by cache-write tokens:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;per_project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;kv&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;kv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])[:&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; written &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  read &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Run it with an optional threshold:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python cache_misses.py          &lt;span class="c"&gt;# default 50,000&lt;/span&gt;
python cache_misses.py 30000    &lt;span class="c"&gt;# lower threshold&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  My results
&lt;/h2&gt;

&lt;p&gt;The data covers 2026-08-31 to 2026-10-01: 2,410 assistant turns across five projects.&lt;/p&gt;

&lt;h3&gt;
  
  
  The misses
&lt;/h3&gt;

&lt;p&gt;Restarts after an idle gap of over one hour, threshold 50k:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Idle gap&lt;/th&gt;
&lt;th&gt;Tokens re-written&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1.1h&lt;/td&gt;
&lt;td&gt;170k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.1h&lt;/td&gt;
&lt;td&gt;80k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.2h&lt;/td&gt;
&lt;td&gt;137k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.2h&lt;/td&gt;
&lt;td&gt;77k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.2h&lt;/td&gt;
&lt;td&gt;193k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.3h&lt;/td&gt;
&lt;td&gt;80k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.4h&lt;/td&gt;
&lt;td&gt;142k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.5h&lt;/td&gt;
&lt;td&gt;58k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.6h&lt;/td&gt;
&lt;td&gt;110k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.6h&lt;/td&gt;
&lt;td&gt;164k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.2h&lt;/td&gt;
&lt;td&gt;148k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.5h&lt;/td&gt;
&lt;td&gt;53k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.6h&lt;/td&gt;
&lt;td&gt;171k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.6h&lt;/td&gt;
&lt;td&gt;103k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.8h&lt;/td&gt;
&lt;td&gt;131k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15.1h&lt;/td&gt;
&lt;td&gt;124k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;18.7h&lt;/td&gt;
&lt;td&gt;233k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;61.8h&lt;/td&gt;
&lt;td&gt;476k&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Idle restarts over 1h&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens re-written cold&lt;/td&gt;
&lt;td&gt;2.65M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mean per restart&lt;/td&gt;
&lt;td&gt;147k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median per restart&lt;/td&gt;
&lt;td&gt;134k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median for gaps of 1.1h–1.6h (10 restarts)&lt;/td&gt;
&lt;td&gt;123k&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The mean is pulled up by one outlier: a session resumed after a 2.5-day weekend that re-wrote 476k tokens in one request. The median is the more honest "cost of a break".&lt;/p&gt;

&lt;h3&gt;
  
  
  The pattern
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Most gaps barely passed the limit.&lt;/strong&gt; 10 of the 18 were between 1.1 and 1.6 hours. These are lunch breaks and meetings, not abandoned sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One stale session cost as much as three or four normal breaks.&lt;/strong&gt; Resuming a huge session after days away is the worst case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session starts are a separate, fixed cost.&lt;/strong&gt; At a 30k threshold the script also shows 14 session starts, each 31k–36k tokens. That is the startup context: system prompt, &lt;code&gt;CLAUDE.md&lt;/code&gt;, memory, skills, and MCP tool definitions. Every new session pays it once, and no timing habit removes it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The cost share
&lt;/h3&gt;

&lt;p&gt;Weighted by relative price across all 2,410 turns:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;th&gt;Weighted share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cache reads&lt;/td&gt;
&lt;td&gt;234M&lt;/td&gt;
&lt;td&gt;55%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache writes, normal incremental&lt;/td&gt;
&lt;td&gt;6.8M&lt;/td&gt;
&lt;td&gt;32%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache writes, idle misses&lt;/td&gt;
&lt;td&gt;2.7M&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Uncached input&lt;/td&gt;
&lt;td&gt;82k&lt;/td&gt;
&lt;td&gt;0.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Idle misses were 13% of the total weighted input. Put another way: without them, input cost would have been 87% of what it was, so the misses added about &lt;strong&gt;15%&lt;/strong&gt; on top.&lt;/p&gt;

&lt;p&gt;The 32% for normal writes is the ordinary cost of new tool output going into the cache each turn. Timing does not change it. Only the 13% slice can be avoided.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Come back within the hour when I can.&lt;/strong&gt; Most of my misses were 5 to 35 minutes past the limit. If I know I will be back soon, I make the next message the first thing I do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Run &lt;code&gt;/compact&lt;/code&gt; before a long break.&lt;/strong&gt; If I will be away for more than an hour and the session is large, I compact while the cache is still warm. The summary is cheap to produce on a warm cache. The restart after the break then re-writes a small summary, not 150k tokens of history, and every turn after that is cheaper too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Do not revive stale sessions.&lt;/strong&gt; For anything older than a day, I start a fresh session with a short brief: goal, current state, key files. That costs about 35k tokens of startup context, not 400k+ of old history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Do not bother with keep-alive pings.&lt;/strong&gt; Sending a "heartbeat" message every hour does keep the cache warm, and each one costs only about 0.1x of the context. On paper it wins for gaps up to about 20 hours (20 pings at 0.1x equal one cold write at 2.0x). In practice it adds noise to the conversation, it stops when the machine sleeps, and it keeps a large context alive that you may not need. &lt;code&gt;/compact&lt;/code&gt; is simpler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plan differences.&lt;/strong&gt; On a subscription plan, tokens count against usage limits, not a bill. The relative weights still show what eats your limits faster, but your plan may not count tokens exactly like API prices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The weights are an estimate.&lt;/strong&gt; Read = 0.1, 1h write = 2.0, uncached = 1.0 follows published API pricing ratios. Output tokens are left out on purpose, because they do not depend on caching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The heuristic is simple.&lt;/strong&gt; "Write &amp;gt; read and write ≥ 50k" catches idle restarts well. It also catches a few non-idle events, such as one write under 5 minutes in my data, so always look at the gap column.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One user, one month.&lt;/strong&gt; These are numbers from my own work: several infrastructure and app repos, long agentic sessions, normal office hours.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it on your own logs
&lt;/h2&gt;

&lt;p&gt;Save the script, run it, and look at two things: the &lt;code&gt;1h-24h&lt;/code&gt; bucket and the &lt;code&gt;&amp;gt;24h&lt;/code&gt; bucket. If they hold a large share of your cache writes, a habit change will save you more than any prompt tweak.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Over to you:&lt;/strong&gt; is saving ~15% of tokens worth changing how you work? Or is the extra attention not worth it? I'd like to know how you handle long sessions and breaks.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;These numbers are averages from one month of my own work. Your numbers could vary.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>productivity</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Typos don't break LLM prompts. One missing quote mark does.</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Sun, 27 Sep 2026 21:02:46 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/typos-dont-break-llm-prompts-one-missing-quote-mark-does-d7d</link>
      <guid>https://dev.to/vadim_albarov/typos-dont-break-llm-prompts-one-missing-quote-mark-does-d7d</guid>
      <description>&lt;p&gt;Everyone who types prompts has wondered the same thing. Does the model care that I wrote "teh" or forgot a comma? Should I clean up my prompt before hitting enter?&lt;/p&gt;

&lt;p&gt;I ran a test to find out. The short answer: stop fixing your spelling. Start checking your quotes and colons.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I tested
&lt;/h2&gt;

&lt;p&gt;I wrote a set of tasks with short, checkable answers. Arithmetic, extraction from a passage, list filtering with "except" and "not", a Python snippet's output, counting words inside quotes, and so on. Then I broke the prompts in controlled ways and sent every version to every model, several times each, in a fresh session every time.&lt;/p&gt;

&lt;p&gt;Two prompt sets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Easy set.&lt;/strong&gt; 12 one-line tasks, each hand-written at five error levels, from clean to "heavy typos plus broken punctuation".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard set.&lt;/strong&gt; 9 long prompts (200 to 450 words) where the question depends on a detail buried in the text. Here I separated the two kinds of damage:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Spelling noise&lt;/em&gt;: a script misspells 35% or 70% of the words, but never touches the words the answer depends on.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Structural break&lt;/em&gt;: the spelling is perfect, but exactly one punctuation mark that carries meaning is wrong. A missing closing quote. A colon dropped before a list. A comma moved around "except". &lt;code&gt;9.45&lt;/code&gt; instead of &lt;code&gt;9:45&lt;/code&gt;. &lt;code&gt;resign&lt;/code&gt; instead of &lt;code&gt;re-sign&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Two realistic profiles on top: &lt;em&gt;non-native grammar&lt;/em&gt; (articles and verb forms off, spelling fine) and &lt;em&gt;voice-to-text&lt;/em&gt; (no punctuation, no capitals, numbers spelled out, "resign" for "re-sign").&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Models: Claude Opus 5, Opus 5.5, Sonnet 5, Haiku 4.5, Fable 5.1 (all at low and medium effort), plus seven local models through Ollama: gemma4, qwen3 14B, phi4 14B, granite 4.2, ornith-1.5 9B, llama 3.1 8B, mistral 7B.&lt;/p&gt;

&lt;p&gt;Total: about 4,900 sessions. An answer counts as correct if the ideal answer appears in the reply, even with an explanation attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: spelling and grammar cost nothing
&lt;/h2&gt;

&lt;p&gt;Here is the hard set. Each cell is the percent of correct answers.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;model&lt;/th&gt;
&lt;th&gt;clean&lt;/th&gt;
&lt;th&gt;70% typos&lt;/th&gt;
&lt;th&gt;non-native grammar&lt;/th&gt;
&lt;th&gt;voice-to-text&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;one punctuation break&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5.5&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fable 5.1&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;92&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 5&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;granite4.2&lt;/td&gt;
&lt;td&gt;93&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;phi4 14B&lt;/td&gt;
&lt;td&gt;74&lt;/td&gt;
&lt;td&gt;74&lt;/td&gt;
&lt;td&gt;85&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma4&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;81&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;td&gt;79&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3 14B&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;74&lt;/td&gt;
&lt;td&gt;85&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ornith-1.5 9B&lt;/td&gt;
&lt;td&gt;74&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;42&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama 3.1 8B&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;63&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mistral 7B&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the Claude rows first. With 70% of the words misspelled, every one of them scored the same as on clean text. Non-native grammar: 100% across the board. This is the prompt that scored 100%:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;shop sellin pens 3dolar each an notbooks 5 dolar,each tom buys 4pen an 2notbook.how much he pay totaly,anser numbr only&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now read the last column. One punctuation mark, with perfect spelling everywhere else, and the same models lose 8 to 23 points.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: which punctuation mark breaks which model
&lt;/h2&gt;

&lt;p&gt;Not every break matters. Most were recovered from context by every Claude model. Two were not.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;break&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;th&gt;Fable&lt;/th&gt;
&lt;th&gt;Haiku&lt;/th&gt;
&lt;th&gt;Sonnet&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;missing closing quote&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;colon dropped before a list&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;comma moved around "except"&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;European decimals (1.500 at 2,50)&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9.45 instead of 9:45&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;managers' → manager's&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;resign instead of re-sign&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;comma dropped in "not closed, and assigned to Lee"&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The missing quote.&lt;/strong&gt; The task: count how many times "report" appears inside a quoted paragraph, not in the text after it. Remove the closing quote and every model counts the three extra "report"s after the paragraph. Answer 8 instead of 5. Opus 5.5 even wrote "the closing quote is missing, so I counted all" and still gave the wrong number. It noticed the problem and followed the broken structure anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The dropped colon.&lt;/strong&gt; "Give the plural form of each word below keeping the order mouse child foot analysis reply with the plurals only comma separated". Sonnet answered &lt;code&gt;mice, children, feet, analyses, replies&lt;/code&gt;. Every local model did the same. The word boundary between instruction and data was gone, so "reply" became data.&lt;/p&gt;

&lt;p&gt;Everything else was recovered, including the genuinely ambiguous one. "List the tickets that are not closed and assigned to Lee" can mean two things once the comma is gone. All five Claude models picked the intended reading every time, because the sentence started with "For Lee's standup".&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3: dictation is a different story
&lt;/h2&gt;

&lt;p&gt;The voice-to-text column is where the small models fall apart. gemma4 goes from 78 to 44. granite from 93 to 56. Two things do the damage: spelled-out numbers ("three hundred forty pallets across twenty two trailers") and, again, the missing colon before a list. Every local model scored 0% on the dictated list task.&lt;/p&gt;

&lt;p&gt;The frontier models mostly shrug it off, with one honest exception. "fifteen hundred units at two fifty each": Sonnet read "two fifty" as $250 in six trials out of six and answered 375,000 instead of 3,750. Fable did it once. Haiku turned "extension two zero one" into 2001 once. These are not model bugs. Spoken numbers are ambiguous, and the model has to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things that did not help
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Effort level.&lt;/strong&gt; Low versus medium made no difference for any model on any task. Structural breaks are not a reasoning problem, so more thinking does not fix them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Warning the model.&lt;/strong&gt; I prepended "the message below may contain spelling, grammar, and punctuation mistakes, answer according to the writer's likely intent". Zero effect on Opus 5, Haiku, and Sonnet. A few points on Opus 5.5 and Fable. Mixed or negative on the local models.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What degrades first: the format, not the answer
&lt;/h2&gt;

&lt;p&gt;Sloppy prompts get slightly sloppier replies before they get wrong ones. Asked for "the number only", Fable and Opus 5.5 add a line of working in about 20% of runs on any prompt. Haiku, on the easy set, drifted into "Yes. If all bloops are razzies..." explanations as the grammar got worse, and at the heaviest level refused four times out of eight because "bloops and razzies are not real words". With clean grammar it had answered the same nonsense-word question fine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Don't fix spelling or grammar. The model does not care, and neither should you.&lt;/li&gt;
&lt;li&gt;Do check the punctuation that carries structure: closing quotes, the colon before a list, commas around "except" and "not", and anything separating your instructions from your data.&lt;/li&gt;
&lt;li&gt;If you dictate prompts, re-read the numbers and the list boundaries before sending.&lt;/li&gt;
&lt;li&gt;Small local models are much less forgiving. If you run a 7B to 14B model, treat every delimiter as load-bearing and avoid dictation.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;p&gt;Twelve short tasks and nine long ones, three trials per cell, one random seed for the typo generator. The voice-to-text prompts are hand-simulated, not real speech engine output. Claude models were called through the Claude Code CLI with a minimal system prompt and tools disabled. Local models ran at Ollama defaults with thinking off. The harness and every prompt and reply are in the repo linked below, so you can rerun it with your own tasks.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Repo with the harness, all prompts, and every raw reply: &lt;a href="https://github.com/valbarov/prompt-noise-test" rel="noopener noreferrer"&gt;github.com/valbarov/prompt-noise-test&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>prompting</category>
      <category>testing</category>
    </item>
    <item>
      <title>12 AWS Defaults That Ship Insecure (and the One Line That Fixes Each)</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Sun, 27 Sep 2026 01:19:02 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/12-aws-defaults-that-ship-insecure-and-the-one-line-that-fixes-each-127f</link>
      <guid>https://dev.to/vadim_albarov/12-aws-defaults-that-ship-insecure-and-the-one-line-that-fixes-each-127f</guid>
      <description>&lt;p&gt;Every AWS default answers exactly one question: will the tutorial work?&lt;/p&gt;

&lt;p&gt;No surprise bill. No failed API call. No "access denied" on step three. That is a great default for a tutorial. It is a terrible default for the thing you spun up "just for staging" that is now production, on the day someone asks who downloaded that bucket in March.&lt;/p&gt;

&lt;p&gt;I run infrastructure for a healthcare company. People with clipboards read my configs. Somewhere around my third Terraform module I noticed my job had a shape, and the shape was not "design clever things". It was flipping the same switches AWS had left in the wrong position, over and over. So I wrote them down.&lt;/p&gt;

&lt;p&gt;Here are the twelve, ranked by how likely each one is to hurt you. One rule of thumb before we start: &lt;code&gt;checkov&lt;/code&gt; and &lt;code&gt;tfsec&lt;/code&gt; check what you wrote. Most of this list is about what you did not write.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. CloudTrail remembers 90 days, then forgets
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; you get "Event history": 90 days of management events, viewable in the console. No trail to S3 unless you create one. And S3 object reads and writes (data events) are not logged at all, even when a trail exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; a credential gets phished in March. In September someone asks what it touched. The window scrolled off in June, and even inside the window "did they download the sensitive bucket" was never recorded. Your answer is a shrug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_cloudtrail"&lt;/span&gt; &lt;span class="s2"&gt;"audit"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;                       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"org-audit"&lt;/span&gt;
  &lt;span class="nx"&gt;s3_bucket_name&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_s3_bucket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;audit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;is_multi_region_trail&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="nx"&gt;enable_log_file_validation&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

  &lt;span class="nx"&gt;event_selector&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;read_write_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"All"&lt;/span&gt;
    &lt;span class="nx"&gt;data_resource&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;type&lt;/span&gt;   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"AWS::S3::Object"&lt;/span&gt;
      &lt;span class="nx"&gt;values&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:s3:::my-sensitive-bucket/"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Data events cost real money at volume. Scope them on purpose, not by forgetting.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. RDS storage is unencrypted, and you cannot fix that later
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; &lt;code&gt;storage_encrypted = false&lt;/code&gt; in Terraform and in the API. The console nudges you. Code does not. And encryption can only be enabled at creation. Later means snapshot, encrypted copy, restore, cutover.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; this is the worst one on the list because it is irreversible in place. Staging quietly becomes production, it always does, and eighteen months later a review finds your primary database in plaintext. The remediation is a maintenance window on your busiest system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;storage_encrypted&lt;/span&gt; &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nx"&gt;kms_key_id&lt;/span&gt;        &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_kms_key&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In my modules this is hardcoded, not a variable. A security invariant that callers can toggle off is not an invariant. It is a default waiting to come back.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Your Postgres accepts plaintext connections
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; on RDS for PostgreSQL 14 and earlier, &lt;code&gt;rds.force_ssl&lt;/code&gt; is 0. The server supports TLS. It also cheerfully accepts connections without it. AWS flipped this to 1 for PostgreSQL 15+, which tells you what AWS thinks of the old default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; any client on &lt;code&gt;sslmode=prefer&lt;/code&gt; (the libpq default) falls back to plaintext if the handshake hiccups. Nothing fails. Nothing logs. Your transmission security now depends on every developer, sidecar, and ad hoc &lt;code&gt;psql&lt;/code&gt; from a bastion remembering a flag.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; enforce it at the engine so client config stops mattering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;parameter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"rds.force_ssl"&lt;/span&gt;
  &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"1"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4. Your load balancer still shakes hands with 2008
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; create an HTTPS listener without a policy and you get &lt;code&gt;ELBSecurityPolicy-2016-08&lt;/code&gt;, which accepts TLS 1.0 and 1.1. Also off by default on the same resource: access logs and deletion protection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; nobody notices, because modern browsers pick 1.2+. Then a customer's security team runs an external scan before signing, and the deal stalls on a finding one attribute would have prevented.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;ssl_policy&lt;/span&gt; &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"ELBSecurityPolicy-TLS13-1-2-2021-06"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And on the &lt;code&gt;aws_lb&lt;/code&gt; itself: &lt;code&gt;access_logs { enabled = true }&lt;/code&gt;, &lt;code&gt;enable_deletion_protection = true&lt;/code&gt;, and &lt;code&gt;drop_invalid_header_fields = true&lt;/code&gt; while you are there.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. VPC Flow Logs do not exist
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; no VPC has flow logs until you create them. No network-level record of anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; GuardDuty flags an instance talking to a known-bad IP. The first question is "since when, and what else did it talk to?" Without flow logs your incident response runs on vibes, and "was anything exfiltrated?" gets answered conservatively, which means expensively.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_flow_log"&lt;/span&gt; &lt;span class="s2"&gt;"vpc"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt;               &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;traffic_type&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"ALL"&lt;/span&gt;
  &lt;span class="nx"&gt;log_destination_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"cloud-watch-logs"&lt;/span&gt;
  &lt;span class="nx"&gt;log_destination&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_cloudwatch_log_group&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;flow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
  &lt;span class="nx"&gt;iam_role_arn&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_iam_role&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;flow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ACCEPT&lt;/code&gt; only or &lt;code&gt;REJECT&lt;/code&gt; only is half an audit trail.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Services create log groups you never asked for
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; log groups default to Never Expire and an AWS-managed key. The sneaky part: RDS log exports, Container Insights, and Lambda create their own log groups on first write, with exactly those defaults, outside your Terraform state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; you enable RDS log exports. RDS creates &lt;code&gt;/aws/rds/instance/.../postgresql&lt;/code&gt; for you. Three years later that orphan holds three years of connection logs plus whatever your app leaked into query errors, retained forever, under a key you did not choose, and invisible to &lt;code&gt;terraform destroy&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; pre-create every log group a service will write to, before the service exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_cloudwatch_log_group"&lt;/span&gt; &lt;span class="s2"&gt;"rds"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;              &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/aws/rds/instance/myapp-db/postgresql"&lt;/span&gt;
  &lt;span class="nx"&gt;retention_in_days&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2192&lt;/span&gt;
  &lt;span class="nx"&gt;kms_key_id&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_kms_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  7. The default security group is a party where everyone knows everyone
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; every VPC ships with a default security group that allows all traffic between members and all outbound. Anything launched without an explicit group lands in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; a contractor spins up a "quick utility box" without thinking about security groups. It joins the default one, alongside everything else that drifted in over the years. That box gets popped, and lateral movement is free, because membership is the authorization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; you cannot delete the default group, but you can strip it bare. A resource with no ingress or egress blocks removes every rule, so anything landing there can talk to precisely nothing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_default_security_group"&lt;/span&gt; &lt;span class="s2"&gt;"this"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;tags&lt;/span&gt;   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"default-DO-NOT-USE"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Loud failure beats silent success.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. RDS backups: one day via the API, zero via Terraform
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; backup retention is one day via API or CLI, seven via the console, and &lt;code&gt;backup_retention_period&lt;/code&gt; defaults to 0 in Terraform. Zero. Automated backups off. Deletion protection is also off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; an instance defined without that line has no backups at all, and it passes plan, apply, and review, because absence does not show up in a diff. You find out during your first real restore attempt, which is the single worst moment available. Bonus: with deletion protection off, one &lt;code&gt;terraform destroy&lt;/code&gt; in the wrong workspace takes the database and its backups together.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;backup_retention_period&lt;/span&gt;   &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;
&lt;span class="nx"&gt;deletion_protection&lt;/span&gt;       &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nx"&gt;skip_final_snapshot&lt;/span&gt;       &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="nx"&gt;final_snapshot_identifier&lt;/span&gt; &lt;span class="err"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"myapp-db-final"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then AWS Backup with a locked vault on top, because backups attached to the instance share the instance's blast radius.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. EBS encryption by default is off, in every region separately
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; account-level EBS encryption by default is disabled, and it is a per-region setting. Turning it on in us-east-1 does nothing for us-west-2.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; your Terraform encrypts every volume it manages. Then someone launches a console instance for a one-off migration, copies a database dump onto it, and that volume is plaintext, because the account default governs ad hoc resources, not your module.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; one resource, once per region you use, including the ones you think you do not use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ebs_encryption_by_default"&lt;/span&gt; &lt;span class="s2"&gt;"this"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;enabled&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  10. SNS stores your alerts in plaintext
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; server-side encryption on SNS topics is off until you set a KMS key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; mostly as an audit finding. Occasionally worse: a team wires appointment reminders through SNS, and now message bodies with patient data sit unencrypted in the messaging layer while every database in the stack is dutifully encrypted. One-line fix or an hours-long finding memo. Pick the line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_sns_topic"&lt;/span&gt; &lt;span class="s2"&gt;"alerts"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;              &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"security-alerts"&lt;/span&gt;
  &lt;span class="nx"&gt;kms_master_key_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_kms_key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Real gotcha: once the topic is encrypted, CloudWatch and EventBridge need &lt;code&gt;kms:Decrypt&lt;/code&gt; and &lt;code&gt;kms:GenerateDataKey*&lt;/code&gt; in the key policy, not in IAM. Otherwise every alarm publish fails silently. Test the path end to end.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. S3 got fixed, which is exactly the problem
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default (current):&lt;/strong&gt; since January 2023 every new object is encrypted with SSE-S3. Since April 2023 new buckets get Block Public Access on and ACLs off. Credit where due: the two most famous S3 footguns are gone for new buckets. Still off: server access logging, versioning, and any customer-managed key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; "S3 encrypts by default now" becomes the reason nobody configures anything further. A sensitive bucket ends up with no access log, no versioning, no policy denying plaintext transport, and a key nobody can audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; for a bucket that matters: SSE-KMS with your own key, versioning on, &lt;code&gt;aws_s3_bucket_logging&lt;/code&gt; to a separate log bucket, and a bucket policy denying &lt;code&gt;aws:SecureTransport = false&lt;/code&gt;. The 2023 change is your floor, not your control.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. ECS Exec gives you a shell, and records nothing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Default:&lt;/strong&gt; ECS Exec logging is &lt;code&gt;DEFAULT&lt;/code&gt;, meaning "whatever awslogs config the task has". If the container has none, session input and output are not logged at all. KMS encryption of the session channel is opt-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it bites:&lt;/strong&gt; an engineer execs into a production task to debug, runs a few queries, pastes results somewhere. Months later an access review needs that session. CloudTrail says a shell was opened at 14:32. That is the entire record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; on the cluster, &lt;code&gt;execute_command_configuration&lt;/code&gt; with &lt;code&gt;logging = "OVERRIDE"&lt;/code&gt;, a dedicated encrypted log group, and a &lt;code&gt;kms_key_id&lt;/code&gt; for the channel. Every session becomes a transcript.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automate it or relive it forever
&lt;/h2&gt;

&lt;p&gt;Twelve items is too many for humans to re-run reliably. That is the real lesson. What works for me:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One baseline module&lt;/strong&gt; for the apply-once account controls: the trail, AWS Config rules like &lt;code&gt;RDS_STORAGE_ENCRYPTED&lt;/code&gt; and &lt;code&gt;VPC_FLOW_LOGS_ENABLED&lt;/code&gt;, EBS default encryption, the account-level S3 public access block.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardcoded invariants&lt;/strong&gt; in workload modules. &lt;code&gt;storage_encrypted&lt;/code&gt;, &lt;code&gt;publicly_accessible = false&lt;/code&gt;, forced TLS, and log validation are not variables in mine. Exposing them as inputs just recreates the permissive default one level up, with your name on it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy as code for the rest.&lt;/strong&gt; &lt;code&gt;checkov&lt;/code&gt; in CI catches regressions in code. AWS Config catches the console-created, script-created, "temporary" resources your code never met. You need both, because the resources that hurt most never saw a pull request.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The pattern across all twelve is the same. AWS optimizes for the first five minutes. Your auditor cares about year five, when someone asks who accessed what and the answer has to exist. Those two goals produce opposite defaults, and closing the gap is, by AWS's own shared responsibility model, your job.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Not legal or compliance advice. I am an infrastructure engineer, and defaults change, sometimes even in the right direction, so check each one against current docs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Which default bit you? I only know the twelve that bit me.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>security</category>
      <category>beginners</category>
      <category>devops</category>
    </item>
    <item>
      <title>Claude Code's Max Effort Cost Me 8x More. It Wrote the Same Code.</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Tue, 15 Sep 2026 21:44:18 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/claude-codes-max-effort-cost-me-8x-more-it-wrote-the-same-code-23hc</link>
      <guid>https://dev.to/vadim_albarov/claude-codes-max-effort-cost-me-8x-more-it-wrote-the-same-code-23hc</guid>
      <description>&lt;p&gt;Claude Code has an &lt;code&gt;/effort&lt;/code&gt; command with five levels: low, medium, high, xhigh, max. The help text for low says "quick, straightforward implementation with minimal overhead." The natural assumption is that the higher levels buy you better code, and that the price is time and tokens.&lt;/p&gt;

&lt;p&gt;I wanted to see that trade-off on a chart. So I built a small harness, gave every level the same task in a clean copy of the same folder, and measured. I did it three times with three different task shapes.&lt;/p&gt;

&lt;p&gt;The chart never appeared. Every level got the same score on every task. Only the clock and the invoice moved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The harness
&lt;/h2&gt;

&lt;p&gt;Nothing clever. One headless Claude Code process per level, running in its own copy of the project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;prompt.txt&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--effort&lt;/span&gt; low &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output-format&lt;/span&gt; stream-json &lt;span class="nt"&gt;--verbose&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--permission-mode&lt;/span&gt; acceptEdits &lt;span class="nt"&gt;--allowedTools&lt;/span&gt; &lt;span class="s2"&gt;"Read,Write,Edit,Bash(python*)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The stream-json output goes to a log file. A small viewer prints readable progress in a visible PowerShell window, so I could watch five terminals work through the same problem one after another. When all runs finish, an analyzer parses the logs for turns, tool calls, cost, and token counts, then runs a hidden test suite against each folder. The child sessions never see the hidden tests.&lt;/p&gt;

&lt;p&gt;I ran each level once. Keep that in mind for every number below. One sample is enough to show a 9x gap. It is not enough to tell a 10 percent difference from noise.&lt;/p&gt;

&lt;p&gt;Model was Fable 5.1 on every run. Claude Code 2.1.272.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark 1: implement from docstrings
&lt;/h2&gt;

&lt;p&gt;A tiny Python order-processing library: money parsing, coupons, tax, inventory with reservations. Nine functions and one class, all stubbed out with &lt;code&gt;NotImplementedError&lt;/code&gt; and detailed docstrings. Three public tests visible. 57 hidden tests.&lt;/p&gt;

&lt;p&gt;The prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Implement every function and method marked &lt;code&gt;# TODO&lt;/code&gt; in the &lt;code&gt;shop&lt;/code&gt; package exactly according to the docstrings. Do not change the docstrings, models.py, or anything in tests/. Run &lt;code&gt;python -m pytest -q&lt;/code&gt; to check your work. When done, reply with a one-line summary.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;effort&lt;/th&gt;
&lt;th&gt;turns&lt;/th&gt;
&lt;th&gt;wall time&lt;/th&gt;
&lt;th&gt;cost&lt;/th&gt;
&lt;th&gt;output tokens&lt;/th&gt;
&lt;th&gt;hidden tests&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;77 s&lt;/td&gt;
&lt;td&gt;$0.59&lt;/td&gt;
&lt;td&gt;5,174&lt;/td&gt;
&lt;td&gt;57/57&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;69 s&lt;/td&gt;
&lt;td&gt;$0.59&lt;/td&gt;
&lt;td&gt;4,995&lt;/td&gt;
&lt;td&gt;57/57&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;174 s&lt;/td&gt;
&lt;td&gt;$1.22&lt;/td&gt;
&lt;td&gt;13,009&lt;/td&gt;
&lt;td&gt;57/57&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xhigh&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;297 s&lt;/td&gt;
&lt;td&gt;$2.04&lt;/td&gt;
&lt;td&gt;24,919&lt;/td&gt;
&lt;td&gt;57/57&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;693 s&lt;/td&gt;
&lt;td&gt;$4.72&lt;/td&gt;
&lt;td&gt;62,196&lt;/td&gt;
&lt;td&gt;57/57&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every level passed everything on the first try. Max took eleven and a half minutes and $4.72 to arrive at the same 57/57 that low reached in 77 seconds for 59 cents. Output tokens grew 12x. Tool calls grew only about 3x, so most of that growth is reasoning, not doing.&lt;/p&gt;

&lt;p&gt;Low and medium were near clones of each other. Same number of turns, same sequence: read, write three files, run pytest, one sanity check, done. The diff between their outputs is 41 lines of reordering.&lt;/p&gt;

&lt;p&gt;The code was not identical across levels, though. Since the tests could not separate them, I diffed the implementations by hand.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Low and medium&lt;/strong&gt; parsed money strings by stripping the sign and the dollar symbol and checking each character. They delete commas, so "1,23.45" is accepted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High&lt;/strong&gt; wrote one verbose regex with named groups that enforces comma groups of three.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Xhigh and max&lt;/strong&gt; factored parsing into helpers. Max also widened &lt;code&gt;Decimal&lt;/code&gt; precision with a &lt;code&gt;localcontext&lt;/code&gt; for very large amounts. No test covers that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tax rounding&lt;/strong&gt;: low and medium used &lt;code&gt;Decimal&lt;/code&gt; with &lt;code&gt;quantize&lt;/code&gt;. High, xhigh and max switched to &lt;code&gt;Fraction&lt;/code&gt; for exact rational math and a hand-written half-up round.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max&lt;/strong&gt; added validation nobody asked for. The inventory constructor rejects negative stock, and reserve rejects quantities below one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So higher effort produced more defensive, more factored code. It did not produce more correct code, because there was nothing left to be more correct about.&lt;/p&gt;

&lt;p&gt;One detail I liked: the only actual spec deviation came from &lt;strong&gt;high&lt;/strong&gt;, not low. The docstring said to raise &lt;code&gt;KeyError&lt;/code&gt; with the coupon code as the message. High raised it with the uppercased code. My test used an all-caps code, so it passed anyway. Effort level did not predict spec fidelity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark 2: port it to Go
&lt;/h2&gt;

&lt;p&gt;The first task was too easy. A port to another language has real traps: Python has &lt;code&gt;Decimal&lt;/code&gt;, &lt;code&gt;Fraction&lt;/code&gt;, arbitrary ints and exceptions. Go has none of those. Errors are values, half-up rounding is manual, and the model has to decide how an &lt;code&gt;OutOfStock&lt;/code&gt; error carries its fields.&lt;/p&gt;

&lt;p&gt;Same harness. The base folder held the Python reference source and an empty Go module. Hidden Go tests import a fixed package path. The analyzer also runs &lt;code&gt;go vet&lt;/code&gt; and &lt;code&gt;gofmt -l&lt;/code&gt; for free quality signals.&lt;/p&gt;

&lt;p&gt;I ran low, medium and high, then stopped to save budget.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;effort&lt;/th&gt;
&lt;th&gt;turns&lt;/th&gt;
&lt;th&gt;wall time&lt;/th&gt;
&lt;th&gt;cost&lt;/th&gt;
&lt;th&gt;output tokens&lt;/th&gt;
&lt;th&gt;hidden tests&lt;/th&gt;
&lt;th&gt;gofmt&lt;/th&gt;
&lt;th&gt;Go lines&lt;/th&gt;
&lt;th&gt;test lines&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;135 s&lt;/td&gt;
&lt;td&gt;$1.04&lt;/td&gt;
&lt;td&gt;11,175&lt;/td&gt;
&lt;td&gt;all pass&lt;/td&gt;
&lt;td&gt;ok&lt;/td&gt;
&lt;td&gt;442&lt;/td&gt;
&lt;td&gt;148&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;401 s&lt;/td&gt;
&lt;td&gt;$2.60&lt;/td&gt;
&lt;td&gt;31,503&lt;/td&gt;
&lt;td&gt;all pass&lt;/td&gt;
&lt;td&gt;test file unformatted&lt;/td&gt;
&lt;td&gt;498&lt;/td&gt;
&lt;td&gt;339&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;470 s&lt;/td&gt;
&lt;td&gt;$3.44&lt;/td&gt;
&lt;td&gt;38,673&lt;/td&gt;
&lt;td&gt;all pass&lt;/td&gt;
&lt;td&gt;ok&lt;/td&gt;
&lt;td&gt;574&lt;/td&gt;
&lt;td&gt;420&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Again, flat on correctness. Cost roughly linear in effort. What each level did with its budget was different, and this part was genuinely interesting to watch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Low&lt;/strong&gt; read the sources, wrote five files, ran vet and test, and stopped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Medium&lt;/strong&gt; wanted to probe the Python original before writing Go, to confirm exact rounding and error behavior. Good instinct. But my tool allowlist for that run did not include Python, so it retried across PowerShell and Bash several times before giving up. That is harness noise, not effort. It also wrote a test suite more than twice the size of low's, and never ran gofmt on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;High&lt;/strong&gt;, with Python allowed, built its own cross-language oracle. It wrote a Python script that ran the original code over many inputs, generated a Go test file from the results, ran that against its port, and then deleted the generated files. That is exactly the grading technique I had planned to build myself. High invented it unprompted, which is the kind of behavior you would hope a higher effort setting buys.&lt;/p&gt;

&lt;p&gt;It just did not change the score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark 3: bug hunt with no execution
&lt;/h2&gt;

&lt;p&gt;Both tasks so far had the same two properties: the spec was complete, and verification was cheap. Any level can iterate against a test suite until green. Extra thinking has nothing to buy.&lt;/p&gt;

&lt;p&gt;So I broke both. A working Go version of the library with 8 planted bugs: a wrong rounding constant, bad comma grouping, a sign error in refund splits, case-sensitive coupon codes, percentage coupons discounting gift cards, tax applied to exempt categories, a non-atomic inventory reserve, and a double-release that inflates stock. The prompt gave 8 vague user reports, one per bug, and one hard rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You cannot run any commands in this session (no go build, no tests). Reason carefully from the code.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tool allowlist: Read, Grep, Glob, Edit. Nothing else.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;effort&lt;/th&gt;
&lt;th&gt;turns&lt;/th&gt;
&lt;th&gt;wall time&lt;/th&gt;
&lt;th&gt;cost&lt;/th&gt;
&lt;th&gt;output tokens&lt;/th&gt;
&lt;th&gt;bugs fixed&lt;/th&gt;
&lt;th&gt;diff&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;52 s&lt;/td&gt;
&lt;td&gt;$0.94&lt;/td&gt;
&lt;td&gt;3,677&lt;/td&gt;
&lt;td&gt;7/8&lt;/td&gt;
&lt;td&gt;+20 / -5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;55 s&lt;/td&gt;
&lt;td&gt;$0.86&lt;/td&gt;
&lt;td&gt;3,585&lt;/td&gt;
&lt;td&gt;7/8&lt;/td&gt;
&lt;td&gt;+16 / -3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;69 s&lt;/td&gt;
&lt;td&gt;$0.94&lt;/td&gt;
&lt;td&gt;5,055&lt;/td&gt;
&lt;td&gt;7/8&lt;/td&gt;
&lt;td&gt;+26 / -4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xhigh&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;125 s&lt;/td&gt;
&lt;td&gt;$1.30&lt;/td&gt;
&lt;td&gt;10,200&lt;/td&gt;
&lt;td&gt;7/8&lt;/td&gt;
&lt;td&gt;+21 / -4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Finally a miss. All four levels missed the same bug, the tax one, and all four made the identical wrong fix. When I looked closely, that was my fault.&lt;/p&gt;

&lt;p&gt;The correct rule was proportional: remove the exempt fraction of the pre-discount subtotal from the post-discount taxable amount. That rule lived only in the Python docstring, which was not in the Go project. From the Go code and the report "grocery and book orders are being taxed," subtracting the exempt amount directly is a completely reasonable fix. Four independent runs converging on it is evidence that the intended rule was not inferable from what they were given. It is not evidence that effort failed.&lt;/p&gt;

&lt;p&gt;So the fair score is 7 of 7 for every level. Flat again.&lt;/p&gt;

&lt;p&gt;What effort did change here was, once more, output tokens: 3.6k, 3.6k, 5k, 10k. Wall time and cost followed. And scope creep showed up at low, high and xhigh, each of which invented a fuzzy gift-card category matcher that handles "Gift Card", "gift_card" and "GIFT-CARD" when the original code used an exact string. Medium alone made the exact minimal change the prompt asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I take from this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Effort controls how much the model deliberates, not what it knows.&lt;/strong&gt; On every task here the answer was fully determined and reachable in one careful pass. Given that, low found it. Max found it too, after thinking about it for eleven minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification loops make effort irrelevant.&lt;/strong&gt; If the session can run tests, any level will iterate until green. The Python and Go benchmarks measured persistence, and persistence was free at every level.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Higher effort spends its budget on self-checks and defensive code.&lt;/strong&gt; Oracles, throwaway verification scripts, extra validation, regexes that reject malformed input the spec never mentioned. Sometimes that is exactly what you want. In a codebase with a scope rule, it is unrequested behavior you now have to review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Medium was the quiet winner on discipline.&lt;/strong&gt; It never won on speed or cost, but on the bug hunt it was the only level that did precisely what the prompt asked and nothing more.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The total bill for this experiment was about $20.&lt;/strong&gt; Roughly $9 for the Python runs, $7 for the Go runs, $4 for the bug hunt. The max Python run alone was almost a quarter of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where effort probably should matter
&lt;/h2&gt;

&lt;p&gt;I have not proven that effort is useless. I have shown three tasks where it did not help, and I think I know why. Effort should matter when deliberation itself is the bottleneck:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two-step inference.&lt;/strong&gt; The symptom is in one file and the cause is in another, and the code near the symptom looks fine. A fix at the symptom passes the reported case and fails hidden siblings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A decoy fix that breaks a hidden invariant.&lt;/strong&gt; An inventory patch that resolves the report and introduces a race under &lt;code&gt;go test -race&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A constraint hidden in a README that conflicts with the obvious fix.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A codebase too large to read.&lt;/strong&gt; Around 2,000 lines, where effort shows in search strategy and in what gets verified before editing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeat runs.&lt;/strong&gt; Three samples per level at minimum. My single-sample wall times have variance comparable to the gaps between adjacent levels.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the next benchmark. Until then, my working rule: for a well-specified implementation task where the session can run tests, low or medium. Reserve high and above for tasks where the model cannot check itself and the answer is not sitting in the docstring.&lt;/p&gt;

&lt;p&gt;And watch the clock. Nothing on max was wrong. It just took nine times longer to be right.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Your AI Knows the Rule. Your Question Decides Whether You Get the Myth</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:30:00 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/your-ai-knows-the-rule-your-question-decides-whether-you-get-the-myth-5544</link>
      <guid>https://dev.to/vadim_albarov/your-ai-knows-the-rule-your-question-decides-whether-you-get-the-myth-5544</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HIPAA sets no retention period for audit logs. The six-year rule covers documentation. Six of the nine results on Google's first page say otherwise, and so did Gemini, eight times out of nine, citing only vendor pages.&lt;/li&gt;
&lt;li&gt;Reworded, the same models were right six times out of seven. Naming the regulation section in the question worked every time, with web search on or off.&lt;/li&gt;
&lt;li&gt;Asking the model to check its own answer made it worse. Each challenge produced a more specific regulator that does not exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If you have never touched healthcare software: HIPAA is the US federal law on health data. Its Security Rule tells hospitals, insurers and every vendor that handles their electronic patient data, called ePHI, what safeguards to have. The HHS Office for Civil Rights enforces it, and the fines are real. The rule itself is short and technology-neutral, which is why so much of what people "know" about it comes from vendors rather than from the text.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;How long does HIPAA require audit logs to be retained?&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everyone in healthcare IT knows the answer: six years. It is in every vendor blog, every compliance checklist, every SIEM sales deck. It is not in the regulation.&lt;/p&gt;

&lt;p&gt;The audit controls standard, &lt;a href="https://www.law.cornell.edu/cfr/text/45/164.312" rel="noopener noreferrer"&gt;45 CFR 164.312(b)&lt;/a&gt;, is one sentence: implement mechanisms that record and examine activity in systems holding ePHI. No period. The only other mention of audit logs, &lt;a href="https://www.law.cornell.edu/cfr/text/45/164.308" rel="noopener noreferrer"&gt;164.308(a)(1)(ii)(D)&lt;/a&gt;, says to review them regularly. No period. Retention is a risk-analysis decision that you write into policy.&lt;/p&gt;

&lt;p&gt;The six years is real but belongs elsewhere. &lt;a href="https://www.law.cornell.edu/cfr/text/45/164.316" rel="noopener noreferrer"&gt;164.316(b)(2)(i)&lt;/a&gt; says to keep &lt;em&gt;documentation&lt;/em&gt;, meaning policies, procedures and written records of required actions, for six years. Two more six-year clocks feed the lore: the &lt;a href="https://www.law.cornell.edu/cfr/text/45/164.528" rel="noopener noreferrer"&gt;accounting of disclosures&lt;/a&gt; lookback, and the &lt;a href="https://www.law.cornell.edu/cfr/text/45/160.414" rel="noopener noreferrer"&gt;six years&lt;/a&gt; OCR has to bring a penalty action. Many organizations pick six years for logs because of that last one. That is prudence, not a mandate.&lt;/p&gt;

&lt;p&gt;Correct answer: HIPAA sets no log retention period, six years covers documentation, decide log retention in your risk analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scoreboard
&lt;/h2&gt;

&lt;p&gt;Fresh chat, default settings, the neutral question, graded against the regulation text. LORE means the answer states the six-year log rule as a HIPAA requirement.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assistant&lt;/th&gt;
&lt;th&gt;Runs&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini web app, 3.6 Flash and 3.1 Pro with reasoning&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;LORE, LORE, LORE, LORE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini API, 3.6 Flash, web search off&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;LORE, LORE, LORE, partial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini API, 3.8 Flash, web search off&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;LORE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT, default model (the UI shows no model label), web search on&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5, high effort&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3 14B, local, no web&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;no myth, but 3 invented authorities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma4 4B, local, no web&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;no myth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Disclosure: Claude Code is my daily tool and got the same rubric. Its answer opened with the bold words "Six years" before correcting itself in the same sentence, and it claimed "regulators treat" logs as documentation, which regulators have never said. That sentence, it turned out, came straight from one of its sources. More on that below.&lt;/p&gt;

&lt;p&gt;The two small local models never stated the myth. A 4B model on my laptop with no internet beat Gemini Pro with reasoning and web search. But qwen3 filled the gap with authorities that do not exist: an HHS recommendation, a NIST document, an OCR position. Right conclusion, invented support. Hold that thought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I changed the question
&lt;/h2&gt;

&lt;p&gt;Two more wordings of the same legal question, on the same Gemini models, the same day.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wording&lt;/th&gt;
&lt;th&gt;Runs&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"How long does HIPAA require..."&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;8 LORE, 1 partial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Does 45 CFR 164.312(b) set a period..."&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4 correct, search on and off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Is it true that HIPAA requires six years..."&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2 correct with search off, 1 LORE with search on&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With search off, the "is it true" runs open with "this is a common misconception" and then explain, correctly, why everyone says six years anyway. The model knew all of it. With search on, the same sentence got "Yes, but with an important legal nuance", followed by "HHS auditors and industry standards universally map audit logs to the six-year documentation rule." They do not.&lt;/p&gt;

&lt;p&gt;The wording that named the section worked everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Fresh chat, default settings, no custom instructions. Run these in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How long does HIPAA require audit logs to be retained?&lt;/li&gt;
&lt;li&gt;Is it true that HIPAA requires six years of audit log retention?&lt;/li&gt;
&lt;li&gt;Does 45 CFR 164.312(b) set a retention period for audit logs?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Right: the audit controls standard sets no period, and six years is documentation under 164.316. Wrong: HIPAA requires six years of audit logs. If your assistant searches the web, expect the second one to still say yes. The third flipped every time for me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the wording matters more than the model
&lt;/h2&gt;

&lt;p&gt;Gemini's answers are grounded on Google Search. Google's own documentation says the model "automatically generates one or multiple search queries and executes them" (&lt;a href="https://ai.google.dev/gemini-api/docs/grounding" rel="noopener noreferrer"&gt;Grounding with Google Search&lt;/a&gt;). So Gemini's sources are whatever Google ranks for the queries the model writes. Here is what each assistant cited for the neutral question.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assistant&lt;/th&gt;
&lt;th&gt;Sources shown under the answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini web, 4 runs&lt;/td&gt;
&lt;td&gt;Pangea (audit-log service), IS Partners (compliance consultancy), Cayosoft (Active Directory audit tool), Aptible (hosting). Vendors only. No primary source in any run that showed sources.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT, search on&lt;/td&gt;
&lt;td&gt;hhs.gov&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5, search on&lt;/td&gt;
&lt;td&gt;USA HIPAA (training and compliance kits), HIPAA Auditors (audit services). Vendors, both accurate.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local models&lt;/td&gt;
&lt;td&gt;none, no web access&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One of Gemini's vendor pages says, word for word, "HIPAA mandates that audit logs must be retained for at least six years, as per 45 C.F.R. § 164.316(b)(2)(i)." Gemini's headline was that sentence with the nouns rearranged. Claude did the same thing with better pages. Its closing line, "HIPAA does not set a single universal log retention period", is nearly word for word from one of its two sources, and the one flaw in its answer, the claim that "regulators treat" logs as documentation, is a sentence from the other. Three grounded assistants, three summaries of vendor pages. The verdict tracked the pages, not the model.&lt;/p&gt;

&lt;p&gt;Then I searched Google for the exact question. Page one, 6 September 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rank&lt;/th&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;What the snippet says&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;HIPAA Journal&lt;/td&gt;
&lt;td&gt;documents must be kept six years (accurate, and about documentation)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Reddit r/cybersecurity&lt;/td&gt;
&lt;td&gt;SIEM log retention thread&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Kiteworks, vendor&lt;/td&gt;
&lt;td&gt;"the HIPAA minimum of six years" for audit logs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Schellman, audit firm&lt;/td&gt;
&lt;td&gt;"all audit logs ... at least 6 years"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Aptible, vendor&lt;/td&gt;
&lt;td&gt;"HIPAA requires six years of audit log retention"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Columbia University policy&lt;/td&gt;
&lt;td&gt;HIPAA documents six years (accurate)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;SecurityMetrics, vendor&lt;/td&gt;
&lt;td&gt;logs "retained for at least six years"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;ChartRequest, vendor&lt;/td&gt;
&lt;td&gt;"Audit logs must be retained for a minimum of six years, as outlined under 45 CFR § 164.316(b)(2)(i)"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;UTMStack, vendor&lt;/td&gt;
&lt;td&gt;"HIPAA's core retention rule is a 6-year minimum"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Six of nine state the six-year log rule as law. Five are vendors whose product benefits from it. Zero primary sources. HHS has no page that ranks here, because the true answer is "there is no number", and nobody writes a ranking page about the absence of a number except the people who would rather you kept six years of logs on their platform. The truth has no marketing budget.&lt;/p&gt;

&lt;p&gt;Two things follow. The myth is in the training data: with search off, the neutral question still produced six years four times out of five, with the right section number attached to the wrong claim. And retrieval enforces it: with search on, the only wording that changed the answer was the one that changed what got retrieved. A question with the section number in it pulls pages about the section, and those pages, vendors included, say it sets no period. "Is it true that HIPAA requires" pulls the same pages as "how long does HIPAA require", and those pages say six years.&lt;/p&gt;

&lt;p&gt;Even the right answers invented authority. One said "HHS has clarified" that raw logs are not documentation. HHS has clarified no such thing. The confident-specifics problem does not go away when the conclusion is right. It stops being visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not ask it to check itself
&lt;/h2&gt;

&lt;p&gt;The obvious objection: push back. So I did, in one Gemini chat with search on.&lt;/p&gt;

&lt;p&gt;First message, the neutral question: six years, plus "HHS interprets audit controls as part of required HIPAA documentation." No such interpretation exists.&lt;/p&gt;

&lt;p&gt;Second message, "is it true": "Yes, it is true." Now "HHS and compliance auditors treat system audit logs as official documentation. Consequently, covered entities must retain these logs for the 6-year period."&lt;/p&gt;

&lt;p&gt;Third message, "check again against the official resources": it quoted both sections verbatim and correctly, wrote that the audit controls standard "contains no timeframe", and one paragraph later: "OCR enforcement applies the 6-year rule to system access logs." No OCR action, guidance or FAQ says that.&lt;/p&gt;

&lt;p&gt;The sources were the same vendor pages each time, so retrieval explains the "yes". It does not explain the escalation. Each challenge made the invented regulator more specific, and the third answer contradicted the text it had just quoted. Once the model has committed, a challenge is a request to justify, not to verify, and it invents the authority it needs. I saw the same pattern with a different assistant last month, when "support this with references" produced a fabricated HHS quote pinned to a real HHS URL. It is not a product quirk.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do now
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Name the section.&lt;/strong&gt; "Does 45 CFR 164.312(b) set a retention period" worked with search on and off, because it changes which pages the model reads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat "is it true" as a partial fix.&lt;/strong&gt; It works when the model answers from memory. With web search on, the vendor pages come back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never ask it to check itself.&lt;/strong&gt; Open a new chat and ask the anchored question there. The chat with the wrong answer in it is poisoned.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the sources first.&lt;/strong&gt; If every link sells the thing the answer says you need, ask again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the section.&lt;/strong&gt; One sentence. Faster than any of the chats.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Post your product, the model label the UI shows, the wording number, and which way it landed. I want to know whether the flip reproduces beyond my machine.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I am an infrastructure engineer, not a lawyer. This is an engineering read of the published regulation, not legal advice. Answers were collected on 6 and 7 September 2026 from consumer web apps, the Gemini API and local models; models and retrieval change constantly, which is part of the point.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Appendix: the pages the assistants cited, with the exact sentence&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The pages that state the six-year log rule are named but not linked. A link is a ranking signal, and ranking is the problem. The four that got it right are linked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cited under the six-year answers&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Page&lt;/th&gt;
&lt;th&gt;What they sell&lt;/th&gt;
&lt;th&gt;Date on page&lt;/th&gt;
&lt;th&gt;Exact sentence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;pangea.cloud, "HIPAA audit log requirements"&lt;/td&gt;
&lt;td&gt;managed audit-log service&lt;/td&gt;
&lt;td&gt;Aug 2024&lt;/td&gt;
&lt;td&gt;"HIPAA mandates that audit logs must be retained for at least six years, as per 45 C.F.R. § 164.316(b)(2)(i)."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cayosoft.com, "HIPAA audit log requirements"&lt;/td&gt;
&lt;td&gt;Active Directory audit software&lt;/td&gt;
&lt;td&gt;Sep 2025&lt;/td&gt;
&lt;td&gt;"HIPAA requires six years of audit log retention." No section cited.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ispartnersllc.com, "HIPAA audit log retention six years"&lt;/td&gt;
&lt;td&gt;HIPAA audits and consulting&lt;/td&gt;
&lt;td&gt;Aug 2022&lt;/td&gt;
&lt;td&gt;Admits "the HHS does not list a set-in-stone rule about the time frame for retaining audit logs", then concludes "you should maintain your audit logs for six years."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aptible.com, "HIPAA audit log retention"&lt;/td&gt;
&lt;td&gt;hosting&lt;/td&gt;
&lt;td&gt;updated Mar 2026&lt;/td&gt;
&lt;td&gt;Headline: "HIPAA requires six years of audit log retention." The body then quotes 164.316 as a documentation rule. Gemini used the headline for a wrong answer and the body for a right one.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Cited under the correct answers&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Page&lt;/th&gt;
&lt;th&gt;What they sell&lt;/th&gt;
&lt;th&gt;Date on page&lt;/th&gt;
&lt;th&gt;Exact sentence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.lakeridge.io/what-are-audit-log-requirements-for-patient-records-164312b" rel="noopener noreferrer"&gt;LakeRidge Technologies&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;compliance services&lt;/td&gt;
&lt;td&gt;Jul 2026&lt;/td&gt;
&lt;td&gt;"HIPAA § 164.312(b) does not state a specific number of days or years that security logs must be retained."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.accountablehq.com/post/hipaa-security-rule-audit-log-retention-period-how-long-to-keep-logs" rel="noopener noreferrer"&gt;Accountable&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;HIPAA compliance software&lt;/td&gt;
&lt;td&gt;Feb 2026&lt;/td&gt;
&lt;td&gt;"The regulation does not prescribe a specific audit log retention period."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://www.usahipaa.com/blog/hipaa-audit-log-retention-requirements" rel="noopener noreferrer"&gt;USA HIPAA&lt;/a&gt; (cited by Claude)&lt;/td&gt;
&lt;td&gt;training, certifications, compliance kits&lt;/td&gt;
&lt;td&gt;Jul 2026&lt;/td&gt;
&lt;td&gt;"HIPAA does not print a retention number inside its audit-controls rule, so teams guess, and most guess too short." Also the source of Claude's "regulators treat that evidence as documentation" sentence.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://hipaaauditors.com/nist-guidelines/au-11-audit-record-retention" rel="noopener noreferrer"&gt;HIPAA Auditors&lt;/a&gt; (cited by Claude)&lt;/td&gt;
&lt;td&gt;audit and certification services&lt;/td&gt;
&lt;td&gt;Sep 2026&lt;/td&gt;
&lt;td&gt;"HIPAA does not set a single universal audit-log day count; define periods risk-based and retain documentation per 164.316."&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four vendors got it right. Their pages rank for the precise question and not for the popular one, which is the whole story in one row.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>discuss</category>
      <category>hipaa</category>
    </item>
    <item>
      <title>How Many Lines of Code Do You Ship to Prod a Week? This Week I Shipped One.</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Wed, 26 Aug 2026 00:18:27 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/how-many-lines-of-code-do-you-ship-to-prod-a-week-this-week-i-shipped-one-140b</link>
      <guid>https://dev.to/vadim_albarov/how-many-lines-of-code-do-you-ship-to-prod-a-week-this-week-i-shipped-one-140b</guid>
      <description>&lt;p&gt;Yes, one. But I am as exhausted as if I had shipped a full feature.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;BackupRetentionPeriod'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="s"&gt;0,&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole fix. It took a week, more Claude Code sessions than I want to count, a &lt;code&gt;/code-review&lt;/code&gt; pass before every pull request, one restart from scratch, and a diff that grew to a couple hundred lines before it shrank back to this. The sentence that made the one-liner possible was fetched in the very first turn. It lost - to a file on my laptop I had not updated in months, and to the model's memory, which was just as old.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Every weekday morning a Step Functions workflow restores an analytics replica from the main database's snapshot, runs some SQL, flips a DNS record, and deletes yesterday's copy. The replica lives one day.&lt;/p&gt;

&lt;p&gt;Except a restored RDS instance inherits its backup retention from the snapshot. So every disposable replica came up with the source's week-long automated backups and started snapshotting itself within minutes. Wasted money, and another copy of PHI nobody asked for.&lt;/p&gt;

&lt;p&gt;I know RDS. Backups are one setting. On a normal instance I would have flipped it without thinking. But this was a restore from a snapshot, and a snapshot carries the source's settings with it - so I had a real doubt: does that one setting still apply at restore time, or does the snapshot win and you have to modify afterwards? Guess: probably one parameter. Doubt: maybe not in this path.&lt;/p&gt;

&lt;p&gt;And the workflow was agent-written. It worked, I had reviewed it, but I had not typed it and did not know every corner of it. Finding the right dict while also settling my doubt felt like a job for the model. So I asked for the change: reliable, no race conditions, check the edge cases, and - I actually wrote this - &lt;em&gt;trivial and bulletproof&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I know vague prompts do not work well. But "trivial and bulletproof" did not feel vague to me. In my head it had a strict meaning: the smallest change that works, plus proof that nothing around it breaks. The model read it differently - as a license to extend the scope. At least that is what I think now; the model did not share its thoughts.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Verified"
&lt;/h2&gt;

&lt;p&gt;The planning turn was impressive. It read the module, queried both production accounts, confirmed the inherited retention, found the first automated backup landing about a minute after each restore. It fetched a stack of AWS docs. Then it wrote:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;And you cannot fix it at restore time. I checked the RDS API model shipped with the local CLI:&lt;br&gt;
&lt;code&gt;RestoreDBInstanceFromDBSnapshot -&amp;gt; BackupRetentionPeriod: False&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That did not fight my belief. It confirmed my doubt, with a checkmark. The thing I was unsure about had been settled by someone who looked.&lt;/p&gt;

&lt;p&gt;The plan that followed was several times the change I had hoped for: modify the instance after restore, absorb the reboot that causes, add IAM, add a retry, widen the timeouts. I asked why. It answered every line - confidently, consistently, with a perfect track record from the values it had just pulled out of production. I had a guess, a doubt, and no working knowledge of the code. It had reasons. I was persuaded, which is not the same as convinced. I said "implement."&lt;/p&gt;

&lt;p&gt;Here is what I missed. In the same response, to show that retention is inherited, it had quoted the RDS User Guide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"When you don't explicitly set this value, the restored database inherits the backup retention period from the source snapshot or instance."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;When you don't explicitly set this value.&lt;/em&gt; The doc was saying the value can be set. The model quoted it for the second half and moved on. Asked about it later, it said it had noticed the conflict and "trusted the stale artifact over the doc."&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop
&lt;/h2&gt;

&lt;p&gt;Every session that week ran the same cycle. &lt;code&gt;/code-review&lt;/code&gt; finds real defects. The model fixes all of them, correctly. At least one fix lands outside the original ask - a cleanup path that was already weak, a timeout that "should" go up while we are here. The next review has new code to examine. Repeat.&lt;/p&gt;

&lt;p&gt;The first review said a mutation did not belong inside a polling Lambda; the fix moved it into its own Step Functions state, with a wait, a three-way Choice, and IAM on a different role. The second review found several bugs in that new surface, and a pre-existing orphan-cleanup problem while it was in there. The third review found bugs in the cleanup rewrite. Each pass was right. Each fix was good. The ask had touched one Lambda; the diff now covered two Lambdas, the state machine, IAM, and the variables file, and had rewritten instance cleanup and database connection handling along the way. None of it was the task.&lt;/p&gt;

&lt;p&gt;Midweek I threw it all away and started from scratch in a fresh session, hoping a clean context would find a better direction. It ran the same check against the same CLI and produced the same design with fresh confidence. Starting over resets the conversation. It does not reset the tools.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/code-review&lt;/code&gt; is good at its job, and that is the problem. It reviews the diff in front of it. It has no way to say the diff should not exist. So every pass certified the direction by omission and handed back a list of things to fix inside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lottery
&lt;/h2&gt;

&lt;p&gt;After the third pass I stopped asking for fixes and asked which findings were real, which were theoretical, and which were irrelevant - and said "none" was a fine answer. That got the most honest reply of the week: a handful real, a few theoretical, one right for the wrong reason, one pure scope creep that had spawned half the last round. "I'd stop here." It even noted the design still lost a race with RDS's first automated snapshot. I asked: commit or rewrite? Commit. I was about to.&lt;/p&gt;

&lt;p&gt;Then I ran &lt;code&gt;/code-review&lt;/code&gt; one last time, as I always did. This one went deep for reasons I still do not know: twenty parallel agents, the maximum, each reading the repo on its own. They used up my almost‑fresh five‑hour session token limit in just 20 minutes. The run died when the limit hit. Three agents came back. One of them said, high severity: the restore API already accepts &lt;code&gt;BackupRetentionPeriod&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One of twenty. One of the three that finished. A minute later on the limit and I would have shipped a couple hundred lines and no budget left to look again.&lt;/p&gt;

&lt;p&gt;The model re-fetched the API reference, found the announcement, confirmed my CLI predated it, and said plainly that its verification had been wrong. Everything came out except the one line.&lt;/p&gt;

&lt;p&gt;One line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where "verified" came from
&lt;/h2&gt;

&lt;p&gt;The CLI on my laptop was many months old. AWS added the parameter after that build, and the CLI never updates its own API model - it is a JSON file frozen into the installer. Every time the agent checked, in every session, it got the same honest wrong answer.&lt;/p&gt;

&lt;p&gt;Half of that is mine. An agent trusts the output of commands on my machine the way I trust my eyes. Stale tools mean stale observations, and the agent cannot tell. Keeping the CLI current used to be housekeeping. With an agent doing the checking, it is ground truth - and when it rots, the agent does not fail. It succeeds at verifying the wrong world.&lt;/p&gt;

&lt;p&gt;The other half: the parameter also postdates the model's training. So the model's memory said no, the local file said no, and I had walked in doubting. Three against one fresh sentence from a live doc. Nothing in the loop weighs a source by its date, so the newest one counted least.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A "verified" that lands on your own doubt is the one to check.&lt;/strong&gt; A claim that contradicts you gets scrutiny for free. A claim that confirms your uncertainty gets none - it feels like closure. It was the one question I had come with, and I let the model close it by assertion.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;"Why?" buys a justification, not a check.&lt;/strong&gt; Every answer was about how to do the thing safely. None was about whether it was needed, because I never asked that. The premise question is: what fact, if wrong, makes this whole plan unnecessary - and where did we verify it?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;"Verified" has a date. So does the model.&lt;/strong&gt; Ask when the tool was current. Then remember the tools on your machine are yours to keep current; an agent inherits every stale binary you leave it and turns it into confident output.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Diff review certifies direction by omission.&lt;/strong&gt; Use &lt;code&gt;/code-review&lt;/code&gt; for what it is - a defect pass. Ask the premise question separately, in plain words, before the second review, not after the fourth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A fresh session is not a second opinion.&lt;/strong&gt; Same model, same tools, same stale CLI, same answer. Independence needs a changed input: a different model, a different source, or a human who knows the system. I was that human, and I had already deferred.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Parallel agents under a hard budget are a lottery.&lt;/strong&gt; Twenty agents share one window. When it hits, the unfinished ones are gone - not returned, not cached. I bought twenty tickets and got to read three.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The uncomfortable version
&lt;/h2&gt;

&lt;p&gt;I did not lose a week because the model was wrong. Models are wrong sometimes; that is priced in. I lost it because the wrong answer matched my own doubt, and a wrong answer that agrees with you does not feel wrong. The model was not lying. It was thorough about the wrong question, standing on a file I had let go stale, and I let its thoroughness stand in for the check I should have made. What finally broke the loop was one lucky agent and one human question - and I would rather not rely on the luck again.&lt;/p&gt;

&lt;p&gt;The one line is committed. One last thing I did check this time, against the botocore changelog rather than a file on my laptop: the Lambda runtime's bundled SDK is newer than the release that added the parameter, so the call will not be rejected client-side. The fix that took a week to find took a few minutes to verify. That ratio is the whole article.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/APIReference/API_RestoreDBInstanceFromDBSnapshot.html" rel="noopener noreferrer"&gt;RestoreDBInstanceFromDBSnapshot - Amazon RDS API Reference&lt;/a&gt; - lists &lt;code&gt;BackupRetentionPeriod&lt;/code&gt; as a request parameter ("Default: Uses existing setting").&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_WorkingWithAutomatedBackups.BackupRetention.html" rel="noopener noreferrer"&gt;Backup retention period - Amazon RDS User Guide&lt;/a&gt; - "During restore operations, you have the option to specify a backup retention period..."&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/02/rds-aurora-backup-configuration-restoring-snapshots" rel="noopener noreferrer"&gt;Amazon RDS now supports backup configuration when restoring snapshots - AWS What's New&lt;/a&gt; - "Previously, restored database instances inherited backup parameter values from snapshot metadata and could only be modified after restore was complete."&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/cli/latest/reference/rds/restore-db-instance-from-db-snapshot.html" rel="noopener noreferrer"&gt;restore-db-instance-from-db-snapshot - AWS CLI Reference&lt;/a&gt; - the current CLI accepts &lt;code&gt;--backup-retention-period&lt;/code&gt;; an older installed CLI does not, because its API model is frozen at build time.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>aws</category>
      <category>claude</category>
      <category>devops</category>
    </item>
    <item>
      <title>Windows Copilot Inverted a HIPAA Rule Three Times - and Cited a Real Paper That Proves the Opposite</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Thu, 20 Aug 2026 03:49:01 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/windows-copilot-inverted-a-hipaa-rule-three-times-and-cited-a-real-paper-that-proves-the-opposite-11ef</link>
      <guid>https://dev.to/vadim_albarov/windows-copilot-inverted-a-hipaa-rule-three-times-and-cited-a-real-paper-that-proves-the-opposite-11ef</guid>
      <description>&lt;p&gt;I have a two-window habit. My main work happens in a Claude Code session, and I don't like burning its context on side questions - so for quick lookups I alt-tab into Windows Copilot and use it as a search-flavored notepad. Ask, skim, close, back to work. It's been part of my routine for a long time, and honestly, I'd never had a problem with its answers before - side questions came back reasonable, and nothing ever sent me down a wrong path. Low stakes, or so I thought.&lt;/p&gt;

&lt;p&gt;This week the side window taught me a HIPAA rule with total confidence. The rule was backwards. Not vague, not incomplete - inverted, in the one direction that would turn a compliance question into a breach.&lt;/p&gt;

&lt;p&gt;I build healthcare infrastructure, so I had the reflexes to check. This is the story of what it said, what the regulation actually says, how the wrong answer survived thinking mode, search mode, and a demand for references - and how the citation trail eventually explained where the inversion came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;I was thinking about deterministic pseudonymization: replacing a patient identifier with &lt;code&gt;hash(key + identifier)&lt;/code&gt; so the same patient always maps to the same token and records stay joinable. The obvious follow-up is what HIPAA thinks about the key. Keep it secret? Share it? Does it matter? Full disclosure: the question I actually typed was leading - I'd already run into the claim that sharing the key was the compliant option somewhere, and I asked about it as if it were established.&lt;/p&gt;

&lt;p&gt;Copilot's answer, summarized:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you keep a &lt;strong&gt;secret key&lt;/strong&gt;, you retain the ability to re-identify. Therefore the data is pseudonymized, still PHI, not de-identified.&lt;/li&gt;
&lt;li&gt;If the key is &lt;strong&gt;public&lt;/strong&gt;, nobody has privileged knowledge. Nobody can reverse a one-way hash. Therefore the data can qualify as de-identified.&lt;/li&gt;
&lt;li&gt;HIPAA, it explained, cares about &lt;em&gt;privileged access&lt;/em&gt;, not cryptographic strength. "You're thinking like an engineer. HIPAA is written by lawyers."&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fllawb1acn7sz56s5fu3r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fllawb1acn7sz56s5fu3r.png" alt="pic 1" width="799" height="218"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgdx9j5ef854i1l3lq6r3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgdx9j5ef854i1l3lq6r3.png" alt="pic 2" width="800" height="204"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's a seductive answer. It has a philosophy. It flatters you for being confused. It reads like someone explaining a genuinely counterintuitive corner of law.&lt;/p&gt;

&lt;p&gt;It is also wrong on both branches, and the second branch is dangerous.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the regulation actually says
&lt;/h2&gt;

&lt;p&gt;The relevant text is &lt;a href="https://www.law.cornell.edu/cfr/text/45/164.514" rel="noopener noreferrer"&gt;45 CFR 164.514&lt;/a&gt;. Safe Harbor's identifier list ends with a catch-all - "any other unique identifying number, characteristic, or code" - with exactly one exception: a re-identification code that satisfies paragraph (c). Paragraph (c) has two prongs:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Derivation.&lt;/strong&gt; The code must &lt;em&gt;not&lt;/em&gt; be derived from or related to information about the individual, and must not be otherwise translatable back to the individual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security.&lt;/strong&gt; The covered entity must not use or disclose the code for other purposes and must not disclose &lt;em&gt;the mechanism for re-identification&lt;/em&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Read those against the two branches:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Secret-key hash.&lt;/strong&gt; &lt;code&gt;HMAC(secret, MRN)&lt;/code&gt; fails Safe Harbor - but not for Copilot's reason. It fails because the token is mathematically &lt;em&gt;derived&lt;/em&gt; from the identifier, and prong 1 prohibits derivation outright, regardless of how well you guard the key. Meanwhile, Copilot's actual claim - "if you can re-identify, it's not de-identified" - contradicts the regulation's text. Paragraph (c) exists precisely so a covered entity &lt;em&gt;can&lt;/em&gt; keep a secret re-identification mechanism while the dataset remains de-identified. Retained re-identification capability is not the disqualifier; derivation and disclosure are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Public-key hash.&lt;/strong&gt; This one fails everything at once. Publishing the key is literally "disclosing the mechanism for re-identification" - prong 2, verbatim. And it makes the data trivially translatable - prong 1 - because health identifiers live in small, enumerable spaces. SSNs are a 10^9 space. Phone numbers, MRNs, emails, name-plus-birthdate: all enumerable. With the key public, anyone hashes every candidate and matches your entire column in seconds on a laptop. One-wayness protects high-entropy inputs; it does nothing for a nine-digit number.&lt;/p&gt;

&lt;p&gt;HHS's own &lt;a href="https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html" rel="noopener noreferrer"&gt;de-identification guidance&lt;/a&gt; closes the loop from both sides. It says a hash &lt;em&gt;without&lt;/em&gt; a secret key counts as an identifying element exactly because recipients can reverse it over the input space. And it says keyed cryptographic hashes are acceptable under the Expert Determination pathway &lt;em&gt;provided the keys are not disclosed, including to the recipients&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;So the real rule is the mirror image of what my sidebar told me:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Construction&lt;/th&gt;
&lt;th&gt;HIPAA status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hash with a &lt;strong&gt;published&lt;/strong&gt; key&lt;/td&gt;
&lt;td&gt;Fails everything - discloses the mechanism, trivially reversible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hash with a &lt;strong&gt;secret&lt;/strong&gt; key&lt;/td&gt;
&lt;td&gt;Fails Safe Harbor (derived code); acceptable under Expert Determination with the key undisclosed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Random token&lt;/strong&gt; + protected mapping table&lt;/td&gt;
&lt;td&gt;The pattern 164.514(c) actually blesses - a random value is not "derived from" anyone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Round two: references make it worse
&lt;/h2&gt;

&lt;p&gt;Maybe I asked badly. I turned on thinking mode, then search mode, and asked it to rethink and support its answer with references.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6ghnafr10ntf2wv1p5js.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6ghnafr10ntf2wv1p5js.png" alt="pic 3" width="800" height="623"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I got the same inverted conclusion, now decorated with emoji section headers, a verdict table, and a source list - and with something worse. The centerpiece was a quoted sentence attributed to HHS, saying a code derived from PHI is an identifier "unless the re-identification key is not retained." I searched for that sentence. Not on HHS.gov, not anywhere I could find. The model composed a plausible-sounding rule that swaps the regulation's actual verb - &lt;em&gt;disclose&lt;/em&gt; - for &lt;em&gt;retain&lt;/em&gt;, wrapped it in quotation marks, and attached the genuine HHS URL to it.&lt;/p&gt;

&lt;p&gt;That's the failure mode that stuck with me: a real link laundering a fake quote. Every reader's citation heuristic - "it links to hhs.gov, so it's grounded" - defeated by construction.&lt;/p&gt;

&lt;p&gt;The funny part is that the correct answer was present in the same response, scattered in the margins. The risks section admitted that a public pepper can be brute-forced over common identifiers - which quietly destroys the headline claim that public-key hashing is irreversible. Another bullet correctly noted that keeping a secret key for linkage requires expert determination - which is the actual rule, and refutes the answer's own bottom line. The model had all the pieces and still shipped the inverted conclusion in the headline, the table, and the summary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round three: it's not me
&lt;/h2&gt;

&lt;p&gt;One hypothesis left: maybe my framing poisoned the well, since my first question presented the counterintuitive claim as a finding. So I opened a fresh chat and asked the neutral, symmetric question - which is better for HIPAA, hash with a private key or a public key - with no premise embedded.&lt;/p&gt;

&lt;p&gt;Verdict, verbatim in spirit: public-key hashing is "far better," private-key hashing is still PHI, and - my favorite line - the public pepper should be "long, random, and not guessable."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4mo6c9r2k935md90o70x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4mo6c9r2k935md90o70x.png" alt="pic 4" width="800" height="372"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A &lt;em&gt;public&lt;/em&gt; value that is &lt;em&gt;not guessable&lt;/em&gt;. It's published. That single phrase is the whole confusion in miniature: the answer needs the key to be simultaneously known to everyone (so nobody has privileged access) and known to no one (so nobody can brute-force). For calibration, I asked ChatGPT's web version the same neutral question. It got it essentially right: public-key hashing rejected for exactly the dictionary-attack reason, keyed HMAC labeled as pseudonymization that doesn't exit PHI obligations by itself, random tokens with a secured mapping table recommended for real de-identification. Not perfect - it never cited the derivation prohibition - but directionally sound everywhere it committed.&lt;/p&gt;

&lt;p&gt;Same question. One product inverted, reproducibly, across independent chats. The other didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The citation trail explains everything
&lt;/h2&gt;

&lt;p&gt;Both of the reference-backed Copilot answers - rounds two and three - leaned on the same academic source: a 2003 AMIA paper by Landi and Rao, &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC1479909/" rel="noopener noreferrer"&gt;"Secure De-identification and Re-identification"&lt;/a&gt;. I looked it up.&lt;/p&gt;

&lt;p&gt;The paper describes an &lt;em&gt;asymmetric encryption&lt;/em&gt; scheme: encrypt patient identifiers with a &lt;strong&gt;public key&lt;/strong&gt; so that only the holder of the matching &lt;strong&gt;private key&lt;/strong&gt; - the data owner - can decrypt and re-identify. "Public key" as in public-key cryptography. One half of a keypair. A system whose entire security rests on the &lt;em&gt;private&lt;/em&gt; key staying secret, and whose explicit purpose is to let the owner retain re-identification capability.&lt;/p&gt;

&lt;p&gt;Now the most plausible failure chain is visible. Retrieval surfaced a paper with "public key" and "de-identification" in close proximity. Summarization collapsed &lt;em&gt;public-key cryptography&lt;/em&gt; into &lt;em&gt;publicly known key&lt;/em&gt;. And then - this is the part I find genuinely instructive - the model didn't just misread a term. It constructed an entire regulatory philosophy around the misreading: the "privileged knowledge" theory of HIPAA, delivered with the confidence of a law professor, appearing in no regulation, contradicted by the very paper being cited. The citation that was supposed to ground the answer was proof of the opposite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model was that, even?
&lt;/h2&gt;

&lt;p&gt;The chat window said "Smart." That's a mode label, not a model name. So I asked Copilot directly which model it runs on. It denied running on any outside lab's models at all - just "proprietary Windows AI technologies" - and said it cannot disclose specifics.&lt;/p&gt;

&lt;p&gt;That answer is worth exactly nothing, and knowing &lt;em&gt;why&lt;/em&gt; it's worth nothing is the useful part. A chatbot's claim about its own identity is generated text like everything else it says - models have no introspective access to the infrastructure serving them, and "I can't disclose" deflections are typically system-prompt policy, not knowledge. The same product that invented an HHS quote is not a reliable witness about its own internals.&lt;/p&gt;

&lt;p&gt;The public reporting says something different and messier: modern assistant products run mixed fleets - frontier models licensed from partner labs alongside the vendor's own in-house models - with a cost-driven router deciding invisibly, per query, which one you get. The lineup shifts between announcements, the routing shifts without any announcement at all, and there's no per-response indicator. The same question may be served by a different model next week, or later today.&lt;/p&gt;

&lt;p&gt;I don't know whether my three inverted answers came from a fast-path model, from the summarization layer garbling search results, or from the persona tuning that opens responses with "let me say this clearly and directly." That's the point: as a user, I &lt;em&gt;can't&lt;/em&gt; know - and asking the product just adds one more unverifiable claim to the pile. When another assistant answers the same question correctly, the difference isn't necessarily raw model capability. It's everything wrapped around the model, and the wrapper is invisible. "Which LLM does this product use" turns out to be the wrong question. The right one is "which pipeline, with which retrieval, which router, and which incentives" - and no consumer product answers it. Not even when you ask it directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm taking away
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Confidence formatting is free.&lt;/strong&gt; Tables, verdict emoji, "let me say this clearly" - none of it correlates with correctness. The most wrong answer in this story was the best-formatted one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Support this with references" does not mean "verify this."&lt;/strong&gt; Once a model has committed to a conclusion, asking for sources can produce &lt;em&gt;justification&lt;/em&gt; instead - up to and including an invented quote pinned to a real government URL. If a quote matters, search for the exact sentence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch for internal contradictions.&lt;/strong&gt; The wrong answers refuted themselves in their own risk sections. A response whose caveats disagree with its headline is telling you which part was retrieved and which part was composed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For anything regulatory, read the primary source.&lt;/strong&gt; 164.514(c) is two paragraphs. It cost five minutes and settled in one reading what three AI answers scrambled. Compliance-by-chatbot is how a "de-identified" dataset ships with a published key and becomes a reportable breach.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The side window deserves the same skepticism as the main one.&lt;/strong&gt; My mistake wasn't using Copilot - it was granting the quick-lookup window a lower evidence bar because the questions felt small. Nothing about the window makes the answers smaller.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'm keeping the two-window setup, but the notepad seat is now vacant. Either something changed in that product - a router, a model swap, a summarization layer - or it could always do this and I simply never caught it; I have no way to tell, and that uncertainty is its own verdict. Either way, a tool I trusted for months just fabricated regulatory quotes with a straight face. So I'm dropping it until it stabilizes, and I'll know it has stabilized the same way I learned it broke: by spot-checking its answers against primary sources. A quick-lookup tool that requires verification of every answer isn't a quick-lookup tool anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Here is the prompt I used for the fresh-chat test, lightly tidied. Run it against your assistant of choice:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In terms of HIPAA compliance, what is better for de-identifying PHI data: HASH(private_key + PHI) or HASH(public_key + PHI)?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The correct answer rejects both as de-identification on their own, flags the public-key variant as trivially reversible by dictionary attack, and mentions that a keyed hash only works under the Expert Determination pathway with the key kept undisclosed. Anything that tells you the public key is the compliant option has inverted 45 CFR 164.514(c).&lt;/p&gt;

&lt;p&gt;Share what you get in the comments - which product, which mode, and which way it landed. I'm genuinely curious whether this reproduces beyond my machine, and a comment thread of timestamped outputs is a better dataset.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclaimer: I'm an infrastructure engineer, not a lawyer; this is an engineering read of published regulations and guidance, not legal advice. The Copilot and ChatGPT responses summarized here were collected in August 2026 from consumer versions of both products; model routing and behavior change constantly, which is rather the point.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hipaa</category>
      <category>security</category>
      <category>llm</category>
    </item>
    <item>
      <title>The De-identification Datasets Problem: i2b2/n2c2 Access Is Broken, and What to Do Instead</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Wed, 19 Aug 2026 03:27:41 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/the-de-identification-datasets-problem-i2b2n2c2-access-is-broken-and-what-to-do-instead-2mo7</link>
      <guid>https://dev.to/vadim_albarov/the-de-identification-datasets-problem-i2b2n2c2-access-is-broken-and-what-to-do-instead-2mo7</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I'm building a PHI de-identification tool, and step one of doing that honestly is benchmarking it against the standard corpora: the i2b2 2006 and 2014 de-identification challenge datasets, now distributed as "n2c2" through Harvard DBMI's data portal. Step one failed. The portal's n2c2 page has said "Temporarily Unavailable" since at least 2026-07-28, registration is closed, and the old i2b2.org dataset page returns HTTP 500. Even when the door is open, access means per-user registration, a data use agreement, and an approval wait - which is why published de-id numbers are so hard to reproduce and why the field's canonical scores are 12-20 years old. I built a synthetic corpus generator instead. Here's the whole investigation, dated, plus what synthetic data honestly can and cannot replace.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step one: you can't get the data
&lt;/h2&gt;

&lt;p&gt;Benchmarking a de-identification tool sounds like the easy part of the project. The hard part is supposed to be the NLP: finding every patient name, date, medical record number, and phone number buried in messy clinical prose, so the notes can be used for research without exposing anyone. Detection is the science. Evaluation is just downloading the test set, right?&lt;/p&gt;

&lt;p&gt;Here is how evaluation actually went for me.&lt;/p&gt;

&lt;p&gt;Every de-identification paper of the last two decades benchmarks against the same corpora: the i2b2 de-identification challenge datasets from 2006 and 2014. If you want your numbers to mean anything to anyone, you report them on i2b2 2014. So on 2026-07-28 I went to get the data from the &lt;a href="https://portal.dbmi.hms.harvard.edu/" rel="noopener noreferrer"&gt;Harvard DBMI Data Portal&lt;/a&gt;, where the corpora live under their post-2018 name, n2c2 (National NLP Clinical Challenges). What I found was a notice: the n2c2 datasets are temporarily unavailable.&lt;/p&gt;

&lt;p&gt;Fine, I thought. Maintenance happens. I built other things and came back.&lt;/p&gt;

&lt;p&gt;As of today, 2026-08-18, nearly three weeks later, here is the exact state of the front door, checked from a fresh session:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The portal itself is up and loads fine.&lt;/li&gt;
&lt;li&gt;The portal's &lt;a href="https://portal.dbmi.hms.harvard.edu/data-sets/" rel="noopener noreferrer"&gt;data sets listing&lt;/a&gt; shows two entries. Neither is an n2c2 NLP corpus (both are 4CE COVID-related sets).&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://portal.dbmi.hms.harvard.edu/projects/n2c2-nlp/" rel="noopener noreferrer"&gt;n2c2 NLP Research Data Sets project page&lt;/a&gt; - the page for the actual challenge corpora, 2006 through 2018 - carries this banner, verbatim: "Temporarily Unavailable. The n2c2 datasets are temporarily unavailable. If you are trying to access data from the 2019 Challenge, tracks 1 (Clinical Semantic Textual Similarity) and 2 (Family History Extraction) are available directly through Mayo Clinic." Below the dataset descriptions: "Registration is not open ... at this time."&lt;/li&gt;
&lt;li&gt;The legacy home of these corpora, &lt;a href="https://www.i2b2.org/NLP/DataSets/" rel="noopener noreferrer"&gt;i2b2.org/NLP/DataSets&lt;/a&gt;, returns HTTP 500.&lt;/li&gt;
&lt;li&gt;The companion &lt;a href="https://n2c2.dbmi.hms.harvard.edu/" rel="noopener noreferrer"&gt;n2c2 informational site&lt;/a&gt; refused my scripted requests outright with an Akamai "Access Denied" page. That one I'll hedge: it may just be bot filtering rather than downtime. But it means I cannot even verify the documentation programmatically.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No ETA. No status page. No mirror. The canonical benchmark for an entire subfield of clinical NLP is a "temporarily unavailable" banner, and "temporarily" has meant at least July 28 through August 18 so far. It may come back tomorrow; the point of this article survives either way, because the outage is only the loudest symptom of a structural problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  What these datasets are, and why everyone needs them
&lt;/h2&gt;

&lt;p&gt;Some history, because the names are confusing. i2b2 (Informatics for Integrating Biology and the Bedside) was an NIH-funded center based at Partners HealthCare that, starting in 2006, ran annual clinical NLP shared tasks on real (de-identified) hospital notes. In 2018 the challenge series was renamed n2c2, and stewardship of the datasets moved to the Department of Biomedical Informatics at Harvard Medical School, distributed via their portal. Same corpora, three names, one door.&lt;/p&gt;

&lt;p&gt;Two of those challenges define de-identification evaluation to this day:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 2006 challenge&lt;/strong&gt; (&lt;a href="https://academic.oup.com/jamia/article/14/5/550/720189" rel="noopener noreferrer"&gt;Uzuner, Luo, and Szolovits, JAMIA 2007&lt;/a&gt;) used hospital discharge summaries in which the authentic PHI had been replaced with synthesized surrogates - including deliberately out-of-vocabulary, made-up names to punish systems that just memorized name lists. Seven teams, sixteen system runs, and the best systems scored above 98% F-measure across PHI categories.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 2014 i2b2/UTHealth challenge&lt;/strong&gt; (&lt;a href="https://www.sciencedirect.com/science/article/pii/S1532046415001823" rel="noopener noreferrer"&gt;Stubbs and Uzuner, JBI 2015&lt;/a&gt;) raised the bar: 1,304 longitudinal records covering 296 patients, over 28,000 annotated PHI instances, annotated under a broad interpretation of HIPAA. Human annotators managed a token-level F1 of 0.927 against the gold standard; the &lt;a href="https://www.researchgate.net/publication/280584382_Automated_systems_for_the_de-identification_of_longitudinal_clinical_narratives_Overview_of_2014_i2b2UTHealth_shared_task_Track_1" rel="noopener noreferrer"&gt;best automated system&lt;/a&gt; hit a strict micro-averaged F1 of 0.936.&lt;/p&gt;

&lt;p&gt;Notice something about both: even the "real" gold standards contain synthetic PHI. The notes are genuine clinical text, but the identifiers in them are surrogates, inserted so the data could be released at all. Keep that in mind for later - the field's ground truth has always been real prose plus fake identifiers.&lt;/p&gt;

&lt;p&gt;That 0.936 from 2014 is, functionally, still the number. When you publish a de-id tool in 2026, reviewers ask how you compare on i2b2 2014. The corpus is twelve years old, drawn from one hospital system, and pre-dates most of what modern EHRs do to note formatting. It is also, at the moment, undownloadable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The access process, even on a good day
&lt;/h2&gt;

&lt;p&gt;Suppose the portal comes back tomorrow. What does access look like then? Per the archived instructions on i2b2.org, the datasets are "freely available" to researchers - subject to a data use agreement, and "each individual user must access the data independently" through the portal. In practice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Register an account on the DBMI portal.&lt;/li&gt;
&lt;li&gt;Sign the Rules of Conduct and the Data Use Agreement (the &lt;a href="https://n2c2.dbmi.hms.harvard.edu/files/n2c2/files/n2c2_2019_-_dua_track_3.pdf" rel="noopener noreferrer"&gt;n2c2 DUAs&lt;/a&gt; are real legal documents, not click-through checkboxes).&lt;/li&gt;
&lt;li&gt;Wait for a human to approve you. There is no published turnaround time.&lt;/li&gt;
&lt;li&gt;Repeat for every individual on your team, because the DUA is per-person, not per-lab.&lt;/li&gt;
&lt;li&gt;Never redistribute the data - which also means never shipping it as a test fixture, never putting it in CI, never publishing your evaluation harness with the inputs included.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is not villainy. The data is real patient prose; a DUA is the legally and ethically appropriate wrapper, and the people who built and maintain these corpora did the field an enormous service. But note the architecture: the standard benchmark for a global research area is administered by one team at one institution, through one portal, with per-user paperwork and no fallback. When that single point of distribution goes down - for maintenance, for a compliance review, for a staffing gap, for whatever is happening right now - the benchmark simply ceases to exist for anyone who doesn't already have a copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this hurts more than my project
&lt;/h2&gt;

&lt;p&gt;Play the incentives forward and the damage compounds:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Published results become unreproducible in practice.&lt;/strong&gt; A paper says "0.94 F1 on i2b2 2014." You cannot check that claim, cannot run the same test set through your own tool, cannot even eyeball the annotation decisions the score depends on. Reproducibility in de-id research is gated on a portal login, and today on a portal banner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;New tools cannot compare against prior art.&lt;/strong&gt; The literature has twenty years of numbers on these corpora. A new open-source tool that cannot access them either skips comparison (and gets dismissed) or quotes other papers' numbers against its own results on different data (and misleads).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The canonical numbers fossilize.&lt;/strong&gt; Because making a new shareable gold standard from real notes is brutally expensive - the 2014 corpus took double annotation, arbitration, and multiple proofreading rounds to reach that 0.927 human F1 - nobody replaces the old benchmarks. The field's reference points are frozen in 2006 and 2014 while clinical documentation, and the models reading it, changed completely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Insiders and outsiders diverge.&lt;/strong&gt; Groups with long-standing access or local hospital data keep publishing; independent developers and open-source maintainers evaluate on whatever they can scrape together. The people most likely to ship a de-id tool you can actually download are the least able to prove it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The alternatives tour
&lt;/h2&gt;

&lt;p&gt;Before building anything, I did the diligence on every other door:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PhysioNet.&lt;/strong&gt; The MIMIC family of ICU databases includes clinical notes, and PhysioNet also hosts a &lt;a href="https://physionet.org/content/deidentifiedmedicaltext/1.0/" rel="noopener noreferrer"&gt;gold standard corpus of 2,434 de-identified nursing notes&lt;/a&gt; with realistic surrogate PHI (Neamatullah et al., 2008). Access requires becoming a credentialed user: identity verification, the CITI "Data or Specimens Only Research" training course, and then a separate DUA per dataset. To PhysioNet's credit, this process is documented, predictable, and actually functioning. But it is still weeks of process, still per-person, still non-redistributable. And MIMIC's notes are already de-identified with placeholders, so to evaluate a de-id tool you must first re-inject surrogate PHI - which is exactly what &lt;a href="https://arxiv.org/pdf/1803.02728" rel="noopener noreferrer"&gt;prior work on synthetically-identified MIMIC notes&lt;/a&gt; does. The "real data" path quietly becomes a synthetic-PHI path anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MTSamples.&lt;/strong&gt; &lt;a href="https://mtsamples.com/" rel="noopener noreferrer"&gt;MTSamples.com&lt;/a&gt; hosts thousands of publicly available transcribed sample medical reports across dozens of specialties. No registration, no DUA, real clinical language structure. The catch: they are samples, so they contain no PHI to find. To make a de-id benchmark out of them you inject fake PHI into the text - an approach with an established lineage in the literature. This is the honest public option, and it is the spirit my workaround follows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Synthea.&lt;/strong&gt; &lt;a href="https://synthea.mitre.org/" rel="noopener noreferrer"&gt;Synthea&lt;/a&gt; generates fully synthetic patients with medically plausible histories - fantastic for structured FHIR data. But its narrative output is template-driven fill-in-the-blank SOAP text linked to the structured record. For de-id evaluation, where the entire game is the messiness of real prose - copy-paste artifacts, headers, inconsistent formatting, abbreviations - templated narrative is the wrong distribution by design.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Newer synthetic clinical text efforts.&lt;/strong&gt; This space is heating up: &lt;a href="https://www.frontiersin.org/journals/digital-health/articles/10.3389/fdgth.2025.1497130/full" rel="noopener noreferrer"&gt;Synthetic4Health&lt;/a&gt; generates annotated synthetic clinical letters; &lt;a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12926592/" rel="noopener noreferrer"&gt;ASQ-PHI&lt;/a&gt; proposes an adversarial synthetic benchmark specifically for de-identification; and there is active work on &lt;a href="https://www.nature.com/articles/s41598-025-86890-3" rel="noopener noreferrer"&gt;whether LLM-generated notes actually match real note distributions&lt;/a&gt; (early answer: imperfectly, and you should measure the gap rather than assume it away). None of these is yet a community-standard replacement for i2b2 2014. All of them are bets on the same thesis I ended up betting on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workaround: a synthetic corpus generator
&lt;/h2&gt;

&lt;p&gt;So I built my own synthetic corpus generator for the PHI de-identification tool I'm building. The design is deliberately boring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Templates styled after real clinical note genres&lt;/strong&gt; - discharge summaries, progress notes, radiology reports - with the section structure, boilerplate, and formatting quirks those genres actually have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Controlled PHI injection.&lt;/strong&gt; Fake names, MRNs, dates, phone numbers, addresses, providers, and facilities generated and inserted at known offsets. Every injected entity is recorded with its exact span and category at generation time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Category balancing.&lt;/strong&gt; Real corpora are dominated by dates and names; rare categories (fax numbers, device IDs, URLs) barely appear. A generator can oversample the rare stuff so your recall numbers on those categories are backed by hundreds of instances instead of four.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The advantages are structural, not incidental:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Labeled by construction.&lt;/strong&gt; No annotation budget, no inter-annotator disagreement, no 0.927 ceiling on ground-truth quality. The generator knows where the PHI is because it put it there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shareable.&lt;/strong&gt; No DUA, because there is no patient. The corpus - and more importantly the generator - can live in a public repo, run in CI, and ship as test fixtures. Anyone can regenerate the exact evaluation set from the code and a seed. That is a property the i2b2 corpora can never have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalable and steerable.&lt;/strong&gt; Need 10,000 more notes with dates in weird formats? That's a parameter, not a grant application.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The honest limits
&lt;/h2&gt;

&lt;p&gt;Here is where I'm obligated to argue against myself, because synthetic evaluation has failure modes that will flatter your tool if you let them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distribution shift.&lt;/strong&gt; My templates are styled after clinical notes; they are not clinical notes. Real notes contain dictation artifacts, OCR junk, mid-sentence copy-paste, and formatting chaos that no template library fully reproduces. A recall number earned on synthetic text is an upper bound, not an estimate, of real-world recall. There is emerging work on &lt;a href="https://link.springer.com/article/10.1186/s44342-026-00072-9" rel="noopener noreferrer"&gt;quantifying exactly this synthetic-to-real gap for PHI taggers&lt;/a&gt;, and the gap is real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overfitting to your own generator.&lt;/strong&gt; This is the insidious one. If the same mental model writes both the fake-name generator and the name-detection logic, the benchmark and the tool share assumptions, and you are grading your own homework. Mitigations exist - independent sources for injection values, formats the detector wasn't designed around, adversarial edge cases - but the risk never reaches zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It complements real-data validation; it does not replace it.&lt;/strong&gt; My position after all this: synthetic corpora are for development, regression testing, category-level diagnostics, and public reproducibility. Before anyone trusts a de-id tool with actual patient data, it needs validation on real notes under proper governance - PhysioNet credentialing, an institutional dataset, or the n2c2 corpora if the door ever reopens. Remember, though, that even those gold standards are real prose with surrogate identifiers. The line between "real benchmark" and "synthetic benchmark" was always a spectrum, not a wall.&lt;/p&gt;

&lt;h2&gt;
  
  
  What better infrastructure would look like
&lt;/h2&gt;

&lt;p&gt;None of this requires new science. It requires plumbing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A status page and a mirror.&lt;/strong&gt; If a dataset is the reference benchmark for a field, "temporarily unavailable" with no ETA on a single portal should be impossible. PhysioNet already demonstrates the model: documented process, predictable credentialing, many datasets under one durable roof.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standing distribution instead of single-lab stewardship.&lt;/strong&gt; Move canonical corpora to infrastructure whose job is distribution, with credentialing handled once per user, not per corpus per portal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish generators, not just corpora.&lt;/strong&gt; A community-maintained synthetic benchmark - generator code plus seeds, calibrated against real data - would give the field something no DUA can: an evaluation anyone can run, extend, and verify. The recent synthetic-benchmark papers are steps in this direction; what's missing is convergence on one that leaderboards accept.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two-track evaluation as the norm.&lt;/strong&gt; Report on the gated real corpus for comparability, and on an open synthetic corpus for reproducibility. Either number alone is half a result.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check the data before you plan the benchmark.&lt;/strong&gt; The canonical de-id corpora are behind a per-user DUA on one portal, and that portal's n2c2 datasets have been "temporarily unavailable" from at least 2026-07-28 through 2026-08-18, with the legacy i2b2.org page throwing HTTP 500. Verify the current state yourself; date what you find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single-institution gatekeeping of a field's benchmark is a reliability bug&lt;/strong&gt;, independent of any outage. Per-user DUAs also mean no CI, no fixtures, no redistribution - reproducibility is structurally capped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The field's reference numbers are from 2006 and 2014.&lt;/strong&gt; Best strict F1 of 0.936 on 1,304 notes from one hospital system is still the bar new tools are measured against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PhysioNet is the functioning real-data path&lt;/strong&gt; - credentialing, CITI training, per-dataset DUA - but its notes need surrogate PHI re-injection for de-id evaluation anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synthetic corpora buy you labels by construction, shareability, and category balance.&lt;/strong&gt; They cost you distribution realism, and they tempt you into grading your own homework. Use them for development and public reproducibility; validate on real data before production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The gold standards were already part synthetic.&lt;/strong&gt; Real notes, surrogate PHI. Synthetic evaluation isn't a betrayal of rigor - unexamined synthetic evaluation is.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;If you work in clinical NLP: how did you get your benchmark data, and how long did access take? And if anyone has current information on when the n2c2 datasets are coming back, the comments are open.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>healthtech</category>
      <category>machinelearning</category>
      <category>nlp</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Is 'HIPAA-Eligible' the Same as 'HIPAA-Compliant'? I Audited AWS's List Against the Security Rule</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Tue, 18 Aug 2026 03:53:27 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/is-hipaa-eligible-the-same-as-hipaa-compliant-i-audited-awss-list-against-the-security-rule-4alf</link>
      <guid>https://dev.to/vadim_albarov/is-hipaa-eligible-the-same-as-hipaa-compliant-i-audited-awss-list-against-the-security-rule-4alf</guid>
      <description>&lt;p&gt;Here is a sentence I have heard, in some form, from three different engineers and one vendor sales deck: "We're fine, we only use HIPAA-eligible AWS services."&lt;/p&gt;

&lt;p&gt;I build healthcare infrastructure for a living, and every time I hear it I do the same mental translation: "We're fine, our data center signed a contract." It's true, it's necessary, and it answers roughly none of the questions a HIPAA auditor will ask you.&lt;/p&gt;

&lt;p&gt;So I decided to make the gap measurable. I took six AWS services that show up in practically every PHI-handling stack - S3, RDS for PostgreSQL, ECS on Fargate, KMS, the CloudTrail/CloudWatch audit pair, and the ALB - and audited each against the Security Rule's technical safeguards (&lt;a href="https://www.law.cornell.edu/cfr/text/45/164.312" rel="noopener noreferrer"&gt;45 CFR 164.312&lt;/a&gt;): what the regulation demands, what the service does by default, and what you must configure yourself.&lt;/p&gt;

&lt;p&gt;Spoiler: of the 35 concrete controls I mapped, the defaults covered 6.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "HIPAA-eligible" actually means
&lt;/h2&gt;

&lt;p&gt;Start with what AWS itself says. The &lt;a href="https://aws.amazon.com/compliance/hipaa-eligible-services-reference/" rel="noopener noreferrer"&gt;HIPAA-eligible services reference&lt;/a&gt; defines eligible services as those that may "create, receive, process, maintain, or transmit" electronic protected health information (ePHI) - &lt;em&gt;provided&lt;/em&gt; you have signed AWS's Business Associate Addendum (BAA) first. And then, in plain sight, the sentence everyone skips: "Customers still must configure these services consistent with HIPAA requirements."&lt;/p&gt;

&lt;p&gt;The fast primer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;covered entity&lt;/strong&gt; (provider, health plan, clearinghouse) or a &lt;strong&gt;business associate&lt;/strong&gt; (anyone handling PHI for them - your SaaS included, if hospitals are your customers) must comply with HIPAA. AWS becomes &lt;em&gt;your&lt;/em&gt; business associate when you sign the BAA - today a self-service click in AWS Artifact.&lt;/li&gt;
&lt;li&gt;The BAA covers AWS's side of the &lt;a href="https://aws.amazon.com/compliance/shared-responsibility-model/" rel="noopener noreferrer"&gt;shared responsibility model&lt;/a&gt;: data centers, hypervisors, service internals - security &lt;em&gt;of&lt;/em&gt; the cloud. Everything you configure - security &lt;em&gt;in&lt;/em&gt; the cloud - stays yours.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;Security Rule&lt;/strong&gt; is the part of HIPAA that talks to engineers. Its technical safeguards (&lt;a href="https://www.law.cornell.edu/cfr/text/45/164.312" rel="noopener noreferrer"&gt;164.312&lt;/a&gt;) name five standards: access control, audit controls, integrity, person or entity authentication, and transmission security.&lt;/li&gt;
&lt;li&gt;Implementation specifications are &lt;strong&gt;required&lt;/strong&gt; or &lt;strong&gt;addressable&lt;/strong&gt;. Addressable does not mean optional - it means implement it, implement a documented equivalent, or document a defensible reason you did neither. Encryption at rest and in transit are both addressable, and every healthcare shop I know treats them as required: "we decided not to encrypt the PHI" is not a paragraph anyone wants to defend in a breach investigation.&lt;/li&gt;
&lt;li&gt;And the kicker, straight from &lt;a href="https://aws.amazon.com/compliance/hipaa-compliance/" rel="noopener noreferrer"&gt;AWS's own HIPAA compliance page&lt;/a&gt;: "There is no HIPAA certification for a cloud service provider (CSP) such as AWS." There is no certification for you either. Only your configuration, your documentation, and eventually someone's audit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So "HIPAA-eligible" means exactly one thing: the service is covered by AWS's BAA, so AWS's side of the split is contractually handled. Whether &lt;em&gt;your&lt;/em&gt; side is handled is the rest of this article.&lt;/p&gt;

&lt;p&gt;Last year I distilled a few years of building this into a set of internal Terraform modules for my team - one hardened block per service, each control traced to its Security Rule citation. The most useful artifact of that project was the diff between "what the module enforces" and "what AWS gave us out of the box." That diff is essentially this article.&lt;/p&gt;

&lt;h2&gt;
  
  
  The audit
&lt;/h2&gt;

&lt;p&gt;Per service: what 164.312 asks, what you get by default, what you must add. Every default claim is sourced; where a default changed recently, I say when.&lt;/p&gt;

&lt;h3&gt;
  
  
  S3 - the poster child of improved-but-insufficient
&lt;/h3&gt;

&lt;p&gt;Credit where due: S3's defaults have genuinely improved. Since &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/01/amazon-s3-automatically-encrypts-new-objects" rel="noopener noreferrer"&gt;January 2023&lt;/a&gt;, every new object is encrypted with SSE-S3, and you cannot turn that off. Since &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/04/amazon-s3-security-best-practices-buckets-default" rel="noopener noreferrer"&gt;April 2023&lt;/a&gt;, new buckets get all four Block Public Access settings enabled and ACLs disabled. Two real 164.312(a) wins, for free. Now the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Transmission security (164.312(e), required standard):&lt;/strong&gt; S3 happily serves plaintext HTTP unless you attach a bucket policy denying &lt;code&gt;aws:SecureTransport = false&lt;/code&gt;. Nothing does this for you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integrity (164.312(c)):&lt;/strong&gt; versioning - your recovery path when an object is improperly altered or deleted - is off by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit controls (164.312(b)):&lt;/strong&gt; server access logging is off by default, and object-level API logging (who read which object) doesn't come from S3 at all - it needs CloudTrail data events, also off by default (see CloudTrail below).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encryption, but accountably:&lt;/strong&gt; SSE-S3 encrypts with a key you can't scope, rotate on your terms, or audit the use of. For PHI you want SSE-KMS with a customer-managed key, plus a policy denying uploads that request the wrong key - so a misconfigured client fails loudly instead of writing PHI under the wrong crypto.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One non-obvious lesson from building the S3 module: keep patient identifiers out of object keys. Keys leak into access logs, CloudTrail events, error messages, and URLs - surfaces your PHI inventory forgot about.&lt;/p&gt;

&lt;h3&gt;
  
  
  RDS for PostgreSQL - encryption is opt-in, still, in 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Encryption at rest (164.312(a)(2)(iv)):&lt;/strong&gt; &lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Overview.Encryption.html" rel="noopener noreferrer"&gt;storage encryption is not enabled by default&lt;/a&gt; at the API level - &lt;code&gt;StorageEncrypted&lt;/code&gt; defaults to false - and can only be set at creation time. Forget it in your Terraform, and the fix is a snapshot-copy-restore migration, not a flag flip.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transmission security (164.312(e)):&lt;/strong&gt; a moving target. For &lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/PostgreSQL.Concepts.General.SSL.html" rel="noopener noreferrer"&gt;RDS PostgreSQL 15 and later, &lt;code&gt;rds.force_ssl&lt;/code&gt; defaults to 1&lt;/a&gt;; for 14 and older it defaults to 0. It also lives in a parameter group, one console edit away from silently becoming 0 again. In my modules it's pinned at the engine, so a client that forgets &lt;code&gt;sslmode=require&lt;/code&gt; gets an error instead of a plaintext session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access control (164.312(a)(1)):&lt;/strong&gt; nothing stops you from making the instance publicly accessible or dropping it in a public subnet. Private placement, security-group-to-security-group ingress (no CIDR allowlists - identities beat address ranges), and &lt;code&gt;publicly_accessible = false&lt;/code&gt; are all decisions you must make and enforce.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integrity (164.312(c)):&lt;/strong&gt; &lt;a href="https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_DeleteInstance.html" rel="noopener noreferrer"&gt;deletion protection is disabled by default in the API and IaC&lt;/a&gt; (the console pre-checks it for production templates - your safety net depends on which tool created the database).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication (164.312(d)) and audit (164.312(b)):&lt;/strong&gt; RDS &lt;em&gt;can&lt;/em&gt; generate and rotate the master password in Secrets Manager so it never touches Terraform state; it can export connection logs to CloudWatch. Both are opt-in. By default the master password is wherever you put it, and connection history is a shrug.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  ECS on Fargate - the good defaults are the invisible ones
&lt;/h3&gt;

&lt;p&gt;Fargate's best HIPAA property doesn't appear in any console checkbox: there are no instances. No AMIs to patch, no SSH, no node agents - a whole slab of the malicious-software surface moves to AWS's side of the BAA. That and IAM's deny-by-default (a fresh task role can do nothing) are real defaults in your favor. Everything else is on you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Access control (164.312(a)(1)):&lt;/strong&gt; nothing prevents tasks in public subnets with public IPs. Private placement is a choice you codify.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication (164.312(d)):&lt;/strong&gt; plain environment variables appear in &lt;code&gt;DescribeTaskDefinition&lt;/code&gt; output and the console. Database credentials must go through the task definition's &lt;code&gt;secrets&lt;/code&gt; integration with Secrets Manager - injected at container start, never written down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unique identity (164.312(a)(2)(i)):&lt;/strong&gt; the execution role (ECS pulling images, injecting secrets) and the task role (your app) are separate identities - if you bother creating them separately instead of reusing one fat role.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit controls (164.312(b)):&lt;/strong&gt; my favorite gap in this audit. ECS Exec - an interactive shell inside a container that may hold PHI in memory - records, by default, only the &lt;code&gt;ExecuteCommand&lt;/code&gt; API call in CloudTrail. What happened &lt;em&gt;inside&lt;/em&gt; the session goes unrecorded unless you configure session I/O logging to an encrypted log group. A shell in a PHI container is an access event; out of the box, it's one with no transcript.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  KMS - the service that only helps if you invite it
&lt;/h3&gt;

&lt;p&gt;KMS is HIPAA-eligible, excellent, and entirely opt-in. Every encryption integration on this page uses KMS &lt;em&gt;if you wire it in&lt;/em&gt;; the eligible list has no opinion about whether you did. Even once you create customer-managed keys:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/kms/latest/developerguide/rotate-keys.html" rel="noopener noreferrer"&gt;Automatic key rotation is disabled by default&lt;/a&gt; for customer-managed keys (AWS-managed keys rotate yearly; the keys you control don't, until you say so).&lt;/li&gt;
&lt;li&gt;The default key policy hands the account root full access and calls it a day. Separating key &lt;em&gt;administration&lt;/em&gt; (policy, rotation, deletion - no decrypt) from key &lt;em&gt;use&lt;/em&gt; (decrypt, no administration) is a policy you write yourself. Worth writing: no single identity can then both reconfigure the crypto and read the data - exactly what 164.312(a)(1) wants from an access control story.&lt;/li&gt;
&lt;li&gt;One default that genuinely protects you: key deletion requires a 7-30 day waiting period. Deleting a key that ever encrypted PHI destroys the data with it, so in my modules the window is pinned to the 30-day maximum and &lt;code&gt;ScheduleKeyDeletion&lt;/code&gt; events page a human.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One design decision I'd defend anywhere: separate keys per purpose - data, logs, backups. Revoking access to PHI data should never break log delivery, and CloudTrail can hold the logs key without gaining any path to the data key.&lt;/p&gt;

&lt;h3&gt;
  
  
  CloudTrail and CloudWatch - the audit trail you think you have
&lt;/h3&gt;

&lt;p&gt;164.312(b) is a required standard with no addressable escape hatch: "implement... mechanisms that record and examine activity in information systems that contain or use ePHI." What you have by default is &lt;a href="https://docs.aws.amazon.com/awscloudtrail/latest/userguide/view-cloudtrail-events.html" rel="noopener noreferrer"&gt;90 days of management-event history&lt;/a&gt; in the console. That's it. Specifically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No durable trail. For records past 90 days you must create a trail delivering to S3. HIPAA's documentation retention requirement (164.316(b)(2)(i)) is six years; most programs apply the same horizon to audit artifacts. 90 days is not in the same universe.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/awscloudtrail/latest/userguide/how-cloudtrail-works.html" rel="noopener noreferrer"&gt;Data events are not logged by default&lt;/a&gt;. Management events tell you someone changed a bucket policy; they do &lt;em&gt;not&lt;/em&gt; tell you someone downloaded 10,000 patient records via &lt;code&gt;GetObject&lt;/code&gt;. Object-level audit is a separate, per-event-billed opt-in - in my modules it's a flag with a cost warning, but a flag someone consciously declines, not a silence nobody noticed.&lt;/li&gt;
&lt;li&gt;Log file validation (the SHA-256 digest chain that makes trail tampering detectable - the 164.312(c) integrity story for the audit trail itself) is a setting, not a given.&lt;/li&gt;
&lt;li&gt;On the CloudWatch side: log groups &lt;a href="https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/Working-with-log-groups-and-streams.html" rel="noopener noreferrer"&gt;never expire by default&lt;/a&gt; - retention is a policy you set deliberately - and while log data is SSE-encrypted at rest, &lt;a href="https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/encrypt-log-data-kms.html" rel="noopener noreferrer"&gt;customer-managed KMS encryption is optional&lt;/a&gt;, attached per log group.&lt;/li&gt;
&lt;li&gt;And the quiet half of 164.312(b): "record &lt;em&gt;and examine&lt;/em&gt;." Metric filters, alarms, anything that turns the trail into a signal a human reviews - none of it exists until you build it. An unread audit trail satisfies nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  ALB - TLS quality depends on which tool you clicked
&lt;/h3&gt;

&lt;p&gt;The load balancer is your transmission security (164.312(e)) chokepoint, with the most path-dependent defaults of the lot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Create an HTTPS listener in the console today and you get a modern TLS 1.3/1.2 policy. Create it through the &lt;a href="https://docs.aws.amazon.com/elasticloadbalancing/latest/application/describe-ssl-policies.html" rel="noopener noreferrer"&gt;CLI or CloudFormation and the default is &lt;code&gt;ELBSecurityPolicy-2016-08&lt;/code&gt;&lt;/a&gt;, which still accepts TLS 1.0 and 1.1. Your cipher floor depends on which tool provisioned the listener. Pin the policy explicitly and this stops being interesting.&lt;/li&gt;
&lt;li&gt;Nothing forces HTTPS to exist at all. A port-80 listener forwarding plaintext is perfectly deployable; the 80-to-443 redirect is a choice.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/elasticloadbalancing/latest/application/load-balancer-access-logs.html" rel="noopener noreferrer"&gt;Access logs are disabled by default&lt;/a&gt; - there goes 164.312(b) at the edge until you wire up the log bucket.&lt;/li&gt;
&lt;li&gt;Deletion protection: off by default. Deleting the front door of a PHI service should take two steps.&lt;/li&gt;
&lt;li&gt;And one my team debated for a week: by default TLS terminates at the ALB, and the hop to your tasks is HTTP inside the private network. A common, defensible posture - but a &lt;em&gt;documented risk decision&lt;/em&gt; your analysis has to own, not something the eligible list settled for you.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The scoreboard
&lt;/h2&gt;

&lt;p&gt;Counting controls involves judgment calls, so here's mine - each row counts the specific 164.312-relevant controls discussed above:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Controls mapped&lt;/th&gt;
&lt;th&gt;Covered by default&lt;/th&gt;
&lt;th&gt;You configure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;S3&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;2 (SSE-S3 since 2023, Block Public Access since 2023)&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RDS Postgres&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1 (forced TLS, PG15+ only)&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECS Fargate&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;2 (no node surface, IAM deny-by-default)&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KMS&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1 (deletion waiting period)&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CloudTrail / CloudWatch&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0 (90-day history is partial credit at best)&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ALB&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;35&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;29&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Roughly one control in six arrives configured. The other 29 are your job, on services that are all on the HIPAA-eligible list. And notice &lt;em&gt;which&lt;/em&gt; six: the freebies cluster in encryption at rest, where AWS has spent a decade raising defaults. Audit controls and transmission security - both required standards - score near zero.&lt;/p&gt;

&lt;p&gt;The floor is genuinely rising: S3 in 2023, PostgreSQL 15's &lt;code&gt;force_ssl&lt;/code&gt;, console TLS policies. But that's also the trap - your actual defaults depend on the year the feature shipped, the engine version, and whether the resource was born in the console or in CloudFormation. A compliance posture made of remembered defaults is a compliance posture made of trivia.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do about it: treat the list as a floor
&lt;/h2&gt;

&lt;p&gt;The eligible list answers one question: may PHI touch this service under our BAA? Everything after that yes is configuration, and configuration you rely on should be code:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Codify the invariants, and don't make them variables.&lt;/strong&gt; The rule that survived every review of our modules: if the Security Rule requires it, there is no toggle. Storage encryption, TLS enforcement, log validation, private placement - hardcoded. Variables exist only for things a compliant deployment may legitimately vary, like retention above the floor. A toggle someone can forget is a finding someone will write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let policy-as-code catch the drift.&lt;/strong&gt; &lt;a href="https://www.checkov.io/" rel="noopener noreferrer"&gt;checkov&lt;/a&gt; and &lt;a href="https://github.com/terraform-linters/tflint" rel="noopener noreferrer"&gt;tflint&lt;/a&gt; in CI flag most of the gaps in this article - unencrypted RDS, missing bucket policies, permissive TLS policies - before they exist. Native &lt;code&gt;terraform test&lt;/code&gt; can pin the invariants so a refactor can't quietly reintroduce a toggle. Where you do deviate, suppress with a written justification inline; that comment &lt;em&gt;is&lt;/em&gt; your addressable-specification documentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Map controls to citations, in the repo.&lt;/strong&gt; Every module of ours carries a table: citation, what the code enforces, and - the column that earns its keep - &lt;em&gt;what remains yours&lt;/em&gt;. No amount of Terraform covers the administrative safeguards: risk analysis, a log review procedure someone actually follows, restore testing, access reviews, BAAs with every other PHI-touching vendor. The code can only make the technical floor solid enough for your people to stand on.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;"HIPAA-eligible" was never a lie. It's a contract term that got promoted, somewhere between the sales deck and the standup, into a security property. The list tells you where PHI may go. The Security Rule tells you what must be true when it gets there. The distance between the two is not covered by anyone's BAA - it's covered by you, ideally in version control.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclaimer: I'm an infrastructure engineer, not a lawyer; this is engineering analysis of published regulations and AWS documentation, not legal advice. Defaults cited were verified in August 2026 and do change (sometimes for the better). Run the details past your compliance officer, who will find at least one thing here that your specific situation makes wrong.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What's the widest eligible-vs-compliant gap you've hit in the wild? I collect these.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>hipaa</category>
      <category>security</category>
      <category>healthtech</category>
    </item>
    <item>
      <title>"Oops, I Forgot to Tell You That's Dangerous": Claude Code Watched Me Wipe Production Redis - Then Helped Carve It Back Off the Disk"</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Wed, 12 Aug 2026 03:47:47 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/oops-i-forgot-to-tell-you-thats-dangerous-claude-code-watched-me-wipe-production-redis-then-300h</link>
      <guid>https://dev.to/vadim_albarov/oops-i-forgot-to-tell-you-thats-dangerous-claude-code-watched-me-wipe-production-redis-then-300h</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I enabled Redis AOF persistence by editing &lt;code&gt;redis.conf&lt;/code&gt; and restarting the service. On the test environment it worked fine. On production, all three nodes came back with &lt;strong&gt;zero keys&lt;/strong&gt;, and the empty dataset overwrote &lt;code&gt;dump.rdb&lt;/code&gt;. The missed step: you must enable AOF on the &lt;strong&gt;running&lt;/strong&gt; instance with &lt;code&gt;CONFIG SET appendonly yes&lt;/code&gt; &lt;em&gt;before&lt;/em&gt; putting it in the config file. We got the data back by carving the unlinked RDB file out of raw disk blocks with &lt;code&gt;dd&lt;/code&gt;. Oh, and the AI assistant that helped me plan the change knew about this trap the whole time - it just didn't mention it until &lt;em&gt;after&lt;/em&gt; the wipe. Here's the full story.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Some context first
&lt;/h2&gt;

&lt;p&gt;The stage for this story is a legacy healthcare project that's been running in production for years. Redis sits at the center of it, wearing two hats: it's the cache in front of SQL Server, and it's also the &lt;strong&gt;feature flag storage&lt;/strong&gt; - the backend seeds flags into Redis at startup, and both backend and client read them from there at runtime. Over the years the flag storage format evolved (more on that in a minute), the cache quietly accumulated datasets that exist &lt;em&gt;only&lt;/em&gt; in Redis, and the Redis setup itself - a bare tarball install with an untouched config - predates everyone's memory of who set it up. In other words: exactly the kind of system where nobody looks at the persistence settings until something forces them to.&lt;/p&gt;

&lt;p&gt;Something did.&lt;/p&gt;

&lt;h2&gt;
  
  
  It started with a feature flag that "flipped itself"
&lt;/h2&gt;

&lt;p&gt;After we published a new release, one of our feature flags apparently turned itself on. Nobody had touched it. The flag had existed for months with a default of &lt;code&gt;false&lt;/code&gt;, and suddenly the client started behaving as if it were &lt;code&gt;true&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I diffed the two release tags first: zero changes to flag definitions or defaults. So I went to where the flags actually live in Redis - and found the flag stored in &lt;strong&gt;two places&lt;/strong&gt;, disagreeing with each other:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;redis-cli GET  FeatureFlag:CrossWindowCommands   → &lt;span class="s2"&gt;"true"&lt;/span&gt;    &lt;span class="c"&gt;# legacy string key&lt;/span&gt;
redis-cli HGET FeatureFlag CrossWindowCommands   → &lt;span class="s2"&gt;"false"&lt;/span&gt;   &lt;span class="c"&gt;# current hash field&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The current implementation stores flags as fields in a &lt;code&gt;FeatureFlag&lt;/code&gt; hash, but there was also a &lt;strong&gt;stale legacy string key&lt;/strong&gt; left over from the old storage format. The &lt;code&gt;false&lt;/code&gt; in the hash, it turned out, wasn't the original state at all - a teammate had dug into the same problem before me and manually set it back to &lt;code&gt;false&lt;/code&gt;. Which made the mismatch the real hint: the legacy key still held &lt;code&gt;true&lt;/code&gt;, untouched, and the values in Redis always win - the startup migration copies legacy keys into the hash with &lt;code&gt;HSETNX&lt;/code&gt;, and the config default only applies when the field doesn't exist yet.&lt;/p&gt;

&lt;p&gt;So the explanation was less dramatic than a flag flipping itself: the flag had been &lt;code&gt;true&lt;/code&gt; in Redis all along. It just didn't matter, because for months no code path ever read it. Then the new client release shipped code that actually &lt;em&gt;did something&lt;/em&gt; with the flag, and a value that had been harmlessly wrong the whole time suddenly had teeth. An unused flag quietly became a used one, and it looked like a spontaneous flip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real finding: our persistence was a time bomb
&lt;/h2&gt;

&lt;p&gt;While investigating, I checked how the Redis instance itself was configured. It's a redis-stack tarball install running natively on RHEL under a custom systemd unit - no Docker, no operator, one master and two replicas with Sentinel. The persistence settings were all defaults - nothing configured at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;appendonly&lt;/span&gt; &lt;span class="n"&gt;no&lt;/span&gt;                    &lt;span class="c"&gt;# default
&lt;/span&gt;&lt;span class="n"&gt;save&lt;/span&gt; &lt;span class="m"&gt;3600&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt; &lt;span class="m"&gt;300&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt; &lt;span class="m"&gt;60&lt;/span&gt; &lt;span class="m"&gt;10000&lt;/span&gt;     &lt;span class="c"&gt;# default snapshot thresholds
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At our write volume, that works out to roughly one snapshot an hour. So every restart could silently roll the dataset back by up to an hour - and worse, a rollback could resurrect old values (like stale legacy flag keys) that had been deleted since the last snapshot. That's exactly the kind of environment where flags "change themselves" and nobody can explain why.&lt;/p&gt;

&lt;p&gt;So we made a decision: park the whodunit, fix the root cause. &lt;strong&gt;Enable AOF&lt;/strong&gt; (append-only file). With &lt;code&gt;aof-timestamp-enabled yes&lt;/code&gt; you even get a timestamped log of every write command - a free audit trail for exactly this class of mystery:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;appendonly&lt;/span&gt; &lt;span class="n"&gt;yes&lt;/span&gt;
&lt;span class="n"&gt;appendfsync&lt;/span&gt; &lt;span class="n"&gt;everysec&lt;/span&gt;
&lt;span class="n"&gt;aof&lt;/span&gt;-&lt;span class="n"&gt;timestamp&lt;/span&gt;-&lt;span class="n"&gt;enabled&lt;/span&gt; &lt;span class="n"&gt;yes&lt;/span&gt;
&lt;span class="n"&gt;auto&lt;/span&gt;-&lt;span class="n"&gt;aof&lt;/span&gt;-&lt;span class="n"&gt;rewrite&lt;/span&gt;-&lt;span class="n"&gt;min&lt;/span&gt;-&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="m"&gt;64&lt;/span&gt;&lt;span class="n"&gt;mb&lt;/span&gt;
&lt;span class="n"&gt;auto&lt;/span&gt;-&lt;span class="n"&gt;aof&lt;/span&gt;-&lt;span class="n"&gt;rewrite&lt;/span&gt;-&lt;span class="n"&gt;percentage&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Honest caveat: it's a &lt;em&gt;what-and-when&lt;/em&gt; audit, not a &lt;em&gt;who&lt;/em&gt; - the AOF doesn't record which client or user issued a command. Since this instance effectively has a single user anyway, that was good enough for a start.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing: everything green
&lt;/h2&gt;

&lt;p&gt;I tested the change end to end. First in a throwaway Docker container: enabled AOF, wrote some flags, and confirmed the AOF captured every command with &lt;code&gt;#TS:&lt;/code&gt; timestamps. I even benchmarked it: 100,000-op &lt;code&gt;redis-benchmark&lt;/code&gt; runs against two identical containers came back at ~181k SET/s &lt;em&gt;with&lt;/em&gt; AOF versus ~142k without. Yes, the AOF run scored higher - which tells you the difference is pure run-to-run noise. With &lt;code&gt;appendfsync everysec&lt;/code&gt;, the write-throughput cost is unmeasurable.&lt;/p&gt;

&lt;p&gt;Then on the test environment cluster. I enabled it &lt;strong&gt;live&lt;/strong&gt; on the running instance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;redis-cli CONFIG SET appendonly &lt;span class="nb"&gt;yes
&lt;/span&gt;redis-cli CONFIG SET aof-timestamp-enabled &lt;span class="nb"&gt;yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I did that live purely to check the concept: right after flipping it I changed a couple of values and watched them show up in the AOF file. A quick sanity check, nothing more - or so I thought. I had no idea this throwaway "test" step was the one doing the heavy lifting.&lt;/p&gt;

&lt;p&gt;I then added the same settings to &lt;code&gt;redis.conf&lt;/code&gt; so they'd survive a restart, rebooted the nodes one by one, and verified. Everything held: &lt;code&gt;aof_enabled:1&lt;/code&gt;, data intact, dummy writes visible in &lt;code&gt;appendonlydir/*.incr.aof&lt;/code&gt;, timestamps and all.&lt;/p&gt;

&lt;p&gt;The test environment was perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production: 0 keys
&lt;/h2&gt;

&lt;p&gt;On production I did what felt like the same change: edited &lt;code&gt;redis.conf&lt;/code&gt; on all three nodes, added &lt;code&gt;appendonly yes&lt;/code&gt;, stopped the sentinels (deliberately, so no surprise failovers mid-maintenance), and rebooted the nodes one by one.&lt;/p&gt;

&lt;p&gt;Why full reboots instead of just restarting the Redis service? Two reasons that felt responsible at the time. All three servers had quietly accumulated &lt;strong&gt;500+ pending OS updates&lt;/strong&gt;, including some severe security patches - so since I was already in a maintenance window, it seemed silly not to apply them and reboot in one go. And a real restart was part of the plan anyway: the whole point of putting &lt;code&gt;appendonly yes&lt;/code&gt; into the config file was to prove the setting survives a node going down, so I wanted to see it hold through a genuine reboot.&lt;/p&gt;

&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;redis-cli DBSIZE
(integer) 0
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Empty. Master and both replicas. And &lt;code&gt;dump.rdb&lt;/code&gt; on disk? Also empty - about 100 bytes of RDB header and nothing else. Roughly 290,000 production keys, gone: cache, feature flags, and one Redis-only ID-mapping dataset that has no SQL fallback at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The missed step
&lt;/h2&gt;

&lt;p&gt;Here's the trap, and if you take one thing from this article, take this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When Redis 7 starts with &lt;code&gt;appendonly yes&lt;/code&gt;, it loads the dataset from the AOF - and ignores &lt;code&gt;dump.rdb&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On the test environment I had run &lt;code&gt;CONFIG SET appendonly yes&lt;/code&gt; on the &lt;em&gt;live&lt;/em&gt; instance first. That triggers an AOF rewrite that builds the AOF base file &lt;strong&gt;from the data currently in memory&lt;/strong&gt;. By the time the test environment restarted, a fully populated AOF existed on disk, and Redis restored from it.&lt;/p&gt;

&lt;p&gt;On production I skipped straight to the config file. So the first startup with &lt;code&gt;appendonly yes&lt;/code&gt; found no AOF containing the dataset, initialized an empty one, and came up blank. From there the damage compounds automatically:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Redis starts empty and creates an empty AOF base.&lt;/li&gt;
&lt;li&gt;The next background save writes the empty dataset &lt;strong&gt;over &lt;code&gt;dump.rdb&lt;/code&gt;&lt;/strong&gt; - and Redis replaces the file via write-temp-then-&lt;code&gt;rename()&lt;/code&gt;, so the original file's inode is unlinked.&lt;/li&gt;
&lt;li&gt;Replicas reconnect and do a full resync &lt;strong&gt;from the empty master&lt;/strong&gt;, wiping themselves too.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Three nodes of redundancy, all faithfully replicating nothing. (Replication protects you from machine failure. It does not protect you from yourself - it &lt;em&gt;distributes&lt;/em&gt; your mistake.)&lt;/p&gt;

&lt;p&gt;Full disclosure: I wasn't doing this alone. I had an AI assistant (Claude) in the loop for the whole change - it designed the test rehearsal, wrote the config, and later drove the recovery. It knew about this trap. It just never said it out loud before the prod run. Here's the moment it understood what had happened, opening with its own mea culpa:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmzw11y6eo3n78bj3b5bf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmzw11y6eo3n78bj3b5bf.png" alt="The moment Claude diagnosed the wipe — and apologized for not flagging the trap before the prod run"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is the classic - and brutal - AOF-enablement trap, and I should have flagged it explicitly for prod: when Redis starts with &lt;code&gt;appendonly yes&lt;/code&gt; in the conf but no AOF files exist yet, it loads from the (nonexistent) AOF and ignores &lt;code&gt;dump.rdb&lt;/code&gt; entirely - it starts empty and creates an empty AOF. On the test environment this didn't bite because you ran &lt;code&gt;CONFIG SET appendonly yes&lt;/code&gt; &lt;em&gt;before&lt;/em&gt; rebooting - that built the AOF from the live dataset, so the reboot loaded it. On prod the conf was edited and rebooted directly - first boot came up empty, the next background save overwrote &lt;code&gt;dump.rdb&lt;/code&gt; with the empty dataset, and each replica that reconnected did a full resync from the empty master and wiped itself too. That's how all three ended up at 0 keys.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A correct diagnosis in one message - six and a half hours too late to be a warning. (To be fair: I never asked "what could go wrong with this rollout?" either. Neither of us rehearsed the failure mode; we only rehearsed success.)&lt;/p&gt;

&lt;p&gt;One honest footnote from the later forensics: block-level evidence showed that at one point during the rollout an AOF base &lt;em&gt;with&lt;/em&gt; the full dataset existed on disk for a while, and the actual wipe most likely happened on a subsequent restart in the sequence. The exact fatal moment is unrecoverable; the end state was unambiguous - empty AOF, empty RDB, empty replicas.&lt;/p&gt;

&lt;p&gt;And of course: no backup copy of &lt;code&gt;dump.rdb&lt;/code&gt; taken before the change, and no off-box backups. (I know. I know.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The recovery: your data is probably still on the disk
&lt;/h2&gt;

&lt;p&gt;Here's the insight that saved us: because Redis replaces &lt;code&gt;dump.rdb&lt;/code&gt; via &lt;code&gt;rename()&lt;/code&gt;, the old file wasn't overwritten in place. Its inode was unlinked, but the &lt;strong&gt;data blocks were still sitting in the free space of the filesystem&lt;/strong&gt;, waiting to be reclaimed.&lt;/p&gt;

&lt;p&gt;So the plan became: stop all writes, and go dig through the raw block device.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1 - freeze everything.&lt;/strong&gt; Stop Redis, don't reboot, minimize writes to the filesystem. Every write is a chance for the filesystem to reclaim the very blocks you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 - try an LVM snapshot.&lt;/strong&gt; The volume group had zero free extents, so no snapshot headroom. Plan B.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 - scan the block device for RDB signatures.&lt;/strong&gt; Every RDB file starts with the magic bytes &lt;code&gt;REDIS00&lt;/code&gt;. The obvious approach dies immediately on a 2 GB RAM box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo grep&lt;/span&gt; &lt;span class="nt"&gt;-abo&lt;/span&gt; &lt;span class="s1"&gt;'REDIS00'&lt;/span&gt; /dev/mapper/rhel-root
&lt;span class="nb"&gt;grep&lt;/span&gt;: memory exhausted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Block devices have no newlines; grep tries to buffer one infinite "line".) So: a tiny Python scanner that reads the device in 32 MB chunks with an overlap of the pattern length, and prints the absolute byte offset of every match:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;DEV&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/mapper/rhel-root&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;PAT&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REDIS00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;CHUNK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DEV&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;buffering&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;buf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;CHUNK&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;buf&lt;/span&gt;
        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PAT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PAT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;prev&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PAT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):]&lt;/span&gt;   &lt;span class="c1"&gt;# overlap so boundary-spanning hits aren't missed
&lt;/span&gt;        &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;CHUNK&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;... scanned &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; GB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it with &lt;code&gt;sudo&lt;/code&gt; and &lt;code&gt;nohup&lt;/code&gt;, go make coffee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4 - triage the hits.&lt;/strong&gt; The scan found 10 candidates. Each got a 128-byte peek, read-only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo dd &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/mapper/rhel-root &lt;span class="nv"&gt;iflag&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;skip_bytes &lt;span class="nv"&gt;skip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$OFFSET&lt;/span&gt; &lt;span class="nv"&gt;bs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;128 &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 | &lt;span class="nb"&gt;od&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two hits were, hilariously, the scanner's own output files (the data directory shared the LV we were scanning - which is also why carve output had to be shipped off-box). Four were post-incident debris: the empty AOF bases and near-empty dumps left behind by the wipe. That left four candidates of 35-38 MB with real data, distinguishable by the timestamps embedded in their headers: the last pre-incident &lt;code&gt;dump.rdb&lt;/code&gt;, and AOF base files from during the incident window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5 - carve the candidates to another machine.&lt;/strong&gt; Never write recovery output to the disk you're recovering from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo dd &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/mapper/rhel-root &lt;span class="nv"&gt;iflag&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;skip_bytes &lt;span class="nv"&gt;skip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;59609546752 &lt;span class="nv"&gt;bs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1M &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;64 &lt;span class="se"&gt;\&lt;/span&gt;
  | ssh user@rescue-box &lt;span class="s1"&gt;'cat &amp;gt; /home/redis/cand_D.rdb'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 6 - validate.&lt;/strong&gt; &lt;code&gt;redis-check-rdb&lt;/code&gt; on the first candidate reported a CRC error - a block near the tail of the file had already been partially reclaimed - but it still parsed all 290,826 keys. Usable in an emergency, so we set it aside and kept going. The next candidate, the freshest one, came back clean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;redis-check-rdb cand_D.rdb
&lt;span class="go"&gt;[offset 39730397] Checksum OK
[offset 39730397] \o/ RDB looks OK! \o/
[info] 290834 keys read
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 7 - verify the contents in a sandbox.&lt;/strong&gt; Never point production at an unverified file. A throwaway instance on the rescue box, isolated port, no config inheritance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;redis-server &lt;span class="nt"&gt;--port&lt;/span&gt; 6391 &lt;span class="nt"&gt;--dir&lt;/span&gt; /home/redis/rescue-test &lt;span class="nt"&gt;--dbfilename&lt;/span&gt; dump.rdb &lt;span class="se"&gt;\&lt;/span&gt;
             &lt;span class="nt"&gt;--appendonly&lt;/span&gt; no &lt;span class="nt"&gt;--daemonize&lt;/span&gt; &lt;span class="nb"&gt;yes
&lt;/span&gt;redis-cli &lt;span class="nt"&gt;-p&lt;/span&gt; 6391 DBSIZE
redis-cli &lt;span class="nt"&gt;-p&lt;/span&gt; 6391 HGETALL FeatureFlag
redis-cli &lt;span class="nt"&gt;-p&lt;/span&gt; 6391 shutdown nosave
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Real keys, real flag values. We had our database back - carved out of unallocated disk blocks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting it back - carefully this time
&lt;/h2&gt;

&lt;p&gt;Restoring into a Sentinel topology has its own trap: if a sentinel promotes a replica that still holds the &lt;em&gt;empty&lt;/em&gt; dataset while you're restoring the master, replication will happily sync the emptiness right back over your restored data. So the order was strict:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Confirm sentinels are stopped on &lt;strong&gt;all&lt;/strong&gt; nodes, then stop replicas, then the master.&lt;/li&gt;
&lt;li&gt;On the master: move the poisoned artifacts aside - never delete evidence mid-incident:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;mv &lt;/span&gt;appendonlydir appendonlydir.bad
   &lt;span class="nb"&gt;mv &lt;/span&gt;dump.rdb dump.rdb.empty
   &lt;span class="nb"&gt;cp &lt;/span&gt;cand_D.rdb dump.rdb
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Set &lt;code&gt;appendonly no&lt;/code&gt; in &lt;code&gt;redis.conf&lt;/code&gt; - the whole point is to force this boot to load &lt;code&gt;dump.rdb&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Start the master and hold your breath:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;   redis-cli DBSIZE
   (integer) 290834
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Start the replicas (they full-resync from the restored master - you do &lt;em&gt;not&lt;/em&gt; restore the RDB onto replicas), then the sentinels, last.&lt;/li&gt;
&lt;li&gt;And only now, enable AOF &lt;strong&gt;the right way&lt;/strong&gt;, on the live instance:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   redis-cli CONFIG SET appendonly &lt;span class="nb"&gt;yes&lt;/span&gt;
   &lt;span class="c"&gt;# wait for: INFO persistence → aof_rewrite_in_progress:0&lt;/span&gt;
   &lt;span class="c"&gt;# verify:   appendonlydir/ base file is megabytes, not ~100 bytes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and &lt;em&gt;then&lt;/em&gt; persist &lt;code&gt;appendonly yes&lt;/code&gt; into &lt;code&gt;redis.conf&lt;/code&gt;. Plus, finally, an immediate &lt;code&gt;scp&lt;/code&gt; of the recovered dump to another machine.&lt;/p&gt;

&lt;p&gt;290,834 keys - eight more than the older candidate, because the winning file was written a couple of minutes later in the timeline. Full recovery.&lt;/p&gt;

&lt;p&gt;One detail I still enjoy: the winning file wasn't the old &lt;code&gt;dump.rdb&lt;/code&gt; at all. It was an &lt;strong&gt;AOF base file&lt;/strong&gt; written &lt;em&gt;during&lt;/em&gt; the incident - the populated base from that brief window when everything still existed. In Redis 7 the AOF base is itself RDB-format, which is why we could rename it to &lt;code&gt;dump.rdb&lt;/code&gt; and boot straight from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Enable AOF on the running instance first.&lt;/strong&gt; &lt;code&gt;CONFIG SET appendonly yes&lt;/code&gt;, wait for the rewrite to finish, verify the AOF base has real size - &lt;em&gt;then&lt;/em&gt; edit the config file. Config-file-then-restart is the data-loss path, because Redis with AOF enabled ignores your RDB at boot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"It worked in test" only counts if test rehearsed the same sequence.&lt;/strong&gt; My test-environment run succeeded &lt;em&gt;because&lt;/em&gt; I happened to run &lt;code&gt;CONFIG SET&lt;/code&gt; first there. Same change, different order of operations, opposite outcome. Rehearse the runbook, not the end state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Copy the data files before touching persistence settings.&lt;/strong&gt; A 30-second &lt;code&gt;cp dump.rdb dump.rdb.$(date +%F)&lt;/code&gt; (and ideally an &lt;code&gt;scp&lt;/code&gt; off-box) would have turned a five-hour incident into a five-minute one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify between steps of a rolling change.&lt;/strong&gt; I rebooted three nodes back to back and only checked &lt;code&gt;DBSIZE&lt;/code&gt; at the end. One check after the first node would have contained the blast radius.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replication is not backup.&lt;/strong&gt; The replicas didn't save the data - they synchronized its destruction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you do lose a file: stop writes immediately.&lt;/strong&gt; Deleted ≠ gone. &lt;code&gt;rename()&lt;/code&gt;-replaced files leave their blocks in free space, and &lt;code&gt;dd&lt;/code&gt; + a signature scan can get them back - but only until something reuses those blocks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loose RDB-only persistence (&lt;code&gt;save 3600 1&lt;/code&gt;, &lt;code&gt;appendonly no&lt;/code&gt;) is its own slow-motion incident.&lt;/strong&gt; Ours had been quietly able to roll back up to an hour of writes on every restart - that's what made a feature flag look haunted in the first place, which is the only reason we went looking.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The irony of the whole story: the change that destroyed the database was the one meant to make it durable. The fix was correct; the &lt;em&gt;order&lt;/em&gt; was fatal. In operations, sequence is part of the change.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Have you ever had a "the fix caused the outage" incident? I'd love to hear about it in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>redis</category>
      <category>postmortem</category>
      <category>devops</category>
    </item>
    <item>
      <title>A Design Flaw in Claude Code's Documentation Skill: One Question, 265k–355k Tokens</title>
      <dc:creator>vadim albarov</dc:creator>
      <pubDate>Sun, 09 Aug 2026 16:34:33 +0000</pubDate>
      <link>https://dev.to/vadim_albarov/a-design-flaw-in-claude-codes-documentation-skill-one-question-265k-355k-tokens-15g7</link>
      <guid>https://dev.to/vadim_albarov/a-design-flaw-in-claude-codes-documentation-skill-one-question-265k-355k-tokens-15g7</guid>
      <description>&lt;p&gt;I typed a one-line question into Claude Code - &lt;code&gt;does fable use api billing?&lt;/code&gt; - and then ran &lt;code&gt;/context&lt;/code&gt; out of habit.&lt;/p&gt;

&lt;p&gt;30% of a 1,000,000-token context window was gone. One question, one answer, 295k tokens used.&lt;/p&gt;

&lt;p&gt;Then I ran the same prompt on a second laptop: &lt;strong&gt;40%&lt;/strong&gt;. Same question, same answer, ~100k tokens more.&lt;/p&gt;

&lt;p&gt;Same CLI version on both machines. I'll call them &lt;strong&gt;laptop 1&lt;/strong&gt; (30%) and &lt;strong&gt;laptop 2&lt;/strong&gt; (40%). This is the story of finding those 100k tokens. Spoiler: every theory I had was wrong, and the root cause turned out to be a design flaw in a single bundled skill - one you can partially work around by &lt;em&gt;adding&lt;/em&gt; files to a folder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The baseline
&lt;/h2&gt;

&lt;p&gt;Here's what &lt;code&gt;/context&lt;/code&gt; showed on laptop 1 after that single exchange:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;System prompt&lt;/td&gt;
&lt;td&gt;5.3k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System tools&lt;/td&gt;
&lt;td&gt;23.7k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP tools (59 tools, deferred)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory files&lt;/td&gt;
&lt;td&gt;318&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills (16 skill descriptions)&lt;/td&gt;
&lt;td&gt;2.2k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Messages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;264.6k&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free space&lt;/td&gt;
&lt;td&gt;703.9k&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things jumped out.&lt;/p&gt;

&lt;p&gt;First, the fixed overhead everyone worries about - MCP servers, skills, memory - is nearly free. 59 MCP tools sat at &lt;strong&gt;0 tokens&lt;/strong&gt; because Claude Code defers their schemas until they're actually used. All 16 skill &lt;em&gt;descriptions&lt;/em&gt; together cost 2.2k tokens.&lt;/p&gt;

&lt;p&gt;Second, the &lt;strong&gt;Messages&lt;/strong&gt; category held 264.6k tokens after a one-line question. The conversation itself was maybe 2k tokens. The rest arrived because my question mentioned a Claude model name, which triggered the built-in &lt;code&gt;claude-api&lt;/code&gt; skill - and a skill trigger doesn't just load instructions. This one injected its entire documentation payload into the conversation as a single message.&lt;/p&gt;

&lt;p&gt;Worth pausing on that: my question was a billing lookup that a single Google search answers in five seconds. And before you conclude "well, agent sessions are just expensive" - they aren't. As a control, I asked two other lookup questions in fresh sessions: &lt;em&gt;"what's the latest Node LTS version?"&lt;/em&gt; and &lt;em&gt;"MIT vs Apache 2.0?"&lt;/em&gt;. Both together cost &lt;strong&gt;7.5k tokens&lt;/strong&gt;. Ordinary questions are cheap.&lt;/p&gt;

&lt;p&gt;The 265k burn has one specific trigger: mentioning a Claude model name. That summons the built-in &lt;code&gt;claude-api&lt;/code&gt; skill, which has no notion of question weight - a casual pricing question gets the exact same multi-hundred-KB documentation payload as "implement a streaming tool-use loop." I paid a quarter of a million tokens for what one web search would have told me.&lt;/p&gt;

&lt;p&gt;On laptop 2, that same category showed &lt;strong&gt;354k tokens&lt;/strong&gt;. The 90k-token mystery lived entirely inside one message.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong theory #1: local CLAUDE.md files
&lt;/h2&gt;

&lt;p&gt;My first guess: laptop 2 has more project instruction files - &lt;code&gt;CLAUDE.md&lt;/code&gt;, memory, rules - quietly injected into context.&lt;/p&gt;

&lt;p&gt;Dead on arrival. Memory files accounted for 318 tokens on laptop 1, and the gap was ~90k tokens ≈ 360 KB of text. A CLAUDE.md would have to be a small book. More importantly, the gap persisted when I ran the prompt from an &lt;strong&gt;empty folder&lt;/strong&gt; on laptop 2 - no project files at all, still ~355k.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong theory #2: skill version
&lt;/h2&gt;

&lt;p&gt;Bundled skills live in a content-addressed cache:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="nv"&gt;%LOCALAPPDATA%&lt;/span&gt;\Temp\claude\bundled&lt;span class="na"&gt;-skills&lt;/span&gt;\&amp;lt;cli&lt;span class="na"&gt;-version&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;\&amp;lt;hash&amp;gt;\claude&lt;span class="na"&gt;-api&lt;/span&gt;\
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both machines ran CLI 2.1.226 with skill version 2.1.226. Same version... but &lt;strong&gt;different hashes&lt;/strong&gt;. Promising! Except when I compared the actual files, every doc folder was byte-identical - 847 KB of assets on both machines. The hash difference was real but, as it turned out later, a symptom rather than a cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong theory #3: the tokenizer
&lt;/h2&gt;

&lt;p&gt;This one almost fooled me, because the arithmetic was beautiful. I had tried Sonnet 5 at some point, and it ships a new tokenizer that produces roughly 30% more tokens for the same text.&lt;/p&gt;

&lt;p&gt;264.6k × 1.3 ≈ &lt;strong&gt;344k&lt;/strong&gt;. Almost exactly laptop 2's number. Same bytes, different ruler!&lt;/p&gt;

&lt;p&gt;Then I ran the control: same model, same effort, both machines. The gap survived. Tokenizer eliminated. (Keep this failure mode in mind though - token-budget intuitions genuinely don't transfer across model families.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Transcript forensics
&lt;/h2&gt;

&lt;p&gt;Same CLI, same skill assets, same model, same prompt - and different token counts. At this point the only honest move was to stop theorizing and diff the actual bytes.&lt;/p&gt;

&lt;p&gt;Claude Code writes every session to a JSONL transcript. Finding the heavy message takes one loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$proj&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Get-ChildItem&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.claude\projects"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Directory&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="n"&gt;Sort-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;LastWriteTime&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Descending&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Select-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-First&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Get-ChildItem&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$proj&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FullName&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Filter&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;*.&lt;/span&gt;&lt;span class="nf"&gt;jsonl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="n"&gt;Sort-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;LastWriteTime&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Descending&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Select-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-First&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;Get-Content&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;FullName&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;ForEach-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="kr"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;$_&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Length&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-gt&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="n"&gt;KB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Line &lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt; : &lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;]::&lt;/span&gt;&lt;span class="n"&gt;Round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;$_&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Length&lt;/span&gt;&lt;span class="n"&gt;/1KB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;) KB"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Laptop 1 transcript: one message of &lt;strong&gt;719 KB&lt;/strong&gt;. Laptop 2 transcript: one message of &lt;strong&gt;957 KB&lt;/strong&gt;. There's the gap - 238 KB of text, ~90k tokens at ~2.6 characters per token.&lt;/p&gt;

&lt;p&gt;The payload is the skill's documentation, embedded as &lt;code&gt;&amp;lt;doc path="..."&amp;gt;&lt;/code&gt; blocks. Extracting the doc lists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;regex&lt;/span&gt;&lt;span class="p"&gt;]::&lt;/span&gt;&lt;span class="n"&gt;Matches&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$line&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'&amp;lt;doc path=\\"([^\\"]+)\\"'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="n"&gt;ForEach-Object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$_&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Groups&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Value&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Laptop 1: 32 docs&lt;/strong&gt; - shared API docs + the &lt;code&gt;python/&lt;/code&gt; folder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Laptop 2: 65 docs&lt;/strong&gt; - shared API docs + &lt;strong&gt;all eight language folders&lt;/strong&gt;: Python, TypeScript, Go, Java, C#, PHP, Ruby, and cURL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All 32 docs the payloads had in common were byte-identical. The laptop 2 payload simply contained 33 extra language docs totaling 237 KB.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reveal
&lt;/h2&gt;

&lt;p&gt;Diffing the instruction text at the top of the two payloads (60 KB each, otherwise identical) surfaced exactly one difference:&lt;/p&gt;

&lt;p&gt;Laptop 1 payload:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;→ Refer to &lt;code&gt;python/claude-api/README.md&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Laptop 2 payload:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;→ Refer to &lt;code&gt;unknown/claude-api/README.md&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;No project language was auto-detected. Ask the user which language they are using, then refer to the matching docs below.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the whole mechanism. When the &lt;code&gt;claude-api&lt;/code&gt; skill triggers, Claude Code detects your project's language from the working directory and injects only that language's documentation. When it detects &lt;em&gt;nothing&lt;/em&gt; - an empty folder, a docs-only repo - the fallback is to inject &lt;strong&gt;documentation for every supported language&lt;/strong&gt;, because the model might need any of them.&lt;/p&gt;

&lt;p&gt;And why did laptop 1 detect Python? The folder I was in had a subdirectory containing a Python project with a &lt;code&gt;.venv&lt;/code&gt; - thousands of &lt;code&gt;.py&lt;/code&gt; files. A stray virtualenv saved me 90k tokens.&lt;/p&gt;

&lt;p&gt;The different cache hashes made sense now too: the skill bundle appears to be cached per &lt;em&gt;rendered variant&lt;/em&gt; - the assets are identical, but the instruction text differs by detection outcome, so each outcome gets its own hash directory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The irony, and the side effect worth knowing
&lt;/h2&gt;

&lt;p&gt;Here's my favorite part. When I got serious about controlling variables, I ran the "clean" experiment: empty folder, fresh session, same prompt. That methodologically pure setup is precisely what &lt;strong&gt;maximizes&lt;/strong&gt; the payload. The controlled experiment created the condition it was measuring.&lt;/p&gt;

&lt;p&gt;The flip side is a genuinely useful, if odd, side effect: &lt;strong&gt;having language context in your working folder reduces token usage.&lt;/strong&gt; Any file that lets Claude Code detect a language - a &lt;code&gt;.py&lt;/code&gt; file, a &lt;code&gt;package.json&lt;/code&gt;, a &lt;code&gt;pyproject.toml&lt;/code&gt;, even a leftover &lt;code&gt;.venv&lt;/code&gt; - pins the skill payload to one language's docs and cuts ~90k tokens (about 25% of the payload) off every &lt;code&gt;claude-api&lt;/code&gt; skill trigger. The final scoreboard, reproduced on both machines:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Working directory&lt;/th&gt;
&lt;th&gt;Docs injected&lt;/th&gt;
&lt;th&gt;Payload&lt;/th&gt;
&lt;th&gt;Messages after one question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Any folder with language markers&lt;/td&gt;
&lt;td&gt;32 (shared + one language)&lt;/td&gt;
&lt;td&gt;719 KB&lt;/td&gt;
&lt;td&gt;~265k tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Empty / language-less folder&lt;/td&gt;
&lt;td&gt;65 (shared + all 8 languages)&lt;/td&gt;
&lt;td&gt;957 KB&lt;/td&gt;
&lt;td&gt;~355-360k tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So if you're about to ask Claude Code API questions from some scratch directory: don't. Run it from a real project - or drop a single &lt;code&gt;pyproject.toml&lt;/code&gt; (or the equivalent for your language) into the scratch folder first. It reads as a joke, but it's a measurable 90k-token difference per session, and on smaller context windows it's not funny at all: &lt;strong&gt;the all-languages payload alone wouldn't fit in a 200k-token context window.&lt;/strong&gt; This entire question is only answerable on a 1M-window model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is a flaw, not a feature
&lt;/h2&gt;

&lt;p&gt;Skills in Claude Code are designed around &lt;em&gt;progressive disclosure&lt;/em&gt;: a one-line description sits in context (all 16 bundled skills together cost 2.2k tokens), and the full instructions load only when triggered. The &lt;code&gt;claude-api&lt;/code&gt; skill follows that pattern for its trigger - and then abandons it entirely for its content: instead of letting the model read the docs it needs on demand, it eagerly injects the whole documentation set as a single message. It's the only bundled skill big enough to need a disk cache at all (847 KB; the other 15 are trivially small).&lt;/p&gt;

&lt;p&gt;The no-language fallback makes it worse in a way that's almost comic. The injected instruction text literally says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No project language was auto-detected. &lt;strong&gt;Ask the user which language they are using&lt;/strong&gt;, then refer to the matching docs below.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;...but by the time the model can ask, all eight languages' documentation has already been paid for. The question comes after the purchase. Any of the obvious designs - ask first and inject one language; inject the shared docs and let the model read language files on demand; scale the payload to the question - would cap the cost at the detected-language level or below.&lt;/p&gt;

&lt;p&gt;To be fair about scope: this is one skill, in one CLI version (2.1.226), and the payload is genuinely useful when you're writing code against the Claude API in a detected-language project. The flaw is the eager all-languages fallback and the trigger's insensitivity to question weight - both fixable upstream without losing what the skill is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Skill payloads dominate context cost.&lt;/strong&gt; Everything people usually blame - MCP servers, memory files, system prompts - added up to ~31k tokens for me. One skill trigger added 265-360k. Audit accordingly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The flaw is specific, not general agent overhead.&lt;/strong&gt; Control questions in fresh sessions ("latest Node LTS?", "MIT vs Apache 2.0?") cost 7.5k tokens combined. The expensive trigger is mentioning a Claude model or API name. For quick Claude pricing/docs lookups, use a web search - inside Claude Code, the skill &lt;em&gt;will&lt;/em&gt; fire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deferred loading works.&lt;/strong&gt; 59 MCP tools at 0 tokens until used. If your setup loads MCP schemas eagerly, that's worth fixing, but it wasn't my problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/context&lt;/code&gt; tells you the category; the transcript tells you the culprit.&lt;/strong&gt; The JSONL line-length trick above takes 30 seconds and points at the exact message.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same CLI version ≠ same context cost.&lt;/strong&gt; The cost depends on runtime conditions - in this case, what's sitting in your working directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Language markers in the folder are a token optimization.&lt;/strong&gt; Unintuitive, but reproducible: give the language detector something to find, and the skill injects one language's docs instead of eight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token intuitions don't transfer across model families.&lt;/strong&gt; My tokenizer theory was wrong &lt;em&gt;this time&lt;/em&gt;, but the ~30% Sonnet 5 difference is real - re-baseline when you switch models.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Everything above was measured on Claude Code 2.1.226 (bundled &lt;code&gt;claude-api&lt;/code&gt; skill 2.1.226) with Fable 5 and Sonnet 5, on two Windows machines. The behavior may well change in future releases - arguably the fallback should ask &lt;em&gt;before&lt;/em&gt; injecting 237 KB of polyglot documentation - but the audit method will keep working regardless.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Reproduce it yourself: ask Claude Code any Claude-API question from an empty folder, run &lt;code&gt;/context&lt;/code&gt;, then do the same from inside a Python or TypeScript repo and compare the Messages category.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;UPD: The 2.1.234 update fixed the issue (at least for me). Now the claude‑api skill is no longer bloated and does not exhaust your limits anymore&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>debugging</category>
    </item>
  </channel>
</rss>
