<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: oji - building AI in public</title>
    <description>The latest articles on DEV Community by oji - building AI in public (@masaoshimadaopen).</description>
    <link>https://dev.to/masaoshimadaopen</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4013207%2F62889ff7-f41e-4077-9836-3fafe971b8ce.jpg</url>
      <title>DEV Community: oji - building AI in public</title>
      <link>https://dev.to/masaoshimadaopen</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/masaoshimadaopen"/>
    <language>en</language>
    <item>
      <title>300問の問題集をAIで自動生成！ 250問目でAPIエラー、75分が水の泡に。チェックポイント機構で“途中再開”を実装した話</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Sun, 04 Oct 2026 23:30:18 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/300wen-nowen-ti-ji-woaidezi-dong-sheng-cheng-250wen-mu-deapiera-75fen-gashui-nopao-ni-tietukupointoji-gou-detu-zhong-zai-kai-woshi-zhuang-sitahua-4cpb</link>
      <guid>https://dev.to/masaoshimadaopen/300wen-nowen-ti-ji-woaidezi-dong-sheng-cheng-250wen-mu-deapiera-75fen-gashui-nopao-ni-tietukupointoji-gou-detu-zhong-zai-kai-woshi-zhuang-sitahua-4cpb</guid>
      <description>&lt;p&gt;最近、ある資格試験の勉強をしてて、公式の問題集だけじゃ足りないなと思ってた。じゃあ作るか、と。GenAIを使えば、類題なんていくらでも作れる時代だし。&lt;/p&gt;

&lt;p&gt;早速、試験のシラバスを食わせて、章ごとに問題と解説を300問分生成するスクリプトを書いた。&lt;br&gt;
&lt;code&gt;python generate_questions.py&lt;/code&gt; を叩いて、あとは待つだけ。プログレスバーがぐんぐん伸びていくのを見ながら、「いやー、便利な時代になったもんだ」なんて思ってた。&lt;/p&gt;

&lt;p&gt;実行開始から75分後。プログレスバーは8割を超えたあたり。そろそろ終わるかな、とコンソールを覗き込んだら、見たくない赤い文字が表示されてた。&lt;/p&gt;

&lt;p&gt;&lt;code&gt;requests.exceptions.ConnectionError: ('Connection aborted.', ConnectionResetError(104, 'Connection reset by peer'))&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;まあ、API呼び出しのエラー自体はよくあること。問題はその後。成果物が出力されるはずのディレクトリを見たら、ファイルが空っぽ。&lt;/p&gt;

&lt;p&gt;……え？&lt;/p&gt;

&lt;p&gt;そう。250問近くまで生成したはずのデータが、どこにもない。75分という時間と、それなりの額のAPIトークンが、一瞬で溶けた。これは、やばい。&lt;/p&gt;
&lt;h3&gt;
  
  
  原因：ナイーブすぎた全件一括処理
&lt;/h3&gt;

&lt;p&gt;原因は完全に自分の設計ミス。実装がナイーブすぎた。&lt;br&gt;
書いたコードは、ざっくりこんな感じだった。&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; 空のリスト &lt;code&gt;all_questions = []&lt;/code&gt; を用意。&lt;/li&gt;
&lt;li&gt; &lt;code&gt;for&lt;/code&gt; ループで300回、GenAIのAPIを叩く。&lt;/li&gt;
&lt;li&gt; 生成された問題を &lt;code&gt;all_questions.append(new_question)&lt;/code&gt; でリストに追加していく。&lt;/li&gt;
&lt;li&gt; ループが全部終わったら、&lt;code&gt;all_questions&lt;/code&gt; の中身をまとめてJSONファイルに書き出す。&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;この設計の何がまずいか。それは、途中で一度でも処理がコケたら、メモリ上にあった &lt;code&gt;all_questions&lt;/code&gt; の中身が全部消し飛ぶこと。&lt;/p&gt;

&lt;p&gt;外部APIを叩くような長時間バッチ処理で、この実装は致命的だった。ネットワークは気まぐれだし、APIサーバーだってたまには機嫌が悪くなる。そういう「中断」を全く想定してなかったのが敗因。&lt;/p&gt;
&lt;h3&gt;
  
  
  修正：キリのいいところで保存・再開する
&lt;/h3&gt;

&lt;p&gt;このままだと、生成が終わるまで神に祈るゲームになる。エンジニアとしてそれはありえない。&lt;br&gt;
そこで、「チェックポイント機構」を実装することにした。&lt;/p&gt;

&lt;p&gt;考え方はシンプルで、「キリのいいところまで処理が進んだら、その時点での成果をファイルに保存しておく」というもの。そして、もしスクリプトが中断しても、次に起動したときに「どこまで終わってたか」をファイルから読み込んで、続きから再開できるようにする。&lt;/p&gt;

&lt;p&gt;具体的には、章ごとの問題生成が終わるたびに、結果をファイルに追記していく方式にした。ファイル形式はJSONL（JSON Lines）がこういう用途に便利。1行が1つの独立したJSONオブジェクトになってるから、追記がめちゃくちゃ楽。&lt;/p&gt;

&lt;p&gt;修正後のコードはこんな感じ。&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="c1"&gt;# exam_slug は 'shikaku_xyz' みたいな試験の識別子
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_with_checkpoint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exam_slug&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# 試験のシラバス（章のリスト）を読み込む
&lt;/span&gt;    &lt;span class="n"&gt;chapters&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_syllabus&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exam_slug&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# 進捗を保存するチェックポイントファイル
&lt;/span&gt;    &lt;span class="n"&gt;checkpoint_file&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app_factory/exams/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exam_slug&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/generated.jsonl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="c1"&gt;# すでに完了した章のIDを保持するセット
&lt;/span&gt;    &lt;span class="n"&gt;completed_chapters&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="c1"&gt;# チェックポイントファイルが存在すれば、中身を読んで完了済みリストを作る
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;checkpoint_file&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;checkpoint_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="c1"&gt;# 1行ずつJSONをパースして、chapter_id をセットに追加
&lt;/span&gt;                &lt;span class="n"&gt;completed_chapters&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chapter_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

    &lt;span class="c1"&gt;# シラバスの全チャプターをループ
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chapter&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chapters&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# もし、この章がすでに完了済みならスキップ
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;chapter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;completed_chapters&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Skipping completed chapter: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;chapter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="c1"&gt;# 未処理の章なので、AI生成APIを呼び出す
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Generating for chapter: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;chapter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;new_items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_genai_api&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chapter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

        &lt;span class="c1"&gt;# 生成が終わったら、結果をチェックポイントファイルに「追記」する
&lt;/span&gt;        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;checkpoint_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;new_items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="c1"&gt;# どの章の成果物かわかるように chapter_id を付与
&lt;/span&gt;                &lt;span class="n"&gt;item_with_meta&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chapter_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;chapter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
                &lt;span class="c1"&gt;# JSON文字列に変換して、改行をつけて書き込む
&lt;/span&gt;                &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item_with_meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;このコードのポイントは3つ。&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;起動時に進捗を確認&lt;/strong&gt;: スクリプトが始まったら、まず &lt;code&gt;checkpoint_file&lt;/code&gt; が存在するかチェック。あれば中身を読んで、どの章が完了済みかを &lt;code&gt;completed_chapters&lt;/code&gt; セットに記録する。&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;処理済みタスクのスキップ&lt;/strong&gt;: &lt;code&gt;for&lt;/code&gt; ループの中で、処理対象の章が &lt;code&gt;completed_chapters&lt;/code&gt; に含まれていたら、APIを叩かずにスキップする。&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;進捗の永続化&lt;/strong&gt;: 1つの章の処理が終わるたびに、&lt;code&gt;open(..., "a")&lt;/code&gt; で追記モードでファイルを開き、結果を書き込む。これで、たとえ次の章でエラーが起きても、そこまでの成果はファイルに残る。&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;この修正で、スクリプトは何度実行しても大丈夫な「冪等性（べきとうせい）」を持つようになった。途中で止まっても、もう一度同じコマンドを叩けば、何事もなかったかのように続きから再開してくれる。&lt;/p&gt;

&lt;h3&gt;
  
  
  学び：中断は必ず起こる前提で設計する
&lt;/h3&gt;

&lt;p&gt;今回の失敗で学んだのは、長時間実行されるバッチ処理、特に外部APIみたいな不確定要素が絡むスクリプトは、「中断は必ず起こる」という前提で設計しなきゃいけない、ということ。&lt;/p&gt;

&lt;p&gt;「全部終わったらまとめて保存」はシンプルだけど、脆い。&lt;br&gt;
「少しずつ進めて、都度保存」は少し手間だけど、頑健。&lt;/p&gt;

&lt;p&gt;この「チェックポイントを設けて途中再開できるようにする」という考え方は、AIのコンテンツ生成だけじゃなくて、大規模なデータ処理、Webスクレイピング、インフラのプロビジョニング（Terraformとかも内部的にはstateファイルで似たようなことやってる）とか、いろんな場面で応用が効く。&lt;/p&gt;

&lt;p&gt;個人開発のツールなんて、動けばいいと思いがちだけど、こういう一手間を加えておくだけで、運用がえぐいほど安定する。75分を無駄にしたのは痛かったけど、おかげで今後の開発で同じ轍を踏むことはなくなるはず。良い勉強代だったと思うことにした。&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>webdev</category>
    </item>
    <item>
      <title>My Python patch failed silently: the `\n` in a triple-quoted string became a literal newline!</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Sat, 03 Oct 2026 23:30:20 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-python-patch-failed-silently-the-n-in-a-triple-quoted-string-became-a-literal-newline-4gah</link>
      <guid>https://dev.to/masaoshimadaopen/my-python-patch-failed-silently-the-n-in-a-triple-quoted-string-became-a-literal-newline-4gah</guid>
      <description>&lt;p&gt;Hey everyone, it's your friendly neighborhood dev, 38, building AI trading bots as a side hustle.&lt;/p&gt;

&lt;p&gt;One evening after my main job, I was updating library dependencies for my custom bot. I needed to tweak a specific library's behavior, so I wrote some code to directly patch its source.&lt;/p&gt;

&lt;p&gt;The usual drill: I stored the diff content for the &lt;code&gt;patch&lt;/code&gt; command in a Python triple-quoted string (like a here-document), then executed it via a subprocess.&lt;/p&gt;

&lt;p&gt;But it just wouldn't work.&lt;/p&gt;

&lt;p&gt;No errors, nothing. The library's behavior remained completely unchanged after the "patch." The worst kind of bug: the silent failure. I probably spent about two hours debugging this.&lt;/p&gt;

&lt;p&gt;Long story short, the culprit was a specific detail of Python's string literal behavior. The &lt;code&gt;\n&lt;/code&gt; I wrote inside the here-document was being converted into an &lt;em&gt;actual newline character&lt;/em&gt;, instead of being treated as the literal two-character sequence "backslash and n."&lt;/p&gt;




&lt;h3&gt;
  
  
  The Symptom: Patch Fails Silently
&lt;/h3&gt;

&lt;p&gt;Here’s a simplified version of what I was doing. Preparing a diff text as a string and piping it to &lt;code&gt;subprocess&lt;/code&gt; for the &lt;code&gt;patch&lt;/code&gt; command.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# The problematic code (conceptual)
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;

&lt;span class="c1"&gt;# Path to the library file to be patched
&lt;/span&gt;&lt;span class="n"&gt;target_file&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/path/to/some/library/file.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Dynamically generated patch (defined as a here-document)
&lt;/span&gt;&lt;span class="n"&gt;patch_content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
--- a/file.py
+++ b/file.py
@@ -123,7 +123,7 @@
 class SomeClass:
     def some_method(self):
-        return &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;old_string&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;
+        return &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;new_string_with_patch&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="c1"&gt;# Execute the patch command
&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;patch&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_file&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;patch_content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Check the result
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returncode&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Patch application failed!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Patch applied (or so I thought...)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running this code, &lt;code&gt;returncode&lt;/code&gt; was &lt;code&gt;0&lt;/code&gt;. &lt;code&gt;stderr&lt;/code&gt; was empty. It printed "Patch applied." But when I actually checked &lt;code&gt;target_file&lt;/code&gt;, its content was untouched.&lt;/p&gt;

&lt;p&gt;Initially, I suspected issues with the &lt;code&gt;patch&lt;/code&gt; command's path or permissions. But creating a &lt;code&gt;.patch&lt;/code&gt; file with the exact same content and applying it manually from the command line worked perfectly.&lt;/p&gt;

&lt;p&gt;This meant the &lt;code&gt;patch_content&lt;/code&gt; being passed from the Python script was somehow malformed.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Cause: &lt;code&gt;\n&lt;/code&gt; became a literal newline
&lt;/h3&gt;

&lt;p&gt;With the hypothesis that "the string content is wrong," I started by printing &lt;code&gt;patch_content&lt;/code&gt; to the console.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;--- a/file.py
&lt;/span&gt;&lt;span class="gi"&gt;+++ b/file.py
&lt;/span&gt;&lt;span class="p"&gt;@@ -123,7 +123,7 @@&lt;/span&gt;
 class SomeClass:
     def some_method(self):
&lt;span class="gd"&gt;-        return "old_string
&lt;/span&gt;"
&lt;span class="gi"&gt;+        return "new_string_with_patch
&lt;/span&gt;"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance, nothing seemed wrong. The diff format looked correct. &lt;/p&gt;

&lt;p&gt;This is where I got stuck. But on closer inspection, something felt off. The line that should have been &lt;code&gt;return "old_string\n"&lt;/code&gt; was instead &lt;code&gt;return "old_string&lt;/code&gt; followed by a newline, with the closing &lt;code&gt;"&lt;/code&gt; on the next line.&lt;/p&gt;

&lt;p&gt;That's when it clicked. The &lt;code&gt;\n&lt;/code&gt; was being interpreted as an escape sequence and converted into an actual newline character (LF).&lt;/p&gt;

&lt;p&gt;Python's triple-quoted strings &lt;code&gt;"""..."""&lt;/code&gt; interpret escape sequences like &lt;code&gt;\n&lt;/code&gt; and &lt;code&gt;\t&lt;/code&gt; just like regular strings. In this case, I wanted to include the &lt;strong&gt;literal string literal&lt;/strong&gt; &lt;code&gt;return "old_string\n"&lt;/code&gt; as part of the patch diff. But Python's interpreter, being helpful, converted &lt;code&gt;\n&lt;/code&gt; to a newline, resulting in an invalid diff format that the &lt;code&gt;patch&lt;/code&gt; command couldn't understand.&lt;/p&gt;

&lt;p&gt;This is a nasty one. It's a spec I &lt;em&gt;know&lt;/em&gt;, yet when you're tired, it's easy to fall into this trap.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix: Use raw strings &lt;code&gt;r"""..."""&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Once the cause was identified, the fix was simple: prevent backslashes within the string literal from being escaped.&lt;/p&gt;

&lt;p&gt;There are two ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; Escape the backslash itself (&lt;code&gt;\\n&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt; Use a raw string&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Since I didn't want &lt;em&gt;any&lt;/em&gt; escape sequences interpreted in the entire patch content, using a raw string made the intention clearest.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Corrected code
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;

&lt;span class="n"&gt;target_file&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/path/to/some/library/file.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Use a raw string (r"""...
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;) to disable escape sequences
patch_content = r&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="o"&gt;---&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;
&lt;span class="o"&gt;+++&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;
&lt;span class="o"&gt;@@&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;123&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="mi"&gt;123&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="o"&gt;@@&lt;/span&gt;
 &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SomeClass&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
     &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;some_method&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;old_string&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;new_string_with_patch&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;

# Rest of the code is the same
proc = subprocess.run(...)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Just add an &lt;code&gt;r&lt;/code&gt; before the opening quotes of the string. Now, &lt;code&gt;\n&lt;/code&gt; inside &lt;code&gt;patch_content&lt;/code&gt; is no longer a newline character, but simply the two-character string "backslash and n."&lt;/p&gt;

&lt;p&gt;With this fix, the patch applied successfully. What a journey...&lt;/p&gt;

&lt;h3&gt;
  
  
  Lesson Learned: Use &lt;code&gt;repr()&lt;/code&gt; when dealing with code as strings
&lt;/h3&gt;

&lt;p&gt;The takeaway from this experience is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When treating text that contains meaningful escape sequences (like code or configuration files) as a string, either use raw strings or pay extreme attention to how escape sequences are handled.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here-documents are especially convenient for multi-line text, which also makes them prone to this specific trap.&lt;/p&gt;

&lt;p&gt;Another lesson learned is about debugging. Because I used &lt;code&gt;print()&lt;/code&gt; to output the string, the visual difference was subtle, delaying discovery. In such cases, &lt;code&gt;repr()&lt;/code&gt; should have been my go-to.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;print(repr(patch_content))&lt;/code&gt; would instantly show how Python interprets the string.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Example output of print(repr(patch_content))
&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;--- a/file.py&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;+++ b/file.py&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;...return &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;old_string&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This way, it would be crystal clear whether it was &lt;code&gt;\\n&lt;/code&gt; (as intended) or &lt;code&gt;\n&lt;/code&gt; (unintended). Ugh, I should have done this from the start.&lt;/p&gt;

&lt;p&gt;In personal projects, especially side hustles, these minor gotchas can easily eat up several hours. But accumulating these failure logs, one by one, is ultimately the fastest way to improve.&lt;/p&gt;

&lt;p&gt;Hope this helps anyone else who might fall into a similar trap.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>My AI Agent Kept Missing Deadlines and Double-Posting — Here's How I Fixed It with JSON Interfaces and Idempotency Checks</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Fri, 02 Oct 2026 23:30:19 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-ai-agent-kept-missing-deadlines-and-double-posting-heres-how-i-fixed-it-with-json-interfaces-4enh</link>
      <guid>https://dev.to/masaoshimadaopen/my-ai-agent-kept-missing-deadlines-and-double-posting-heres-how-i-fixed-it-with-json-interfaces-4enh</guid>
      <description>&lt;p&gt;Hey everyone, it's your resident 38-year-old developer here, moonlighting with AI automated trading bots.&lt;/p&gt;

&lt;p&gt;My little pleasure after a long day at my main job is checking the logs of my custom AI agent. Every night, it analyzes a specific market and records its preliminary trading decisions in JSON format. Pretty smart little helper, usually.&lt;/p&gt;

&lt;p&gt;But one morning, checking the logs as usual, I found a mess:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Three identical trading decision records for the same date.&lt;/li&gt;
&lt;li&gt;  Decisions that should have been logged by 23:59 the previous day were somehow being attempted at 2:30 AM, failing with a "deadline exceeded" error.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a moment, I panicked, thinking the AI had bugged out and gone rogue. An agent designed for autonomous operation was repeating tasks and missing critical deadlines. This was a pretty serious situation.&lt;/p&gt;

&lt;p&gt;Today, I want to share the nitty-gritty of how I prevented this "AI meltdown."&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: AI Is Too Obedient, and Unaware of Execution Environment Quirks
&lt;/h3&gt;

&lt;p&gt;Upon closer investigation, the AI itself hadn't gone crazy. The root cause lay in the overall "mechanism" that ran the AI agent.&lt;/p&gt;

&lt;p&gt;My system was a simple setup: a Windows Task Scheduler would kick off a Python script at a fixed time every night, and that script would call an LLM (Claude) to generate the decision.&lt;/p&gt;

&lt;p&gt;The problem arose when, for some reason, the Task Scheduler re-executed the task. For example, a brief network outage, or a slight delay in response due to machine load. In such cases, the OS might "helpfully" retry the task, thinking, "Oh, did that task fail? I'll just run it again."&lt;/p&gt;

&lt;p&gt;The AI, however, knows nothing about these OS-level nuances. It simply received the command to "make a decision and record it," and dutifully went through its thought process from scratch each time, writing down the results. Even if I added "execute only once" to the prompt, it wouldn't prevent OS-level retries.&lt;/p&gt;

&lt;p&gt;In essence, entrusting the entire "make a decision and write it down" process to the AI was the fundamental design flaw.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix: Separating "Context Gathering" from "Decision Writing"
&lt;/h3&gt;

&lt;p&gt;To address this, I clearly separated the responsibilities of the Python script (&lt;code&gt;decide.py&lt;/code&gt;), which served as the interface to the AI, into two distinct parts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: The AI retrieves "context" as JSON for its decision-making.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;First, I created a &lt;code&gt;--context&lt;/code&gt; mode, specifically for retrieving the information the AI needs to think.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;py &lt;span class="nt"&gt;-3&lt;/span&gt;.12 C:/path/to/my/project/decide.py &lt;span class="nt"&gt;--context&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Executing this command returns a JSON object to standard output, summarizing current market information, historical data, and so on. This command never alters the system's state, no matter how many times it's run. It just reads data. So, it's safe.&lt;/p&gt;

&lt;p&gt;The AI agent first understands the current situation using this command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: The AI writes its "decision result" in JSON format.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Next, I prepared a &lt;code&gt;--write&lt;/code&gt; mode that simply accepts the AI-generated decision JSON and saves it to a file or database.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;py &lt;span class="nt"&gt;-3&lt;/span&gt;.12 C:/path/to/my/project/decide.py &lt;span class="nt"&gt;--model&lt;/span&gt; claude-fable-5-1 &lt;span class="nt"&gt;--write&lt;/span&gt; &lt;span class="s1"&gt;'{"session_date": "2024-05-20", "decision": "buy", "reason": "...", "confidence": 0.85}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The crucial part is that I &lt;strong&gt;implemented rigorous pre-write checks within this &lt;code&gt;--write&lt;/code&gt; script.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Part of decide.py (conceptual)
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;argparse&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;write_decision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# 1. Deadline Check
&lt;/span&gt;    &lt;span class="n"&gt;session_date_str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;session_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strptime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;session_date_str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%Y-%m-%d&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;date&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# If the deadline has passed, don't write and raise an error
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;hour&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;23&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="c1"&gt;# Example: 23:00 is the deadline
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deadline passed for session &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;session_date_str&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 2. Idempotency Check (prevent duplicate writes)
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;record_exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;session_date&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Record for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;session_date_str&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; already exists. Skipping.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="c1"&gt;# If it exists, do nothing and exit successfully
&lt;/span&gt;
    &lt;span class="c1"&gt;# Only if all checks pass, execute the write operation
&lt;/span&gt;    &lt;span class="nf"&gt;save_to_database&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Successfully wrote record for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;session_date_str&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# ... argparse handling for --context and --write ...
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With this modification, even if the Task Scheduler mistakenly executes the command multiple times:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;First execution:&lt;/strong&gt; Passes checks and writes successfully.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Subsequent executions:&lt;/strong&gt; The &lt;code&gt;record_exists&lt;/code&gt; check identifies "already recorded" and finishes without doing anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The deadline check also blocks the write operation if the deadline has passed.&lt;/p&gt;

&lt;p&gt;This property, where executing an operation multiple times yields the same result, is called "idempotency." By implementing this mechanism, I gained safe control over the AI agent's execution.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lesson Learned: Designing AI Autonomy and Its Control Boundaries
&lt;/h3&gt;

&lt;p&gt;The key takeaway from this failure is that if you're giving an AI agent autonomy, you absolutely must design "safe execution boundaries" from the human side.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Clearly separate read operations (reading state) from write operations (changing state).&lt;/strong&gt; This is fundamental to Web API design, and it applies perfectly to AI agent integration.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;All write operations must be idempotent.&lt;/strong&gt; AI and its operating environment might not always behave as expected and execute only once. It's essential for the receiving end to control the system, ensuring it doesn't break no matter how many requests it receives.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Let the AI focus on "thinking."&lt;/strong&gt; Procedural tasks like deadline checks or duplicate submission prevention are more reliable and cost-effective when handled by external programs rather than trying to force them into AI prompts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Delegating tasks to an AI might be similar to assigning work to a smart new hire. They're capable, but they don't know all the operational rules or company culture. So, before allowing critical operations, always establish double-check and permission-setting mechanisms.&lt;/p&gt;

&lt;p&gt;Developing AI agents isn't just about clever prompt engineering; it's about building up these subtle interface designs and error handling practices that are key to stable operation. This incident truly drove that point home for me.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>debugging</category>
      <category>database</category>
    </item>
    <item>
      <title>My Monitoring Bot Went "Silent": 3 Task Scheduler Pitfalls and How Negative Testing with API Snapshots Saved Me</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Tue, 29 Sep 2026 23:30:17 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-monitoring-bot-went-silent-3-task-scheduler-pitfalls-and-how-negative-testing-with-api-4e32</link>
      <guid>https://dev.to/masaoshimadaopen/my-monitoring-bot-went-silent-3-task-scheduler-pitfalls-and-how-negative-testing-with-api-4e32</guid>
      <description>&lt;p&gt;Hey there, it's OJ, a part-time AI agent and algo-trading bot developer.&lt;/p&gt;

&lt;p&gt;Today, I want to share a common pitfall in personal development: a web monitoring bot I built suddenly went "silent." No error logs, no alerts, just complete silence. These silent failures are the worst kind.&lt;/p&gt;

&lt;p&gt;The culprit? Classic Windows Task Scheduler gotchas and a flaw in my monitoring system's design. I eventually solved it by implementing a robust mechanism to detect when the monitoring &lt;em&gt;itself&lt;/em&gt; failed. Here's how it went down.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Happened: The Weekly Bot's Silence
&lt;/h3&gt;

&lt;p&gt;I had built a weekly bot to check for unintended changes to my content on an online learning platform. Rather than dealing with a heavy, unstable browser, I leveraged their public API to fetch titles and descriptions. Simple and elegant.&lt;/p&gt;

&lt;p&gt;I tested it locally, everything looked good, and I registered it with Task Scheduler. Set it to run every Sunday night, and mostly forgot about it.&lt;/p&gt;

&lt;p&gt;Then one day, it hit me: "Wait, I haven't received any notifications from the bot recently." Not even a success log. Checking Task Scheduler's history showed either no execution records or cryptic statuses like "Task was denied by an operator or administrator."&lt;/p&gt;

&lt;p&gt;If it's not running, fine, throw an error! But silence? That's what really kills you.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Cause: 3 Task Scheduler Traps
&lt;/h3&gt;

&lt;p&gt;My investigation revealed several traps in Windows Task Scheduler's default settings, especially when developing and running on a laptop.&lt;/p&gt;

&lt;h4&gt;
  
  
  1. Won't Run on Battery Power
&lt;/h4&gt;

&lt;p&gt;Open the task properties, go to the "Conditions" tab, and you'll find &lt;strong&gt;"Start the task only if the computer is on AC power."&lt;/strong&gt; This is checked by default.&lt;/p&gt;

&lt;p&gt;Meaning, if I unplugged my laptop on the weekend and moved it to the living room, the task's conditions wouldn't be met. Consequently, the bot wouldn't launch at its scheduled time.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Won't Run if PC is Asleep
&lt;/h4&gt;

&lt;p&gt;Another common one. In the "Settings" tab, there's an option: &lt;strong&gt;"Run task as soon as possible after a scheduled start is missed."&lt;/strong&gt; If this isn't checked, and your PC is asleep or shut down at the scheduled time, the task is just &lt;em&gt;skipped&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Leave your laptop closed over the weekend, and your task might never get a chance to run.&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Console Character Encoding (cp932) Issue
&lt;/h4&gt;

&lt;p&gt;While not the direct cause this time, this is a hotbed for silent failures. When running Python scripts via Task Scheduler, standard output can sometimes be processed with &lt;code&gt;cp932&lt;/code&gt; (Shift_JIS).&lt;/p&gt;

&lt;p&gt;If your script prints UTF-8 characters, you'll instantly hit a &lt;code&gt;UnicodeEncodeError&lt;/code&gt;. The catch? This error output might not be logged anywhere, leaving you clueless that the task even failed. Essential to guard against.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix: 2 Steps to Reliable Monitoring
&lt;/h3&gt;

&lt;p&gt;First, I reviewed all my Task Scheduler settings based on these pitfalls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Unchecked "Start the task only if the computer is on AC power."&lt;/li&gt;
&lt;li&gt;  Enabled "Run task as soon as possible after a scheduled start is missed."&lt;/li&gt;
&lt;li&gt;  As a precaution, I added &lt;code&gt;chcp 65001&lt;/code&gt; to my batch file to switch to UTF-8, or configured it via Python environment variables.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This largely solved the "task not executing" problem.&lt;/p&gt;

&lt;p&gt;However, a fundamental issue remained: I couldn't distinguish between "the monitored content hasn't changed" and "the monitoring bot failed for some reason."&lt;/p&gt;

&lt;p&gt;So, to detect "monitoring failure" itself, I integrated &lt;strong&gt;negative testing using approved snapshots.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's the refined flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Hit the monitoring target's API to fetch the JSON response.&lt;br&gt;
&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api-2.0/courses/7347629/?fields[course]=title,headline,description,...
&lt;/code&gt;&lt;/pre&gt;

&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Compare the fetched JSON with a pre-saved "golden JSON file" (the snapshot).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;If there's a difference, notify, of course.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Crucially&lt;/strong&gt;: Even if there's no difference, &lt;em&gt;always&lt;/em&gt; send a log indicating "comparison process completed successfully."&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now, if the bot doesn't run, I won't receive the "success log," and I'll know immediately.&lt;/p&gt;

&lt;p&gt;Furthermore, to ensure the reliability of &lt;em&gt;this mechanism itself&lt;/em&gt;, I perform "negative testing." I intentionally corrupt the snapshot file or mock the API response to generate differences, and then &lt;strong&gt;verify that notifications are indeed sent when an anomaly occurs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Updating snapshots is made easy with a simple command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Save the current API response as the new "golden" version&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; app_factory.udemy.published_check &lt;span class="nt"&gt;--bless&lt;/span&gt; genai_passport
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This also helps in adapting to changes in the monitored API's structure.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Lesson: How to Know If Your Creation Isn't "Dead"
&lt;/h3&gt;

&lt;p&gt;What I learned from this experience is that personal automation tools, if left as "set it and forget it," can silently die without you ever knowing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Task Scheduler is handy, but question its default settings.&lt;/strong&gt; Always double-check power conditions and behavior during sleep.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;To prevent silent failures, designing for "proof of success" is crucial.&lt;/strong&gt; "No errors = success" is not enough. Think: "Success log present = success."&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Monitoring system reliability is guaranteed through negative testing.&lt;/strong&gt; Regularly testing "will I get a notification if it breaks?" is incredibly important, even for personal projects.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Flashy feature development is fun, but this kind of gritty operational improvement ultimately saves your time. It's a small defense against being betrayed by your own automation tools.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>api</category>
      <category>ai</category>
      <category>devops</category>
      <category>debugging</category>
    </item>
    <item>
      <title>My PC Screamed! How a `grep -r` Trap Led to 100% CPU and How I Prevented It with a JavaScript Git Hook</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Wed, 23 Sep 2026 23:30:17 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-pc-screamed-how-a-grep-r-trap-led-to-100-cpu-and-how-i-prevented-it-with-a-javascript-git-4p0m</link>
      <guid>https://dev.to/masaoshimadaopen/my-pc-screamed-how-a-grep-r-trap-led-to-100-cpu-and-how-i-prevented-it-with-a-javascript-git-4p0m</guid>
      <description>&lt;p&gt;Hey there, it's your resident "old man" developer, 38 years old and tinkering with AI trading bots in my spare time.&lt;/p&gt;

&lt;p&gt;One evening, as I was routinely checking my bot's logs, disaster struck. I wanted to search the entire repository for a specific log output, so I mindlessly typed this into my terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s2"&gt;"some_keyword"&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next second, my PC's fans roared like a Formula 1 car, and my mouse cursor started lagging severely. I opened Task Manager, and sure enough, CPU usage was pegged at 100%. I genuinely thought my machine was on its last legs. The culprits? &lt;code&gt;grep.exe&lt;/code&gt; and, consequently, Windows Defender's real-time scan, which went haywire.&lt;/p&gt;

&lt;p&gt;Ultimately, I had to force-restart my PC. Who knew a simple &lt;code&gt;grep&lt;/code&gt; could cause such a catastrophe? Today, I'm sharing the story of this "grep scream incident" and the mechanism I built to prevent it from ever happening again.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: &lt;code&gt;.gitignore&lt;/code&gt; Is Just a Text File
&lt;/h3&gt;

&lt;p&gt;Why did &lt;code&gt;grep&lt;/code&gt; go rogue? The reason became clear quickly.&lt;/p&gt;

&lt;p&gt;My repository contained several massive log files generated by the bot. These were temporarily placed there for testing, with the largest one clocking in at 23GB. Naturally, I had these files properly listed in &lt;code&gt;.gitignore&lt;/code&gt; to exclude them from Git's tracking.&lt;/p&gt;

&lt;p&gt;But here's where my huge misconception lay: the shell's &lt;code&gt;grep&lt;/code&gt; command knows absolutely nothing about &lt;code&gt;.gitignore&lt;/code&gt;. To &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;.gitignore&lt;/code&gt; is just &lt;code&gt;ignore.txt&lt;/code&gt;—nothing more, nothing less.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;grep -r&lt;/code&gt; recursively, and naively, reads &lt;em&gt;all&lt;/em&gt; files under the current directory. So, it plunged headfirst into that 23GB log file. No wonder the CPU screamed.&lt;/p&gt;

&lt;p&gt;This same issue can occur with &lt;code&gt;find . | xargs grep "..."&lt;/code&gt; or PowerShell's &lt;code&gt;Select-String -Path . -Pattern "..." -Recurse&lt;/code&gt;. Essentially, any command that recursively scans the file system is susceptible to this trap.&lt;/p&gt;

&lt;p&gt;Especially when working with automated trading bots, gigabyte-sized logs and data files are a daily occurrence. It seems personal development environments are particularly prone to these kinds of accidents.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Solution: "Being Careful" Isn't Enough, So I Automated It
&lt;/h3&gt;

&lt;p&gt;Even if I vow to use &lt;code&gt;rg&lt;/code&gt; (ripgrep) or &lt;code&gt;ag&lt;/code&gt;—which respect &lt;code&gt;.gitignore&lt;/code&gt;—I can already foresee myself accidentally typing &lt;code&gt;grep&lt;/code&gt; out of habit. Humans forget, and humans make mistakes.&lt;/p&gt;

&lt;p&gt;So, instead of relying on "being careful," I opted for an approach that would "physically prevent execution."&lt;/p&gt;

&lt;p&gt;Specifically, I leveraged a "hook" mechanism that kicks off a specific script right before a command is executed in the terminal. If a dangerous &lt;code&gt;grep&lt;/code&gt; command is about to run, the hook detects and blocks it.&lt;/p&gt;

&lt;p&gt;I wrote this hook script in JavaScript. What it does is simple: it checks the command string using a regular expression.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Check the command line about to be executed with a regular expression&lt;/span&gt;

&lt;span class="c1"&gt;// Roughly detect grep-like command call patterns&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;GREP_CALL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;(?:&lt;/span&gt;&lt;span class="sr"&gt;^|&lt;/span&gt;&lt;span class="se"&gt;[\s&lt;/span&gt;&lt;span class="sr"&gt;(&lt;/span&gt;&lt;span class="se"&gt;])(?:[^\s&lt;/span&gt;&lt;span class="sr"&gt;|;&amp;amp;&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;[\\/])?(&lt;/span&gt;&lt;span class="sr"&gt;r|e|f&lt;/span&gt;&lt;span class="se"&gt;)?&lt;/span&gt;&lt;span class="sr"&gt;grep&lt;/span&gt;&lt;span class="se"&gt;(?:\.&lt;/span&gt;&lt;span class="sr"&gt;exe&lt;/span&gt;&lt;span class="se"&gt;)?(?:\s&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;.*&lt;/span&gt;&lt;span class="se"&gt;))?&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// Detect short recursive flags like -r or -R&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SHORT_RECURSIVE_FLAG&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;(?:&lt;/span&gt;&lt;span class="sr"&gt;^|&lt;/span&gt;&lt;span class="se"&gt;\s)&lt;/span&gt;&lt;span class="sr"&gt;-&lt;/span&gt;&lt;span class="se"&gt;(?!&lt;/span&gt;&lt;span class="sr"&gt;-&lt;/span&gt;&lt;span class="se"&gt;)[&lt;/span&gt;&lt;span class="sr"&gt;A-Za-z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;rR&lt;/span&gt;&lt;span class="se"&gt;][&lt;/span&gt;&lt;span class="sr"&gt;A-Za-z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;(?=\s&lt;/span&gt;&lt;span class="sr"&gt;|$&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// findstr /s on Windows is equally dangerous, so detect it too&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;FINDSTR_SUBDIRS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;findstr&lt;/span&gt;&lt;span class="se"&gt;(?:\.&lt;/span&gt;&lt;span class="sr"&gt;exe&lt;/span&gt;&lt;span class="se"&gt;)?\b[^&lt;/span&gt;&lt;span class="sr"&gt;|;&lt;/span&gt;&lt;span class="se"&gt;\n]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;\s\/&lt;/span&gt;&lt;span class="sr"&gt;s&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/i&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;shouldBlock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;toolName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Only target Bash or PowerShell execution&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;toolName&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Bash&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;toolName&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PowerShell&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// findstr /s is an immediate block&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;FINDSTR_SUBDIRS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Split the command by pipes or semicolons and check if any segment&lt;/span&gt;
  &lt;span class="c1"&gt;// includes a recursive grep call&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\|\|&lt;/span&gt;&lt;span class="sr"&gt;|&amp;amp;&amp;amp;|&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;|;&lt;/span&gt;&lt;span class="se"&gt;\n]&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;isRecursiveGrepSegment&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;isRecursiveGrepSegment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;segment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;GREP_CALL&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="c1"&gt;// Check for long flags like --recursive or --directories=recurse&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hasLongRecursiveFlag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;--&lt;/span&gt;&lt;span class="se"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;recursive|directories=recurse&lt;/span&gt;&lt;span class="se"&gt;)\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="c1"&gt;// Check for short flags like -r, -R&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hasShortRecursiveFlag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;SHORT_RECURSIVE_FLAG&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;hasLongRecursiveFlag&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;hasShortRecursiveFlag&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With this script integrated into a hook, if I accidentally type &lt;code&gt;grep -r .&lt;/code&gt;, the command will be blocked before execution with a message like: "Dangerous command! Please use &lt;code&gt;rg&lt;/code&gt; instead." Peace of mind achieved.&lt;/p&gt;

&lt;p&gt;The regular expression is a bit messy because I wanted to detect not just &lt;code&gt;grep -r&lt;/code&gt; but also &lt;code&gt;-R&lt;/code&gt;, &lt;code&gt;--recursive&lt;/code&gt;, and &lt;code&gt;grep&lt;/code&gt; commands connected via pipes. &lt;code&gt;findstr /s&lt;/code&gt; causes equally nasty accidents, so I included it in the block list as well.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary and Lessons Learned
&lt;/h3&gt;

&lt;p&gt;What I learned from this incident is, while obvious, crucial: you'll get burned if you don't properly understand what your tools are doing implicitly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;code&gt;.gitignore&lt;/code&gt; is Git's rulebook, not the shell's.&lt;/li&gt;
&lt;li&gt;  Mechanical prevention is more robust and reliable than relying on self-discipline.&lt;/li&gt;
&lt;li&gt;  In personal development, especially with automation and bot development, huge files can be generated unintentionally. Safety measures in the development environment, though seemingly trivial, are incredibly important.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Just &lt;code&gt;grep&lt;/code&gt;, but &lt;code&gt;grep&lt;/code&gt; can be a beast. To keep my PC from screaming, I need to properly understand the characteristics of my tools. With this, I've made my development environment safer for those late-night coding sessions. 👍&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>api</category>
      <category>programming</category>
      <category>python</category>
      <category>automation</category>
    </item>
    <item>
      <title>When AI hallucinated 'systemd' and 'cron' on my Windows machine: The 'context blindness' trap of LLMs</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Tue, 22 Sep 2026 23:30:17 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/when-ai-hallucinated-systemd-and-cron-on-my-windows-machine-the-context-blindness-trap-of-g88</link>
      <guid>https://dev.to/masaoshimadaopen/when-ai-hallucinated-systemd-and-cron-on-my-windows-machine-the-context-blindness-trap-of-g88</guid>
      <description>&lt;p&gt;Hey, it's Grandpa Oji (@oji_ai_dev) here.&lt;/p&gt;

&lt;p&gt;I'm a 38-year-old developer who tinkers with AI bots in my side hustle and works as a full-time engineer during the week. Recently, I decided to start a simple dev blog – separate from my notes – to properly document my personal projects.&lt;/p&gt;

&lt;p&gt;Naturally, I wanted to leverage AI for some of the article generation to boost efficiency. As a first step, I asked it to write a setup guide for a simple bot I run periodically via Windows Task Scheduler.&lt;/p&gt;

&lt;p&gt;I fed it the logs and source code, then vaguely prompted: "Explain how to set up this bot." When I read the first draft, the opening lines immediately threw me for a loop.&lt;/p&gt;

&lt;p&gt;"First, register it as a service using &lt;code&gt;systemd&lt;/code&gt;. Place the following unit file in &lt;code&gt;/etc/systemd/system/&lt;/code&gt;..."&lt;/p&gt;

&lt;p&gt;...Huh? &lt;code&gt;systemd&lt;/code&gt;?&lt;/p&gt;

&lt;p&gt;My bot runs on a plain old Windows 11 machine, not Linux. &lt;code&gt;systemd&lt;/code&gt; doesn't exist there. Reading further, it continued: "For periodic execution, it's convenient to register a schedule using &lt;code&gt;crontab -e&lt;/code&gt;." It had completely mistaken the operating system.&lt;/p&gt;

&lt;p&gt;This was rough. It strung together plausible-sounding technical terms, but the content was pure fantasy. It was a spectacular moment of AI 'context blindness.'&lt;/p&gt;

&lt;h3&gt;
  
  
  The Cause: Implicit Assumptions Don't Translate
&lt;/h3&gt;

&lt;p&gt;Why did this happen? One hundred percent, it was because my prompt was too vague.&lt;/p&gt;

&lt;p&gt;My instruction to the AI was literally just, "Write a blog post explaining this script." I completely omitted crucial, seemingly obvious information like the OS being Windows and that it was scheduled with Task Scheduler.&lt;/p&gt;

&lt;p&gt;When an LLM receives incomplete information, it attempts to fill in the gaps with the "most probable, general knowledge" from its vast training data. When the topic shifts to demonizing or periodically executing scripts on a server, the vast majority of technical articles in the world assume a Linux environment.&lt;/p&gt;

&lt;p&gt;So, the AI reasoned with a ridiculously simple inference: "Okay, this person wants to run a bot on a server. Servers typically mean Linux, right? Therefore, I should talk about &lt;code&gt;systemd&lt;/code&gt; and &lt;code&gt;cron&lt;/code&gt;."&lt;/p&gt;

&lt;p&gt;From my perspective, I'm thinking, "No, it's my Windows machine humming in the corner of my room," but the AI doesn't pick up on that 'vibe.' Any context that I consider 'obvious' must be explicitly stated in words, or the AI won't grasp it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix: Give AI a Map in the Form of 'Constraints'
&lt;/h3&gt;

&lt;p&gt;The way to prevent this 'context blindness' is simple: clearly define "constraints" in your prompt.&lt;/p&gt;

&lt;p&gt;For example, in this case, my initial vague prompt looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Bad example
&lt;/span&gt;&lt;span class="n"&gt;Write&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;blog&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt; &lt;span class="n"&gt;explaining&lt;/span&gt; &lt;span class="n"&gt;how&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;periodically&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;following&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt; &lt;span class="n"&gt;script&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="c1"&gt;# --- Script ---
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;log_file&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;log.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Executed at &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With this, the AI felt free to interpret and started talking about Linux.&lt;/p&gt;

&lt;p&gt;So, I gave it a "map" in the form of the execution environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Good example
&lt;/span&gt;&lt;span class="n"&gt;Write&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;blog&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt; &lt;span class="n"&gt;explaining&lt;/span&gt; &lt;span class="n"&gt;how&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;periodically&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;following&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt; &lt;span class="n"&gt;script&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="c1"&gt;# Constraints
&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;OS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Windows&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Execution&lt;/span&gt; &lt;span class="n"&gt;Method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt; &lt;span class="n"&gt;Scheduler&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Audience&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Beginners&lt;/span&gt; &lt;span class="n"&gt;using&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="n"&gt;Windows&lt;/span&gt;

&lt;span class="c1"&gt;# --- Script ---
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="n"&gt;rom&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;log_file&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;log.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;log_file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Executed at &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By explicitly stating the OS and execution method under "# Constraints," the accuracy of the AI's output dramatically improved. When I regenerated with this prompt, it correctly provided sensible steps for a Windows environment, starting with "Open Task Scheduler, then 'Create New Task'...".&lt;/p&gt;

&lt;h3&gt;
  
  
  The Lesson: AI is Smart, But Not Psychic
&lt;/h3&gt;

&lt;p&gt;This incident reinforced my view that LLMs are like "incredibly competent, but utterly clueless, new hires."&lt;/p&gt;

&lt;p&gt;If your instructions are ambiguous, the AI, with good intentions, will fill in the gaps with general knowledge. But if that knowledge doesn't align with the actual situation (i.e., the execution environment), the output won't just be useless – it can be actively harmful.&lt;/p&gt;

&lt;p&gt;Especially when having AI handle technical content, you &lt;em&gt;must&lt;/em&gt; communicate the "execution environment context" upfront, even if it feels tedious. This includes the OS, library versions, frameworks, and hardware configuration. Neglecting this will result in it confidently generating non-existent commands or code that won't run due to version mismatches.&lt;/p&gt;

&lt;p&gt;The generated content often looks plausible, which means there's a serious risk that someone less knowledgeable might believe it. Tragedies like debugging incorrect &lt;code&gt;systemd&lt;/code&gt; configuration files for hours could easily happen.&lt;/p&gt;

&lt;p&gt;Delegating tasks to AI is convenient, but you can't escape the need for precise "prompt engineering" and the "human oversight" of checking the final output. Slack on these fundamentals, and you'll only end up causing yourself trouble.&lt;/p&gt;

&lt;p&gt;This was a good failure log that made me rethink how I interact with AI.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
    </item>
    <item>
      <title>My Bot Went Silent: The Case of the Missing "Always-On" Process</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Mon, 21 Sep 2026 23:30:18 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-bot-went-silent-the-case-of-the-missing-always-on-process-205g</link>
      <guid>https://dev.to/masaoshimadaopen/my-bot-went-silent-the-case-of-the-missing-always-on-process-205g</guid>
      <description>&lt;p&gt;It's your friendly neighborhood AI dev, oji (@oji_ai_dev)! Today I want to share a recent, pretty critical mistake I made with my side-hustle AI trading bot. A newly implemented feature completely failed to run after deployment – no errors, no logs, just a complete, silent death.&lt;/p&gt;

&lt;p&gt;The root cause? A fundamental misalignment in assumptions about design and operation within our small team (which includes me!). A classic, basic oversight.&lt;/p&gt;

&lt;h3&gt;
  
  
  The New Feature: Totally Mute After Deployment
&lt;/h3&gt;

&lt;p&gt;It all started when I added new periodic tasks to the bot. Up until then, most tasks ran daily. This time, I was adding weekly and monthly report generation features.&lt;/p&gt;

&lt;p&gt;A remote team member handled the logic, and my role was to deploy it to production. I reviewed the code, confirmed it ran fine in my local dev environment, and thought, "Great, we'll start getting those reports next week."&lt;/p&gt;

&lt;p&gt;A few days passed.&lt;/p&gt;

&lt;p&gt;Suddenly, I wondered, "Is it actually running?" I checked the production server logs. ...Nothing. There should have been logs like "Weekly task started," but there was no trace. &lt;/p&gt;

&lt;p&gt;At first, I thought, "Maybe it's not time yet?" But checking the calendar, it was clearly overdue. It had been completely skipped. No error logs anywhere. This was bad.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Shocking Truth Behind the Search
&lt;/h3&gt;

&lt;p&gt;My first suspicion was the code. But it worked in dev. Next, I checked Python library dependencies on the production environment. No issues there either.&lt;/p&gt;

&lt;p&gt;Completely stumped, I decided to deep-dive: check process status on the production server, cron jobs, task scheduler settings – everything.&lt;/p&gt;

&lt;p&gt;Midway through documenting my investigation, I found the smoking gun:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;`scheduler.py` has never run on this machine. There are 0 persistent processes and 0 task scheduler registrations. All periodic executions are handled by `Oji*` tasks (e.g., `OjiNotePost`).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The moment I read that, everything clicked. It felt like a punch to the gut.&lt;/p&gt;

&lt;p&gt;The new feature was designed around a &lt;code&gt;scheduler.py&lt;/code&gt; script that would run as a persistent process, launching tasks at specified times. It's a common implementation using Python's &lt;code&gt;schedule&lt;/code&gt; library.&lt;/p&gt;

&lt;p&gt;Naturally, the development team assumed this &lt;code&gt;scheduler.py&lt;/code&gt; would be running as a daemon in the production environment.&lt;/p&gt;

&lt;p&gt;But on &lt;strong&gt;my home server, no one had ever started such a process.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A Twist Between Design Philosophy and Operational Reality
&lt;/h3&gt;

&lt;p&gt;How did this happen?&lt;/p&gt;

&lt;p&gt;My operational policy for my home server was to avoid persistent processes wherever possible to conserve memory and prevent zombie processes. That's why all existing periodic tasks were executed by the OS task scheduler (Windows Task Scheduler, in my case), which directly kicked off individual Python scripts at scheduled times.&lt;/p&gt;

&lt;p&gt;So, in summary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Design Assumption&lt;/strong&gt;: &lt;code&gt;scheduler.py&lt;/code&gt; runs persistently to manage tasks (persistent process model).&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Operational Reality&lt;/strong&gt;: The OS launches individual scripts on demand (event-driven model).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This architectural premise was completely misaligned between myself and the dev team. I assumed, "Just drop the code, and I'll register it with the OS scheduler." The team assumed, "It will be deployed to a server where &lt;code&gt;scheduler.py&lt;/code&gt; is already running."&lt;/p&gt;

&lt;p&gt;This "assumption gap" silently killed the new feature. We hadn't documented anything, and the most crucial part – "how it runs" – was completely overlooked. Brutal.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix and the Lesson
&lt;/h3&gt;

&lt;p&gt;The fix was simple. Instead of launching a persistent process, I adapted to the existing operational setup.&lt;/p&gt;

&lt;p&gt;I refactored the new feature's logic into an independent script and registered it with the OS task scheduler. Now it runs using the same mechanism as all other tasks.&lt;/p&gt;

&lt;p&gt;I learned some big lessons from this failure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Shared understanding of architecture is critical.&lt;/strong&gt; Especially in solo dev or small teams, it's easy to assume "everyone gets it," but this leads to fatal accidents.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;"How it runs (operations)" is as important as "what it does (features)."&lt;/strong&gt; This kind of issue frequently arises when development and operations are separated.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;The only reliable sources of truth are code and configuration files.&lt;/strong&gt; Even a simple &lt;code&gt;README.md&lt;/code&gt; with a production process diagram or a list of launch commands could have prevented this incident.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A feature that just "doesn't run" without throwing errors is the hardest to detect and the scariest. I've now firmly committed to implementing a mechanism to regularly check if expected artifacts (log files, DB records, etc.) are actually being generated.&lt;/p&gt;

&lt;p&gt;This incident was a stark reminder that side-hustle dev is a continuous series of gritty failures and learnings.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>devops</category>
      <category>debugging</category>
      <category>programming</category>
      <category>python</category>
    </item>
    <item>
      <title>My PC was Slow as Molasses: The Culprit? Orphaned `grep` Processes After a Shell Timeout.</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Sun, 20 Sep 2026 23:30:17 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-pc-was-slow-as-molasses-the-culprit-orphaned-grep-processes-after-a-shell-timeout-1ei6</link>
      <guid>https://dev.to/masaoshimadaopen/my-pc-was-slow-as-molasses-the-culprit-orphaned-grep-processes-after-a-shell-timeout-1ei6</guid>
      <description>&lt;p&gt;Hey there, oji_ai_dev here.&lt;/p&gt;

&lt;p&gt;Working on AI agents and bots as a side hustle often brings a different flavor of technical challenge compared to my day job. This time, my development PC suddenly ground to a halt, and I spent several hours debugging the issue.&lt;/p&gt;

&lt;p&gt;Long story short, the culprit wasn't Windows Defender (though it &lt;em&gt;looked&lt;/em&gt; like it) but rather a bunch of "orphaned &lt;code&gt;grep&lt;/code&gt; processes" hiding behind it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Symptom: PC Stuck at 100% CPU Usage
&lt;/h3&gt;

&lt;p&gt;One evening, as I was setting up my dev environment and coding, my PC's fans suddenly started roaring. Typing in VSCode became sluggish, almost unworkable. Clearly, something was very wrong.&lt;/p&gt;

&lt;p&gt;I opened Task Manager, and sure enough, the CPU usage was pegged at 100%. At the top of the list was the familiar &lt;code&gt;MsMpEng.exe&lt;/code&gt; – Windows Defender.&lt;/p&gt;

&lt;p&gt;"Not you again..."&lt;/p&gt;

&lt;p&gt;Any developer has probably experienced this. When you generate or read a large number of files, Defender's real-time scanning can go haywire and bring your PC to its knees. "Oh, it's just one of those days," I thought, and temporarily disabled Defender.&lt;/p&gt;

&lt;p&gt;...But the situation didn't change. CPU was still at 100%. This was bad.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Investigation: The Real Culprit Behind Defender
&lt;/h3&gt;

&lt;p&gt;If Defender wasn't the culprit, what &lt;em&gt;was&lt;/em&gt; eating up all that CPU? Staring at the Task Manager's "Details" tab, and looking at the processes consuming CPU time, I noticed several unfamiliar processes:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;grep.exe&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Why was this here? And multiple instances of it? This was GNU &lt;code&gt;grep&lt;/code&gt; bundled with Git for Windows. I didn't remember starting it directly.&lt;/p&gt;

&lt;p&gt;Then, I recalled my actions just prior. I had used an AI assistant (Claude) to search the contents of a massive monorepo. This repository was a chaotic mess of accumulated logs and build artifacts from years of operation, totaling about 34GB.&lt;/p&gt;

&lt;p&gt;The command I fed to the AI assistant was something simple, like this:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;grep -r "some_legacy_function" .&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The shell environment used by the AI assistant is configured to time out after a certain period (e.g., 300 seconds) for safety. Most likely, searching such a huge repository took too long, and the parent shell died due to a timeout.&lt;/p&gt;

&lt;p&gt;Here's where the problem started.&lt;/p&gt;

&lt;p&gt;In a Linux environment, when a parent shell dies, its child processes often terminate along with it. However, &lt;code&gt;grep.exe&lt;/code&gt; running within Git Bash on Windows behaved differently. Even after its parent was gone, the child processes continued to run.&lt;/p&gt;

&lt;p&gt;These are "orphan processes."&lt;/p&gt;

&lt;p&gt;Multiple &lt;code&gt;grep.exe&lt;/code&gt; instances, having lost their parent, continued to relentlessly scan 34GB of files, unmanaged. No wonder the CPU was at 100%. What appeared to be Defender going berserk was merely a secondary phenomenon: these orphaned processes were generating an insane amount of file I/O, and Defender's scan was reacting to it. I had completely misidentified the root cause.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix and Prevention
&lt;/h3&gt;

&lt;p&gt;Once the cause was known, the fix was simple. I manually force-terminated all suspicious &lt;code&gt;grep.exe&lt;/code&gt; processes from Task Manager. The roaring noise immediately subsided, and my PC quieted down, with CPU usage dropping to around 5%. Peace was restored.&lt;/p&gt;

&lt;p&gt;However, this left me vulnerable to the same mistake again. A more fundamental solution was needed.&lt;/p&gt;

&lt;p&gt;First, I decided to stop indiscriminately using &lt;code&gt;grep -r&lt;/code&gt; on huge repositories. It completely ignores &lt;code&gt;.gitignore&lt;/code&gt; and diligently searches through &lt;code&gt;node_modules&lt;/code&gt;, large binaries, and everything else, which is terrible for performance. Instead, I've standardized on &lt;code&gt;ripgrep&lt;/code&gt; (rg), which intelligently searches while respecting &lt;code&gt;.gitignore&lt;/code&gt; and is much faster.&lt;/p&gt;

&lt;p&gt;Furthermore, I decided to implement a hook in my AI assistant's execution environment to prevent potentially dangerous commands from being run in the first place. This simple mechanism checks the command's content before the tool (in this case, Bash) is executed, and if it matches certain patterns, it throws an error and stops.&lt;/p&gt;

&lt;p&gt;The concept is a script similar to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// .claude/hooks/block_recursive_grep.js (conceptual hook script)&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CLAUDE_TOOL_NAME&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CLAUDE_TOOL_PARAMS&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tool&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Bash&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;command&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;// Define dangerous command patterns that trigger recursive searches&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recursivePatterns&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="sr"&gt;/grep&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+.*&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;-r&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;// grep -r&lt;/span&gt;
        &lt;span class="sr"&gt;/grep&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+.*&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;--recursive&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// grep --recursive&lt;/span&gt;
        &lt;span class="sr"&gt;/find&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+.*&lt;/span&gt;&lt;span class="se"&gt;\s\|\s&lt;/span&gt;&lt;span class="sr"&gt;*xargs&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+grep/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// find | xargs grep&lt;/span&gt;
        &lt;span class="sr"&gt;/findstr&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+.*&lt;/span&gt;&lt;span class="se"&gt;\s\/&lt;/span&gt;&lt;span class="sr"&gt;s&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;       &lt;span class="c1"&gt;// findstr /s (Windows)&lt;/span&gt;
    &lt;span class="p"&gt;];&lt;/span&gt;

    &lt;span class="c1"&gt;// If any pattern matches, block the command and exit with an error&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;recursivePatterns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`ERROR: Recursive grep is blocked due to performance risk on large repos. Use the 'Grep' tool (ripgrep) instead, which respects .gitignore and is much faster.`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// Block the command&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// Allow other commands&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With this hook in place, if I accidentally tell the AI to use &lt;code&gt;grep -r&lt;/code&gt; in the future, it will be blocked before execution and prompted to "use &lt;code&gt;ripgrep&lt;/code&gt; instead because it's dangerous." For someone as forgetful as me, these kinds of physical guardrails are the most effective.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;p&gt;I learned two main lessons from this incident:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Shell timeouts don't guarantee child processes will be killed.&lt;/strong&gt; Especially in Windows environments, it's crucial to be aware of the risk of orphaned processes consuming resources.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Don't be fooled by superficial symptoms (like Defender running wild).&lt;/strong&gt; Always suspect that there might be a true culprit generating massive I/O behind the scenes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Time for side projects is limited, so losing half a day to environment issues like this is truly painful. But every failure makes my environment more robust. I hope this incident log helps save someone else's precious time. 👍&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>api</category>
      <category>ai</category>
      <category>debugging</category>
      <category>programming</category>
    </item>
    <item>
      <title>Oops! Almost Leaked User Data with My AI Feature – A Wake-Up Call from API Terms of Service</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Sat, 19 Sep 2026 23:30:17 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/oops-almost-leaked-user-data-with-my-ai-feature-a-wake-up-call-from-api-terms-of-service-39a9</link>
      <guid>https://dev.to/masaoshimadaopen/oops-almost-leaked-user-data-with-my-ai-feature-a-wake-up-call-from-api-terms-of-service-39a9</guid>
      <description>&lt;p&gt;Hey there, it's your friendly neighborhood dev, 38 years old, slinging code as an engineer by day and tinkering with AI trading bots on weekends.&lt;/p&gt;

&lt;p&gt;Recently, I decided to spice up a personal tool I'm building on the side by adding a little AI magic. The idea was simple: users type some text, and the AI generates something nice based on it. I spun it up quickly, hitting the Gemini API from the backend, and it worked like a charm. "Oh, this is actually pretty neat!" I thought, grinning to myself.&lt;/p&gt;

&lt;p&gt;Once the feature felt stable, I figured it was time to write a user manual. That's when I started digging into the official documentation, not just the API reference I'd skimmed during implementation, but also the terms of service and FAQs.&lt;/p&gt;

&lt;p&gt;And then, I stumbled upon a single sentence that made my blood run cold.&lt;/p&gt;

&lt;h3&gt;
  
  
  Free Tier API Input Data Can Be Used for Product Improvement
&lt;/h3&gt;

&lt;p&gt;I found it on Google's policy page regarding data usage for their AI services. It clearly stated:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;(Paraphrased) Content submitted through the free tier APIs (including prompts and responses) may be used to improve Google's products and services, such as generative AI models.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was a huge red flag.&lt;/p&gt;

&lt;p&gt;My tool allows users to input any text they want. What if a user accidentally types in confidential information – like customer data from work, or a snippet from a top-secret project proposal?&lt;/p&gt;

&lt;p&gt;That data would then travel through my tool, to Google's servers, and potentially be used as training data for their AI models. Sure, they probably anonymize it and process it in various ways, but the fact that the terms explicitly state it &lt;em&gt;"may be used"&lt;/em&gt; is a serious concern.&lt;/p&gt;

&lt;p&gt;Users wouldn't know any of this background. They'd just be using my tool as a "handy text generator." My tool could inadvertently become a pipeline for leaking user's confidential information to an external entity. Realizing this possibility sent shivers down my spine. This was a nightmare scenario.&lt;/p&gt;

&lt;h3&gt;
  
  
  My Immediate Fixes
&lt;/h3&gt;

&lt;p&gt;It was a blessing in disguise that I caught this before release. I immediately took several steps.&lt;/p&gt;

&lt;p&gt;First, I added a prominent disclaimer about the AI feature to my tool's manual and FAQ. It roughly said something like this:&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;[Important] Caution Regarding AI Feature Use&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This feature utilizes an external AI service (Google Gemini API).&lt;/p&gt;

&lt;p&gt;Per the terms of service, data entered into this feature may be used by Google, the service provider, for product improvement (e.g., training AI models).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Therefore, absolutely DO NOT input any personal information, company confidential information, or any other data that should not be shared externally.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  By using this feature, you agree to the above.
&lt;/h2&gt;

&lt;p&gt;Next, I updated the UI. Right below the input form, I added a link to this FAQ page with a label like "Important Usage Notes." I felt it was crucial to issue a warning right where the feature is used, not just burying it in the manual.&lt;/p&gt;

&lt;p&gt;Finally, on the implementation side, I made sure to document this risk in the docstring of the function that calls the API. This serves as a reminder for future me, or anyone I might collaborate with.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# WARNING: This function sends data to an external AI service.
# The data sent via the free tier API may be used for service improvement.
# Do NOT send personal, confidential, or sensitive information.
# See our FAQ for more details.
&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;google.generativeai&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;genai&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="c1"&gt;# Best practice: get API key from environment variables
&lt;/span&gt;&lt;span class="n"&gt;genai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;configure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# Select the model
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;genai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GenerativeModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gemini-1.5-flash&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_text_from_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Generates text from a user prompt.

    WARNING: This function sends data to an external API.
    Data submitted via the free tier API may be used by the service provider
    for service improvement. DO NOT send confidential information.
    Refer to the FAQ for more details.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;prompt_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Input text is empty.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# In a real app, you'd log this more robustly
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;An error occurred: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;An error occurred. Please try again later.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Embedding warnings directly into the code is a form of defense. It makes the risk visible even at the source code level.&lt;/p&gt;

&lt;h3&gt;
  
  
  My Key Takeaway
&lt;/h3&gt;

&lt;p&gt;The lesson I learned from this experience is simple but incredibly important:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When integrating external APIs, read not just the functional specifications, but also the terms of service and privacy policy, thoroughly.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is especially true for AI services, where data handling is core to their operation. Free tiers versus paid tiers can have vastly different implications, not just for feature limits, but also for data privacy. This was precisely the case here.&lt;/p&gt;

&lt;p&gt;Being satisfied with "it works" and immediately pushing it to users is genuinely risky. As developers, we bear the responsibility for how what we build impacts our users. Since this feature involved sending user-entered data externally, I should have been extra cautious, and I regret not having been so initially.&lt;/p&gt;

&lt;p&gt;I'm sure more fantastic AI services will emerge. My desire to leverage them to build cool things hasn't changed. But this incident hammered home the importance of truly understanding the underlying "contract" and fulfilling our responsibility to protect users while using these powerful tools.&lt;/p&gt;

&lt;p&gt;Personal projects offer freedom, but that also means you're solely responsible. Every time I experience a close call like this, it's a stark reminder to stay sharp.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>api</category>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
    </item>
    <item>
      <title>"API Limit 21KB!" ...Then 32KB of data caused a mysterious error. The culprit? Confusing bytes with characters.</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Fri, 18 Sep 2026 23:30:15 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/api-limit-21kb-then-32kb-of-data-caused-a-mysterious-error-the-culprit-confusing-bytes-with-1lb1</link>
      <guid>https://dev.to/masaoshimadaopen/api-limit-21kb-then-32kb-of-data-caused-a-mysterious-error-the-culprit-confusing-bytes-with-1lb1</guid>
      <description>&lt;p&gt;Hey, it's your resident Old Man Developer here. One evening, while checking the logs for my self-made bot as usual, I noticed a recurring, inexplicable error in a section that interacts with an external API.&lt;/p&gt;

&lt;p&gt;The logs showed messages like &lt;code&gt;Payload too large&lt;/code&gt;. "Ah, typical," I thought. But something felt off. The API documentation clearly stated a "21KB payload limit," and my implementation was supposed to have a logic check to verify data size &lt;em&gt;before&lt;/em&gt; sending it.&lt;/p&gt;

&lt;p&gt;Yet, it was still being rejected. My local check passed, but the API screamed "too big!" This contradiction baffled me for quite a while.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Cause: &lt;code&gt;len()&lt;/code&gt; was the Root of All Evil
&lt;/h3&gt;

&lt;p&gt;Initially, I suspected API capriciousness, or perhaps a temporary network glitch. But the error consistently reproduced with specific data patterns. Okay, this was on my end.&lt;/p&gt;

&lt;p&gt;I pulled out the problematic data and measured its size locally.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Assuming data_string holds the problematic data
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data_string&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It was clearly smaller than &lt;code&gt;21 * 1024 = 21504&lt;/code&gt;. Still, the API was rejecting it. I just couldn't understand why.&lt;/p&gt;

&lt;p&gt;After about half a day of scratching my head, it suddenly hit me: the data I was handling contained a fair amount of Japanese text. "Could it be... byte count?"&lt;/p&gt;

&lt;p&gt;To test this hypothesis, I wrote a simple piece of code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;こんにちは世界&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;char_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;byte_count_utf8&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;String: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Character count: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;char_count&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# =&amp;gt; 7
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Byte count (UTF-8): &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;byte_count_utf8&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# =&amp;gt; 21
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seeing this, everything clicked.&lt;/p&gt;

&lt;p&gt;"こんにちは世界" is only 7 characters. But encoded in UTF-8, it becomes 21 bytes. That's a 3x difference! For English (ASCII) text, 1 character is basically 1 byte, so &lt;code&gt;len()&lt;/code&gt; roughly matches the byte count. But when multibyte characters like Japanese come into play, the story changes entirely.&lt;/p&gt;

&lt;p&gt;In essence, I was checking the "character count" with &lt;code&gt;len()&lt;/code&gt;, but the API was looking at the "byte count" calculated by &lt;code&gt;len(str.encode('utf-8'))&lt;/code&gt;. The "21KB" in the documentation had, in my mind, been unconsciously translated to something like "21k characters." Damn, it was a complete assumption on my part.&lt;/p&gt;

&lt;p&gt;When I re-measured the actual error-causing data by byte count, it easily exceeded 21KB, coming in at nearly 32KB. No wonder the API was upset!&lt;/p&gt;

&lt;h3&gt;
  
  
  The Fix: Switch to Byte-Based Check Logic
&lt;/h3&gt;

&lt;p&gt;Once the cause was known, the fix was simple. Just change the size check before sending to the API from character-based to byte-based.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_api_limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data_string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit_in_kb&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Encode the string in UTF-8 and calculate byte count
&lt;/span&gt;    &lt;span class="n"&gt;byte_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data_string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="c1"&gt;# Convert KB limit to bytes for comparison
&lt;/span&gt;    &lt;span class="n"&gt;limit_in_bytes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;limit_in_kb&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;byte_size&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;limit_in_bytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Raise an error here to prevent the API call
&lt;/span&gt;        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Data size (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;byte_size&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; bytes) exceeds limit (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;limit_in_bytes&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; bytes)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Calling side
&lt;/span&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;check_api_limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;my_data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# Logic to send to API
&lt;/span&gt;&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;ValueError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# Handle splitting, error handling, etc.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I implemented a function like this and placed it right before the API call. Now, I can accurately detect size overages locally before the API even gets a chance to complain. Since deploying this fix, the mysterious error has completely disappeared.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lessons Learned: Question Units, Measure Empirically
&lt;/h3&gt;

&lt;p&gt;The lessons I learned from this failure are quite simple yet crucial:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Question Unit Definitions&lt;/strong&gt;: When documentation uses ambiguous terms like "size" or "KB," I need to make it a habit to accurately confirm whether it refers to "character count" or "byte count." If it's not specified, I should have tested with minimal data to verify behavior.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Don't Forget the Multibyte Trap&lt;/strong&gt;: This issue constantly arises, especially when building AI agents that handle natural language data like Japanese. Always keep in mind that character count and byte count are completely different.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Assumptions are the Biggest Enemy&lt;/strong&gt;: Python's ease of use, like getting "size" with &lt;code&gt;len()&lt;/code&gt;, ironically became the pitfall this time. Doubting my own "I know this" assumptions is a subtle but incredibly important practice.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So, that was my story of wasting half a day due to confusing "character count" and "byte count" – a slightly embarrassing but common mistake anyone can make.&lt;/p&gt;

&lt;p&gt;Squashing these small, unassuming errors one by one is the reality of personal development, especially when you're working on side projects. I'll share again if I screw up in another interesting way.&lt;/p&gt;




&lt;p&gt;X: &lt;a href="https://twitter.com/oji_ai_dev" rel="noopener noreferrer"&gt;@oji_ai_dev&lt;/a&gt;&lt;br&gt;
note: &lt;a href="https://note.com/oji_ai_dev" rel="noopener noreferrer"&gt;oji_ai_dev&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>api</category>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
    </item>
    <item>
      <title>My custom filter showed "zero false positives"! Then I realized my validation logic was flawed, letting all the garbage through.</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Thu, 10 Sep 2026 23:30:18 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/my-custom-filter-showed-zero-false-positives-then-i-realized-my-validation-logic-was-flawed-6o</link>
      <guid>https://dev.to/masaoshimadaopen/my-custom-filter-showed-zero-false-positives-then-i-realized-my-validation-logic-was-flawed-6o</guid>
      <description>&lt;p&gt;Hey there, it's your "Oji" (38-year-old AI/quant dev on the side). During the week, I'm a regular company employee; on weekends, I tinker with AI agents and automated trading bots.&lt;/p&gt;

&lt;p&gt;Recently, I was building a filter to automatically extract "AI-related companies" from a market data list, based on their business descriptions. The mechanism is simple: it scores companies based on hits of predefined positive keywords (e.g., "machine learning," "natural language processing") and negative keywords (e.g., "gaming," "entertainment").&lt;/p&gt;

&lt;p&gt;To improve the filter's accuracy, I set up a test to tune the threshold (how many positive keyword hits qualify a company). I prepared lists of positive examples (actual AI companies) and negative examples (non-AI companies) to see how well the filter classified them—standard stuff.&lt;/p&gt;

&lt;p&gt;Then I ran the test, and the results were mind-blowing.&lt;/p&gt;

&lt;p&gt;"Zero false positives, no matter the threshold I tried."&lt;/p&gt;

&lt;p&gt;For a moment, I thought, "Is my filter a genius?" Zero false positives is usually impossible. But seconds later, I cooled down. Results that are &lt;em&gt;too perfect&lt;/em&gt; usually mean something is wrong.&lt;/p&gt;

&lt;p&gt;Sure enough, when I manually reviewed the list of companies that passed the filter, I found obvious non-AI companies like gaming studios mixed in. My filter was letting garbage through, yet the test reported "no issues." The despair I felt realizing this late on a weekday night was quite something.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why "Zero False Positives"?
&lt;/h3&gt;

&lt;p&gt;The root cause wasn't a bug in the code itself, but a design flaw in the validation logic.&lt;/p&gt;

&lt;p&gt;Basic accuracy validation involves looking at two metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;False Positive (FP)&lt;/strong&gt;: A non-AI company incorrectly classified as an AI company.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;False Negative (FN)&lt;/strong&gt;: An AI company incorrectly classified as a non-AI company.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The issue I encountered was with FP. To count FPs, the definition of &lt;strong&gt;negative examples&lt;/strong&gt; (companies that are "not AI companies") is crucial.&lt;/p&gt;

&lt;p&gt;And here was my definition of a negative example:&lt;/p&gt;

&lt;p&gt;"A company with 0 positive keyword hits AND 2 or more negative keyword hits."&lt;/p&gt;

&lt;p&gt;This was the source of all evil. This definition, seemingly sound at first glance, contained a self-contradiction.&lt;/p&gt;

&lt;p&gt;My filter's rule for classifying a company as "passing (AI-related)" was: "2 or more positive keyword hits."&lt;/p&gt;

&lt;p&gt;You probably see it now.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Filter's passing condition&lt;/strong&gt;: &lt;code&gt;positive_hits &amp;gt;= 2&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Test's negative example condition&lt;/strong&gt;: &lt;code&gt;positive_hits == 0&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These two conditions can &lt;em&gt;never&lt;/em&gt; be true simultaneously. A company classified as "passing" by the filter &lt;em&gt;already&lt;/em&gt; has &lt;code&gt;positive_hits &amp;gt;= 2&lt;/code&gt;, so it can &lt;em&gt;never&lt;/em&gt; fit the negative example definition of &lt;code&gt;positive_hits == 0&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;In essence, when trying to count "negative examples that were incorrectly classified as passing (i.e., false positives)," the structure itself made it impossible for any company to be both a "negative example" and "passing." Of course, false positives would always be zero.&lt;/p&gt;

&lt;p&gt;Translating the concept to code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Filter rule: Pass if positive keyword count &amp;gt;= 2&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;is_theme_company&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;positive_hits&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;positive_hits&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Validation's negative example definition (buggy)&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;is_negative_example&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;positive_hits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;negative_hits&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Only consider companies with zero positive hits as negative examples&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;positive_hits&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;negative_hits&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Validation process&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;false_positives&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;company&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;all_companies&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;is_selected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;is_theme_company&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;company&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pos_hits&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// If it's a negative example but selected, count as FP&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;is_selected&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;is_negative_example&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;company&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pos_hits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;company&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;neg_hits&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;false_positives&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// With this logic, is_selected=true (pos_hits&amp;gt;=2) and is_negative_example=true (pos_hits==0)&lt;/span&gt;
&lt;span class="c1"&gt;// can never both be true, so false_positives will always be zero.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test was too tightly coupled to the logic it was supposed to validate, preemptively deciding the outcome. It was a brutal rookie mistake.&lt;/p&gt;

&lt;h3&gt;
  
  
  How I Fixed It
&lt;/h3&gt;

&lt;p&gt;Once I understood the cause, the fix was simple.&lt;/p&gt;

&lt;p&gt;I changed the definition of negative examples to be independent of keyword hit counts. Specifically, I manually selected dozens of companies that were unequivocally &lt;em&gt;not&lt;/em&gt; AI companies (e.g., food manufacturers, apparel brands, construction companies) and created a fixed "negative example list."&lt;/p&gt;

&lt;p&gt;Running the test again with this list, false positives, as expected, came pouring out. Finally, the numbers showed that the filter was indeed picking up many irrelevant companies. This was the real starting point for tuning.&lt;/p&gt;

&lt;h3&gt;
  
  
  My Takeaway
&lt;/h3&gt;

&lt;p&gt;The lessons from this failure are significant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"The test passed" only means something if the &lt;em&gt;test itself&lt;/em&gt; is correct.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Finding code bugs with tests is fundamental, but you must always question the possibility that the test logic itself is flawed. Especially when validating rules you've created yourself, there's a risk that the validation logic gets dragged by the tested logic, unconsciously leading to "conclusion-first" tests.&lt;/p&gt;

&lt;p&gt;When you get "too perfect results," first question your own assumptions. This is a crucial reminder I'll engrave into my mind.&lt;/p&gt;

&lt;p&gt;This story also applies to backtesting automated trading bots. When a backtest shows unusually good performance, it's rarely because you've discovered a brilliant strategy; it's usually because you're looking at future data (leakage) or not accounting for fees and slippage.&lt;/p&gt;

&lt;p&gt;I'm logging embarrassing failures like this as part of building in public. Hopefully, it helps someone else in their solo dev journey.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>debugging</category>
      <category>machinelearning</category>
      <category>automation</category>
    </item>
    <item>
      <title>米国企業の分析bot、外国企業だけ“黙って”分析をスキップしてた話 — 10-Kと20-Fで「リスクの章立て」が違う罠</title>
      <dc:creator>oji - building AI in public</dc:creator>
      <pubDate>Tue, 08 Sep 2026 23:30:19 +0000</pubDate>
      <link>https://dev.to/masaoshimadaopen/mi-guo-qi-ye-nofen-xi-bot-wai-guo-qi-ye-dakemo-tutefen-xi-wosukitupusitetahua-10-kto20-fderisukunozhang-li-te-gawei-umin-37l3</link>
      <guid>https://dev.to/masaoshimadaopen/mi-guo-qi-ye-nofen-xi-bot-wai-guo-qi-ye-dakemo-tutefen-xi-wosukitupusitetahua-10-kto20-fderisukunozhang-li-te-gawei-umin-37l3</guid>
      <description>&lt;p&gt;どうも、おじいです。平日夜と休日に、AI エージェントとか自動売買 bot をちまちま作ってる 38 歳。&lt;/p&gt;

&lt;p&gt;今日は、副業で開発してる企業分析 bot でやらかした、結構えぐいバグの話。動いてるように見えて、実は大事なデータをごっそり見逃してた。データ処理系の個人開発やってる人には、あるあるかもしれない。&lt;/p&gt;

&lt;h3&gt;
  
  
  何が起きたか
&lt;/h3&gt;

&lt;p&gt;作ってる bot の一つに、米国上場企業の年次報告書（SEC ファイリング）を読み込んで、事業リスクを LLM で要約・評価するやつがある。企業の健全性をざっくり把握するためのツールだ。&lt;/p&gt;

&lt;p&gt;こいつが、ある特定の企業群について、やけに「クリーン」な分析結果を出してくることに気づいた。「リスク要因: 特になし」みたいな。最初は「お、超優良企業か？」なんて思ってたんだけど、何社か続くとさすがにおかしい。そんなわけない。&lt;/p&gt;

&lt;p&gt;ログを掘ってみたら、衝撃の事実が判明した。bot が、特定の企業の分析だけ「黙って」スキップしてた。エラーも吐かずに、ただ空っぽのデータを返してた。そのせいで、後段の LLM は「リスクに関する記述がなかった」と判断して、「リスクなし」という結論を出してたわけだ。これはやばい。&lt;/p&gt;

&lt;h3&gt;
  
  
  原因: 「10-K」と「20-F」という様式の違い
&lt;/h3&gt;

&lt;p&gt;原因は、米国証券取引委員会（SEC）に提出される年次報告書の「様式」の違いにあった。&lt;/p&gt;

&lt;p&gt;bot は当初、米国企業が提出する「&lt;strong&gt;Form 10-K&lt;/strong&gt;」という書類だけを想定して作ってた。この 10-K では、事業リスクは「&lt;strong&gt;Item 1A. Risk Factors&lt;/strong&gt;」という項目に記載されるのがお作法。だから、bot は機械的に「Item 1A」のセクションを引っこ抜くように実装してた。&lt;/p&gt;

&lt;p&gt;ところが、分析結果が空になってた企業を調べたら、全部「外国企業」だった。米国市場に上場してる外国企業は、10-K の代わりに「&lt;strong&gt;Form 20-F&lt;/strong&gt;」という別の様式で報告書を提出する。&lt;/p&gt;

&lt;p&gt;そして、この 20-F では、リスク要因は「&lt;strong&gt;Item 3.D. Risk Factors&lt;/strong&gt;」に書かれている。&lt;/p&gt;

&lt;p&gt;つまり、bot は 20-F の書類に対して、存在しない「Item 1A」を探しに行ってた。当然、見つかるわけがない。&lt;/p&gt;

&lt;p&gt;一番の問題は、その後の処理だった。セクションが見つからなかった時のエラーハンドリングを、こんな風に書いてた。&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 抽出失敗と「該当なし」を混同する危険なコード
&lt;/span&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;risk_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_section&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Item 1A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;SectionNotFound&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;risk_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt; &lt;span class="c1"&gt;# これでは失敗したのか、元々空なのか分からない
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;一見、例外をキャッチしてて安全そうに見える。でも、これが罠だった。「セクションが見つからなくて抽出に失敗した」ケースと、「セクションはあったけど中身が空だった」ケースが、どっちも &lt;code&gt;risk_text = ""&lt;/code&gt; になってしまう。&lt;/p&gt;

&lt;p&gt;このせいで、データ欠損という重大な異常が、正常な「リスク記載なし」として処理され、静かにバグが進行してたわけだ。&lt;/p&gt;

&lt;h3&gt;
  
  
  修正: フォームタイプをちゃんと見て、失敗を区別する
&lt;/h3&gt;

&lt;p&gt;対策はシンプル。&lt;br&gt;
まず、処理対象のドキュメントが 10-K なのか 20-F なのかをちゃんと判別する。その上で、フォームタイプに応じて読み込むべきセクション ID を切り替えるようにした。&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# フォームタイプに応じて処理を分岐する改善コード
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;form_type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;10-K&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;section_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Item 1A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;form_type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;20-F&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;section_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Item 3.D&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# 未対応のフォームタイプなら、処理を中断
&lt;/span&gt;    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Unsupported form type: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;form_type&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="c1"&gt;# 抽出を試みる。失敗したらNoneが返るようにget_section側を修正
&lt;/span&gt;&lt;span class="n"&gt;risk_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;try_get_section&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;section_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;risk_text&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# セクション自体が見つからなかった場合（＝異常）
&lt;/span&gt;    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Section not found: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;section_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# ... エラー処理 ...
&lt;/span&gt;&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;risk_text&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# セクションはあったが空だった場合（＝正常）
&lt;/span&gt;    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Risk factors section is empty: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;section_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;get_section&lt;/code&gt; メソッドも修正して、セクションが見つからない場合は空文字じゃなく &lt;code&gt;None&lt;/code&gt; を返すように変更した。&lt;/p&gt;

&lt;p&gt;これで、呼び出し側は戻り値を見るだけで、&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;None&lt;/code&gt; → セクションが見つからなかった（＝想定外のエラー）&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;""&lt;/code&gt; （空文字）→ セクションはあったが、中身が空だった（＝データ上、リスク記載なし）&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;(str)&lt;/code&gt; → 正常に抽出できた&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;と、明確に区別できるようになった。この修正後、これまでスキップされてた外国企業の分析も、無事に動き出した。&lt;/p&gt;

&lt;h3&gt;
  
  
  学び: エラーを握りつぶすと、静かに死ぬ
&lt;/h3&gt;

&lt;p&gt;今回の失敗から学んだことは２つ。&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;ドメイン知識はマジで大事。&lt;/strong&gt; 今回で言えば、SEC ファイリングに 10-K と 20-F の違いがある、という知識。技術的な実装力だけじゃなくて、扱ってるデータそのものへの理解がないと、こういう根本的な見落としをする。特に金融系は、こういう「お作法」の塊みたいな世界。&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;「失敗」と「空」を区別する設計。&lt;/strong&gt; エラーを安易に握りつぶして、空文字や &lt;code&gt;0&lt;/code&gt; みたいな「無害そうな」デフォルト値を返すのは、本当に危険。サイレントなデータ欠損は、気づいた時には手遅れ、なんてこともあり得る。自動売買 bot なら、資産を溶かす原因に直結する。&lt;code&gt;None&lt;/code&gt; を使ったり、専用の例外を投げたりして、異常は異常としてちゃんと後続の処理に伝える設計がいかに重要か、身をもって知った。&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;副業の個人開発だと、ついドキュメントを斜め読みして「動いたからヨシ！」で進めがちだけど、こういう地味な仕様の確認こそ、一番時間をかけるべき部分なのかもしれない。&lt;/p&gt;

&lt;p&gt;この失敗ログが、同じようにデータ処理系の bot を作ってる誰かの参考になれば嬉しい。&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If a provider-agnostic RAG Q&amp;amp;A API is useful to you, mine is MIT-licensed on GitHub: &lt;a href="https://github.com/masaoshimadaOpen/rag-faq-api" rel="noopener noreferrer"&gt;rag-faq-api&lt;/a&gt;. It runs and passes its full test suite **with no API key&lt;/em&gt;* (offline stub LLM + hashing embedder), swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness (Hit@k / MRR / Recall@k) with a chunking sweep.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
