<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Venkat Girish</title>
    <description>The latest articles on DEV Community by Venkat Girish (@venkat_girish_d0d414e01e5).</description>
    <link>https://dev.to/venkat_girish_d0d414e01e5</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149745%2Ff5cbeb05-4ae5-41de-99b8-a8b1e63d9daf.png</url>
      <title>DEV Community: Venkat Girish</title>
      <link>https://dev.to/venkat_girish_d0d414e01e5</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/venkat_girish_d0d414e01e5"/>
    <language>en</language>
    <item>
      <title>I Built a Machine-Memory Agent and Found Out I Wasn't Using Memory at All</title>
      <dc:creator>Venkat Girish</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:01:20 +0000</pubDate>
      <link>https://dev.to/venkat_girish_d0d414e01e5/i-built-a-machine-memory-agent-and-found-out-i-wasnt-using-memory-at-all-2l7b</link>
      <guid>https://dev.to/venkat_girish_d0d414e01e5/i-built-a-machine-memory-agent-and-found-out-i-wasnt-using-memory-at-all-2l7b</guid>
      <description>&lt;p&gt;My README said the agent used &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;'s Reflect operation to reason over a machine's failure history. My code called Recall, stuffed the results into a prompt, and asked an LLM to do the reasoning itself. Nobody caught it until I went looking for a bug in something else entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this thing actually does
&lt;/h2&gt;

&lt;p&gt;The idea is simple and, I think, underused: every industrial machine should have its own memory. A technician fixes a pump today; six months from now, a different technician sees the same vibration pattern on the same pump and has no idea it happened before. The service report is in a binder somewhere, or in the first technician's head, and neither is searchable at 2am when the line is down.&lt;/p&gt;

&lt;p&gt;So I built a small system — a JSON file of synthetic maintenance records, a Flask UI, and &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; doing the memory work — that lets a technician describe what they're seeing and get back an answer grounded in what actually happened last time, including the fixes that &lt;em&gt;didn't&lt;/em&gt; work. Not "here's a generic troubleshooting guide," but "R. Bhatt tried re-greasing this exact bearing in February, it didn't help, here's what did."&lt;/p&gt;

&lt;p&gt;The architecture is intentionally small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Technician input (symptoms / sensor readings)
              │
              ▼
        demo_agent.py
              │
      ┌───────┴───────┐
      │               │
   RECALL          REFLECT
      │               │
      └───────┬───────┘
              ▼
        Hindsight Memory
              │
              ▼
         Groq LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two modes, two different jobs: a side-by-side "with vs. without memory" comparison uses Recall — grab relevant fragments, hand them to an LLM, let it reason. A pre-repair check uses Reflect, which is supposed to do the reasoning itself, grounded directly in memory. That distinction is the whole story of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moment I found out my "Reflect" wasn't reflecting
&lt;/h2&gt;

&lt;p&gt;I was doing an honest self-review before treating this as done, and I asked myself a simple question: does the code actually call &lt;code&gt;reflect()&lt;/code&gt; anywhere? I grepped. It didn't.&lt;/p&gt;

&lt;p&gt;Here's what &lt;code&gt;check_before_repair&lt;/code&gt; looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_before_repair&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;English&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="n"&gt;texts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;machine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are the Chief Reliability Engineer for a plant. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;SCHEMA&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json_mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recall, then a generic chat completion doing the actual judgment call. It worked fine in a demo. It was also not what I'd told anyone it was. The difference matters more than it sounds: Recall's job is retrieval — hand back relevant fragments and get out of the way. Reflect's job is synthesis — decide whether this matches a known failure pattern, weigh conflicting evidence, and answer the actual judgment question. I was doing the second thing by hand, with a hand-rolled prompt, instead of asking the memory system to do it.&lt;/p&gt;

&lt;p&gt;The fix was to call Hindsight's &lt;code&gt;reflect()&lt;/code&gt; directly, with structured output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reflect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;hs&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;reflect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tags_match&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;any&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;response_schema&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;response_schema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;include_facts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;friendly_error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;response_schema&lt;/code&gt; gives me a JSON Schema and Hindsight returns &lt;code&gt;structured_output&lt;/code&gt; matching it — no more asking an LLM to "please respond only with valid JSON" and hoping. &lt;code&gt;include_facts=True&lt;/code&gt; is the part that actually changed the architecture: it returns a &lt;code&gt;based_on&lt;/code&gt; field naming the exact memories, mental models, and directives Hindsight used to produce the answer. My "why does the agent think this" panel in the UI now renders &lt;em&gt;that&lt;/em&gt; — not a list of things I recalled and hoped the model used, but a list of things the model reports it actually used.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotcha that cost me an afternoon: IDs live in &lt;code&gt;context&lt;/code&gt;, not &lt;code&gt;text&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Once Reflect was wired in, my citation-checking code broke. I'd built a guard that only lets the model cite a record ID if that ID appears somewhere in the retrieved text — a cheap, effective hallucination check. It started rejecting every citation.&lt;/p&gt;

&lt;p&gt;Turned out Hindsight's fact extraction paraphrases content into atomic statements, and the &lt;code&gt;[MAINT-2025-001]&lt;/code&gt;-style tag I'd embedded in my retained text doesn't survive that paraphrasing. But the &lt;code&gt;context&lt;/code&gt; string I'd also passed at retain time — &lt;code&gt;"maintenance record MAINT-2025-001 | machine MCH-017 | tags: ..."&lt;/code&gt; — comes back untouched on every fact:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;id_blob&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;texts&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;facts&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;conf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id_blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;srcs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sources_from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id_blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Lesson, in one line: if you need something to survive Hindsight's extraction pass verbatim, put it in &lt;code&gt;context&lt;/code&gt;, not in the content you're asking it to summarize.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bug that only exists because of async
&lt;/h2&gt;

&lt;p&gt;The other real bug: a cached Hindsight client, reused across requests, would occasionally throw &lt;code&gt;Timeout context manager should be used inside a task&lt;/code&gt; — a classic symptom of an aiohttp session created in one asyncio event loop being reused in another. Flask's dev server spins up a new thread per request; the sync wrapper spins up a new event loop per call. Cache the client across both and eventually they collide.&lt;/p&gt;

&lt;p&gt;The fix was cheap once I understood it: stop caching, build a fresh client per call, and retry once on anything that looks transient:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;hs&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;hindsight_client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Hindsight&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                      &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.2&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;_is_transient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I only found this because I tested concurrent requests on purpose. It would have shown up live, in front of the people I most wanted it not to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What memory that grows actually looks like
&lt;/h2&gt;

&lt;p&gt;The part that convinced me this wasn't just retrieval-with-extra-steps was the closed loop. Log a failed fix, and the very next diagnosis for that machine changes — not because I re-ran a script, but because the memory itself grew:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before logging feedback:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Extending lubrication intervals from monthly to quarterly failed twice, once due to a manual policy change and once due to a CMMS software reset."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;After logging "re-greased the bearing housing — didn't work":&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Re-greasing the bearing housing has been attempted and resulted in failure, as noted in log LIVE-001. Discontinue all attempts to re-grease; replace the drive-end bearing with SKF 6309 directly."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same machine, same underlying question, a genuinely different answer because a technician's outcome became a memory a few seconds earlier. I also gave each machine a "digital twin" — a Hindsight mental model whose &lt;code&gt;source_query&lt;/code&gt; asks it to synthesize that machine's full history into a living profile, with &lt;code&gt;trigger={"refresh_after_consolidation": True}&lt;/code&gt; so it updates itself as new incidents come in, rather than a summary I'd have to regenerate by hand. And I moved a couple of standing rules — lockout-tagout before touching rotating equipment, never invent a technician name or a date — out of my prompt text and into &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Hindsight directives&lt;/a&gt;, so they're enforced by the memory system on every request instead of by convention in a string I control.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell someone starting this tomorrow
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grep your own claims.&lt;/strong&gt; If your docs say you use an operation, verify the code path actually calls it. I didn't, for weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide what needs to survive verbatim, and put it somewhere that isn't summarized.&lt;/strong&gt; &lt;code&gt;context&lt;/code&gt; is not &lt;code&gt;content&lt;/code&gt;; treat them differently on purpose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test concurrency before your users do.&lt;/strong&gt; The scariest bug I found only existed under simultaneous requests, and I only found it by deliberately generating them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A system's memory should get smarter without you touching code.&lt;/strong&gt; If the only way your agent "learns" is by editing a seed file, it's not really learning — it's re-authoring.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval and reasoning are different jobs.&lt;/strong&gt; Recall for "what's relevant." Reflect for "what does this mean." Conflating them is an easy trap, and I fell into it first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this needed a bigger model. It needed the memory system to actually be doing the job I'd told everyone it was doing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
