<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Palle Poojitha </title>
    <description>The latest articles on DEV Community by Palle Poojitha  (@pallepoojith).</description>
    <link>https://dev.to/pallepoojith</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147761%2Fdac124f5-5b5e-4120-a82a-35e9a53bb783.png</url>
      <title>DEV Community: Palle Poojitha </title>
      <link>https://dev.to/pallepoojith</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pallepoojith"/>
    <language>en</language>
    <item>
      <title>My Hindsight Agent Started Citing Its Own Guesses as Evidence</title>
      <dc:creator>Palle Poojitha </dc:creator>
      <pubDate>Tue, 29 Sep 2026 01:24:51 +0000</pubDate>
      <link>https://dev.to/pallepoojith/my-hindsight-agent-started-citing-its-own-guesses-as-evidence-2542</link>
      <guid>https://dev.to/pallepoojith/my-hindsight-agent-started-citing-its-own-guesses-as-evidence-2542</guid>
      <description>&lt;p&gt;The answer looked great. Loop, the social media agent I built on &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight agent memory&lt;/a&gt;, recommended a carousel at 10 AM for a café's new single-origin coffee, and it backed the plan with a confident justification: "We're using the carousel layout and 10 AM slot that were recommended in our Sep 28 2026 plan for the Araku coffee series."&lt;/p&gt;

&lt;p&gt;The only plan from Sep 28 was one Loop had written itself, a few minutes earlier, with nothing useful in memory. The café's real history pointed the other way: its Friday-evening reels mostly reached 4,700 to 7,000 people, and its static posts never broke 900.&lt;/p&gt;

&lt;p&gt;My agent had started agreeing with itself. This post covers how that happened, why it's an easy mistake with any long-term memory, and what the fix looks like in code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Loop does
&lt;/h2&gt;

&lt;p&gt;Loop is a strategist for one brand at a time. You describe the brand in a few taps: what you post about, where, how you sound, and any hard rules. Then you ask it things like "write an Instagram post for our new Araku coffee" or "why do my discount posts get so little reach?" When a post goes live, you log how it did: reach, likes, comments, shares, and a note such as "lots of comments asked where the beans are from."&lt;/p&gt;

&lt;p&gt;The stack is deliberately small. A FastAPI app (&lt;code&gt;main.py&lt;/code&gt;) serves a single-page UI, Groq handles generation (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt;, with &lt;code&gt;qwen/qwen3-32b&lt;/code&gt; as a fallback), and Hindsight holds everything the agent knows. Every brand gets its own memory bank, created with a short mission statement that tells Hindsight what the memory is for: learning which formats, topics, tones and posting times work for this audience, and updating beliefs when new results contradict old ones.&lt;/p&gt;

&lt;p&gt;Every chat message runs the same three steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Recall&lt;/strong&gt; what Hindsight knows that's relevant to the request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate&lt;/strong&gt; a reply grounded in those memories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retain&lt;/strong&gt; something, so the next reply is better.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Post results go through &lt;code&gt;retain&lt;/code&gt; too, with their real timestamps. Hindsight extracts facts from what it's given and consolidates them in the background into observations, beliefs like "reels featuring the café's craft drive the highest engagement." Those observations fill a "What Loop has learned" panel beside the chat, and &lt;code&gt;reflect&lt;/code&gt; turns the whole history into a weekly readout. If you haven't used Hindsight's &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;retain, recall and reflect APIs&lt;/a&gt;, that is the entire surface area. The rest of this post is about the places where it was easy to hold them wrong.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkx1h8d0qohbhljelx1ch.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkx1h8d0qohbhljelx1ch.png" alt="How Loop uses Hindsight: recall before every reply, retain only owner messages and measured results" width="800" height="460"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug: step three stored the wrong thing
&lt;/h2&gt;

&lt;p&gt;My first version of step three looked reasonable. It stored the exchange: &lt;code&gt;content=f"The user asked: {req.message}\nLoop replied: {reply[:600]}"&lt;/code&gt;. Store the conversation and the agent remembers the conversation. That's what memory means, right?&lt;/p&gt;

&lt;p&gt;Here's what actually happened. The first time I asked about the new coffee, recall came back empty (more on why in the lessons), so the model did what models do: it produced a plausible generic plan, a carousel at 10 AM. That reply went straight into memory. Hindsight did exactly its job. It extracted a clean, dated fact: "Loop created a carousel Instagram post for Araku coffee featuring farm, roasting, brewing, and tasting notes."&lt;/p&gt;

&lt;p&gt;The next time I asked, recall surfaced that fact right next to the real post history. To a language model, a dated statement about a carousel plan for Araku coffee looks exactly like evidence, so it cited it. The memory layer wasn't wrong. I had fed it the agent's own speculation and filed it as experience.&lt;/p&gt;

&lt;p&gt;I think this generalizes. Any agent that writes its outputs into the same store it reads evidence from will drift toward agreeing with itself. The first guess gets cited, the citation makes the second answer more confident, and the second answer gets stored too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: retain inputs, not outputs
&lt;/h2&gt;

&lt;p&gt;The fix was two changes. First, step three now stores only what the owner said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# RETAIN what the owner said (requests, preferences, feedback) in the background so the user
# isn't kept waiting. Loop's own reply is deliberately not stored: otherwise its past guesses
# come back later as "evidence" and the agent starts agreeing with itself.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;hs_run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;HS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The brand owner told Loop: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;owner message in chat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;retain_async&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Owner messages are real signal: "we're launching Araku coffee," "never mention competitors," "stop using emojis." The agent's replies are hypotheses, and a hypothesis only becomes evidence after it has been tested. That's exactly what the "Log results" flow is for. A draft that gets posted and measured comes back as a result with numbers attached. A draft nobody posted shouldn't come back at all.&lt;/p&gt;

&lt;p&gt;Second, the system prompt now says so explicitly, because older banks can still hold self-generated facts: "Measured post results and owner statements are evidence. Loop's own earlier drafts are NOT evidence: never justify a choice by saying you suggested it before."&lt;/p&gt;

&lt;p&gt;&lt;code&gt;retain_async=True&lt;/code&gt; turned out to be a nice side effect. Retention now happens in the background, so the user gets the reply without waiting on fact extraction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall: two queries, not one
&lt;/h2&gt;

&lt;p&gt;The other half of grounding is what you ask memory for. A single recall on the user's message works for specific questions and fails for vague ones. "What should I post this weekend?" doesn't look anything like "the café is closed on Mondays" or "never use more than 4 hashtags," but both matter. So every chat runs two recalls and merges them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Two searches: the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s actual request + a standing &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;what do we know&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; query.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;queries&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;brand voice, owner rules, audience preferences, best and worst performing posts and timing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;merged&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;q&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;queries&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;recall_texts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;q&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
                &lt;span class="n"&gt;merged&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;merged&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The standing query is boring on purpose. It's the "things you must never forget about this brand" query, and it means the owner's rules land in context even when the question doesn't mention them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like now
&lt;/h2&gt;

&lt;p&gt;To make the effect of memory visible, the UI has a Memory switch. Off means no recall and no retain: the model gets a generic system prompt and its own separate conversation history. That separate history was its own bug. My first version shared one history across both modes, which quietly leaked brand details into the "generic" answers and made the comparison meaningless.&lt;/p&gt;

&lt;p&gt;To test Loop, I seeded a bank with 24 posts of history for a café brand, plus five owner rules. Then I asked the same question in both modes: "Write an Instagram post for our new Araku single-origin coffee."&lt;/p&gt;

&lt;p&gt;With memory off, Loop wrote a competent caption for a generic coffee brand and finished with eight hashtags: &lt;code&gt;#ArakuCoffee #SingleOrigin #SpecialtyCoffee #CoffeeLovers #FreshRoast #FromFarmToCup #MorningBoost #CoffeeCulture&lt;/code&gt;. The owner's very first rule is "never more than 4 hashtags."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq7ekxm7tff04tw67tmwd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq7ekxm7tff04tw67tmwd.png" alt="Memory off: a generic caption that breaks the owner's hashtag rule" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;With memory on, it suggested a 15-second reel of the first pour, and its "Why" line used numbers it hadn't invented: "The Sep 25 Reel of the first Araku pour pulled 6,880 reach, 905 likes, 184 comments, and 132 shares—our strongest recent performance, so we'll replicate that format and timing while staying under the 4-hashtag limit."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6y9cggf3lrmgocih38bk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6y9cggf3lrmgocih38bk.png" alt="Memory on: 14 memories recalled, and a draft grounded in a real post's numbers" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then the loop closed. I logged a result for that reel: 7,120 reach, 948 likes, 201 comments, and a note that lots of commenters asked where the beans came from. In a later session, with a fresh page and an empty chat history, the new draft invited followers to ask about the beans' journey, and its "Why" line cited the new result: "A Reel on 28 Sept 2026 about Araku beans reached 7,120 people, earned 948 likes and sparked many 'origin?' comments."&lt;/p&gt;

&lt;p&gt;The weekly readout, written by &lt;code&gt;reflect&lt;/code&gt;, turned that one note into a recommendation: test a post "dedicated exclusively to the 'story behind the bean'". It also flagged the 15%-off announcement that reached 570 people as the thing that wasn't working.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0lqgut7usi3ol37pl0q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0lqgut7usi3ol37pl0q.png" alt="The weekly readout from reflect, next to the beliefs Hindsight consolidated" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;None of that lives in the prompt or the chat history. It lives in the bank.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Retain inputs and measured outcomes, never raw outputs.&lt;/strong&gt; If your agent writes into the store it reads evidence from, it will eventually cite itself. Outputs should earn their way into memory by being tested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Give every memory a stable identity.&lt;/strong&gt; Sample posts and the brand profile use fixed document IDs, so saving a profile twice replaces it instead of stacking two contradictory versions. Hindsight enforces this for async batches: it rejects a batch whose items share a document ID, which is how I found out my seeding code was wrong. Each item now carries its own ID and its real timestamp:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;_post_sentence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;post performance&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;_post_time&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;document_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sample-post-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;02&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Timestamp everything that happened at a specific time.&lt;/strong&gt; "What's working lately?" is a temporal question. Results carry their real posting time through &lt;code&gt;timestamp=&lt;/code&gt;, so memory can tell an August post from one logged yesterday. The readout's talk of "the latest reel" depends on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Know your client's concurrency model.&lt;/strong&gt; Remember the empty recall that started all this? The Python client I used wraps async calls in a sync API, and its HTTP session is tied to the event loop of the first thread that uses it. FastAPI runs sync routes on a thread pool, so calls began failing at random with "Timeout context manager should be used inside a task." It looked like flaky memory, but it was a threading problem on my side. The smallest fix was one worker thread that owns every Hindsight call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;_hs_thread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_workers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;thread_name_prefix&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hindsight&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;hs_run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;_hs_thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;submit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;result&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cleaner long-term version is async routes calling the client's async methods directly. The single thread was the smallest change that made recall reliable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Show the memory.&lt;/strong&gt; The "What Loop has learned" panel and the "What it remembered" button did more for trust than any prompt tweak. When an answer cites a number, you can open the exact memories it came from. When something's wrong, you can see why. That's how I spotted my own agent's carousel plan sitting in its recall results, dressed up as history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;The model didn't get smarter between the eight-hashtag answer and the one that cited a real reel. What changed was what it was allowed to remember, and what I stopped letting it remember. If you're adding long-term memory to an agent, it's worth reading up on &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;what agent memory is and how it differs from stuffing context&lt;/a&gt; before you decide what to write into it. In Loop, the most important line of code is the one that doesn't retain the reply.&lt;br&gt;
Project code: &lt;a href="https://github.com/pallepoojith/Apex-agentic-Hindsight" rel="noopener noreferrer"&gt;https://github.com/pallepoojith/Apex-agentic-Hindsight&lt;/a&gt;&lt;br&gt;
Shared with the Code.in community.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>python</category>
      <category>agents</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
