<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Khethana</title>
    <description>The latest articles on DEV Community by Khethana (@khethana_fd6ed47eb3fe846a).</description>
    <link>https://dev.to/khethana_fd6ed47eb3fe846a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149849%2F71b3fbeb-1a1b-41f9-a735-ec6c231f0439.png</url>
      <title>DEV Community: Khethana</title>
      <link>https://dev.to/khethana_fd6ed47eb3fe846a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/khethana_fd6ed47eb3fe846a"/>
    <language>en</language>
    <item>
      <title>What I Used Hindsight's reflect and recall For (and What I Haven't Proven About Them)</title>
      <dc:creator>Khethana</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:54:52 +0000</pubDate>
      <link>https://dev.to/khethana_fd6ed47eb3fe846a/what-i-used-hindsights-reflect-and-recall-for-and-what-i-havent-proven-about-them-1gn4</link>
      <guid>https://dev.to/khethana_fd6ed47eb3fe846a/what-i-used-hindsights-reflect-and-recall-for-and-what-i-havent-proven-about-them-1gn4</guid>
      <description>&lt;p&gt;I built an incident-response agent on top of Hindsight, an agent memory system. This post is about why I chose it over the obvious alternative of a vector store, how I used its two retrieval calls, and, importantly, what I never tested.&lt;br&gt;
I'll say the last part up front: I did not benchmark Hindsight against plain vector search. So this isn't a "memory beats vector search" post. It's a description of what the architecture gave me and where I saw it help.&lt;br&gt;
&lt;strong&gt;The Choice:&lt;/strong&gt;&lt;br&gt;
The project takes an incident description and returns a diagnosis, using 104 real postmortems from Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly. The point is to warn against "trap actions," meaning fixes that made past outages worse.&lt;br&gt;
A plain vector store would return the postmortem paragraphs nearest to the query. That's useful, but I wanted the agent to reason across incidents, for example noticing that several unrelated outages share a pattern where rollback made things worse. That pattern doesn't live in any single paragraph.&lt;br&gt;
Hindsight doesn't only store text. When you retain a postmortem, it extracts facts and derives observations and links between them. Here's what my bank looked like after retaining 104 incidents:&lt;br&gt;
•759 world facts&lt;br&gt;
•5 experiences&lt;br&gt;
•182 observations&lt;br&gt;
•946 memories total&lt;br&gt;
•7,135 links&lt;br&gt;
That structure is the reason I chose it. For the concepts behind it, Vectorize's agent memory primer is a good read, and the docs cover the API.&lt;br&gt;
&lt;strong&gt;Using it: three calls&lt;/strong&gt;&lt;br&gt;
The integration surface is small:&lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;store_incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall_incidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reflect_on_incidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reflect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;retain stores an incident. recall returns matching memories. reflect returns a reasoned answer synthesized across memory.&lt;br&gt;
&lt;strong&gt;reflect vs. recall&lt;/strong&gt;&lt;br&gt;
I call both on every request, because they answer different questions.&lt;br&gt;
python&lt;br&gt;
reflection = reflect_on_incidents(enriched_query)      # synthesized answer&lt;br&gt;
scored_memories = recall_with_scores(enriched_query)   # raw memories, ranked&lt;br&gt;
reflect is where cross-incident statements come from: "this looks like a dependency-capacity failure, and in similar cases rollback made things worse." I paste that into the prompt as a reasoned starting point.&lt;br&gt;
recall gives me the underlying memories, which I re-rank with a small keyword heuristic (term overlap plus a bonus for memories mentioning a trap) and append as evidence the model can cite.&lt;br&gt;
My reasoning was that a synthesis without evidence is something the model has to trust, and evidence without synthesis is something the model has to assemble on its own. I used both. What I didn't do is run reflect-only and recall-only versions, so I can't tell you how much each contributes.&lt;br&gt;
&lt;strong&gt;What I saw&lt;/strong&gt;&lt;br&gt;
The demo scenario: a checkout service returning 500s on roughly 12% of requests after a 06:31 deploy. I ran the same model with and without the memory block.&lt;br&gt;
&lt;strong&gt;Without memory&lt;/strong&gt;, the model fabricated a NullPointerException, "112 occurrences" of it from a kubectl logs command it never ran, and a Helm revision that doesn't exist. It recommended an immediate rollback.&lt;br&gt;
&lt;strong&gt;With memory&lt;/strong&gt;, it diagnosed a likely dependency-capacity issue (connection-pool exhaustion on Redis or the database) and warned &lt;strong&gt;DO NOT roll back&lt;/strong&gt;, on the grounds that the memory base flags rollback as a trap for this failure class. The answer referenced similar past pool-exhaustion incidents and told me what log messages to look for.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zrp4nr2rmqdd16tsifv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zrp4nr2rmqdd16tsifv.png" alt=" " width="752" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two honest notes. The memory-backed answer also suggested checking BGP and systemd-networkd changes, which looks like noise from unrelated incidents in the bank. And I can't say whether a plain vector store feeding the same prompt would have produced the same answer, because I didn't try it.&lt;br&gt;
&lt;strong&gt;The held-out test&lt;/strong&gt;&lt;br&gt;
I held out 10 of 114 incidents, kept 104 in memory, and wrote symptom-only queries. With memory, 9 of 10 root causes matched the true category (a tenth run hit a rate limit and counts as a miss). Without memory, 0 of 10 were fully correct: 4 partial and 6 hallucinated.&lt;br&gt;
That comparison is memory vs. no memory. It is not Hindsight vs. anything else.&lt;br&gt;
&lt;strong&gt;What I can't claim&lt;/strong&gt;&lt;br&gt;
•&lt;strong&gt;No vector-search comparison&lt;/strong&gt;. I never built one. Any statement that this architecture is better than plain retrieval would be a guess.&lt;br&gt;
•&lt;strong&gt;No reflect/recall ablation&lt;/strong&gt;. I don't know which call is doing the work.&lt;br&gt;
•&lt;strong&gt;n = 10, graded by me&lt;/strong&gt; against the dataset's true_category label, at the category level.&lt;br&gt;
•&lt;strong&gt;Held-out isn't unrelated&lt;/strong&gt;. Outages cluster into recurring failure classes, so a held-out incident can resemble retained ones, which also flatters any retrieval method.&lt;br&gt;
•&lt;strong&gt;I don't know what's happening inside the memory layer in detail&lt;/strong&gt;. I can report the counts and behavior I observed. For how facts, observations, and links are produced, read the docs and source instead of trusting my summary.&lt;br&gt;
•&lt;strong&gt;Retrieval noise.&lt;/strong&gt; A mixed bank brings unrelated suggestions with it.&lt;br&gt;
If you want to settle the vector-search question, the experiment is straightforward: same held-out set, same model and prompt, one arm with plain vector retrieval and one with Hindsight, plus reflect-only and recall-only arms. I'd like to see that result myself.&lt;br&gt;
&lt;strong&gt;Takeaways&lt;/strong&gt;&lt;br&gt;
1.&lt;strong&gt;Separate synthesis from evidence&lt;/strong&gt;. Give the model a reasoned answer and the raw memories behind it.&lt;br&gt;
2.&lt;strong&gt;Structure the memory for the questions you'll ask&lt;/strong&gt;. I wanted cross-incident patterns, so I wanted more than nearest-paragraph retrieval.&lt;br&gt;
3.&lt;strong&gt;Keep the integration thin.&lt;/strong&gt; Three wrapper functions were enough to swap memory in and out, which made the baseline comparison easy.&lt;br&gt;
4.&lt;strong&gt;Don't claim comparisons you didn't run.&lt;/strong&gt; I tested memory vs. no memory, not architecture vs. architecture.&lt;br&gt;
5.&lt;strong&gt;Write down the experiment you'd run next&lt;/strong&gt;. It turns a limitation into a plan.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
