<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sahasra</title>
    <description>The latest articles on DEV Community by Sahasra (@sahasra).</description>
    <link>https://dev.to/sahasra</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149793%2Ff24d1722-a352-4897-97b5-d4da9331f44e.png</url>
      <title>DEV Community: Sahasra</title>
      <link>https://dev.to/sahasra</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sahasra"/>
    <language>en</language>
    <item>
      <title>I made Hindsight Memory Survive a New Interaction</title>
      <dc:creator>Sahasra</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:13:09 +0000</pubDate>
      <link>https://dev.to/sahasra/i-made-hindsight-memory-survive-a-new-interaction-3p6c</link>
      <guid>https://dev.to/sahasra/i-made-hindsight-memory-survive-a-new-interaction-3p6c</guid>
      <description>&lt;p&gt;I Let the Current Issue Shape Hindsight Recall&lt;br&gt;
Persistent memory is useful only when it is relevant to the question in front of the agent. A customer may have years of support history, but a new request about billing does not need every note about a router, password reset, and delivery delay. I made the current issue part of the Hindsight query so retrieval has a task to answer, while a separate customer tag limits whose history can be considered.&lt;br&gt;
That separation is the design decision I keep coming back to. The current issue shapes relevance. The customer tag shapes scope. When those are mixed into one natural-language sentence, it is harder to see which part of the system is responsible for which guarantee.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/..%2Farticle-assets%2Fsupport-memory-architecture.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/..%2Farticle-assets%2Fsupport-memory-architecture.svg" alt="Support Memory Agent architecture from customer issue through Hindsight recall and Groq response generation" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
The API route uses the issue to query Hindsight, carries the selected memory into Groq, and stores the response for the next interaction.&lt;br&gt;
Search for a task, not a transcript&lt;br&gt;
The route receives a customer name and issue from the Next.js page. It derives the customer's tag and uses the issue to ask Hindsight for previous support issues, resolutions, and preferences related to that problem.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;taggedMemory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;bankId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Recall previous support issues, resolutions, and preferences related to: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;customerTag&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="na"&gt;tagsMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;any_strict&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I prefer this to asking for a broad customer summary on every request. A broad summary may be useful for an account overview, but support responses are usually about an immediate task. Querying with the issue gives retrieval a concrete anchor: prior attempts, similar incidents, known preferences, and resolutions related to what the customer just said.&lt;br&gt;
The strict tag filter handles a different question: whose information is in scope? Without it, a semantically similar issue from another customer could rank well. That might make the search look productive while quietly violating customer separation. The API uses the tag as an eligibility condition and the issue as the relevance query.&lt;br&gt;
I keep those dimensions separate when debugging too. If Hindsight returns the wrong customer's record, changing the wording of the issue query is the wrong first move; I need to inspect the tag and matching mode. If the tag is correct but the result is irrelevant, then the retrieval query, result count, and memory representation are the places to investigate. Separating scope from relevance gives me a smaller search space when behavior is surprising.&lt;br&gt;
This is one of the practical benefits of Hindsight's memory system: the application can ask a natural-language question while also applying structured recall options. The Hindsight documentation describes the API surface; the agent memory overview explains why memory is meant to retrieve useful context for a task, not just append more transcript text.&lt;br&gt;
How much memory should reach Groq?&lt;br&gt;
The current route takes the first recall result. If there is no result, it uses an explicit no-memory message. Then it sends that text and the current issue to Groq.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;previousMemory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="nx"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`No previous memory found for &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This choice keeps the prompt small and the user-visible “Previous Memory” section easy to understand. There is one displayed memory and one generated answer. It also creates an obvious future tuning point: should this support task use one result, a short set of ranked results, or an explicit Hindsight-generated synthesis?&lt;br&gt;
I would not decide that by stuffing every result into the prompt. More context can mean more irrelevant details, more tokens, and more opportunities for the model to combine separate incidents incorrectly. The right amount depends on the support task and the quality of the retrieved result. For a single recurring issue, one prior resolution may be enough. For a complicated case spanning several sessions, a short set of linked events may be necessary.&lt;br&gt;
The interface reflects the same tradeoff by showing one “Previous Memory” item. That is easy for an operator to scan and easy to connect to the answer, but it should not be mistaken for a complete customer timeline. Hindsight can contain more than the first result. If the support process needs multiple incidents, I would show a small, ranked set with dates and source context rather than silently passing a long list to Groq and hiding it from the person reviewing the answer.&lt;br&gt;
The generation instruction is deliberately plain. It asks the model to use memory when available, avoid inventing history, avoid mentioning another customer's information, handle the current issue when memory is empty, and stay under 120 words. The current issue and selected memory are separate fields in the user prompt, which helps the model understand which is present context and which is historical context.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`
You are a professional customer support agent.

Use previous customer memory when available.

Rules:
- Be friendly and concise.
- Never invent customer history.
- Do not mention information belonging to another customer.
- If there is no previous memory, simply handle the current issue.
- Keep the response under 120 words.
`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`
Customer: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

Current issue:
&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

Previous memory:
&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;previousMemory&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

Write a personalized support response.
`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example: same customer, different issue&lt;br&gt;
Imagine a customer previously contacted support because a router restarted overnight. The support agent walked through a power-adapter replacement, and that resolved the issue. Months later, the customer reports an intermittent connection after moving the router to another room.&lt;br&gt;
The query includes the new symptom and the Hindsight tag selects the same customer's record. A relevant prior resolution can then shape the Groq response: the agent can acknowledge the previous power issue and ask whether the adapter or outlet changed, instead of repeating the first generic troubleshooting script. If there is no relevant memory, the prompt explicitly tells Groq to address the current issue without manufacturing continuity.&lt;br&gt;
That example is not a benchmark claim. It describes the intended data path: issue-conditioned recall, selected history in the prompt, and a response grounded in both. The code does not establish that every returned memory is relevant. I still need to judge retrieval quality separately from answer quality.&lt;br&gt;
The questions I use to evaluate recall&lt;br&gt;
There are at least three independent checks for this flow:&lt;br&gt;
Scope: did Hindsight return records tagged for the right customer?&lt;br&gt;
Relevance: did the current issue retrieve history that actually helps with this problem?&lt;br&gt;
Use: did Groq apply that history correctly without inventing a resolution or preference?&lt;br&gt;
A single “memory found” boolean collapses these checks. The UI already shows the recalled text separately, which gives the operator some evidence for scope and relevance. For a production workflow, I would retain those distinctions in logs and evaluations as well: a correct customer with irrelevant memory is different from a memory-service failure, and both are different from a response that ignored useful history.&lt;br&gt;
The lesson for me is simple: retrieval is a product decision as much as a model decision. Hindsight needs a precise question and a precise customer scope. Groq needs only the context that survived those choices. If I make that boundary visible in the code and interface, I can improve the retrieval strategy without asking the model to compensate for a vague memory query.&lt;br&gt;
There is another reason to keep the query and prompt stages legible: the customer issue is user-supplied text. It should be treated as data, not instructions to the model or the memory service. The route places it in the issue slot and retains it as part of the interaction. In a deployed support workflow, I would test adversarial issue text explicitly and delimit historical memory from the current request so neither can quietly rewrite the support rules. Roles and prompt wording help, but they do not replace input handling and evaluation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>hindsight</category>
    </item>
  </channel>
</rss>
