<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manasvi Pinnamaneni</title>
    <description>The latest articles on DEV Community by Manasvi Pinnamaneni (@manasvi_pinnamaneni_30da1).</description>
    <link>https://dev.to/manasvi_pinnamaneni_30da1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149416%2F46c72116-da03-40a8-8306-20cf195c34e2.jpg</url>
      <title>DEV Community: Manasvi Pinnamaneni</title>
      <link>https://dev.to/manasvi_pinnamaneni_30da1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manasvi_pinnamaneni_30da1"/>
    <language>en</language>
    <item>
      <title>I Let Hindsight Teach Our Invoice Agent From Feedback</title>
      <dc:creator>Manasvi Pinnamaneni</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:53:36 +0000</pubDate>
      <link>https://dev.to/manasvi_pinnamaneni_30da1/i-let-hindsight-teach-our-invoice-agent-from-feedback-4moi</link>
      <guid>https://dev.to/manasvi_pinnamaneni_30da1/i-let-hindsight-teach-our-invoice-agent-from-feedback-4moi</guid>
      <description>&lt;p&gt;The first version of my invoice agent could make a reasonable decision. The harder problem was making the next decision with everything we had already learned.&lt;/p&gt;

&lt;p&gt;I was building an Accounts Payable system that takes an invoice, looks at vendor history, and returns an APPROVE, REVIEW, or HOLD recommendation. The missing piece was continuity: when a human reviewer corrected the agent, that correction needed to become useful context for a future invoice.&lt;/p&gt;

&lt;p&gt;That is where Hindsight changed the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;The application has a React and Vite frontend, a Node.js and Express backend, Groq for LLM-based invoice analysis, Hindsight for agent memory, and MySQL for persistent invoice history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The request path is straightforward:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6nyeuvziswprw146a8e4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6nyeuvziswprw146a8e4.png" alt=" " width="800" height="529"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjljfzfeo9bgxe5w63ex0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjljfzfeo9bgxe5w63ex0.png" alt=" " width="800" height="342"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I use &lt;strong&gt;MySQL&lt;/strong&gt; alongside that flow, but for a different reason. Hindsight provides contextual memory for the agent. MySQL provides structured application history: &lt;u&gt;invoice amounts, decisions, confidence, reasons, recommendations, and timestamps&lt;/u&gt;.&lt;/p&gt;

&lt;p&gt;The distinction matters. I don't want an agent-memory system to become my transactional database, and I don't want a relational table to pretend it is contextual memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem I actually had to solve
&lt;/h2&gt;

&lt;p&gt;The interesting problem wasn't generating an invoice recommendation.&lt;/p&gt;

&lt;p&gt;It was handling the second invoice.&lt;/p&gt;

&lt;p&gt;Suppose a vendor normally sends invoices between ₹60,000 and ₹80,000, with shipping between ₹3,000 and ₹5,000. An invoice for ₹70,000 plus ₹4,000 shipping is ordinary.&lt;/p&gt;

&lt;p&gt;Now suppose a different invoice is outside those ranges.&lt;/p&gt;

&lt;p&gt;A stateless request can see the current numbers, but it doesn't necessarily know that a previous exception was verified against a purchase order, or that a human reviewer previously approved a similar case.&lt;/p&gt;

&lt;p&gt;I could have kept adding rules to the prompt. That would have moved the problem around without solving it.&lt;/p&gt;

&lt;p&gt;Instead, I treated previous decisions as part of the agent's working context.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;Current invoice&lt;br&gt;
     +&lt;br&gt;
Relevant experience&lt;br&gt;
     +&lt;br&gt;
Human corrections&lt;br&gt;
     =&lt;br&gt;
Current decision&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is the core reason I integrated Hindsight agent memory on GitHub rather than building another collection of ad-hoc history queries.&lt;/p&gt;
&lt;h2&gt;
  
  
  I retrieve two kinds of memory
&lt;/h2&gt;

&lt;p&gt;The first implementation decision was not to retrieve everything about the vendor.&lt;/p&gt;

&lt;p&gt;I use two separate recall queries. The first asks for general vendor context:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s7cou7gu3tedi4s7jbb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s7cou7gu3tedi4s7jbb.png" alt=" " width="787" height="388"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The second query is deliberately narrower. It asks Hindsight specifically for previous human decisions, purchase-order verification, exceptions, and reviewer instructions.&lt;/p&gt;

&lt;p&gt;This separation was useful because not all historical information has the same value.&lt;/p&gt;

&lt;p&gt;This separation was useful because not all historical information has the same value.&lt;/p&gt;

&lt;p&gt;A previous invoice analysis is useful. A previous human decision about a similar exception is often more useful. A verified purchase order can be more useful still.&lt;/p&gt;

&lt;p&gt;The Hindsight documentation describes memory as something an agent can retain and retrieve over time. In this application, that capability maps naturally onto vendor behavior and human review history.&lt;/p&gt;
&lt;h2&gt;
  
  
  More memory created a new problem
&lt;/h2&gt;

&lt;p&gt;Once retrieval worked, I ran into the next problem: too much context.&lt;/p&gt;

&lt;p&gt;Two recall queries can return overlapping memories. Some memories are long. Some are less relevant than a human correction. Sending everything to the LLM makes the reasoning context harder to control.&lt;/p&gt;

&lt;p&gt;So I added a small memory preparation stage. I combine the results, remove duplicates, prioritize human decisions and purchase-order verification, and keep only a small amount of text:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdzkpn3xozmw2o5zdl2r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdzkpn3xozmw2o5zdl2r.png" alt=" " width="646" height="482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is not a sophisticated ranking system. That is intentional. I wanted a small, inspectable retrieval policy before introducing more machinery.&lt;/p&gt;

&lt;p&gt;Hindsight remains the source of contextual memory while the application decides how much of that memory should enter the reasoning step.&lt;/p&gt;
&lt;h2&gt;
  
  
  Groq sees the invoice and its history
&lt;/h2&gt;

&lt;p&gt;After memory preparation, the backend passes the current invoice, vendor profile, and compact memory set to Groq. The response contains the decision, confidence, reason, and recommendation. The application then records the analysis back into Hindsight so later requests have another piece of history available.&lt;/p&gt;

&lt;p&gt;This is the part of agent memory explained by Vectorize that I found most useful conceptually: memory is not just storage. It changes what the agent can take into account on a later interaction.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fonnr4npobabmz4jj6q5i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fonnr4npobabmz4jj6q5i.png" alt=" " width="800" height="381"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhq0jh1sv4mrizx32nlz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhq0jh1sv4mrizx32nlz.png" alt=" " width="800" height="271"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The important memory is often the human correction
&lt;/h2&gt;

&lt;p&gt;The most interesting endpoint in the backend is the feedback path.&lt;/p&gt;

&lt;p&gt;After an invoice is analyzed, a human reviewer can accept or request further review. That decision is written to Hindsight with the reasoning supplied by the reviewer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;async function saveFeedback(feedback) {
    await remember(

        Human feedback for ${feedback.vendor_name}.

        Invoice amount:
        ₹${feedback.total_amount}

        Agent decision:
        ${feedback.agent_decision}

        Human decision:
        ${feedback.human_decision}

        Human feedback:
        ${feedback.feedback}
        ,
        {
            type: "human_feedback",
            vendor: feedback.vendor_name,
            importance: "high"
        }
    );
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail is what gets stored.&lt;/p&gt;

&lt;p&gt;I don't save only:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;APPROVE&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I save the relationship between the agent's decision, the human decision, and the reason.&lt;/p&gt;

&lt;p&gt;That gives a later retrieval query something much more useful than a bare status value.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete invoice flow
&lt;/h2&gt;

&lt;p&gt;Consider ABC Industrial Supplies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Its configured profile contains:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Typical invoices:&lt;/strong&gt;    ₹60K–₹80K&lt;br&gt;
&lt;strong&gt;Typical shipping:&lt;/strong&gt;    ₹3K–₹5K&lt;br&gt;
&lt;strong&gt;Payment terms:&lt;/strong&gt;       Net 30&lt;br&gt;
&lt;strong&gt;Approval threshold:&lt;/strong&gt;  ₹75K&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A normal invoice might be:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invoice amount:&lt;/strong&gt; ₹70,000&lt;br&gt;
&lt;strong&gt;Shipping:&lt;/strong&gt;       ₹4,000&lt;br&gt;
&lt;strong&gt;Total:&lt;/strong&gt;          ₹74,000&lt;/p&gt;

&lt;p&gt;The agent can retrieve the vendor's historical pattern, compare the current invoice with that context, and return an approval recommendation.&lt;/p&gt;

&lt;p&gt;Now consider an invoice that exceeds the normal threshold.&lt;/p&gt;

&lt;p&gt;The first response can be &lt;strong&gt;REVIEW&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A human reviewer then checks the purchase order and supporting documents and approves it.&lt;/p&gt;

&lt;p&gt;That human decision is remembered.&lt;/p&gt;

&lt;p&gt;When another high-value invoice from the same vendor arrives, the feedback retrieval query specifically looks for previous human decisions and purchase-order verification. The agent can therefore distinguish between:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"This is outside the normal range"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"We have previously seen this type of exception&lt;br&gt;
and verified it."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is a much more useful form of learning than simply increasing the amount of historical data available to the model.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why I kept MySQL
&lt;/h2&gt;

&lt;p&gt;Hindsight made the agent stateful, but I still needed conventional persistence.&lt;/p&gt;

&lt;p&gt;The invoice history table stores structured fields such as invoice number, vendor, amounts, AI decision, confidence, reason, recommendation, human feedback, and timestamps.&lt;/p&gt;

&lt;p&gt;The frontend can load this persistent history from the backend instead of relying only on browser state.&lt;/p&gt;

&lt;p&gt;This gives me two different ways to answer two different questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MySQL answers:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What decisions did the system record?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight answers:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What previous experience is relevant to this decision?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those are related questions, but they aren't the same question.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdybt70mge6gopt85z9x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdybt70mge6gopt85z9x.png" alt=" " width="800" height="553"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Human feedback is better memory than generic history&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A human correction contains both a decision and a reason. Treating feedback as first-class memory made Hindsight substantially more useful than simply storing previous invoices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Retrieval quality matters as much as memory storage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I had to remove duplicates, prioritize human feedback, and constrain the amount of memory passed to the LLM. More context is not automatically better context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Agent memory and application persistence should stay separate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MySQL gives me predictable structured history and auditability. Hindsight gives the agent contextual recall. Keeping those responsibilities separate makes the architecture easier to reason about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Memory should affect behavior&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A UI label saying "Hindsight active" doesn't prove much. The useful test is behavioral: a later invoice should be analyzed with relevant previous decisions available to the reasoning step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. A small retrieval policy is a good starting point&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I started with two targeted recall queries, duplicate removal, prioritization, a four-memory limit, and a 500-character text limit. That made the system inspectable before introducing more complex ranking.&lt;/p&gt;
&lt;h2&gt;
  
  
  The architecture I ended up with
&lt;/h2&gt;

&lt;p&gt;The part I care most about is the loop from human feedback back into Hindsight.&lt;/p&gt;

&lt;p&gt;It changes the system from an invoice classifier into an application that can accumulate experience.&lt;/p&gt;

&lt;p&gt;I don't think memory removes the need for human review. In Accounts Payable, the opposite is more useful: human review becomes another source of context.&lt;/p&gt;

&lt;p&gt;The agent does not need to remember everything.&lt;/p&gt;

&lt;p&gt;It needs to remember the things that can change the next decision.&lt;br&gt;
The first version of my invoice agent could make a reasonable decision. The harder problem was making the next decision with everything we had already learned.&lt;/p&gt;

&lt;p&gt;I was building an Accounts Payable system that takes an invoice, looks at vendor history, and returns an &lt;code&gt;APPROVE&lt;/code&gt;, &lt;code&gt;REVIEW&lt;/code&gt;, or &lt;code&gt;HOLD&lt;/code&gt; recommendation. The missing piece was continuity: when a human reviewer corrected the agent, that correction needed to become useful context for a future invoice.&lt;/p&gt;

&lt;p&gt;That is where Hindsight changed the design.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;The application has a React and Vite frontend, a Node.js and Express backend, Groq for LLM-based invoice analysis, Hindsight for agent memory, and MySQL for persistent invoice history.&lt;/p&gt;

&lt;p&gt;The request path is deliberately straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;React + Vite
      |
      v
Node.js + Express
      |
      +------&amp;gt; Hindsight
      |          |
      |          +--&amp;gt; vendor history
      |          +--&amp;gt; previous decisions
      |          +--&amp;gt; human feedback
      |
      +------&amp;gt; Groq
      |          |
      |          +--&amp;gt; decision
      |          +--&amp;gt; confidence
      |          +--&amp;gt; reason
      |          +--&amp;gt; recommendation
      |
      v
Human reviewer
      |
      v
Hindsight
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I use MySQL alongside that flow, but for a different reason. Hindsight provides contextual memory for the agent. MySQL provides structured application history: invoice amounts, decisions, confidence, reasons, recommendations, and timestamps.&lt;/p&gt;

&lt;p&gt;The distinction matters. I don't want an agent-memory system to become my transactional database, and I don't want a relational table to pretend it is contextual memory.&lt;/p&gt;

&lt;p&gt;The frontend exposes the workflow as an invoice analysis screen. A user selects a vendor, enters the invoice and shipping amounts, and submits the invoice. The backend validates the request, finds the vendor profile, recalls relevant memories, sends the current context to Groq, stores the resulting analysis in Hindsight, and persists the decision in MySQL.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem I actually had to solve
&lt;/h2&gt;

&lt;p&gt;The interesting problem wasn't generating an invoice recommendation.&lt;/p&gt;

&lt;p&gt;It was handling the second invoice.&lt;/p&gt;

&lt;p&gt;Suppose a vendor normally sends invoices between ₹60,000 and ₹80,000, with shipping between ₹3,000 and ₹5,000. An invoice for ₹70,000 plus ₹4,000 shipping is ordinary.&lt;/p&gt;

&lt;p&gt;Now suppose a different invoice is outside those ranges.&lt;/p&gt;

&lt;p&gt;A stateless request can see the current numbers, but it doesn't necessarily know that a previous exception was verified against a purchase order, or that a human reviewer previously approved a similar case.&lt;/p&gt;

&lt;p&gt;I could have kept adding rules to the prompt. That would have moved the problem around without solving it.&lt;/p&gt;

&lt;p&gt;Instead, I treated previous decisions as part of the agent's working context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The idea became:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current invoice
     +
Relevant experience
     +
Human corrections
     =
Current decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the core reason I integrated &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight agent memory on GitHub&lt;/a&gt; rather than building another collection of ad-hoc history queries.&lt;/p&gt;

&lt;p&gt;The first implementation decision was not to retrieve "everything about the vendor."&lt;/p&gt;

&lt;p&gt;I use two separate recall queries.&lt;/p&gt;

&lt;p&gt;The first asks for general vendor context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;vendorMemoryQuery&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`
Analyze invoice from &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vendor_name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.

Find important historical information about:
vendor invoice ranges,
shipping patterns,
previous discrepancies,
previous invoice resolutions,
and previous approval decisions.

Focus only on information relevant to this vendor.
`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;vendorMemories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;vendorMemoryQuery&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second query is deliberately narrower. It asks Hindsight for previous human decisions and corrections:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;feedbackMemoryQuery&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`
Find previous human feedback and human decisions
involving &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vendor_name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.

Look specifically for:
human approval decisions,
human review decisions,
purchase order verification,
supporting document verification,
previous invoice exceptions,
previous resolutions,
and instructions given by human reviewers.

Return memories that can help decide how to handle
the current invoice.
`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;feedbackMemories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;feedbackMemoryQuery&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation was useful because not all historical information has the same value.&lt;/p&gt;

&lt;p&gt;A previous invoice analysis is useful.&lt;/p&gt;

&lt;p&gt;A previous human decision about a similar exception is often more useful.&lt;/p&gt;

&lt;p&gt;A verified purchase order can be more useful still.&lt;/p&gt;

&lt;p&gt;I wanted the retrieval step to make that distinction explicit instead of leaving the LLM with an undifferentiated historical dump.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; describes memory as something an agent can retain and retrieve over time. In this application, that capability maps naturally onto vendor behavior and human review history.&lt;/p&gt;

&lt;h2&gt;
  
  
  More memory created a new problem
&lt;/h2&gt;

&lt;p&gt;Once retrieval worked, I ran into the next problem: too much context.&lt;/p&gt;

&lt;p&gt;Two recall queries can return overlapping memories. Some memories are long. Some are less relevant than a human correction. Sending everything to the LLM is an easy way to make the context harder to reason about.&lt;/p&gt;

&lt;p&gt;So I added a small memory preparation stage.&lt;/p&gt;

&lt;p&gt;First I combine the two result sets and remove exact duplicates. Then I prioritize memories containing human decisions, purchase-order verification, or related feedback. Finally, I keep only a small amount of text:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fis4b12zvmbzntwfzupxd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fis4b12zvmbzntwfzupxd.png" alt=" " width="612" height="492"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is not a sophisticated ranking system. That is intentional.&lt;/p&gt;

&lt;p&gt;I wanted a small, inspectable retrieval policy before introducing more machinery.&lt;/p&gt;

&lt;p&gt;The important part is that Hindsight remains the source of contextual memory while the application decides how much of that memory should enter the reasoning step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Groq sees the invoice and its history
&lt;/h2&gt;

&lt;p&gt;After memory preparation, the backend passes the current invoice, vendor profile, and compact memory set to the LLM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;analyzeWithLLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;compactMemories&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response contains the fields the UI needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;decision
confidence
reason
recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application then records the analysis back into Hindsight:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;remember&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`
    Invoice analysis:

    Vendor: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vendor_name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Invoice amount: ₹&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Shipping: ₹&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shipping&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Total amount: ₹&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;total_amount&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Agent decision:
    &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Confidence:
    &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Reason:
    &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

    Recommendation:
    &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;recommendation&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;
    `&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;invoice_analysis&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;vendor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vendor_name&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates a feedback loop around the reasoning step.&lt;/p&gt;

&lt;p&gt;The agent recalls previous experience before making a decision, then records the new analysis so that later requests have another piece of history available.&lt;/p&gt;

&lt;p&gt;This is the part of &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory explained by Vectorize&lt;/a&gt; that I found most useful conceptually: memory is not just storage. It changes what the agent can take into account on a later interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The important memory is often the human correction
&lt;/h2&gt;

&lt;p&gt;The most interesting endpoint in the backend is the feedback path.&lt;/p&gt;

&lt;p&gt;After an invoice is analyzed, a human reviewer can accept or request further review. That decision is written to Hindsight with the reasoning supplied by the reviewer:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F392a1owpsenl9b4cu5az.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F392a1owpsenl9b4cu5az.png" alt=" " width="670" height="568"&gt;&lt;/a&gt;&lt;br&gt;
The important detail is what gets stored.&lt;/p&gt;

&lt;p&gt;I don't save only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;APPROVE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I save the relationship between the agent's decision, the human decision, and the reason.&lt;/p&gt;

&lt;p&gt;That gives a later retrieval query something much more useful than a bare status value.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete invoice flow
&lt;/h2&gt;

&lt;p&gt;Consider ABC Industrial Supplies.&lt;/p&gt;

&lt;p&gt;Its configured profile contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Typical invoices:    ₹60K–₹80K
Typical shipping:    ₹3K–₹5K
Payment terms:       Net 30
Approval threshold:  ₹75K
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A normal invoice might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice amount: ₹70,000
Shipping:       ₹4,000
Total:          ₹74,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent can retrieve the vendor's historical pattern, compare the current invoice with that context, and return an approval recommendation.&lt;/p&gt;

&lt;p&gt;Now consider an invoice that exceeds the normal threshold.&lt;/p&gt;

&lt;p&gt;The first response can be &lt;code&gt;REVIEW&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A human reviewer then checks the purchase order and supporting documents and approves it.&lt;/p&gt;

&lt;p&gt;That human decision is remembered.&lt;/p&gt;

&lt;p&gt;When another high-value invoice from the same vendor arrives, the feedback retrieval query specifically looks for previous human decisions and purchase-order verification. The agent can therefore distinguish between "this is outside the normal range" and "we have previously seen this kind of exception and verified it."&lt;/p&gt;

&lt;p&gt;That is a much more useful form of learning than simply increasing the amount of historical data available to the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I kept MySQL
&lt;/h2&gt;

&lt;p&gt;Hindsight made the agent stateful, but I still needed conventional persistence.&lt;/p&gt;

&lt;p&gt;The invoice history table contains structured fields such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;invoice_number
vendor_name
invoice_amount
shipping
total_amount
ai_decision
confidence
ai_reason
recommendation
human_decision
feedback
created_at
updated_at
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The backend stores the AI result through a dedicated service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;databaseId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;saveInvoiceDecision&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;invoice_number&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;invoice_number&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;vendor_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vendor_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;shipping&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shipping&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;total_amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;total_amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="na"&gt;recommendation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;recommendation&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The frontend can then load persistent invoice history from the backend instead of relying only on browser state.&lt;/p&gt;

&lt;p&gt;This separation gives me two different ways to answer two different questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MySQL answers:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What decisions did the system record?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Hindsight answers:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What previous experience is relevant to this decision?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are related questions, but they aren't the same question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in the application
&lt;/h2&gt;

&lt;p&gt;The visible UI is intentionally simple.&lt;/p&gt;

&lt;p&gt;The dashboard exposes invoice analysis, vendor intelligence, the current decision, retrieved memories, and human feedback. The user doesn't have to interact with a separate chatbot or manually construct a memory query.&lt;/p&gt;

&lt;p&gt;The interesting state change happens underneath the interface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before memory:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice A → decision
Invoice B → decision
Invoice C → decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each request is effectively isolated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With Hindsight:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice A
   ↓
Decision
   ↓
Remember

Invoice B
   ↓
Recall A
   ↓
Decision
   ↓
Remember

Invoice C
   ↓
Recall A + B + human feedback
   ↓
Decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the behavior I was trying to build.&lt;/p&gt;

&lt;p&gt;The goal wasn't to make the agent remember every interaction. It was to make previous interactions available when they were relevant.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Human feedback is better memory than generic history
&lt;/h3&gt;

&lt;p&gt;A human correction contains a decision and a reason.&lt;/p&gt;

&lt;p&gt;That is exactly the kind of context that can help with a future exception. Treating feedback as first-class memory made the Hindsight integration substantially more useful than simply storing previous invoices.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Retrieval quality matters as much as memory storage
&lt;/h3&gt;

&lt;p&gt;Once an agent has memory, the next problem is deciding what to retrieve.&lt;/p&gt;

&lt;p&gt;I had to remove duplicates, prioritize human feedback, and constrain the amount of memory passed to the LLM.&lt;/p&gt;

&lt;p&gt;More context is not automatically better context.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Agent memory and application persistence should stay separate
&lt;/h3&gt;

&lt;p&gt;MySQL gives me predictable structured history and auditability.&lt;/p&gt;

&lt;p&gt;Hindsight gives the agent contextual recall.&lt;/p&gt;

&lt;p&gt;Trying to make one system perform both jobs would make the architecture harder to reason about.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Memory should affect behavior, not just appear in a dashboard
&lt;/h3&gt;

&lt;p&gt;A UI label saying "Hindsight active" doesn't prove much.&lt;/p&gt;

&lt;p&gt;The useful test is behavioral: a later invoice should be analyzed with relevant previous decisions available to the reasoning step.&lt;/p&gt;

&lt;p&gt;That is why the human-feedback loop is more important to me than the memory indicator in the UI.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. A small retrieval policy is a good starting point
&lt;/h3&gt;

&lt;p&gt;I didn't start with a complex ranking pipeline.&lt;/p&gt;

&lt;p&gt;I started with two targeted recall queries, duplicate removal, prioritization, a four-memory limit, and a 500-character text limit.&lt;/p&gt;

&lt;p&gt;That made the system inspectable.&lt;/p&gt;

&lt;p&gt;I can make retrieval more sophisticated later without first having to untangle an opaque memory pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture I ended up with
&lt;/h2&gt;

&lt;p&gt;The final design is deliberately modest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌───────────────┐
                    │ React + Vite  │
                    └───────┬───────┘
                            │
                            ▼
                  ┌──────────────────┐
                  │ Node + Express   │
                  └───────┬──────────┘
                          │
             ┌────────────┼────────────┐
             │            │            │
             ▼            ▼            ▼
        Hindsight       Groq         MySQL
        Memory          LLM          History
             │            │            │
             └──────┬─────┘            │
                    ▼                  │
              AI decision              │
                    │                  │
                    ▼                  │
             Human feedback            │
                    │                  │
                    └──────► Hindsight │
                                       │
                              Persistent records
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The part I care most about is the loop from human feedback back into Hindsight.&lt;/p&gt;

&lt;p&gt;It changes the system from an invoice classifier into an application that can accumulate experience.&lt;/p&gt;

&lt;p&gt;I don't think memory removes the need for human review. In Accounts Payable, the opposite is more useful: human review becomes another source of context.&lt;/p&gt;

&lt;p&gt;That is the design I would carry forward as the system grows.&lt;/p&gt;

&lt;p&gt;The agent does not need to remember everything.&lt;/p&gt;

&lt;p&gt;It needs to remember the things that can change the next decision.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>database</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
