DEV Community

Manasvi Pinnamaneni
Manasvi Pinnamaneni

Posted on

I Let Hindsight Teach Our Invoice Agent From Feedback

The first version of my invoice agent could make a reasonable decision. The harder problem was making the next decision with everything we had already learned.

I was building an Accounts Payable system that takes an invoice, looks at vendor history, and returns an APPROVE, REVIEW, or HOLD recommendation. The missing piece was continuity: when a human reviewer corrected the agent, that correction needed to become useful context for a future invoice.

That is where Hindsight changed the design.

What the system does

The application has a React and Vite frontend, a Node.js and Express backend, Groq for LLM-based invoice analysis, Hindsight for agent memory, and MySQL for persistent invoice history.

The request path is straightforward:

I use MySQL alongside that flow, but for a different reason. Hindsight provides contextual memory for the agent. MySQL provides structured application history: invoice amounts, decisions, confidence, reasons, recommendations, and timestamps.

The distinction matters. I don't want an agent-memory system to become my transactional database, and I don't want a relational table to pretend it is contextual memory.

The problem I actually had to solve

The interesting problem wasn't generating an invoice recommendation.

It was handling the second invoice.

Suppose a vendor normally sends invoices between ₹60,000 and ₹80,000, with shipping between ₹3,000 and ₹5,000. An invoice for ₹70,000 plus ₹4,000 shipping is ordinary.

Now suppose a different invoice is outside those ranges.

A stateless request can see the current numbers, but it doesn't necessarily know that a previous exception was verified against a purchase order, or that a human reviewer previously approved a similar case.

I could have kept adding rules to the prompt. That would have moved the problem around without solving it.

Instead, I treated previous decisions as part of the agent's working context.

Current invoice
+
Relevant experience
+
Human corrections
=
Current decision

That is the core reason I integrated Hindsight agent memory on GitHub rather than building another collection of ad-hoc history queries.

I retrieve two kinds of memory

The first implementation decision was not to retrieve everything about the vendor.

I use two separate recall queries. The first asks for general vendor context:

The second query is deliberately narrower. It asks Hindsight specifically for previous human decisions, purchase-order verification, exceptions, and reviewer instructions.

This separation was useful because not all historical information has the same value.

This separation was useful because not all historical information has the same value.

A previous invoice analysis is useful. A previous human decision about a similar exception is often more useful. A verified purchase order can be more useful still.

The Hindsight documentation describes memory as something an agent can retain and retrieve over time. In this application, that capability maps naturally onto vendor behavior and human review history.

More memory created a new problem

Once retrieval worked, I ran into the next problem: too much context.

Two recall queries can return overlapping memories. Some memories are long. Some are less relevant than a human correction. Sending everything to the LLM makes the reasoning context harder to control.

So I added a small memory preparation stage. I combine the results, remove duplicates, prioritize human decisions and purchase-order verification, and keep only a small amount of text:

This is not a sophisticated ranking system. That is intentional. I wanted a small, inspectable retrieval policy before introducing more machinery.

Hindsight remains the source of contextual memory while the application decides how much of that memory should enter the reasoning step.

Groq sees the invoice and its history

After memory preparation, the backend passes the current invoice, vendor profile, and compact memory set to Groq. The response contains the decision, confidence, reason, and recommendation. The application then records the analysis back into Hindsight so later requests have another piece of history available.

This is the part of agent memory explained by Vectorize that I found most useful conceptually: memory is not just storage. It changes what the agent can take into account on a later interaction.

The important memory is often the human correction

The most interesting endpoint in the backend is the feedback path.

After an invoice is analyzed, a human reviewer can accept or request further review. That decision is written to Hindsight with the reasoning supplied by the reviewer:

async function saveFeedback(feedback) {
    await remember(

        Human feedback for ${feedback.vendor_name}.

        Invoice amount:
        ₹${feedback.total_amount}

        Agent decision:
        ${feedback.agent_decision}

        Human decision:
        ${feedback.human_decision}

        Human feedback:
        ${feedback.feedback}
        ,
        {
            type: "human_feedback",
            vendor: feedback.vendor_name,
            importance: "high"
        }
    );
}
Enter fullscreen mode Exit fullscreen mode

The important detail is what gets stored.

I don't save only:

APPROVE

I save the relationship between the agent's decision, the human decision, and the reason.

That gives a later retrieval query something much more useful than a bare status value.

A concrete invoice flow

Consider ABC Industrial Supplies.

Its configured profile contains:

Typical invoices: ₹60K–₹80K
Typical shipping: ₹3K–₹5K
Payment terms: Net 30
Approval threshold: ₹75K

A normal invoice might be:

Invoice amount: ₹70,000
Shipping: ₹4,000
Total: ₹74,000

The agent can retrieve the vendor's historical pattern, compare the current invoice with that context, and return an approval recommendation.

Now consider an invoice that exceeds the normal threshold.

The first response can be REVIEW.

A human reviewer then checks the purchase order and supporting documents and approves it.

That human decision is remembered.

When another high-value invoice from the same vendor arrives, the feedback retrieval query specifically looks for previous human decisions and purchase-order verification. The agent can therefore distinguish between:

"This is outside the normal range"

and:

"We have previously seen this type of exception
and verified it."

That is a much more useful form of learning than simply increasing the amount of historical data available to the model.

Why I kept MySQL

Hindsight made the agent stateful, but I still needed conventional persistence.

The invoice history table stores structured fields such as invoice number, vendor, amounts, AI decision, confidence, reason, recommendation, human feedback, and timestamps.

The frontend can load this persistent history from the backend instead of relying only on browser state.

This gives me two different ways to answer two different questions.

MySQL answers:

What decisions did the system record?

Hindsight answers:

What previous experience is relevant to this decision?

Those are related questions, but they aren't the same question.

What I learned

1. Human feedback is better memory than generic history

A human correction contains both a decision and a reason. Treating feedback as first-class memory made Hindsight substantially more useful than simply storing previous invoices.

2. Retrieval quality matters as much as memory storage

I had to remove duplicates, prioritize human feedback, and constrain the amount of memory passed to the LLM. More context is not automatically better context.

3. Agent memory and application persistence should stay separate

MySQL gives me predictable structured history and auditability. Hindsight gives the agent contextual recall. Keeping those responsibilities separate makes the architecture easier to reason about.

4. Memory should affect behavior

A UI label saying "Hindsight active" doesn't prove much. The useful test is behavioral: a later invoice should be analyzed with relevant previous decisions available to the reasoning step.

5. A small retrieval policy is a good starting point

I started with two targeted recall queries, duplicate removal, prioritization, a four-memory limit, and a 500-character text limit. That made the system inspectable before introducing more complex ranking.

The architecture I ended up with

The part I care most about is the loop from human feedback back into Hindsight.

It changes the system from an invoice classifier into an application that can accumulate experience.

I don't think memory removes the need for human review. In Accounts Payable, the opposite is more useful: human review becomes another source of context.

The agent does not need to remember everything.

It needs to remember the things that can change the next decision.
The first version of my invoice agent could make a reasonable decision. The harder problem was making the next decision with everything we had already learned.

I was building an Accounts Payable system that takes an invoice, looks at vendor history, and returns an APPROVE, REVIEW, or HOLD recommendation. The missing piece was continuity: when a human reviewer corrected the agent, that correction needed to become useful context for a future invoice.

That is where Hindsight changed the design.

What the system does

The application has a React and Vite frontend, a Node.js and Express backend, Groq for LLM-based invoice analysis, Hindsight for agent memory, and MySQL for persistent invoice history.

The request path is deliberately straightforward:

React + Vite
      |
      v
Node.js + Express
      |
      +------> Hindsight
      |          |
      |          +--> vendor history
      |          +--> previous decisions
      |          +--> human feedback
      |
      +------> Groq
      |          |
      |          +--> decision
      |          +--> confidence
      |          +--> reason
      |          +--> recommendation
      |
      v
Human reviewer
      |
      v
Hindsight
Enter fullscreen mode Exit fullscreen mode

I use MySQL alongside that flow, but for a different reason. Hindsight provides contextual memory for the agent. MySQL provides structured application history: invoice amounts, decisions, confidence, reasons, recommendations, and timestamps.

The distinction matters. I don't want an agent-memory system to become my transactional database, and I don't want a relational table to pretend it is contextual memory.

The frontend exposes the workflow as an invoice analysis screen. A user selects a vendor, enters the invoice and shipping amounts, and submits the invoice. The backend validates the request, finds the vendor profile, recalls relevant memories, sends the current context to Groq, stores the resulting analysis in Hindsight, and persists the decision in MySQL.

The problem I actually had to solve

The interesting problem wasn't generating an invoice recommendation.

It was handling the second invoice.

Suppose a vendor normally sends invoices between ₹60,000 and ₹80,000, with shipping between ₹3,000 and ₹5,000. An invoice for ₹70,000 plus ₹4,000 shipping is ordinary.

Now suppose a different invoice is outside those ranges.

A stateless request can see the current numbers, but it doesn't necessarily know that a previous exception was verified against a purchase order, or that a human reviewer previously approved a similar case.

I could have kept adding rules to the prompt. That would have moved the problem around without solving it.

Instead, I treated previous decisions as part of the agent's working context.

The idea became:

Current invoice
     +
Relevant experience
     +
Human corrections
     =
Current decision
Enter fullscreen mode Exit fullscreen mode

That is the core reason I integrated Hindsight agent memory on GitHub rather than building another collection of ad-hoc history queries.

The first implementation decision was not to retrieve "everything about the vendor."

I use two separate recall queries.

The first asks for general vendor context:

const vendorMemoryQuery = `
Analyze invoice from ${invoice.vendor_name}.

Find important historical information about:
vendor invoice ranges,
shipping patterns,
previous discrepancies,
previous invoice resolutions,
and previous approval decisions.

Focus only on information relevant to this vendor.
`;

const vendorMemories = await recall(
    vendorMemoryQuery
);
Enter fullscreen mode Exit fullscreen mode

The second query is deliberately narrower. It asks Hindsight for previous human decisions and corrections:

const feedbackMemoryQuery = `
Find previous human feedback and human decisions
involving ${invoice.vendor_name}.

Look specifically for:
human approval decisions,
human review decisions,
purchase order verification,
supporting document verification,
previous invoice exceptions,
previous resolutions,
and instructions given by human reviewers.

Return memories that can help decide how to handle
the current invoice.
`;

const feedbackMemories = await recall(
    feedbackMemoryQuery
);
Enter fullscreen mode Exit fullscreen mode

This separation was useful because not all historical information has the same value.

A previous invoice analysis is useful.

A previous human decision about a similar exception is often more useful.

A verified purchase order can be more useful still.

I wanted the retrieval step to make that distinction explicit instead of leaving the LLM with an undifferentiated historical dump.

The Hindsight documentation describes memory as something an agent can retain and retrieve over time. In this application, that capability maps naturally onto vendor behavior and human review history.

More memory created a new problem

Once retrieval worked, I ran into the next problem: too much context.

Two recall queries can return overlapping memories. Some memories are long. Some are less relevant than a human correction. Sending everything to the LLM is an easy way to make the context harder to reason about.

So I added a small memory preparation stage.

First I combine the two result sets and remove exact duplicates. Then I prioritize memories containing human decisions, purchase-order verification, or related feedback. Finally, I keep only a small amount of text:

This is not a sophisticated ranking system. That is intentional.

I wanted a small, inspectable retrieval policy before introducing more machinery.

The important part is that Hindsight remains the source of contextual memory while the application decides how much of that memory should enter the reasoning step.

Groq sees the invoice and its history

After memory preparation, the backend passes the current invoice, vendor profile, and compact memory set to the LLM:

const decision = await analyzeWithLLM(
    invoice,
    vendor,
    compactMemories
);
Enter fullscreen mode Exit fullscreen mode

The response contains the fields the UI needs:

decision
confidence
reason
recommendation
Enter fullscreen mode Exit fullscreen mode

The application then records the analysis back into Hindsight:

await remember(
    `
    Invoice analysis:

    Vendor: ${invoice.vendor_name}

    Invoice amount: ₹${invoice.amount}

    Shipping: ₹${invoice.shipping}

    Total amount: ₹${invoice.total_amount}

    Agent decision:
    ${decision.decision}

    Confidence:
    ${decision.confidence}

    Reason:
    ${decision.reason}

    Recommendation:
    ${decision.recommendation}
    `,
    {
        type: "invoice_analysis",
        vendor: invoice.vendor_name
    }
);
Enter fullscreen mode Exit fullscreen mode

That creates a feedback loop around the reasoning step.

The agent recalls previous experience before making a decision, then records the new analysis so that later requests have another piece of history available.

This is the part of agent memory explained by Vectorize that I found most useful conceptually: memory is not just storage. It changes what the agent can take into account on a later interaction.

The important memory is often the human correction

The most interesting endpoint in the backend is the feedback path.

After an invoice is analyzed, a human reviewer can accept or request further review. That decision is written to Hindsight with the reasoning supplied by the reviewer:


The important detail is what gets stored.

I don't save only:

APPROVE
Enter fullscreen mode Exit fullscreen mode

I save the relationship between the agent's decision, the human decision, and the reason.

That gives a later retrieval query something much more useful than a bare status value.

A concrete invoice flow

Consider ABC Industrial Supplies.

Its configured profile contains:

Typical invoices:    ₹60K–₹80K
Typical shipping:    ₹3K–₹5K
Payment terms:       Net 30
Approval threshold:  ₹75K
Enter fullscreen mode Exit fullscreen mode

A normal invoice might be:

Invoice amount: ₹70,000
Shipping:       ₹4,000
Total:          ₹74,000
Enter fullscreen mode Exit fullscreen mode

The agent can retrieve the vendor's historical pattern, compare the current invoice with that context, and return an approval recommendation.

Now consider an invoice that exceeds the normal threshold.

The first response can be REVIEW.

A human reviewer then checks the purchase order and supporting documents and approves it.

That human decision is remembered.

When another high-value invoice from the same vendor arrives, the feedback retrieval query specifically looks for previous human decisions and purchase-order verification. The agent can therefore distinguish between "this is outside the normal range" and "we have previously seen this kind of exception and verified it."

That is a much more useful form of learning than simply increasing the amount of historical data available to the model.

Why I kept MySQL

Hindsight made the agent stateful, but I still needed conventional persistence.

The invoice history table contains structured fields such as:

invoice_number
vendor_name
invoice_amount
shipping
total_amount
ai_decision
confidence
ai_reason
recommendation
human_decision
feedback
created_at
updated_at
Enter fullscreen mode Exit fullscreen mode

The backend stores the AI result through a dedicated service:

const databaseId = await saveInvoiceDecision({
    invoice_number:
        invoice.invoice_number || null,

    vendor_name:
        invoice.vendor_name,

    amount:
        invoice.amount,

    shipping:
        invoice.shipping,

    total_amount:
        invoice.total_amount,

    decision:
        decision.decision,

    confidence:
        decision.confidence,

    reason:
        decision.reason,

    recommendation:
        decision.recommendation
});
Enter fullscreen mode Exit fullscreen mode

The frontend can then load persistent invoice history from the backend instead of relying only on browser state.

This separation gives me two different ways to answer two different questions.

MySQL answers:

What decisions did the system record?

Hindsight answers:

What previous experience is relevant to this decision?

Those are related questions, but they aren't the same question.

What changed in the application

The visible UI is intentionally simple.

The dashboard exposes invoice analysis, vendor intelligence, the current decision, retrieved memories, and human feedback. The user doesn't have to interact with a separate chatbot or manually construct a memory query.

The interesting state change happens underneath the interface.

Before memory:

Invoice A → decision
Invoice B → decision
Invoice C → decision
Enter fullscreen mode Exit fullscreen mode

Each request is effectively isolated.

With Hindsight:

Invoice A
   ↓
Decision
   ↓
Remember

Invoice B
   ↓
Recall A
   ↓
Decision
   ↓
Remember

Invoice C
   ↓
Recall A + B + human feedback
   ↓
Decision
Enter fullscreen mode Exit fullscreen mode

That is the behavior I was trying to build.

The goal wasn't to make the agent remember every interaction. It was to make previous interactions available when they were relevant.

What I learned

1. Human feedback is better memory than generic history

A human correction contains a decision and a reason.

That is exactly the kind of context that can help with a future exception. Treating feedback as first-class memory made the Hindsight integration substantially more useful than simply storing previous invoices.

2. Retrieval quality matters as much as memory storage

Once an agent has memory, the next problem is deciding what to retrieve.

I had to remove duplicates, prioritize human feedback, and constrain the amount of memory passed to the LLM.

More context is not automatically better context.

3. Agent memory and application persistence should stay separate

MySQL gives me predictable structured history and auditability.

Hindsight gives the agent contextual recall.

Trying to make one system perform both jobs would make the architecture harder to reason about.

4. Memory should affect behavior, not just appear in a dashboard

A UI label saying "Hindsight active" doesn't prove much.

The useful test is behavioral: a later invoice should be analyzed with relevant previous decisions available to the reasoning step.

That is why the human-feedback loop is more important to me than the memory indicator in the UI.

5. A small retrieval policy is a good starting point

I didn't start with a complex ranking pipeline.

I started with two targeted recall queries, duplicate removal, prioritization, a four-memory limit, and a 500-character text limit.

That made the system inspectable.

I can make retrieval more sophisticated later without first having to untangle an opaque memory pipeline.

The architecture I ended up with

The final design is deliberately modest:

                    ┌───────────────┐
                    │ React + Vite  │
                    └───────┬───────┘
                            │
                            ▼
                  ┌──────────────────┐
                  │ Node + Express   │
                  └───────┬──────────┘
                          │
             ┌────────────┼────────────┐
             │            │            │
             ▼            ▼            ▼
        Hindsight       Groq         MySQL
        Memory          LLM          History
             │            │            │
             └──────┬─────┘            │
                    ▼                  │
              AI decision              │
                    │                  │
                    ▼                  │
             Human feedback            │
                    │                  │
                    └──────► Hindsight │
                                       │
                              Persistent records
Enter fullscreen mode Exit fullscreen mode

The part I care most about is the loop from human feedback back into Hindsight.

It changes the system from an invoice classifier into an application that can accumulate experience.

I don't think memory removes the need for human review. In Accounts Payable, the opposite is more useful: human review becomes another source of context.

That is the design I would carry forward as the system grows.

The agent does not need to remember everything.

It needs to remember the things that can change the next decision.

Top comments (0)