Remembering Isn't Learning: Building Memory That Changes a Procurement Recommendation
An agent remembering something is not the same as an agent learning from it.
I kept returning to that distinction while building VendorPulse, an AI procurement decision-support application. It is easy to bolt a memory store onto an LLM app, save a few notes, and call the result "an agent with memory." If the stored notes never change what the system recommends, nothing has been learned. You have a database with a nicer interface.
So I designed VendorPulse around one principle: memory should change future behavior. Everything below follows from it.
Why stateless evaluation is not enough
A stateless evaluator sees a purchase request and a list of candidate vendors. It can reason about the request in general terms, but it has no idea that one vendor was 22 days late last July, or that another has been consistently boring and reliable. Each evaluation starts from zero.
Procurement is a domain where history matters. Vendor behavior is a pattern over time: late deliveries, rejection rates, partial recoveries, and inconsistent quality. A single snapshot doesn't capture it.
There is also a middle option that looks like a fix but isn't: retrieval alone. If the system fetches old records and shows them next to the recommendation, a human might benefit, but the agent hasn't changed. Learning requires a loop:
Memory is retrieved.
Memory enters the reasoning.
The reasoning affects a recommendation.
The real outcome is fed back into memory.
The loop continues.
Break any link and you're back to remembering without learning.
VendorPulse architecture
VendorPulse combines five pieces:
SQLite for structured procurement history and vendor statistics
Hindsight for persistent memory of vendor experiences
Groq as the LLM evaluation layer
FastAPI for the backend workflow
React/TypeScript for the frontend
The workflow runs in nine steps:
Create a procurement request.
Fetch structured vendor history.
Recall relevant Hindsight memories.
Evaluate vendors using baseline information plus recalled memory.
Generate a risk assessment and recommendation.
A human selects and approves a vendor.
Record the actual outcome.
Retain the outcome in Hindsight.
Use that memory in future evaluations.
I split structured data and memory deliberately. SQLite answers precise questions: what orders exist, what were the delays, what are the aggregate stats. Hindsight holds the experiential layer, the narrative of what happened with a vendor and what it implied. They complement each other, and the evaluator gets both.
The demo scenario
The running example: a request from Apex Manufacturing for 10,000 kg of structural steel, delivery target 19 October 2026, budget ₹10,00,000, priority High.
The history:
Vendor
July
August
October
SteelCore
+22 days late, 12% rejected
+18 days late, 9% rejected
on time, 3% rejected
MetalWorks
on time, 2% rejected
+2 days late, 3% rejected
on time, 2% rejected
PrimeSteel
+5 days late, 5% rejected
on time, 4% rejected
—
Baseline evaluation
The baseline evaluator gets the request and the structured statistics, with no recalled memory. From the table you can compute averages. SteelCore averages about 13 days late across three orders and an 8% rejection rate, MetalWorks under a day late and about 2.3%, PrimeSteel 2.5 days and 4.5%.
That is useful, and a decent baseline will rank MetalWorks well. But aggregates flatten the story. SteelCore's most recent order was on time with a low rejection rate. An average treats that as one data point among three. A baseline has no way to represent what kind of failure this was or how much weight the recovery deserves.
The baseline is also the control. Having it lets me see what memory actually adds instead of assuming it helps.
Hindsight recall
Before evaluation, the backend recalls memories relevant to the request. A representative sketch (not an exact excerpt from the repository):
python
async def recall_vendor_context(request, vendors):
query = (
f"Vendor delivery and quality experience for {request.material}, "
f"priority {request.priority}, vendors: {', '.join(vendors)}"
)
memories = await hindsight.recall(bank_id="vendorpulse", query=query)
return [m.text for m in memories]
I picked Hindsight because it's built as a memory layer for agents, not as generic document retrieval. I wanted retain and recall to be first-class operations I could drop into a procurement loop, without building my own memory infrastructure around a vector store. It fit the architecture cleanly: the backend retains at one point in the workflow and recalls at another, and the memory bank persists between sessions. If you want background, the Hindsight GitHub, the Hindsight documentation, and What is Agent Memory cover the model well.
Memory-informed reasoning
Recall by itself doesn't change anything. The step that matters is putting memory into the evaluation prompt in a way the model has to reason with:
python
def build_evaluation_prompt(request, vendor_stats, memories):
return f"""
Evaluate vendors for this purchase request.
REQUEST: {request.summary()}
STRUCTURED VENDOR STATS: {vendor_stats}
RECALLED EXPERIENCES: {memories}
For each vendor, give a risk assessment. Explicitly state where
recalled experiences change your assessment relative to the
statistics alone, especially given the delivery deadline and priority.
Return JSON: vendor, risk_level, reasoning, recommendation.
"""
The instruction to say where memory changes the assessment is deliberate. It makes the influence visible and checkable, and it stops the model from treating recalled context as decoration.
In the demo, the kind of reasoning this enables is the difference between "SteelCore has a mediocre average" and something closer to: SteelCore was badly late twice and has one good recent order, which is encouraging but thin evidence, and this is a high-priority order with a fixed date. MetalWorks has been consistently on time with low rejections. That is a judgment about pattern and stakes, not just arithmetic.
Recommendation and human decision
The evaluator produces a risk assessment and recommendation per vendor. It doesn't place an order.
The procurement manager remains the decision-maker. They see the recommendation and the reasoning, including which memories shaped it, and then select and approve a vendor. They may know things the system doesn't: a relationship, a pricing conversation, a capacity constraint. VendorPulse is decision support, not autonomous procurement. That boundary is intentional, and it also makes the loop more honest, because the recorded outcome reflects a real decision by a person.
Outcome retention
After delivery, the actual outcome is recorded: delivery date versus target, rejection rate, and any notes. The record is written to SQLite and retained in Hindsight as a natural-language experience:
python
async def retain_outcome(order, outcome):
text = (
f"Order {order.id}: {order.vendor} supplied {order.quantity} kg of "
f"{order.material}. Delivered {outcome.days_late:+d} days vs target; "
f"{outcome.rejection_pct}% rejected. Priority was {order.priority}."
)
await hindsight.retain(bank_id="vendorpulse", content=text)
This is the step that closes the loop. Without it, the system would recall a frozen past. With it, every completed order becomes context for the next one.
The next evaluation
Suppose the manager chose a vendor for the Apex order and the outcome came back. Now that experience is in memory. The next request for structural steel, perhaps from a different customer, recalls it alongside the earlier history.
If SteelCore delivers on time again, the system now has two consecutive good orders after two bad ones, and the evaluator can reasonably soften its concern. If it slips, the earlier pattern is reinforced. Either way, the same evaluation code produces different reasoning because the memory it recalls has changed. That is the behavior I was aiming for: not a smarter model, but a system whose context improves with experience.
Lessons I learned
Retrieval is the easy half. Getting memories back is straightforward. Making them alter a recommendation requires deliberate prompt and workflow design.
Keep a baseline. Without a memory-free evaluation to compare against, you can't tell what memory contributes.
Structured data and memory do different jobs. Statistics answer "how many"; memory carries "what happened and why it mattered." Merging them into one store would have weakened both.
Make influence visible. Asking the model to state where memory changed its view made the system easier to inspect and trust appropriately.
Keep the human in the loop. Recorded outcomes are only meaningful when they reflect real decisions, and a decision-support tool should support decisions.
Conclusion
Storing experiences is the beginning of agent memory, not the end. What matters is whether a recalled experience enters the reasoning, shapes a recommendation, and gets replaced by a richer experience once the real outcome is known.
VendorPulse is my attempt to build that full loop in a small, concrete domain: recall, reason, recommend, decide, record, repeat. The design goal is simple to state and demanding to implement: an agent that has remembered something should behave differently because of it.
Top comments (0)