An agent can recall exactly what happened and still make the same mistake. I built VendorPulse because that gap bothered me.
Retrieval is easy. You store a document, you fetch it, you show it in a sidebar. The user sees the old information and does the mental work. The agent itself hasn't learned anything. Its next decision is unchanged.
Learning, for an agent, has a stricter definition. Memory must enter reasoning. Reasoning must change a recommendation or an action. The outcome of that action must be fed back into memory. The loop must continue. If any link is missing, you have a search engine, not a memory-augmented agent.
VendorPulse is my attempt to implement that full loop for procurement.
Why Stateless Evaluation Fails
Procurement looks stateless on paper. You have a request for material, quantity, delivery date, budget, priority. You have a vendor list with price, lead time, certification. You can write a scoring function and rank vendors.
The problem is that the scoring function never saw July.
In manufacturing, vendor performance is seasonal, contextual, and excuse-driven. A vendor who is perfect in October can be 22 days late in July monsoon season with 12% rejection. An ERP shows their average rating as 4.2. It doesn't remember that last July they said "truck breakdown" and this July they said "truck breakdown" again.
I wanted an evaluation that gets better because it lived through those months. Not because I added more rules, but because it accumulated experiences.
VendorPulse Architecture
I kept the architecture intentionally narrow.
Frontend: React / TypeScript. A single workflow for creating a procurement request, viewing baseline vs memory-informed evaluations, and recording an outcome.
Backend: FastAPI. Handles the procurement workflow, talks to SQLite for structured history, talks to the LLM layer, and talks to Hindsight.
Structured layer: SQLite. PO history, vendor master, delivery records, rejection rates. Deterministic facts.
Memory layer: Hindsight. Persistent, semantic memory for episodic experiences. This is where the unstructured, contextual memory lives.
LLM layer: Groq for evaluation. I chose Groq for low-latency reasoning, because the evaluation step needs to compare baseline stats with several recalled memories in one prompt.
Central principle: Memory should change future behavior. So Hindsight sits in the middle of the architecture diagram, not as a logging sink. Every evaluation must pass through a recall step, and every outcome must pass through a retain step.
If you are new to this layer, start here:
Hindsight GitHub
Hindsight documentation
What is Agent Memory
I selected Hindsight because I needed a memory system that supports semantic recall on messy, human-like experiences, not just key-value lookup. Procurement excuses don't have a schema. "Truck breakdown during monsoon" and "vehicle failure in heavy rain" should match. Hindsight handles that.
Baseline Evaluation: What Happens Without Memory
Take the demo I used to validate the design:
Apex Manufacturing, Structural Steel, 10,000 kg, delivery by 19 October 2026, budget ₹10,00,000, high priority.
Three vendors in SQLite:
SteelCore: Cheapest. Avg 4.1 rating.
MetalWorks: Mid-price. Avg 4.3 rating.
PrimeSteel: Premium. Avg 4.5 rating.
A baseline evaluation without memory does what most procurement tools do. It fetches structured stats, calculates a weighted score from price, average delay, and average rejection rate, and recommends the cheapest compliant vendor.
baseline evaluation - no memory context
def evaluate_baseline(request, vendors):
scores = []
for v in vendors:
avg_delay = get_avg_delay(v.id)
avg_reject = get_avg_reject(v.id)
price_score = 1.0 - (v.price / request.budget)
score = 0.5 * price_score + 0.3 * (1 - avg_delay/30) + 0.2 * (1 - avg_reject)
scores.append((v, score))
return sorted(scores, key=lambda x: x[1], reverse=True)
In this scenario, SteelCore wins on price. The baseline says "Meets budget, certified, recommended." It is not wrong. It is just amnesic. It averaged July, August, and October together and lost the seasonal pattern.
This is the core failure of stateless evaluation. It compresses history into an average.
Hindsight Recall: Retrieving Relevant Experiences
Before evaluation, VendorPulse queries Hindsight with the current request context.
The query is not just "SteelCore". It is "Structural Steel, October delivery, high priority, 10,000 kg, monsoon tail-end, Apex Manufacturing". Hindsight returns semantically relevant episodes, not just exact matches.
For this request, recall returns:
SteelCore: July +22 days late, 12% rejected. Reason: logistics delay during monsoon.
SteelCore: August +18 days late, 9% rejected. Reason: same transporter issue.
SteelCore: October on time, 3% rejected.
MetalWorks: July on time, 2% rejected.
MetalWorks: August +2 days late, 3% rejected.
PrimeSteel: July +5 days late, 5% rejected. August on time, 4% rejected.
These are not in SQLite as neat rows. They are retained as natural language experiences with context, outcome, and decision.
recall relevant vendor experiences
memories = hindsight.recall(
query=f"{request.material} {request.delivery_month} {request.priority}",
filters={"material": "Structural Steel", "customer": "Apex Manufacturing"},
top_k=8
)
memories = list of episodic experiences with vendor, delay, rejection, excuse, season
Retrieval alone is not the point. I could have displayed these in a side panel and stopped. The evaluation would still be stateless. The manager would do the learning, not the agent.
Memory-Informed Reasoning: Where Memory Changes the Score
The difference is in the evaluation prompt. I pass both the structured stats and the recalled memories to the LLM as separate contexts, and I instruct it to treat memories as contextual overrides, not just additional facts.
Baseline context: avg delay, avg rejection, price.
Memory context: seasonal pattern, repeated excuse, recency, October performance.
The agent's reasoning changes:
Without memory: "SteelCore is cheapest, average delay 13 days, acceptable for October."
With memory: "SteelCore is cheapest, but its two recent monsoon deliveries were 18+ days late with the same transporter excuse and high rejection. October was on time, suggesting seasonality. However, delivery target is Oct 19 - still within monsoon tail in this region. MetalWorks was on time in July and only +2 days in August for same material. Risk-adjusted, MetalWorks has lower true cost when delay penalty is factored."
That is not just showing old tickets. That is using them to change the recommendation.
memory-informed evaluation
def evaluate_with_memory(request, vendors, memories, stats):
prompt = build_prompt(
request=request,
structured_stats=stats,
episodic_memories=memories,
instruction="Use episodic memories to adjust risk. If a vendor shows repeated seasonal delays with same excuse, increase risk. If October history is clean, note seasonality. Recommend based on true cost, not just quoted price."
)
evaluation = groq_llm.generate(prompt, tools=[calculate_true_cost])
return evaluation # includes reasoning, risk_score, recommendation
Memory has now entered reasoning. The next step is that reasoning must affect action.
Recommendation and Human Decision
VendorPulse does not auto-purchase. This was a deliberate design choice.
The agent generates a risk assessment and a ranked recommendation with evidence:
Recommended: MetalWorks. Reason: Clean July/August history for Structural Steel, low rejection, higher price but lower expected delay cost. Evidence: July on time 2% reject, August +2 days 3% reject.
Risky: SteelCore. Reason: Cheapest but pattern of 18-22 day delays in monsoon months with repeat excuse. October was on time, but target date still in high-risk window. Evidence: July +22d 12%, August +18d 9%.
The procurement manager sees both baseline and memory-informed views side by side. The difference is the demonstration of learning. The manager remains the decision-maker and selects the vendor.
For the article scenario, the manager selects MetalWorks despite higher quoted price.
Outcome Retention: Closing the Loop
This is where most memory demos stop. They recall, they recommend, they finish. The loop is open. The system will make the same recommendation next month because it never learned what happened after the recommendation.
VendorPulse requires an outcome step.
After delivery, the manager records the actual outcome: actual delivery date, actual rejection rate, reason, additional cost.
This outcome is then retained in Hindsight as a new experience.
retain outcome as new episodic memory
experience = f"""
Material: {request.material} 10,000 kg for Apex Manufacturing
Vendor: {selected_vendor.name}
Requested: {request.delivery_target}
Actual: Delivered {actual_delay} days late, {actual_rejection}% rejected
Context: {request.priority} priority, {request.delivery_month} season
Outcome: {"On time" if actual_delay <= 2 else "Delayed"} - {outcome_notes}
Lesson: {lesson_learned}
"""
hindsight.retain(
text=experience,
metadata={
"vendor": selected_vendor.name,
"material": request.material,
"month": request.delivery_month,
"delay": actual_delay,
"rejection": actual_rejection
}
)
Now the memory store has grown by one. Not a log row in SQLite, but a contextual experience that will be semantically recalled next time someone orders Structural Steel in October.
How the Next Evaluation Uses That Experience
On the next request for Structural Steel in October, recall will return the new experience plus the older ones. If MetalWorks delivered on time with 2% rejection, its risk score drops further. If it was late, its risk score increases.
Over 10-20 procurements, the system develops a per-vendor, per-material, per-season memory that no average can capture. It learns that SteelCore fails in July-August but is fine in October. It learns that MetalWorks handles monsoon better. It learns which excuses repeat.
That is the behavior change I wanted to prove. Same request, different month, different recommendation, because memory changed.
Lessons I Learned Building It
Memory schema matters more than LLM prompt.
I started with free-form memories and got noisy recall. I got much better results when I enforced a consistent structure for retain: material + vendor + month + delay + rejection + excuse + lesson. Still natural language, but predictable enough for semantic search to work.You need two evaluations to prove memory works.
I built a split view from day one: baseline from SQLite alone vs memory-informed from SQLite + Hindsight. Without that, you cannot show that memory changed behavior. For engineers, show your work. The side-by-side is the demo.Human-in-the-loop is a feature, not a limitation.
Early on I was tempted to make it auto-select. I removed that. Procurement managers don't want autonomous purchasing. They want better evidence. Keeping the manager as the decision-maker made the product more credible and the memory loop more honest - the outcome is real human-verified data, not an agent's assumption.Retention must be explicit and immediate.
If you wait to retain until end of day or batch job, memories are lost. I made outcome recording part of the core workflow, not an optional step. Create request -> Evaluate -> Select -> Record outcome -> Retain. If you break that chain, learning stops.
Conclusion
An agent remembering something is not the same as an agent learning from it.
Learning requires that the memory be recalled in context, that it enters reasoning, that reasoning changes a recommendation, that a human acts on it, that the real outcome is captured, and that the outcome becomes a new memory.
VendorPulse implements that loop for a very specific domain where seasonal, vendor-specific patterns actually matter. The technology choices - SQLite for facts, Hindsight for experiences, Groq for reasoning, FastAPI and React for the loop - were all in service of keeping that loop tight and demonstrable.
The system doesn't know everything. It just doesn't forget what it already learned.


Top comments (0)