ARTICLE 3 - For Agent / LLM Teammate
Title: The Evaluation Layer: Where Recall Has to Become Reasoning
I built the evaluation layer for VendorPulse - the part where recalled memories actually change a recommendation. This is where most memory demos fail.
Retrieving from Hindsight is easy. You call recall() and you get relevant experiences. The hard part is making those experiences affect reasoning.
A naive approach I tried first: concatenate memories into the prompt and ask Groq to "consider them". Result: LLM summarized memories but still recommended the cheapest vendor. It treated memories as trivia.
The fix: Separate baseline reasoning from memory-informed reasoning.
I implemented two distinct evaluation paths in FastAPI:python# evaluation.py
def baseline_evaluate(request):
stats = get_vendor_stats(request.material)
# only uses avg delay, avg rejection, price
return groq.evaluate(f"Rank vendors by price and avg stats: {stats}")
def memory_informed_evaluate(request, memories, stats):
# stats + episodic memories
return groq.evaluate(
system="You are a procurement risk analyst. Memories are contextual overrides.",
context={
"structured": stats,
"episodic": memories, # from Hindsight
"rule": "If vendor shows repeat excuse + repeat delay in same season, increase risk score. Calculate true cost = quoted price + (expected delay * penalty)"
}
)Concrete example that made it click:
Request: Structural Steel, 10,000 kg, Oct 19, ₹10L, High priority
Memories recalled:
SteelCore July +22d 12% reject, August +18d 9% reject, October on time 3%MetalWorks July on time 2%, August +2d 3%Baseline says: SteelCore cheapest, recommend.
Memory-informed says:
"SteelCore: Pattern of monsoon failures with repeat transporter excuse. October history clean, but Oct 19 still in monsoon tail. Risk-adjusted true cost = ₹10L + 15 days * ₹5000 penalty = ₹10,75,000 with high variance.
MetalWorks: Stable July-August for same material, low rejection. True cost = ₹10,50,000 with low variance. Recommend MetalWorks despite higher quoted price."
That is learning. Same data, different reasoning, different recommendation.
Top comments (0)