DEV Community

Eranki Giridhar Goud
Eranki Giridhar Goud

Posted on

I Built an Agent With Persistent Memory. It Immediately Recommended the Wrong Supplier.

I built SupplyMind over a few focused sessions: a small agent that watches for supply chain disruptions — a cyclone closing a port, a supplier failing a quality check — and recommends what to do next based on what actually happened last time. Not a chatbot with a system prompt full of company trivia. An agent backed by Hindsight, an open-source memory layer that retains, recalls, and reflects on real incidents the way a person with institutional memory would.

The first time I ran it end-to-end, it recommended a supplier based on exactly one matching memory — a copper wire deal that had nothing to do with the cyclone I'd just described. The system was working exactly as built. My assumptions about what "remembering" would look like were wrong. That gap was the most useful part of building this.

What it does
SupplyMind takes a plain-English disruption report — "Cyclone hits Visakhapatnam port, our steel coil shipment is stuck" — and turns it into a cited, evidence-backed recommendation in three steps:

Normalize. An LLM call (via Groq) turns free text into structured fields: disruption type, location, materials, cause.
Recall. That structured signal becomes a Hindsight query against a bank of past incidents. Hindsight runs semantic, keyword, and temporal retrieval in parallel and returns the memories it thinks are relevant.
Score and recommend. A deterministic Python scorer — no LLM involved — ranks candidate suppliers by success rate, response speed, and cost from the recalled evidence, and the system writes a recommendation that cites the specific memory IDs behind it.
A manager accepts, modifies, or rejects the recommendation. Whatever they decide, and whatever actually happens afterward, gets written back into Hindsight as a new memory. The next disruption benefits from it.

The part that broke: memory doesn't store what you wrote
I seeded the bank with twenty synthetic disruption records, each with a stable ID tag like [D07], written in a specific sentence structure so I could parse it back out later with a regex. My first version of the evidence parser looked for that literal bracketed ID and specific phrases like cost: 18 lakh.

It found nothing. Every recall came back empty of citations.

The reason is worth understanding if you're building on agent memory rather than a plain vector store: Hindsight doesn't hand back your original text. It extracts facts from what you retain and reconstructs a natural-language summary at recall time. My [D07] tag was gone. "Cost: 18 lakh INR" had become "cost of 18 lakh INR" in one memory and "18 lakh in costs" in another. A regex tuned to my exact phrasing was matching nothing, because there was no "exact phrasing" to match against anymore — the memory had been paraphrased.

That's a feature, not a bug. It's what makes recall work across "typhoon" and "cyclone" and "severe weather" without three separate keyword rules. But it means you can't treat retained text as a database row you'll read back verbatim.

The fix was to stop relying on anything I'd embedded in the text and instead match on something durable and unlikely to be rephrased: the date. Every seeded memory carries its original event date somewhere in the paraphrase, so I built a small lookup from date back to my source CSV row, and pulled supplier names, response times, and outcomes with looser, more forgiving regex patterns instead of exact-phrase matches:

def parse_evidence(memory_texts: list[str]) -> list[dict]:
records = []
for text in memory_texts:
rec = {"raw": text}
m = re.search(r"(\d{4}-\d{2}-\d{2})", text)
rec["id"] = DATE_TO_ID.get(m.group(1)) if m else None

    rec["alternate_supplier"] = "Unknown"
    for name in SUPPLIERS:
        if name in text:
            rec["alternate_supplier"] = name
            break

    m = re.search(r"cost(?:\s+of)?\s+([\d.]+)\s*lakh", text, re.IGNORECASE)
    rec["cost_lakh"] = float(m.group(1)) if m else 10.0
    records.append(rec)
return records
Enter fullscreen mode Exit fullscreen mode

Once that landed, citations populated correctly and the scorer's numbers stopped defaulting to placeholder values. Recommendations started reading like this:

Activate Sahyadri Forgings. Expected delay ~8.8h response, ~10.9 lakh INR cost.
Confidence: Medium (based on 13 case(s)). Cited: D02, D07, D15, D08, D12, D10.

Every number in that sentence traces back to a real memory ID, not a guess.

The other lesson: memory has rate limits too
Retaining twenty records the naive way — one call after another, as fast as Python could fire them — hit Groq's free-tier throughput limit within seconds. Then, after switching to on_demand service tier to fix that, I hit a second, stricter limit: the 120-billion-parameter model's daily quota, with a ProviderRateLimitResetError and a retry timestamp four minutes out.

The fix was two changes, not one: move to the smaller gpt-oss-20b model, which Hindsight's own docs recommend for Groq's free tier, and add exponential backoff with jitter between retains instead of firing them back to back.
for attempt in range(6):

try:

    retain(text, context=f"disruption incident: {r['disruption_type']}", timestamp=...)

    break

except Exception as e:

    wait = 30 * (attempt + 1)

    time.sleep(wait)
Enter fullscreen mode Exit fullscreen mode

time.sleep(12)

It's not elegant, but seeding twenty records went from repeatedly crashing to reliably finishing in under ten minutes. If you're prototyping against a free-tier LLM behind your memory layer, budget for this. It's not a Hindsight problem, it's an underlying-provider problem that any agent doing bulk writes will hit.

What the grounded pipeline actually buys you
The scorer never sees raw LLM output for its numbers — it computes success rate, average response time, and average cost directly from parsed evidence. The recommendation-writing step is only allowed to cite memory IDs that were actually recalled; anything else gets rejected. That split matters more than it sounds like on paper. It means "Activate Sahyadri Forgings" isn't a plausible-sounding guess — it's a sentence built entirely out of arithmetic on real records, with an audit trail back to which thirteen incidents produced that number.

When I fed the system a signal with genuinely thin history — one unrelated copper-wire memory instead of a matching port closure — it surfaced that single weak match with a "Low" confidence label rather than confidently making something up. That's the behavior you want from something making cost and timeline claims: visible uncertainty when the evidence is thin, not synthetic confidence.

Takeaways
Retained text is not stored text. Design your parsing around durable facts (dates, named entities) instead of exact phrasing, because a memory system that paraphrases for better recall will paraphrase away your parsing hooks too.
Score deterministically, write with an LLM, and validate the seam between them. Keep the arithmetic out of the model's hands and the citations checked against what was actually retrieved.
Free-tier LLM quotas are a real constraint on bulk memory writes, not just on live inference. Budget retries and backoff from the start.
Confidence should be able to say "I don't know much here." A system that hedges on thin evidence is more trustworthy than one that always sounds sure.
The code is on GitHub if you want to see the full pipeline, including the FastAPI backend and the scorer: github.com/vectorize-io/hindsight for the memory layer itself, and Hindsight's own docs at hindsight.vectorize.io if you're setting up retain/recall/reflect for the first time.

Top comments (0)