In my test data, a security-first proposal won with Meridian Valley Health System and lost with Cedar Ridge Medical Group. (All 23 RFPs in this project are synthetic. I wrote them to build the test.)
The outcomes flipped for a reason. Meridian's timeline was flexible. Cedar Ridge demanded a hard 90-day go-live, and my 7-month security-first plan couldn't meet it.
Why retrieval isn't enough
Put both proposals in a vector database and search "security-first," and you get both back. That's correct, and it's also useless, because the real question is left to you: which lesson applies to the bid I'm writing today?
I wanted an agent that holds the contradiction instead of returning it. It should form a belief about why we win or lose, and revise that belief when new evidence arrives.
So I built MEMORA, an RFP memory agent on Hindsight, using Groq (openai/gpt-oss-120b) for reasoning and Python around it. If agent memory is new to you, this overview is a good primer, and the Hindsight docs cover the API I used.
How it works
MEMORA ingests past proposals into a Hindsight memory bank. It deliberately does not ingest the lessons_learned field. If I fed it my own conclusions, it would only be retrieving them. A regression test enforces this, so that field can never reach memory.
Before I write a new response, premortem() recalls relevant history and returns risks with evidence tags, plus a block of known conflicts. A third bid, Bayshore Regional Health, resolved the Cedar Ridge conflict: security-first won again when delivered as a phased 60-day secure core. The warning now reads roughly: security-first is risky under hard deadlines, unless it's phased. The resolution is stored next to the conflict it resolves.
The before/after
The strongest proof is a three-stage script, learning_demo.py. It uses a separate bank and reveals the history one piece at a time.
Stage 0 has Cedar Ridge and Bayshore hidden. The pre-mortem returns only generic risks, citing unrelated proposals like RFP-002, RFP-009 and RFP-023.
Stage 1 retains Cedar Ridge (RFP-011) and reruns the same question. A new risk appears immediately, citing RFP-011 and echoing the real situation: new clinics, a 90-day go-live.
Stage 2 retains Bayshore (RFP-019). The risk text evolves: it now ties security-first directly to the timeline risk and cites RFP-003 and RFP-011 together.
The agent code is identical across stages. Only the memory changes.
One reliability detail cost me real time. Hindsight banks persist indefinitely, so an RFP retained in an earlier run silently contaminates every later "before" baseline. The demo now resets itself at the start of every run:
# learning_demo.py
try:
memory.client.delete_bank(bank_id=DEMO_BANK)
except Exception:
pass # the bank may not exist on the very first run
A demo that only works once isn't a demo.
The bug that taught me the most
For a while, Stage 0 kept failing. The agent cited RFP-011 even when I had never put it in memory.
I blamed Hindsight. I was wrong. My own prompt's example JSON used real IDs like "RFP-011" as formatting placeholders, and the model treated them as real evidence. It wasn't remembering anything. It was copying my example.
The fix was obviously fake placeholders:
# agent.py, inside PROMPT
"evidence": ["RFP-XXX"]
"evidence_for": ["RFP-XXX"], "evidence_against": ["RFP-YYY"],
"resolution": "RFP-ZZZ or null if unresolved"
I also added a unit test that fails if a real ID ever appears in the prompt's example, so this can't return silently. The lesson: an LLM can't tell a formatting example from real evidence, so never put real data in your examples. I only found this because I built a before/after harness strict enough to expose it.
Honest limitations
The conflict wiring is plain engineering, not memory magic. Some recalled memories came back without attribution, so I couldn't rely on recall alone to say which RFP a belief came from. I wrote _known_conflicts() to read the win/loss conflicts and their resolutions deterministically from my data and pass them to the model. That part would work with any database. What memory contributes is the evolving risk analysis across the three stages.
The demo stages the conflict block too. To simulate "not yet learned," the demo filters hidden RFPs out of that deterministic block. So Stages 1 and 2 show memory changing the risk analysis, but the conflict lines themselves are controlled by my script.
The data is synthetic. Twenty-three hand-written RFPs prove the mechanism, not real-world accuracy. Real win/loss reasons are rarely as clean as "the timeline was too tight."
The agent learns only what I tell it. If I log an outcome with the wrong reason, it will confidently learn the wrong belief.
What's next
The natural next step is a closed loop: after each real bid I log the outcome, and the agent revises its beliefs. The three-stage demo is that loop compressed into one script.
The code is at github.com/manimounika-gubbala/MEMORA. If you build with agent memory, test it the way I ended up testing mine: hide the evidence, run, reveal it, run again. If the output doesn't change, your agent isn't learning.

Top comments (0)