Retaining RFPs in Hindsight: What My Proposal Assistant Actually Remembers
I built ProposalMind to draft responses to Requests for Proposal (RFPs), and I used Hindsight as its persistent memory layer. Before writing about the project, I reread every file and compared the README with what the code does. The wiring is simple. The more interesting question turned out to be what the application actually stores, and that is the subject of this article. For background on the concepts, see the Hindsight documentation and Vectorize's overview of agent memory.
What ProposalMind is
The repo has two scripts. ui.py is a Streamlit app: you paste an RFP, fill in a company name, contact details and a client name, and click Generate. It saves the RFP to memory, recalls from memory, asks an LLM for a nine-section proposal, shows a keyword-based coverage score, and offers a PDF download built with ReportLab. app.py is a separate command-line script that stores a hardcoded sample proposal, reads an RFP from input(), recalls, generates a shorter four-section response and prints it. The two share no code, and the README documents only ui.py.
Generation goes through Groq, using the OpenAI Python SDK:
groq_client = OpenAI(
api_key=os.getenv("GROQ_API_KEY"),
base_url="https://api.groq.com/openai/v1",
)
The model is openai/gpt-oss-120b in both files, sent as a single user message with no system prompt and no extra parameters.
The Hindsight calls
Hindsight is used in two calls against one hardcoded bank named ProposalMind. I do not use reflect(), tags or metadata, and nothing in the repo creates or configures the bank. In ui.py, both calls run when Generate is clicked with a non-empty RFP:
hindsight.retain(
bank_id=BANK_ID,
content=f"New RFP received:\n{rfp}",
context="current RFP",
)
memories = hindsight.recall(
bank_id=BANK_ID,
query=rfp,
)
The pasted RFP is stored with a short prefix, and then the entire RFP text is used as the search query. My reading is that this avoids writing any query logic, though the code doesn't say why it was chosen. app.py takes a different route: its recall query is a fixed question about what enterprise clients care about in data analytics proposals, and the new RFP never reaches the search.
How recalled text reaches the model
Recalled memories become plain text inside the prompt, under the heading RELEVANT PAST EXPERIENCE. The two files build that text differently:
# app.py: pull out each memory's text
memories = "\n".join(
f"- {memory.text}"
for memory in memory_result.results
)
# ui.py: convert the whole response object
memory_text = str(memories)
The ui.py version converts the entire recall response with str() and pastes the result into both the prompt and the "View recalled experience" expander. The repository shows this conversion, but not what it produces. It could be readable text or an object representation. I have not seen a live Hindsight response, so that needs checking.
The rest of the ui.py prompt includes the client name, company details, the raw RFP, nine required section headings, and instructions not to use placeholders, invent contact details, or invent past results that aren't in the provided experience.
A walkthrough, and what I actually tested
Suppose I paste: "We need Power BI reporting, strong data security, and a fast implementation with clear project milestones."
-
Retain. Hindsight receives
New RFP received:followed by that text, with contextcurrent RFP. -
Recall. The same text is the query, and the result is converted with
str(). - Generation. The prompt combines the RFP, company details and recalled text, and Groq returns the proposal.
- Requirement detection and coverage. These render below the proposal. They are not passed to the LLM.
I could not run this against live services, because the environment where I did the checking had no network access. Instead I ran the real ui.py with the Hindsight, OpenAI and Streamlit libraries replaced by stubs that record their inputs, using proposals I wrote by hand. That verified the exact retain() and recall() arguments, the Groq base URL and model, the single-message prompt, and the deterministic behavior below. It says nothing about what Hindsight returns or what the model would write.
Requirement detection is substring matching on the RFP for four items: Power BI, security, fast/rapid/quick, and timeline/milestone/schedule. Coverage is a separate check against four fixed categories, each with a keyword list searched in the generated proposal:
"Fast Implementation": [
"fast implementation",
"rapid",
"quick",
"agile",
],
The categories are not derived from the RFP; an rfp_lower variable is computed and never used. In my stubbed runs, a proposal that followed the nine template headings scored 75%. Fast Implementation failed, even though the RFP asked for it, because none of its keywords appeared. Proposals without matching keywords scored 0% to 25%. The score can't reach 100% unless one of those four Fast Implementation keywords appears. The page also shows 0% and four warnings before anything has been generated.
What the memory actually contains
Here is the point I would have missed without rereading. In ui.py, the only thing sent to retain() is the incoming RFP text. Generated proposals, outcomes, feedback and user edits are never stored. The only "successful proposal" in the repository is the hardcoded Acme Retail sample in app.py. The UI never writes it, so it would reach the bank only if someone ran that script against the same bank.
The README says recalled experience can inform topics such as security requirements, implementation strategies and project milestones. In the code, what comes back depends on what was stored, and the UI stores requests. So the design is closer to retaining incoming requests than to building a history of successful proposals. I have not compared proposals with and without memory, so I make no claim about their quality.
Two details I have not verified. Retain runs just before recall, so recall may or may not include the RFP just submitted. And every Generate click re-submits the same RFP to retain(). How Hindsight handles repeated content, including whether it deduplicates, was not tested.
Smaller rough edges
-
PDF text. Non-ASCII characters are dropped. In my test, "Café" became "Caf" and Japanese text vanished. Smart quotes, dashes and bullets are converted to ASCII, and Markdown bold markers are stripped without producing bold text. Lines starting with
1.through9.become headings, including body lines like1.5 weeks of discovery. Star bullets and table rows come through as plain text. -
Dependencies.
requirements.txtlistshindsight-client,openaiandpython-dotenv, butui.pyalso importsstreamlitandreportlab, which the README lists in its stack. I treat that as a dependency gap visible in the repository. I did not test a fresh install.
What I would change next
- Retain generated proposals, and outcomes where they exist, alongside the incoming RFP.
- Use the text of each recalled result in a structured form, as
app.pydoes, instead of stringifying the whole response. - Derive the coverage checklist from the actual RFP.
- Add
streamlitandreportlabtorequirements.txt.
Still to verify against a live Hindsight
What str(memories) actually looks like, whether the current RFP shows up in its own recall, how repeated retains behave, and what recall returns after app.py has seeded the sample.
Takeaway
The Hindsight integration is two calls. What deserves attention is the input: right now, what the app sends to memory is what clients asked for, not what worked.
Top comments (0)