<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Victor Isaac</title>
    <description>The latest articles on DEV Community by Victor Isaac (@victorisaacchintaeng).</description>
    <link>https://dev.to/victorisaacchintaeng</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148108%2Fdf50d365-c40c-4b5c-86c5-bc7d5219a8fd.jpg</url>
      <title>DEV Community: Victor Isaac</title>
      <link>https://dev.to/victorisaacchintaeng</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/victorisaacchintaeng"/>
    <language>en</language>
    <item>
      <title>My Agent Kept Forgetting Our Company, So I Gave It Hindsight</title>
      <dc:creator>Victor Isaac</dc:creator>
      <pubDate>Mon, 28 Sep 2026 22:01:47 +0000</pubDate>
      <link>https://dev.to/victorisaacchintaeng/my-agent-kept-forgetting-our-company-so-i-gave-it-hindsight-37p3</link>
      <guid>https://dev.to/victorisaacchintaeng/my-agent-kept-forgetting-our-company-so-i-gave-it-hindsight-37p3</guid>
      <description>&lt;p&gt;In February, a lender sent my father's interior fit-out company a vendor pre-qualification questionnaire. Forty-odd questions, a folder of supporting documents, and a deadline of the next afternoon. The company was qualified. It had already built six gold-loan branches for another lender. Five months and two reminder emails later, the questionnaire was still not submitted.&lt;/p&gt;

&lt;p&gt;Nobody was lazy. The answers existed. They were just scattered across old email threads, previous submissions and one person's head. Every new questionnaire started from zero.&lt;/p&gt;

&lt;p&gt;That is the problem we built Prequal to solve: an agent that answers vendor questionnaires from a permanent memory of one company, shows the evidence behind every answer, and gets better every time someone corrects it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a normal LLM agent fails here
&lt;/h2&gt;

&lt;p&gt;Our first instinct was the obvious one: paste the company profile into a prompt and let a model fill in the questionnaire. It fails in three ways.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It forgets. Each questionnaire is a fresh session, so last month's corrections are gone.&lt;/li&gt;
&lt;li&gt;It guesses. Ask a model "Do you hold ISO 9001?" with thin context and it will happily produce a confident, plausible, wrong answer.&lt;/li&gt;
&lt;li&gt;It can't show its work. A reviewer has no idea which fact an answer came from.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For vendor paperwork, the second failure is the dangerous one. A fabricated certification on a pre-qualification form is worse than no answer at all.&lt;/p&gt;

&lt;p&gt;So the design goal became: the agent should only say what it can prove, and its memory should outlive any single session. That's where &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight agent memory&lt;/a&gt; came in.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the system hangs together
&lt;/h2&gt;

&lt;p&gt;Prequal is a small FastAPI backend, a single-page review screen, SQLite for questionnaire state, Groq for the language model, and Hindsight as the only place company knowledge lives.&lt;/p&gt;

&lt;p&gt;Every question goes through the same pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Classify.&lt;/strong&gt; Each question has a type: exact field, document request, yes/no, numeric by year, or free text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic answers first.&lt;/strong&gt; GSTIN, PAN, entity type, address: these are looked up directly. No model call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document check.&lt;/strong&gt; "Attach your ISO 9001 certificate" is checked against the company's document inventory. On file, or not held. Still no model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall.&lt;/strong&gt; Everything else pulls evidence from Hindsight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Draft with tools.&lt;/strong&gt; The model must call &lt;code&gt;submit_answer&lt;/code&gt; with the evidence IDs it used, or &lt;code&gt;flag_gap&lt;/code&gt;. Plain text is never treated as an answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails.&lt;/strong&gt; Low confidence or uncited drafts go to human review. Any positive claim about a certification the company doesn't hold is blocked.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Automation before AI. The model only sees the questions that actually need judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core story: memory is the product
&lt;/h2&gt;

&lt;p&gt;Here is the part that surprised us. Without memory, Prequal is useless. With an empty memory bank, every single question comes back as "needs review". The agent has nothing to stand on, so it says so.&lt;/p&gt;

&lt;p&gt;That's exactly the behaviour we wanted, and it's only possible because the agent treats &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; as its source of truth rather than as a nice-to-have.&lt;/p&gt;

&lt;p&gt;We give each company its own memory bank with a mission and three directives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;DIRECTIVES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no-fabrication&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Never state a certification, registration, licence or figure that is not in memory. If it is absent, say it is absent.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;When a fact has more than one value over time, report the timeline, most recent first, and name the dates.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reviewer-wins&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Prefer the reviewer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s corrected answers over the agent&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s own earlier drafts.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On day one we retain three kinds of memory: company facts, one fact per document (on file or not held), and every answer the company has given on past questionnaires. Turnover is retained once per financial year, with a timestamp:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;fy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;turnover_by_fy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;w&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s annual turnover for FY &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fy&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; was &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;label&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (audited).&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;annual turnover from audited accounts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="n"&gt;ts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;_fy_end&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fy&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;profile&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;financial&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That timestamp matters. When a questionnaire asks for "turnover for the last three years", recall returns dated facts, and the review screen renders them as a timeline, newest first. When a new year's figure is retained, it simply appears at the top. No overwriting, no stale numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The learning loop
&lt;/h2&gt;

&lt;p&gt;The part engineers ask about most is how the agent improves. It's almost boring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retain_review&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;final_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reviewer-approved answer on the &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;client&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; questionnaire &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
               &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;to the question &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;final_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; Reviewer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s note: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reviewer correction for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;client&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="n"&gt;timestamp&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;utcnow&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every approval and every edit becomes a memory, with the reviewer's reason attached. On the next questionnaire, a similar question recalls the corrected answer and the model cites it. If a reviewer says "always lead with the six-branch project", that preference is in memory for every future questionnaire, from every client.&lt;/p&gt;

&lt;p&gt;This is the difference between an agent with a longer prompt and an agent with memory. A prompt resets. Memory accumulates.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like in use
&lt;/h2&gt;

&lt;p&gt;A reviewer opens a new questionnaire and clicks run. Questions stream in with one of three statuses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Auto-filled:&lt;/strong&gt; answered from an exact field, a document on file, or cited evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Needs you:&lt;/strong&gt; thin evidence or low confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gap:&lt;/strong&gt; nothing in memory answers it, or the document isn't held.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Click any answer and the evidence panel shows exactly which memories it came from, with the cited ones marked. A live log on the right shows every retain, recall and reflect call as it happens. Before answering, one &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;reflect&lt;/a&gt; call produces a short brief: which past projects matter for this client, which documents are missing.&lt;/p&gt;

&lt;p&gt;The moment that sold us on the approach: asked "Do you hold ISO 9001 certification?", the agent answers "No. We do not hold an ISO 9001 certificate at present" and marks it as a gap for the reviewer to decide. It doesn't invent a certificate number. It tells you what's missing.&lt;/p&gt;

&lt;p&gt;On a full run against a 43-question questionnaire from an NBFC, with real Hindsight memory and Groq's gpt-oss-120b drafting, Prequal answered 38 questions, sent 1 to a human for review, and flagged 4 as gaps. All four gaps were documents the company does not hold. The company in the demo is fictional, modelled on a real firm's structure with synthetic figures.&lt;/p&gt;

&lt;p&gt;Here's a two-minute walkthrough:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/cmx4w6JXkEw" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Deterministic first, model second.&lt;/strong&gt; In our 43-question sample, 15 questions are exact fields and 13 are document checks. Sending those to a model adds cost, latency and risk for nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Force the model to cite, then check the citations.&lt;/strong&gt; Tool calls with evidence IDs, validated after the fact, turned "trust me" answers into auditable ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Timestamps are underrated.&lt;/strong&gt; Retaining facts with the date they became true made "what changed" a query instead of a data-cleaning job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Make the empty state honest.&lt;/strong&gt; An agent that says "I don't know this company yet" is more trustworthy than one that fills the silence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Function calling fails. Plan for it.&lt;/strong&gt; We validate every tool argument, retry once with the schema echoed back, fall back to a second model, and treat anything else as "needs review". A crash is never an answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Parsing Excel and PDF questionnaires directly, writing answers back into the client's own template, and document expiry alerts so a solvency certificate never lapses silently.&lt;/p&gt;

&lt;p&gt;The bigger lesson is about &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;what agent memory actually is&lt;/a&gt;. It isn't a bigger context window. It's a record that outlives the session, knows when things were true, and gets better every time a human corrects it. For a small company answering the same forty questions over and over, that's the whole product.&lt;/p&gt;

&lt;p&gt;Code: &lt;a href="https://github.com/victorisaacchinta-eng/prequal" rel="noopener noreferrer"&gt;https://github.com/victorisaacchinta-eng/prequal&lt;/a&gt;&lt;br&gt;
Demo: &lt;a href="https://youtu.be/cmx4w6JXkEw" rel="noopener noreferrer"&gt;https://youtu.be/cmx4w6JXkEw&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
