<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mani Mounika Gubbala</title>
    <description>The latest articles on DEV Community by Mani Mounika Gubbala (@manimounikagubbala).</description>
    <link>https://dev.to/manimounikagubbala</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150334%2F37fc642d-94c5-4005-983c-54899afdc770.png</url>
      <title>DEV Community: Mani Mounika Gubbala</title>
      <link>https://dev.to/manimounikagubbala</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manimounikagubbala"/>
    <language>en</language>
    <item>
      <title>Same Strategy, Two Hospitals, One Win and One Loss: Teaching an Agent Why</title>
      <dc:creator>Mani Mounika Gubbala</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:34:37 +0000</pubDate>
      <link>https://dev.to/manimounikagubbala/same-strategy-two-hospitals-one-win-and-one-loss-teaching-an-agent-why-19dd</link>
      <guid>https://dev.to/manimounikagubbala/same-strategy-two-hospitals-one-win-and-one-loss-teaching-an-agent-why-19dd</guid>
      <description>&lt;p&gt;In my test data, a security-first proposal won with Meridian Valley Health System and lost with Cedar Ridge Medical Group. (All 23 RFPs in this project are synthetic. I wrote them to build the test.)&lt;/p&gt;

&lt;p&gt;The outcomes flipped for a reason. Meridian's timeline was flexible. Cedar Ridge demanded a hard 90-day go-live, and my 7-month security-first plan couldn't meet it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why retrieval isn't enough
&lt;/h2&gt;

&lt;p&gt;Put both proposals in a vector database and search "security-first," and you get both back. That's correct, and it's also useless, because the real question is left to you: &lt;em&gt;which lesson applies to the bid I'm writing today?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I wanted an agent that holds the contradiction instead of returning it. It should form a belief about why we win or lose, and revise that belief when new evidence arrives.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;MEMORA&lt;/strong&gt;, an RFP memory agent on &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;, using Groq (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt;) for reasoning and Python around it. If agent memory is new to you, &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;this overview&lt;/a&gt; is a good primer, and the &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight docs&lt;/a&gt; cover the API I used.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;MEMORA ingests past proposals into a Hindsight memory bank. It deliberately does &lt;strong&gt;not&lt;/strong&gt; ingest the &lt;code&gt;lessons_learned&lt;/code&gt; field. If I fed it my own conclusions, it would only be retrieving them. A regression test enforces this, so that field can never reach memory.&lt;/p&gt;

&lt;p&gt;Before I write a new response, &lt;code&gt;premortem()&lt;/code&gt; recalls relevant history and returns risks with evidence tags, plus a block of known conflicts. A third bid, Bayshore Regional Health, resolved the Cedar Ridge conflict: security-first won again when delivered as a phased 60-day secure core. The warning now reads roughly: &lt;em&gt;security-first is risky under hard deadlines, unless it's phased.&lt;/em&gt; The resolution is stored next to the conflict it resolves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The before/after
&lt;/h2&gt;

&lt;p&gt;The strongest proof is a three-stage script, &lt;code&gt;learning_demo.py&lt;/code&gt;. It uses a separate bank and reveals the history one piece at a time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 0&lt;/strong&gt; has Cedar Ridge and Bayshore hidden. The pre-mortem returns only generic risks, citing unrelated proposals like RFP-002, RFP-009 and RFP-023.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 1&lt;/strong&gt; retains Cedar Ridge (RFP-011) and reruns the same question. A new risk appears immediately, citing RFP-011 and echoing the real situation: new clinics, a 90-day go-live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 2&lt;/strong&gt; retains Bayshore (RFP-019). The risk text evolves: it now ties security-first directly to the timeline risk and cites RFP-003 and RFP-011 together.&lt;/p&gt;

&lt;p&gt;The agent code is identical across stages. Only the memory changes.&lt;/p&gt;

&lt;p&gt;One reliability detail cost me real time. Hindsight banks persist indefinitely, so an RFP retained in an earlier run silently contaminates every later "before" baseline. The demo now resets itself at the start of every run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# learning_demo.py
&lt;/span&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;delete_bank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DEMO_BANK&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;pass&lt;/span&gt;  &lt;span class="c1"&gt;# the bank may not exist on the very first run
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A demo that only works once isn't a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug that taught me the most
&lt;/h2&gt;

&lt;p&gt;For a while, Stage 0 kept failing. The agent cited RFP-011 even when I had never put it in memory.&lt;/p&gt;

&lt;p&gt;I blamed Hindsight. I was wrong. My own prompt's example JSON used real IDs like &lt;code&gt;"RFP-011"&lt;/code&gt; as formatting placeholders, and the model treated them as real evidence. It wasn't remembering anything. It was copying my example.&lt;/p&gt;

&lt;p&gt;The fix was obviously fake placeholders:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# agent.py, inside PROMPT
&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RFP-XXX&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence_for&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RFP-XXX&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence_against&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RFP-YYY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resolution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RFP-ZZZ or null if unresolved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also added a unit test that fails if a real ID ever appears in the prompt's example, so this can't return silently. The lesson: &lt;strong&gt;an LLM can't tell a formatting example from real evidence, so never put real data in your examples.&lt;/strong&gt; I only found this because I built a before/after harness strict enough to expose it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The conflict wiring is plain engineering, not memory magic.&lt;/strong&gt; Some recalled memories came back without attribution, so I couldn't rely on recall alone to say which RFP a belief came from. I wrote &lt;code&gt;_known_conflicts()&lt;/code&gt; to read the win/loss conflicts and their resolutions deterministically from my data and pass them to the model. That part would work with any database. What memory contributes is the evolving risk analysis across the three stages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The demo stages the conflict block too.&lt;/strong&gt; To simulate "not yet learned," the demo filters hidden RFPs out of that deterministic block. So Stages 1 and 2 show memory changing the risk analysis, but the conflict lines themselves are controlled by my script.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The data is synthetic.&lt;/strong&gt; Twenty-three hand-written RFPs prove the mechanism, not real-world accuracy. Real win/loss reasons are rarely as clean as "the timeline was too tight."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent learns only what I tell it.&lt;/strong&gt; If I log an outcome with the wrong reason, it will confidently learn the wrong belief.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The natural next step is a closed loop: after each real bid I log the outcome, and the agent revises its beliefs. The three-stage demo is that loop compressed into one script.&lt;/p&gt;

&lt;p&gt;The code is at &lt;a href="https://github.com/manimounika-gubbala/MEMORA" rel="noopener noreferrer"&gt;github.com/manimounika-gubbala/MEMORA&lt;/a&gt;. If you build with agent memory, test it the way I ended up testing mine: hide the evidence, run, reveal it, run again. If the output doesn't change, your agent isn't learning.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljg3q5smoyiop3a36pbz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljg3q5smoyiop3a36pbz.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
