<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Seeramreddi Praveen</title>
    <description>The latest articles on DEV Community by Seeramreddi Praveen (@praveeen77955).</description>
    <link>https://dev.to/praveeen77955</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150390%2Fed0ac62b-c975-4db2-b699-a74c4bbb933a.png</url>
      <title>DEV Community: Seeramreddi Praveen</title>
      <link>https://dev.to/praveeen77955</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/praveeen77955"/>
    <language>en</language>
    <item>
      <title>Memory-Grounded Prompting - How Retrieved Deal History Changes What an LLM Actually Generates</title>
      <dc:creator>Seeramreddi Praveen</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:28:16 +0000</pubDate>
      <link>https://dev.to/praveeen77955/memory-grounded-prompting-how-retrieved-deal-history-changes-what-an-llm-actually-generates-4mbj</link>
      <guid>https://dev.to/praveeen77955/memory-grounded-prompting-how-retrieved-deal-history-changes-what-an-llm-actually-generates-4mbj</guid>
      <description>&lt;p&gt;Role: AI/ML Engineer&lt;/p&gt;




&lt;p&gt;There's a misconception I kept running into when we started this project. People treat retrieval-augmented generation like the retrieval step is just a fancier search engine, and the LLM does all the real work.&lt;/p&gt;

&lt;p&gt;After building DealMind's coaching engine, I can tell you it's the other way around. The quality of what gets retrieved almost entirely determines the quality of what gets generated. The LLM is pattern completion - it completes the pattern the retrieved context sets up. Give it a vague pattern, you get vague output. Give it a specific, narrative-rich memory, and the output gets specific and narrative-rich too.&lt;/p&gt;

&lt;p&gt;Here's what I learned building this.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Prompt Structure
&lt;/h2&gt;

&lt;p&gt;Every coaching turn in DealMind fires one Groq call. The system prompt defines the agent's role and locks the output format:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
system = (&lt;br&gt;
    "You are Synapse, an AI sales coach. You watch live sales conversations "&lt;br&gt;
    "and give the sales rep real-time coaching based on patterns from "&lt;br&gt;
    "past won and lost deals.\n\n"&lt;br&gt;
    "Return ONLY valid JSON - no markdown:\n"&lt;br&gt;
    '{"synapse_type": "warning" or "success", '&lt;br&gt;
    '"reference_deal": "", '&lt;br&gt;
    '"pattern": "", '&lt;br&gt;
    '"suggestion": "", '&lt;br&gt;
    '"coached_reply": ""}'&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The user prompt is where the retrieved memories go in:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
user = (&lt;br&gt;
    f"Customer just said: \"{body.customer_message}\"\n\n"&lt;br&gt;
    f"Conversation so far:\n{conv_text}\n\n"&lt;br&gt;
    f"Historical deal patterns (won/lost):\n{history_context}\n\n"&lt;br&gt;
    f"Current deal memories:\n{deal_context}\n\n"&lt;br&gt;
    "Based on historical patterns, provide coaching for the sales rep."&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;history_context is the concatenated text of the top 3 recalled historical deal documents. deal_context comes from the current deal's own memory bank. Both are labelled clearly in the prompt - that labelling matters more than you'd expect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2cxuieywf96wii7864p.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2cxuieywf96wii7864p.jpeg" alt=" " width="799" height="423"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lines 759-776: The system prompt locks Groq into returning structured JSON with 5 fields. The user prompt injects the customer message, conversation history, recalled historical deals, and current deal memories - all in one call.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Memory Actually Changes the Output
&lt;/h2&gt;

&lt;p&gt;This is the thing worth paying attention to. Without retrieved memory in the prompt, Groq generates generic coaching:&lt;/p&gt;

&lt;p&gt;json&lt;br&gt;
{&lt;br&gt;
  "synapse_type": "warning",&lt;br&gt;
  "reference_deal": "General best practice",&lt;br&gt;
  "pattern": "Price negotiations should involve all stakeholders",&lt;br&gt;
  "suggestion": "Make sure InfoSec is aligned before discussing price",&lt;br&gt;
  "coached_reply": "Let me make sure we have full alignment before discussing pricing..."&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Fine. Forgettable. Useless in a live negotiation.&lt;/p&gt;

&lt;p&gt;With the Meridian Corp memory recalled - $380K lost because the rep agreed to a discount before getting InfoSec sign-off - Groq generates something genuinely useful:&lt;/p&gt;

&lt;p&gt;json&lt;br&gt;
{&lt;br&gt;
  "synapse_type": "warning",&lt;br&gt;
  "reference_deal": "Meridian Corp - LOST $380K",&lt;br&gt;
  "pattern": "Rep agreed to a 15% discount in Week 3 before InfoSec gatekeeper Elena Vasquez had been engaged. When Elena joined in Week 5, she blocked the multi-tenant architecture. No leverage left.",&lt;br&gt;
  "suggestion": "Do not discuss pricing until Priya Sharma (InfoSec) has formally signed off. Lock compliance first, price second.",&lt;br&gt;
  "coached_reply": "Robert, before we talk numbers I want to make sure Priya has everything she needs from our compliance team. Once InfoSec is locked, we can look at commercial terms together."&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Same model. Same output schema. Completely different output. The model is completing the pattern that the Meridian Corp transcript sets up. Without that memory, it has nothing specific to pattern-match against, so it falls back to generic advice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh28opgzwd9mcfrb9x7h3.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh28opgzwd9mcfrb9x7h3.jpeg" alt=" " width="799" height="423"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Turn 1: Customer asks about multi-tenant architecture. Hindsight recalls DataFlow Inc (LOST $290K) and Meridian Corp (LOST $380K). Groq generates a Pattern Warning with the exact coached reply - all in one API call.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Dual-Bank Strategy
&lt;/h2&gt;

&lt;p&gt;We use two separate Hindsight banks. One for all historical deals, one per active deal.&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
history_memories = await recall(&lt;br&gt;
    bank_id="historical-deals",&lt;br&gt;
    query=body.customer_message,&lt;br&gt;
    budget="high",&lt;br&gt;
    max_results=3,&lt;br&gt;
)&lt;br&gt;
Same model. Same output schema. Completely different output. The model is completing the pattern that the Meridian Corp transcript sets up. Without that memory, it has nothing specific to pattern-match against, so it falls back to generic advice.&lt;br&gt;
deal_memories = await recall(&lt;br&gt;
    bank_id=body.deal_id,&lt;br&gt;
    query=body.customer_message,&lt;br&gt;
    budget="mid",&lt;br&gt;
    max_results=3,&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;budget="high" for historical - we want thorough cross-deal search. budget="mid" for current deal - faster, we just need recent call notes from this specific deal. The results from both go into clearly labelled sections of the user prompt.&lt;/p&gt;

&lt;p&gt;The labelling is not optional. Early on we just dumped both into one block. The model would mix up which context was historical and which was current-deal, and the coaching became inconsistent. Adding explicit headers - "Historical deal patterns (won/lost)" and "Current deal memories" - fixed it almost immediately.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Thing That Surprised Me Most
&lt;/h2&gt;

&lt;p&gt;I expected prompt engineering to be the hard part. Finding the right system prompt wording, getting the JSON schema to work reliably, handling edge cases in the output parsing.&lt;/p&gt;

&lt;p&gt;That stuff was hard, but it wasn't the hardest part. The hardest part was realising how much the quality of the prompt depends on the quality of what's in memory. A beautifully engineered prompt fed garbage memories still produces garbage coaching. You can't engineer your way out of bad data upstream.&lt;/p&gt;

&lt;p&gt;If the Meridian Corp transcript just said "lost $380K - compliance issues," no prompt in the world would generate the specific, actionable coaching we needed. The transcript had to include Elena Vasquez's name, the Week 5 timeline, the specific sequence of mistakes. That narrative richness is what the model latches onto.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;GitHub:&lt;/em&gt; &lt;a href="https://github.com/grsanudeep42-cmd/dealmind" rel="noopener noreferrer"&gt;https://github.com/grsanudeep42-cmd/dealmind&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Demo video:&lt;/em&gt; &lt;a href="https://youtu.be/pxUxM-SIMSk" rel="noopener noreferrer"&gt;https://youtu.be/pxUxM-SIMSk&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Resources on Hindsight and agent memory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;https://github.com/vectorize-io/hindsight&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;https://hindsight.vectorize.io/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;https://vectorize.io/what-is-agent-memory&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Shoutout to &lt;a href="https://code.in" rel="noopener noreferrer"&gt;@Code.in&lt;/a&gt; for running this challenge.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>security</category>
    </item>
  </channel>
</rss>
