<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Charan teja</title>
    <description>The latest articles on DEV Community by Charan teja (@charan_teja_7e5307e5250b8).</description>
    <link>https://dev.to/charan_teja_7e5307e5250b8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150818%2F763e2839-d3b0-46a6-8e70-6068a5208759.jpg</url>
      <title>DEV Community: Charan teja</title>
      <link>https://dev.to/charan_teja_7e5307e5250b8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/charan_teja_7e5307e5250b8"/>
    <language>en</language>
    <item>
      <title>Making Agent Memory Visible</title>
      <dc:creator>Charan teja</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:27:06 +0000</pubDate>
      <link>https://dev.to/charan_teja_7e5307e5250b8/making-agent-memory-visible-34j1</link>
      <guid>https://dev.to/charan_teja_7e5307e5250b8/making-agent-memory-visible-34j1</guid>
      <description>&lt;p&gt;ARTICLE 3 - For Agent / LLM Teammate&lt;br&gt;
Title: The Evaluation Layer: Where Recall Has to Become Reasoning&lt;br&gt;
I built the evaluation layer for VendorPulse - the part where recalled memories actually change a recommendation. This is where most memory demos fail.&lt;br&gt;
Retrieving from Hindsight is easy. You call recall() and you get relevant experiences. The hard part is making those experiences affect reasoning.&lt;br&gt;
A naive approach I tried first: concatenate memories into the prompt and ask Groq to "consider them". Result: LLM summarized memories but still recommended the cheapest vendor. It treated memories as trivia.&lt;br&gt;
The fix: Separate baseline reasoning from memory-informed reasoning.&lt;br&gt;
I implemented two distinct evaluation paths in FastAPI:python# evaluation.py&lt;br&gt;
def baseline_evaluate(request):&lt;br&gt;
    stats = get_vendor_stats(request.material)&lt;br&gt;
    # only uses avg delay, avg rejection, price&lt;br&gt;
    return groq.evaluate(f"Rank vendors by price and avg stats: {stats}")&lt;/p&gt;

&lt;p&gt;def memory_informed_evaluate(request, memories, stats):&lt;br&gt;
    # stats + episodic memories&lt;br&gt;
    return groq.evaluate(&lt;br&gt;
        system="You are a procurement risk analyst. Memories are contextual overrides.",&lt;br&gt;
        context={&lt;br&gt;
            "structured": stats,&lt;br&gt;
            "episodic": memories, # from Hindsight&lt;br&gt;
            "rule": "If vendor shows repeat excuse + repeat delay in same season, increase risk score. Calculate true cost = quoted price + (expected delay * penalty)"&lt;br&gt;
        }&lt;br&gt;
    )Concrete example that made it click:&lt;br&gt;
Request: Structural Steel, 10,000 kg, Oct 19, ₹10L, High priority&lt;br&gt;
Memories recalled:&lt;br&gt;
SteelCore July +22d 12% reject, August +18d 9% reject, October on time 3%MetalWorks July on time 2%, August +2d 3%Baseline says: SteelCore cheapest, recommend.&lt;br&gt;
Memory-informed says: &lt;br&gt;
"SteelCore: Pattern of monsoon failures with repeat transporter excuse. October history clean, but Oct 19 still in monsoon tail. Risk-adjusted true cost = ₹10L + 15 days * ₹5000 penalty = ₹10,75,000 with high variance.&lt;br&gt;
MetalWorks: Stable July-August for same material, low rejection. True cost = ₹10,50,000 with low variance. Recommend MetalWorks despite higher quoted price."&lt;br&gt;
That is learning. Same data, different reasoning, different recommendation.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>fastapi</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
