<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Akshith Bijigiri</title>
    <description>The latest articles on DEV Community by Akshith Bijigiri (@akshith_bijigiri).</description>
    <link>https://dev.to/akshith_bijigiri</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148994%2F09050cc9-34a8-41bc-991b-1dbcf1869a2a.png</url>
      <title>DEV Community: Akshith Bijigiri</title>
      <link>https://dev.to/akshith_bijigiri</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/akshith_bijigiri"/>
    <language>en</language>
    <item>
      <title>Why I Stopped Using Generic LLM Wrappers for My Agent</title>
      <dc:creator>Akshith Bijigiri</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:07:08 +0000</pubDate>
      <link>https://dev.to/akshith_bijigiri/why-i-stopped-using-generic-llm-wrappers-for-my-agent-5ga9</link>
      <guid>https://dev.to/akshith_bijigiri/why-i-stopped-using-generic-llm-wrappers-for-my-agent-5ga9</guid>
      <description>&lt;h1&gt;
  
  
  &lt;strong&gt;Why I Stopped Using Generic LLM Wrappers for My Agent&lt;/strong&gt;
&lt;/h1&gt;

&lt;p&gt;Building an AI agent is easy when you only need it to generate a response.&lt;/p&gt;

&lt;p&gt;Building one that &lt;strong&gt;remembers what happened five calls ago&lt;/strong&gt; is a different problem.&lt;/p&gt;

&lt;p&gt;While building &lt;strong&gt;DealMemory&lt;/strong&gt;, a sales intelligence agent with persistent memory, I ran into a surprisingly simple bug:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.text vs .content
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That small issue cost me hours of debugging and taught me an important lesson about working with LLMs, memory systems, and agent frameworks.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: LLMs Don't Remember Your Deals
&lt;/h2&gt;

&lt;p&gt;Imagine a sales representative has spoken with Acme Corp five times.&lt;/p&gt;

&lt;p&gt;During those calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The CTO raised API latency concerns.&lt;/li&gt;
&lt;li&gt;The CFO pushed back on pricing three times.&lt;/li&gt;
&lt;li&gt;Salesforce was mentioned as a competitor.&lt;/li&gt;
&lt;li&gt;The security team requested a SOC 2 report.&lt;/li&gt;
&lt;li&gt;Legal was discussing a 10% volume discount.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All this information might exist in CRM notes, but a normal LLM doesn't automatically know it.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Brief me on Acme Corp."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and you might get:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Review the stakeholder map, identify objections, and prepare for pricing discussions."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Technically correct, but not very useful.&lt;/p&gt;

&lt;p&gt;That's what we wanted to solve with &lt;strong&gt;DealMemory&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Our Approach
&lt;/h2&gt;

&lt;p&gt;DealMemory gives every deal its own memory bank.&lt;/p&gt;

&lt;p&gt;We built it using:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hindsight&lt;/strong&gt; — persistent memory&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Groq&lt;/strong&gt; — LLM inference&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streamlit&lt;/strong&gt; — user interface&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python&lt;/strong&gt; — application logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8c2nu3d52gtt6bvtqf1y.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8c2nu3d52gtt6bvtqf1y.jpeg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The basic flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Streamlit UI
     ↓
Agent Layer
     ↓
Hindsight Memory
     ↓
Relevant Deal History
     ↓
Groq LLM
     ↓
Grounded Sales Brief
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is that the LLM doesn't start with an empty context.&lt;/p&gt;

&lt;p&gt;It receives relevant information retrieved from the deal's history.&lt;/p&gt;




&lt;h2&gt;
  
  
  Hindsight: Retain, Recall, Reflect
&lt;/h2&gt;

&lt;p&gt;Hindsight gave us three important operations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retain
&lt;/h3&gt;

&lt;p&gt;Store every call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;store_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acme_corp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Call 2: CFO pushed on pricing. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ROI framing improved engagement.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;call_log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each deal gets its own memory bank, so Acme's information doesn't mix with Globex or Initech.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recall
&lt;/h3&gt;

&lt;p&gt;When the rep asks for a briefing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;recall_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acme_corp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What objections have been raised?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where I ran into the bug.&lt;/p&gt;




&lt;h1&gt;
  
  
  The &lt;code&gt;.text&lt;/code&gt; vs &lt;code&gt;.content&lt;/code&gt; Bug
&lt;/h1&gt;

&lt;p&gt;I initially assumed the retrieved memory would behave like a typical LLM response.&lt;/p&gt;

&lt;p&gt;I tried accessing the result using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the Hindsight recall result exposed the actual stored memory through:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the context needed to be constructed like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;memories&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The frustrating part was that &lt;code&gt;.content&lt;/code&gt; is something we commonly encounter when working with LLM responses.&lt;/p&gt;

&lt;p&gt;But a memory retrieval result isn't necessarily an LLM response.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Different layers of an AI application can have completely different object interfaces.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When debugging these systems, checking the actual object is often more useful than guessing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That simple step would have saved me a lot of time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpenmp45klb9zrgiv5ney.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpenmp45klb9zrgiv5ney.png" alt=" " width="800" height="566"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Then Came &lt;code&gt;reflect()&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Recall gives us the relevant information.&lt;/p&gt;

&lt;p&gt;But Hindsight also provides &lt;code&gt;reflect()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The difference is roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recall   → What happened?
Reflect  → What does it mean?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example, recall might show:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Call 1: CFO questioned pricing
Call 2: CFO questioned pricing
Call 4: CFO questioned pricing
ROI discussion improved engagement
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reflection can turn that history into a useful pattern:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The CFO consistently shows pricing sensitivity, but ROI-based positioning has improved engagement.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's much closer to a real sales copilot.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Before and After
&lt;/h2&gt;

&lt;p&gt;Without memory:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Schedule a discovery call and prepare for pricing objections."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With DealMemory:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Stakeholder:&lt;/strong&gt; CTO raised API latency concerns.&lt;br&gt;
&lt;strong&gt;Objection:&lt;/strong&gt; CFO pushed on pricing three times.&lt;br&gt;
&lt;strong&gt;Competitor:&lt;/strong&gt; Salesforce mentioned during discovery.&lt;br&gt;
&lt;strong&gt;Open items:&lt;/strong&gt; SOC 2 report and 10% volume discount.&lt;br&gt;
&lt;strong&gt;Next action:&lt;/strong&gt; Send an updated proposal using ROI framing and follow up with legal.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The LLM didn't become smarter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The context became better.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6423uvji31rpj2rc9efu.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6423uvji31rpj2rc9efu.jpeg" alt=" " width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  One More Debugging Detail: Async Indexing
&lt;/h2&gt;

&lt;p&gt;There was another small issue we had to handle.&lt;/p&gt;

&lt;p&gt;Hindsight's memory retention is asynchronous, so newly stored information may take a short time before it becomes searchable.&lt;/p&gt;

&lt;p&gt;Our flow therefore waits briefly after logging a call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Log Call
   ↓
Retain
   ↓
Wait for indexing
   ↓
Recall
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prevents the confusing situation where you store information and immediately wonder:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Why can't my agent find it?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;The biggest lesson wasn't simply to use &lt;code&gt;.text&lt;/code&gt; instead of &lt;code&gt;.content&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It was this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't assume every component in an AI pipeline follows the same response format.&lt;/strong&gt;&lt;br&gt;
Memory systems, LLM APIs, retrieval systems, and tools can all return different objects.&lt;br&gt;
Inspect the actual data before building assumptions around it.&lt;br&gt;
More importantly, building DealMemory changed how I think about AI agents.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How do I make my LLM smarter?"&lt;br&gt;
sometimes the better question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"How do I give my LLM better context?"&lt;/strong&gt;&lt;br&gt;
That's what DealMemory is built around:&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;text&lt;br&gt;
Generic&lt;br&gt;
   ↓&lt;br&gt;
Specific&lt;br&gt;
   ↓&lt;br&gt;
Pattern-aware&lt;/p&gt;

&lt;p&gt;And that small &lt;code&gt;.text&lt;/code&gt; bug was one of the debugging lessons that helped us get there.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
