<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rama krishna</title>
    <description>The latest articles on DEV Community by Rama krishna (@rama_krishna_c048c2d91677).</description>
    <link>https://dev.to/rama_krishna_c048c2d91677</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3796199%2F85ca672a-04cb-423a-94ff-d2c7f967bcdf.jpg</url>
      <title>DEV Community: Rama krishna</title>
      <link>https://dev.to/rama_krishna_c048c2d91677</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rama_krishna_c048c2d91677"/>
    <language>en</language>
    <item>
      <title>My Agent Remembered the Fix and Was Wrong</title>
      <dc:creator>Rama krishna</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:05:34 +0000</pubDate>
      <link>https://dev.to/rama_krishna_c048c2d91677/my-agent-remembered-the-fix-and-was-wrong-410i</link>
      <guid>https://dev.to/rama_krishna_c048c2d91677/my-agent-remembered-the-fix-and-was-wrong-410i</guid>
      <description>&lt;p&gt;Last month my incident-recovery agent looked at a failing search query, remembered it had fixed the same symptom before, and proposed the same fix. It would have been the wrong one.&lt;/p&gt;

&lt;p&gt;The bug was a search gateway returning results from the wrong corpus version. The first time, an alias was pointing at the old collection. This time the alias was already correct, and the real problem was somewhere else.&lt;/p&gt;

&lt;p&gt;What AfterTrace does&lt;/p&gt;

&lt;p&gt;AfterTrace is a command-line tool that helps recover from incidents. It follows one loop: detect, diagnose, propose, approve, fix, verify, retain.&lt;/p&gt;

&lt;p&gt;The rule I built around is: memory proposes, live evidence disposes. The agent can remember, but it never acts on memory alone.&lt;/p&gt;

&lt;p&gt;It's built from:&lt;/p&gt;

&lt;p&gt;Qdrant Cloud for the vector data the search gateway queries&lt;br&gt;
Hindsight for the agent's long-term memory&lt;br&gt;
SQLite (one file) to log incidents, events, aliases, and cache state&lt;br&gt;
A Python CLI that runs the loop and asks a human to approve every write&lt;/p&gt;

&lt;p&gt;You can see the full project here:&lt;br&gt;
&lt;a href="https://github.com/Ramakrishna1967/AfterTrace" rel="noopener noreferrer"&gt;https://github.com/Ramakrishna1967/AfterTrace&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The idea&lt;/p&gt;

&lt;p&gt;Most talk about agent memory is about the upside: the agent remembers, so it gets better. But a memory that makes an agent quicker at being right also makes it quicker at being wrong when things change.&lt;/p&gt;

&lt;p&gt;So I gave Hindsight recall one small job: decide what to check first. A recalled fix never gets applied on its own. Before any change, the agent re-checks the live state that fix depends on.&lt;/p&gt;

&lt;p&gt;I tested this with three scenarios, each run in a fresh process so the only thing carried over is what's stored in memory.&lt;/p&gt;

&lt;p&gt;Scenario 1: cold start&lt;/p&gt;

&lt;p&gt;A new corpus, B, is loaded correctly, but the alias still points at the old one, A. A query that should return B returns A.&lt;/p&gt;

&lt;p&gt;With no memory, the agent checks that B is fine, sees the alias points at A, proposes switching it, and waits for approval. After the change it re-queries B, confirms it works, and saves the real incident to Hindsight. I only save what actually happened, never a made-up summary.&lt;/p&gt;

&lt;p&gt;Scenario 2: transfer&lt;/p&gt;

&lt;p&gt;A different corpus, same kind of bug. A fresh process calls recall through the Hindsight API, gets the experience from Scenario 1, and checks the alias first. That's the payoff of persistent memory: it went straight to the likeliest cause. It still checked live state before writing anything, and the query passed.&lt;/p&gt;

&lt;p&gt;Scenario 3: the important one&lt;/p&gt;

&lt;p&gt;Same symptom, but now the alias already points at B. The real cause is a stale cache serving old results.&lt;/p&gt;

&lt;p&gt;Recall returns the old alias fix. The agent re-checks the live alias, sees the fix no longer applies, and prints:&lt;/p&gt;

&lt;p&gt;REJECTED recalled alias fix&lt;/p&gt;

&lt;p&gt;It then looks for another cause, finds the stale cache, and clears only the cache. The alias is left alone.&lt;/p&gt;

&lt;p&gt;Without the check, Scenario 3 ends with the agent rewriting a healthy alias while the real fault stays. With it, memory decides what to look at and live evidence decides what to change.&lt;/p&gt;

&lt;p&gt;What I learned&lt;br&gt;
Use recall to prioritize, not to authorize. Ranking suspects by past incidents helps a lot. Acting on them blindly is the risk.&lt;br&gt;
Re-check what a recalled fix depends on. A remembered fix assumes the world is unchanged. Test that.&lt;br&gt;
Only save real incidents. Made-up memories poison future recalls.&lt;br&gt;
Test the failure case first. The scenario where memory is wrong shows whether your design is safe.&lt;br&gt;
Fresh processes show what memory really does. Restarting between scenarios made it clear what came from Hindsight and what was just in RAM.&lt;/p&gt;

&lt;p&gt;What's next&lt;/p&gt;

&lt;p&gt;Right now the checks cover alias and cache faults. Other fault types need their own checks. I also want to save rejected recalls, so the rejection itself becomes something the agent learns from.&lt;/p&gt;

&lt;p&gt;If you're building an agent with persistent memory, Hindsight is worth trying. The biggest lesson for me was deciding exactly what memory is allowed to decide.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I'm a CS Sophomore and I Built an Open-Source Observability Platform for AI Agents</title>
      <dc:creator>Rama krishna</dc:creator>
      <pubDate>Fri, 27 Feb 2026 09:30:11 +0000</pubDate>
      <link>https://dev.to/rama_krishna_c048c2d91677/im-a-cs-sophomore-and-i-built-an-open-source-observability-platform-for-ai-agents-12j7</link>
      <guid>https://dev.to/rama_krishna_c048c2d91677/im-a-cs-sophomore-and-i-built-an-open-source-observability-platform-for-ai-agents-12j7</guid>
      <description>&lt;p&gt;Three months ago I was debugging a LangGraph agent that kept failing in production. Token costs were spiking, I had no idea which tool call was causing it, and every time I tried to reproduce the bug the agent did something completely different. That's when I realized: AI agents need their own observability tooling. So I built it.&lt;br&gt;
The Problem with Current Tools&lt;br&gt;
Web observability tools like Datadog and New Relic are great — for web services. But AI agents are fundamentally different. They run for minutes or hours, not milliseconds. Their costs are variable and per-token. They're non-deterministic, so you can't reproduce failures the same way. And they process sensitive data — full conversation text — that needs to be handled carefully.&lt;br&gt;
What AgentStack Does&lt;br&gt;
AgentStack is a free, self-hostable observability platform built specifically for AI agents. The entire SDK surface is one decorator — @observe — and that's intentional. I didn't want developers to rewrite their agents or learn a new paradigm. You instrument what you already have, and AgentStack does the rest.&lt;br&gt;
The moment you add @observe to your agent function, every execution starts producing full structured traces — arguments, return values, timing, exceptions, token counts, and cost. It captures sync functions, async functions, and deeply nested call chains automatically. Agent calls a tool, tool calls an LLM, LLM calls another tool — AgentStack links all of it into a clean parent-child span tree.&lt;br&gt;
The Time Machine is the feature I'm most proud of. AI agents are non-deterministic — you can't just run the same input again and expect the same bug to appear. Time Machine lets you step through any past execution, span by span, and see exactly what every LLM returned, what every tool did, and which decision path the agent took. It's like a debugger, but for agent runs that already happened.&lt;br&gt;
The Security Engine runs in real time. Every span is analyzed for prompt injection patterns, PII leakage, token explosions, and anomalous latency. If something suspicious happens, an alert surfaces in the dashboard within seconds. And before any span ever hits storage, AgentStack automatically scrubs SSNs, credit card numbers, emails, phone numbers, and API keys — no configuration required.&lt;br&gt;
Cost Analytics was something I personally needed badly. When you're running GPT-4 agents in production, costs can spiral fast and silently. AgentStack tracks per-model token usage and calculates USD cost for every single span, with timeseries charts broken down by hour, day, or week across GPT-4, Claude, Gemini, and any other provider.&lt;br&gt;
The whole thing is fully self-hostable with one Docker Compose command. Redis, ClickHouse, the collector, the API, the workers, and the React dashboard — all spin up together. Your traces never leave your infrastructure. No SaaS subscription. No data sharing. Just your agents, fully visible, on your own servers&lt;br&gt;
If this resonates with you, I'd love a star on GitHub — it's my first open-source launch and every star genuinely helps others find the project.&lt;br&gt;
👉 &lt;a href="https://github.com/Ramakrishna1967/AgentStack" rel="noopener noreferrer"&gt;https://github.com/Ramakrishna1967/AgentStack&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
