<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Akkireddy Hasini</title>
    <description>The latest articles on DEV Community by Akkireddy Hasini (@akkireddy_hasini_bfb7cd39).</description>
    <link>https://dev.to/akkireddy_hasini_bfb7cd39</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150626%2Fbd114a82-c389-4419-b4e7-a70e64a26e48.png</url>
      <title>DEV Community: Akkireddy Hasini</title>
      <link>https://dev.to/akkireddy_hasini_bfb7cd39</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/akkireddy_hasini_bfb7cd39"/>
    <language>en</language>
    <item>
      <title>A Memory ON/OFF Toggle Was My Best Hindsight Debugging Tool</title>
      <dc:creator>Akkireddy Hasini</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:01:32 +0000</pubDate>
      <link>https://dev.to/akkireddy_hasini_bfb7cd39/a-memory-onoff-toggle-was-my-best-hindsight-debugging-tool-3ofa</link>
      <guid>https://dev.to/akkireddy_hasini_bfb7cd39/a-memory-onoff-toggle-was-my-best-hindsight-debugging-tool-3ofa</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mxh89v7xdn61nhqqr6h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mxh89v7xdn61nhqqr6h.png" alt=" " width="799" height="368"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8uh96tndgw9xuvfnk4i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8uh96tndgw9xuvfnk4i.png" alt=" " width="800" height="369"&gt;&lt;/a&gt;The plan looked great, and I couldn't tell whether it was using any of our incident history.&lt;br&gt;
That was the real problem when I built Incident Copilot. Getting an LLM to write a fluent incident response plan is easy. Knowing whether that plan came from our past outages or from the model's general training is hard. A confident answer proves nothing, so I added a switch.&lt;br&gt;
What the system does&lt;br&gt;
Incident Copilot is a small FastAPI service with a single-page UI. You paste an alert and symptoms, and it returns a ranked plan. Beside the plan, a "Memories used" panel shows the exact records the plan was built from.&lt;/p&gt;

&lt;p&gt;The flow is one recall followed by one LLM call. There is no planner loop and no vector store I maintain. The model is openai/gpt-oss-120b on Groq, and the memory layer is Hindsight, the open-source agent memory system. Four endpoints do the work:&lt;br&gt;
• POST /plan recalls memories and returns a structured plan.&lt;br&gt;
• POST /action records a step's outcome (worked, failed, harmful) mid-incident.&lt;br&gt;
• POST /close retains the full postmortem.&lt;br&gt;
• GET /patterns asks Hindsight to reflect on recurring root causes.&lt;/p&gt;

&lt;p&gt;If you're new to the idea, Vectorize has a good overview of what agent memory is, and the Hindsight docs cover retain, recall, and reflect, the three operations I use.&lt;br&gt;
Why I needed an off switch&lt;br&gt;
Without memory, an LLM gives you the textbook plan: check the pool, look at recent deploys, restart the pods, scale up. Those are sensible in isolation, which is the trap. If recall silently breaks (empty results, wrong bank, a bad query), the plan still reads well and nothing tells you.&lt;br&gt;
So /plan takes a use_memory flag. Same model, same prompt, same alert, and the only variable is whether recall runs. The same flag drives my evaluation script:&lt;br&gt;
python&lt;br&gt;
for mode in ("off", "on"):&lt;br&gt;
    p = agent.plan(inc["alert"], inc["symptoms"], use_memory=(mode == "on"))&lt;br&gt;
    row[mode] = grade(p, inc)&lt;br&gt;
Once that flag existed, every question became an A/B test instead of a hunch. The toggle turned "does memory help?" into something I could answer per alert in about ten seconds.&lt;br&gt;
What I put in memory: failures with outcomes&lt;br&gt;
A toggle only shows a difference if there is something worth recalling. My seed data is in data/incidents.json, and each incident records every action taken and how it turned out:&lt;br&gt;
json&lt;br&gt;
{&lt;br&gt;
  "id": "INC-101",&lt;br&gt;
  "service": "checkout-service",&lt;br&gt;
  "alert": "checkout-service p99 latency &amp;gt; 4s for 5m",&lt;br&gt;
  "actions": [&lt;br&gt;
    { "action": "Restarted checkout pods", "outcome": "harmful",&lt;br&gt;
      "note": "Latency got worse; the reconnect storm exhausted the pool..." },&lt;br&gt;
    { "action": "Scaled checkout to 12 replicas", "outcome": "harmful",&lt;br&gt;
      "note": "More replicas opened more DB connections and hit max_connections..." }&lt;br&gt;
  ]&lt;br&gt;
}&lt;br&gt;
Restarting and scaling are the two things everyone reaches for at 3 a.m., and in this failure family both make it worse. At retain time I render each outcome into the text as an uppercase word (Outcome: HARMFUL), because Hindsight extracts memories from natural language and I wanted "restarting pods was harmful" to survive as a fact, not as a dropped metadata field.&lt;/p&gt;

&lt;p&gt;Reading the toggle: the same alert, two ways&lt;br&gt;
I ran an alert that isn't in the seed data: "Kafka consumer lag growing on the orders topic," with the symptom "Consumer group keeps rebalancing every few minutes."&lt;br&gt;
Memory ON: the plan's first step was to roll back config deploy #5102, which lowered the orders-api connection pool size and idle timeout. It cited INC-104, where the same kind of settings change caused instability. A yellow banner warned that several referenced incidents were older than six months. In the Memories used panel, one recalled record reads: "Restarting orders-api pods worsened the incident by triggering a reconnect storm that exhausted the connection pool."&lt;/p&gt;

&lt;p&gt;The panel is what made this a debugging tool and not just a demo. Each recalled memory is tagged world, experience, or observation, and dated. When a plan looked wrong, I could see in one glance whether retrieval had fetched the wrong record or the model had misused the right one.&lt;br&gt;
The loop that makes it learn&lt;br&gt;
Every step in a plan has Worked, Failed, and Harmful buttons. Clicking one calls /action, which retains a short note right away, keyed to a stable document_id so edits replace records instead of duplicating them:&lt;br&gt;
python&lt;br&gt;
_retain(&lt;br&gt;
    bank_id=BANK,&lt;br&gt;
    content=f"During incident {incident_id}, attempted: {action}. Outcome: {outcome.upper()}. {note}",&lt;br&gt;
    context="live incident action",&lt;br&gt;
    document_id=f"incident-{incident_id}-action-{n}",&lt;br&gt;
)&lt;br&gt;
I marked a pool-sizing step as worked for incident LIVE-001 and the UI showed "Worked (saved)." A later alert on payments-api (HikariPool-1 - Connection is not available, request timed out after 30000ms) got a plan that cited LIVE-001 directly, saying a recent deploy had altered pool settings and caused the same timeout. An incident I had just closed was already shaping the next recommendation. The plan also cited INC-108 for the rollback step.&lt;br&gt;
The "Show patterns" button calls Hindsight's reflect operation and returns a "Recurring Root Causes and Failed Fixes" summary across the whole bank, starting with families like expired mTLS certificates. I use it between incidents more than during them.&lt;/p&gt;

&lt;p&gt;A choice that mattered: recall across services&lt;br&gt;
I don't filter recall by service:&lt;br&gt;
python&lt;br&gt;
resp = client.recall(bank_id=BANK, query=f"{alert} {symptoms}".strip(),&lt;br&gt;
                     budget="high", max_tokens=3000)&lt;br&gt;
A config change that shrinks a connection pool looks the same on checkout-service, orders-api, and payments-api. In my data the same root cause appears across several services, and a per-service index would have found nothing for a new one. The Kafka run above worked because it recalled an orders-api incident for a Kafka consumer symptom, which shows the flip side too: recall keyed on symptoms can pull a plan toward a root cause that doesn't fit. Nothing in the Kafka alert mentioned a config change, yet the plan led with one.&lt;br&gt;
Lessons learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Build the off switch first. It is cheaper than any evaluation framework, and without it you can't tell what memory contributes.&lt;/li&gt;
&lt;li&gt;Put the outcome in the text. "Harmful" as a searchable word in the memory, not just a JSON field, is what let recall surface warnings.&lt;/li&gt;
&lt;li&gt;Show the recalled memories next to the answer. It separates retrieval bugs from generation bugs faster than anything else I tried.&lt;/li&gt;
&lt;li&gt;Real timestamps matter. I retain each incident with its resolution date, so the agent can reason about staleness. The six-month banner in the Kafka run comes from that.&lt;/li&gt;
&lt;li&gt;Expect over-anchoring. Cross-service recall is powerful and occasionally too confident. &lt;/li&gt;
&lt;li&gt;Give memory time to settle. Retained content is consolidated in the background, so recall right after seeding can look thin. I lost time to that once before I read the seed script's reminder.
Where this goes next
The code is at &lt;a href="https://github.com/Hasinireddy2407/incident-copilot" rel="noopener noreferrer"&gt;https://github.com/Hasinireddy2407/incident-copilot&lt;/a&gt;. If you want to try the memory layer yourself, start with the Hindsight GitHub repository.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>backend</category>
      <category>llm</category>
      <category>python</category>
    </item>
  </channel>
</rss>
