<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rahul Kalakoti</title>
    <description>The latest articles on DEV Community by Rahul Kalakoti (@rahul_kalakoti_34d0f44c70).</description>
    <link>https://dev.to/rahul_kalakoti_34d0f44c70</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147101%2Fa7954cf8-9efd-4f6e-8476-b3d957764835.png</url>
      <title>DEV Community: Rahul Kalakoti</title>
      <link>https://dev.to/rahul_kalakoti_34d0f44c70</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rahul_kalakoti_34d0f44c70"/>
    <language>en</language>
    <item>
      <title>My On-Call Agent Remembered the Fix That Took Down Checkout</title>
      <dc:creator>Rahul Kalakoti</dc:creator>
      <pubDate>Mon, 28 Sep 2026 12:05:58 +0000</pubDate>
      <link>https://dev.to/rahul_kalakoti_34d0f44c70/my-on-call-agent-remembered-the-fix-that-took-down-checkout-1de2</link>
      <guid>https://dev.to/rahul_kalakoti_34d0f44c70/my-on-call-agent-remembered-the-fix-that-took-down-checkout-1de2</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2jhzgbuk2ax30f8rkfd4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2jhzgbuk2ax30f8rkfd4.png" alt=" " width="800" height="336"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4ykxmz1phsjowt3g4f8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4ykxmz1phsjowt3g4f8.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;02:06 UTC checkout-api started timing out. I pasted the alert into a capable LLM and asked what to do. It gave me a sensible, textbook answer: raise the connection pool size, rolling-restart the pods, and if that fails, consider a database failover.&lt;/p&gt;

&lt;p&gt;Every step of that answer had already been tried on this team. Raising the pool in April made things worse, because Postgres hit max_connections. The failover in May turned a SEV2 into a four-minute full checkout outage. The model wasn't wrong in general. It simply didn't know our history, and in incident response, knowing the history is most of the job.&lt;/p&gt;

&lt;p&gt;So I built Deja, an incident-response agent that remembers. It retains every postmortem, every fix that failed, and every time an engineer told it "you were wrong." When the pager fires it recalls what happened last time before it says anything. This post covers how it works, where Hindsight agent memory fits in, and what I got wrong along the way.&lt;/p&gt;

&lt;p&gt;What Deja does&lt;br&gt;
The loop is small on purpose:&lt;/p&gt;

&lt;p&gt;An alert fires (service, alert text, a few log lines).&lt;br&gt;
Deja recalls similar incidents from a Hindsight memory bank and loads the team's hard rules.&lt;br&gt;
It answers twice. The same LLM triages the alert once with memory and once without. The UI shows them side by side.&lt;br&gt;
An engineer resolves the incident and grades Deja's triage: spot on, partly right, or wrong.&lt;br&gt;
Deja retains the postmortem and the grade. The next similar alert benefits.&lt;br&gt;
The side-by-side view is the whole argument for memory. On the 02:06 checkout alert the stateless agent says "increase the pool." Deja says:&lt;/p&gt;

&lt;p&gt;Seen before: INC-2031 (Apr 8) and INC-2064 (May 14). The nightly bulk-export cron is hitting the Postgres primary. Kill the job. Do NOT raise the pool size (made it worse in INC-2031). Do NOT fail over the primary (full outage in INC-2064). Page Priya Raman.&lt;/p&gt;

&lt;p&gt;Both answers come from the same model on the same alert. The only difference is memory.&lt;/p&gt;

&lt;p&gt;Why I didn't just use RAG over postmortems&lt;br&gt;
The obvious approach is to embed the postmortem docs and do vector search. That breaks down quickly in on-call work:&lt;/p&gt;

&lt;p&gt;Alerts and postmortems use different words. The alert says HikariPool-1 - Connection is not available. The postmortem says "bulk-export starved checkout." Pure semantic search misses the exact error string, and pure keyword search misses the paraphrase.&lt;br&gt;
Time matters. "Starts at 02:06" is a strong signal when the culprit is a 02:00 cron.&lt;br&gt;
Entities matter across services. A CoreDNS problem that hit payments last month is relevant to search-indexer today.&lt;br&gt;
Some knowledge is a rule, not a fact. "Never fail over the primary for pool exhaustion" should always apply, not only when retrieval happens to rank it highly.&lt;br&gt;
Hindsight covers each of these with a separate primitive, so I didn't have to build that layer myself. Recall runs semantic, BM25, graph and temporal retrieval in parallel and reranks the fused results. Hard rules are directives. Summaries that should stay current are mental models.&lt;/p&gt;

&lt;p&gt;The memory bank is configured for SRE work&lt;br&gt;
The bank is created once at startup. create_bank is an upsert, so it runs on every boot:&lt;/p&gt;

&lt;p&gt;self.client.create_bank(&lt;br&gt;
    bank_id=self.bank,&lt;br&gt;
    name="Deja on-call memory",&lt;br&gt;
    mission=BANK_MISSION,&lt;br&gt;
    retain_mission=(&lt;br&gt;
        "Extract operational facts: service names, symptoms and error signatures, root causes, "&lt;br&gt;
        "remediation steps that worked, remediation steps that failed or made things worse, "&lt;br&gt;
        "people involved, time-to-resolve, recurring schedules (cron jobs, deploys)..."&lt;br&gt;
    ),&lt;br&gt;
    reflect_mission=REFLECT_MISSION,&lt;br&gt;
    disposition_skepticism=4,&lt;br&gt;
    disposition_literalism=4,&lt;br&gt;
    disposition_empathy=2,&lt;br&gt;
)&lt;br&gt;
The retain mission mattered more than I expected. Without it, fact extraction kept the narrative and dropped the one sentence that matters most at 2am: what did NOT work. With it, "failover caused a full outage" becomes a first-class fact that recall can surface.&lt;/p&gt;

&lt;p&gt;Every resolution becomes memory&lt;br&gt;
When an engineer closes an incident, Deja retains a postmortem with its real timestamp, tags, and structured metadata. If the engineer graded the triage, Deja also retains a separate feedback memory:&lt;/p&gt;

&lt;p&gt;items = [{&lt;br&gt;
    "content": postmortem_text(record),&lt;br&gt;
    "context": "incident postmortem",&lt;br&gt;
    "timestamp": inc["started_at"],&lt;br&gt;
    "document_id": inc["id"],&lt;br&gt;
    "metadata": postmortem_metadata(record),   # root_cause, fix, failed, resolved_by, ttr...&lt;br&gt;
    "tags": [f"service:{svc}", f"severity:{sev}", "type:postmortem"],&lt;br&gt;
}]&lt;br&gt;
if verdict:&lt;br&gt;
    items.append({&lt;br&gt;
        "content": f"Engineer feedback on Deja's triage of {inc_id}: verdict={verdict}. "&lt;br&gt;
                   f"Deja predicted: {predicted}. Actual root cause: {actual}. Correction: {notes}",&lt;br&gt;
        "context": "engineer feedback",&lt;br&gt;
        ...&lt;br&gt;
    })&lt;br&gt;
self.memory.retain_many(items)&lt;br&gt;
self.memory.refresh_runbook(svc)&lt;br&gt;
The feedback memory is what makes this learning rather than just search. When Deja confidently blames the database and the real cause was an inactive Debezium replication slot, that correction is retained. The next time a similar alert fires, recall surfaces both the incident and the fact that Deja got it wrong last time.&lt;/p&gt;

&lt;p&gt;Triage: recall, then reason, with a baseline for honesty&lt;br&gt;
memories = self.memory.recall(alert_query(inc))          # Hindsight recall&lt;br&gt;
rules    = self.memory.directives()                        # hard team rules&lt;br&gt;
deja     = llm.json(DEJA_SYSTEM, rules + memories + incident)&lt;br&gt;
baseline = llm.json(BASELINE_SYSTEM, incident)             # same model, no memory&lt;br&gt;
The system prompt has a few strict requirements: team-specific claims must cite an incident ID, anything a memory marks as failed must appear in do_not, and if nothing relevant was recalled the agent must say so and cap its confidence at 35%. That last rule is what makes the first occurrence of a new failure mode honest. Deja says "I've never seen this," gives general advice, and makes no attempt to pattern-match its way into a confident wrong answer.&lt;/p&gt;

&lt;p&gt;Runbooks nobody has to write&lt;br&gt;
Each service gets a mental model in Hindsight. Deja creates it with a source query ("known failure modes, the fix that worked, actions that did NOT work, who to page") and a trigger to refresh after consolidation:&lt;/p&gt;

&lt;p&gt;self.client.create_mental_model(&lt;br&gt;
    bank_id=self.bank, id=f"runbook-{service}", name=f"Runbook: {service}",&lt;br&gt;
    source_query=f"Write the operational runbook for {service}: known failure modes ...",&lt;br&gt;
    tags=[f"service:{service}"],&lt;br&gt;
    trigger={"refresh_after_consolidation": True},&lt;br&gt;
)&lt;br&gt;
After every resolved incident the runbook is rebuilt from memory. Nobody schedules a "runbook cleanup sprint" anymore, because the runbook is a side effect of doing on-call.&lt;/p&gt;

&lt;p&gt;Things that were painful&lt;br&gt;
Open-weight models and JSON. Running gpt-oss-120b on Groq is fast, but I still got fenced JSON,  blocks, and the occasional json_validate_failed. The LLM wrapper now tries JSON mode, then plain mode, then a fallback model (qwen3-32b), extracts the outermost object, and returns None rather than raising. A normaliser coerces whatever comes back into one shape the UI trusts.&lt;br&gt;
Retain is not instant. Fact extraction and consolidation take time. For seeding history I use async retain and let the UI show the memory count climbing. For a live resolution I retain synchronously, because the demo is "resolve it, fire it again, watch Deja know."&lt;br&gt;
Rules don't belong in retrieval. Stored as an ordinary memory, "never fail over the primary" only shows up when recall happens to rank it highly. As a Hindsight directive it applies on every triage, no matter what recall returns.&lt;br&gt;
Lessons&lt;br&gt;
Store what failed, explicitly. A postmortem's most valuable line is the one listing what didn't work. Tell the memory layer to extract it.&lt;br&gt;
Show the counterfactual. A side-by-side against the same model without memory is the most convincing evidence I've found.&lt;br&gt;
Make the agent admit ignorance. Cap confidence when recall comes back empty. A wrong answer delivered confidently at 2am costs more than no answer.&lt;br&gt;
Close the loop with feedback. Grading the agent takes one click and turns every incident into training data, with no fine-tuning involved.&lt;br&gt;
Separate facts from rules. Facts go through recall. Non-negotiables are directives.&lt;br&gt;
The code is on GitHub: &lt;a href="https://github.com/Rahulkalakoti45/deja-oncall" rel="noopener noreferrer"&gt;https://github.com/Rahulkalakoti45/deja-oncall&lt;/a&gt;. If you run on-call for anything, try pointing your last ten postmortems at it and see what it remembers. Read more about what agent memory is and the Hindsight docs.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
