<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: YELUBANDI HARI DURGA RAMAKRISHNA</title>
    <description>The latest articles on DEV Community by YELUBANDI HARI DURGA RAMAKRISHNA (@ramakrishna_yeluband).</description>
    <link>https://dev.to/ramakrishna_yeluband</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150190%2F6835eb41-8f3b-4d73-8eb6-598f194470f9.jpg</url>
      <title>DEV Community: YELUBANDI HARI DURGA RAMAKRISHNA</title>
      <link>https://dev.to/ramakrishna_yeluband</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ramakrishna_yeluband"/>
    <language>en</language>
    <item>
      <title>How I built a runbook that learns using Hindsight</title>
      <dc:creator>YELUBANDI HARI DURGA RAMAKRISHNA</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:45:26 +0000</pubDate>
      <link>https://dev.to/ramakrishna_yeluband/how-i-built-a-runbook-that-learns-using-hindsight-3kij</link>
      <guid>https://dev.to/ramakrishna_yeluband/how-i-built-a-runbook-that-learns-using-hindsight-3kij</guid>
      <description>&lt;p&gt;Every runbook I have ever used was out of date the day after it was written. The one for our payment service still said "restart the pods" long after two separate on-call engineers had proven that restarting the pods makes the problem worse. So I stopped writing runbooks and built one that updates itself from what actually happened.&lt;/p&gt;

&lt;p&gt;What the system does&lt;br&gt;
The Hindsight Incident Responder is an on-call agent for production engineering teams. When an alert arrives — PaymentService 5xx rate &amp;gt; 5%, a Postgres log line about exhausted connection slots, a deploy twenty minutes ago — it answers three questions:&lt;/p&gt;

&lt;p&gt;What is most likely wrong, and how sure are we?&lt;br&gt;
What should I do, in order, with proven fixes first?&lt;br&gt;
What should I not bother trying, because it already failed here before?&lt;br&gt;
Everything in those answers comes from memory. Each resolved incident is retained into Hindsight, an open-source agent memory system from Vectorize. When a new alert comes in, the agent recalls similar incidents and grounds its plan in them, citing incident IDs. After the incident is closed, the engineer marks which suggested steps worked or failed and writes a short post-mortem, which is structured, reviewed, and retained. A separate reflect call reasons across the whole bank to answer questions like "why does payment-service keep going down?"&lt;/p&gt;

&lt;p&gt;The LLM never changes between the first incident and the hundredth. The runbook it produces does, because the memory does. To make that visible rather than something I ask people to trust, the UI answers every alert twice — once with no memory, once with Hindsight — side by side.&lt;/p&gt;

&lt;p&gt;Same alert, same model: without memory on the left, with Hindsight on the right&lt;/p&gt;

&lt;p&gt;How it hangs together&lt;br&gt;
The stack is small on purpose. A Streamlit console renders an alert feed and the results. agent.py orchestrates triage, post-mortem structuring, follow-up questions and pattern queries. memory.py is the only module that knows Hindsight exists; it exposes retain, recall and reflect, and appends every call to an in-process trace the sidebar renders live. llm.py is an OpenAI-compatible client, so switching between Gemini and Groq is a three-line change in .env.&lt;/p&gt;

&lt;p&gt;Where Hindsight sits&lt;/p&gt;

&lt;p&gt;Hindsight runs as a hosted service. Our memory bank holds the team's incident history: what the alerts looked like, what the root cause turned out to be, what fixed it, what didn't, and what we learned. The Hindsight documentation covers the retrieval model in depth; the short version is that it extracts facts and entities from what you retain and combines semantic, keyword, entity and temporal signals when you recall, which is what lets a query phrased nothing like the original post-mortem still find the right incident.&lt;/p&gt;

&lt;p&gt;The core problem: a runbook is a frozen answer&lt;br&gt;
A traditional runbook encodes one engineer's understanding at one point in time. It can't distinguish two incidents with identical alerts but different causes, it never records the twelve minutes someone wasted on the wrong fix, and nobody updates it at 2 AM. I wanted the runbook to be a derived artifact: regenerated at triage time from every relevant incident the team has ever closed, including the failures.&lt;/p&gt;

&lt;p&gt;That reframing drove three design decisions.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An incident becomes two memories, written for the retriever
My first attempt retained each post-mortem as one block of prose. Recall was mediocre. The alert says remaining connection slots are reserved for non-replication superuser connections; the post-mortem said "pool exhaustion after v2.31." Those are the same event, but the link was buried.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So each incident is split into a diagnosis memory and a resolution memory, and both are dense with the things a future alert will contain — exact error strings, config keys, service names:&lt;/p&gt;

&lt;p&gt;Python&lt;/p&gt;

&lt;p&gt;def incident_to_memories(inc: Dict) -&amp;gt; List[Dict]:&lt;br&gt;
    header = (f"INCIDENT {inc['id']} | {inc['date']} | {inc['severity']} | "&lt;br&gt;
              f"service: {inc['service']} | {inc['title']}")&lt;br&gt;
    tags = ", ".join(inc.get("tags", []))&lt;br&gt;
    steps = " ".join(f"{i+1}) {s}" for i, s in enumerate(inc.get("resolution_steps", [])))&lt;br&gt;
    failed = "; ".join(inc.get("failed_attempts", [])) or "none recorded"&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;diagnosis = (f"{header}\nSymptoms: {inc['symptoms']}\n"
             f"Root cause: {inc['root_cause']}\nTags: {tags}")
resolution = (f"{header}\nResolution steps that WORKED: {steps} "
              f"Runbook used: {inc.get('runbook', 'n/a')}. "
              f"Resolved in {inc.get('time_to_resolve_min', '?')} minutes.\n"
              f"Attempts that FAILED (do not repeat): {failed}\n"
              f"Lesson: {inc.get('lesson', '')}\nTags: {tags}")

meta = {"incident_id": inc["id"], "service": inc["service"],
        "severity": inc["severity"], "tags": tags}
ts = f"{inc['date']}T00:00:00Z"
return [
    {"text": diagnosis,  "metadata": {**meta, "kind": "diagnosis"},  "timestamp": ts},
    {"text": resolution, "metadata": {**meta, "kind": "resolution"}, "timestamp": ts},
]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The phrase Attempts that FAILED (do not repeat) is the most valuable line in the system. It is the part of a post-mortem that never makes it into a static runbook.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Triage recalls first, and the prompt is not allowed to improvise
Triage is deliberately boring. The raw alert text is the recall query. The returned memories are formatted with their incident IDs from metadata — I learned that Hindsight sometimes returns an extracted fact rather than the verbatim text, so relying on the ID appearing in the sentence was fragile — and the LLM is asked for JSON:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Python&lt;/p&gt;

&lt;p&gt;def triage(alert: str, use_memory: bool = True) -&amp;gt; Dict:&lt;br&gt;
    mems = memory.recall(alert, max_tokens=6000, top_k=12) if use_memory else []&lt;br&gt;
    system = prompts.TRIAGE_WITH_MEMORY if use_memory else prompts.TRIAGE_AMNESIA&lt;br&gt;
    user = prompts.TRIAGE_USER.format(alert=alert, memories=_format_memories(mems))&lt;br&gt;
    plan = llm.chat_json(system, user)&lt;br&gt;
    return {"memories": mems, "plan": plan, "mode": "hindsight" if use_memory else "amnesia"}&lt;br&gt;
The prompt rules are where the "runbook" character comes from. Fixes are ranked by what worked. Every step cites an incident. A recurring failure class gets a final step tagged pattern describing the permanent prevention the memories recommend. And the "don't bother trying" list has a hard constraint:&lt;/p&gt;

&lt;p&gt;text&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;do_not_try must contain ONLY failed attempts that are literally stated in the
memories, each with the incident ID where it is stated. Never invent one.&lt;/li&gt;
&lt;li&gt;Never attribute a fact, step, or failure to an incident ID unless that memory
actually contains it.
That constraint was not in the first version. I added it after watching the model confidently attribute a made-up failure to a real incident number. A runbook that cites sources is only useful if the citations are real.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Nothing enters memory unreviewed
The learning loop is what makes this a runbook that learns rather than one that merely searches. After the incident, the engineer marks each suggested step as worked or failed, writes what happened in plain words, and the agent turns that into the same structured shape:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Python&lt;/p&gt;

&lt;p&gt;def structure_postmortem(free_text, step_feedback=None, alert="", next_id=None) -&amp;gt; Dict:&lt;br&gt;
    feedback_txt = ""&lt;br&gt;
    if step_feedback:&lt;br&gt;
        worked = [f["step"] for f in step_feedback if f["result"] == "worked"]&lt;br&gt;
        failed = [f["step"] for f in step_feedback if f["result"] == "failed"]&lt;br&gt;
        feedback_txt = (f"\n\nOriginal alert: {alert}\nSteps that WORKED: {worked}\n"&lt;br&gt;
                        f"Steps that FAILED: {failed}")&lt;br&gt;
    hint = f"\nUse id {next_id}. Today's date is {date.today().isoformat()}."&lt;br&gt;
    structured = llm.chat_json(prompts.STRUCTURE_POSTMORTEM + hint, free_text + feedback_txt)&lt;br&gt;
    ...&lt;br&gt;
    return structured   # shown to the engineer; retained only after confirmation&lt;br&gt;
The record is displayed, and only a human click sends it to memory.retain_incident. That gate exists because of the worst bug I hit, described below.&lt;/p&gt;

&lt;p&gt;Review the structured record before it is retained&lt;/p&gt;

&lt;p&gt;What it does in practice&lt;br&gt;
The bank holds 22 incidents across nine services, seeded from our post-mortem history, including three deliberate patterns: payment-service connection-pool exhaustion recurring three times, two auth-service Redis incidents with the same missing-TTL cause, and two pairs of look-alike alerts with different root causes.&lt;/p&gt;

&lt;p&gt;Take the payment alert above. Without memory, the model says: roll back the v2.52 deploy, terminate idle Postgres connections. Sensible, generic, and blind to this team's history.&lt;/p&gt;

&lt;p&gt;With Hindsight, the same model matches INC-017, INC-031 and INC-051 and explains why each matches. The first step is DB_POOL_MAX=60 followed by kubectl rollout restart deploy/payment-service, sourced to INC-017. A final step tagged pattern recommends computing the connection budget from maxReplicas and adding the CI check the team wrote after the second occurrence. Under "already tried before — didn't work": restarting pods without adjusting the pool size — failed in INC-017, the pool refilled and exhausted again within three minutes. The UI also shows that the matched incidents were resolved in about fifteen minutes on average, computed from their recorded times, against a clearly labelled assumed baseline.&lt;/p&gt;

&lt;p&gt;The case I care most about is the look-alike. Same service, same 502 alert, but the log says x509: certificate has expired or is not yet valid calling psp-gateway.internal. The agent matches INC-044 only, gives the certificate fix — delete the psp-gateway-tls secret so cert-manager reissues it, restart the gateway — and warns that the database runbook was misapplied for twelve minutes the last time this exact confusion happened. A static runbook indexed by alert name cannot do that. Memory that includes root causes and dead ends can.&lt;/p&gt;

&lt;p&gt;After closing an incident and retaining the post-mortem, re-running the same alert cites the new incident within seconds. The runbook has updated itself.&lt;/p&gt;

&lt;p&gt;A real recall call in the live trace&lt;/p&gt;

&lt;p&gt;The bug that changed the design&lt;br&gt;
Early on I clicked "save" with the post-mortem template still empty. The structuring model did what a language model does with a blank page: it invented a plausible incident. Root cause "connection leak," a runbook called RB-DB-CONN-01 that has never existed, and a failed attempt — "waiting for traffic to subside" — that nobody ever tried. That fiction was retained as INC-100.&lt;/p&gt;

&lt;p&gt;Then everything got worse. INC-100 outranked three real incidents in recall, so the agent stopped citing the actual config fix. Its invented failed attempts leaked into recommendations for the unrelated certificate alert. One fabricated memory, and every answer touching payment-service degraded.&lt;/p&gt;

&lt;p&gt;This is the property of agent memory that deserves more attention than it gets: it is permanent by design. Garbage in is not garbage out once; it is garbage out until someone notices and purges it. Vectorize's overview of agent memory describes memory as what lets an agent improve across sessions. The corollary is that it also lets an agent degrade across sessions if the ingestion path is careless.&lt;/p&gt;

&lt;p&gt;The fix was the review gate above, plus a structuring prompt that is forbidden from filling gaps: if the notes don't say what fixed it, resolution_steps is empty and root_cause is "unknown"; if no failed attempt is mentioned, failed_attempts is empty. I retested with deliberately useless notes and got exactly that — an honest, empty record I could discard.&lt;/p&gt;

&lt;p&gt;A smaller pain worth recording: the Hindsight Python client uses an async HTTP session internally, and Streamlit re-runs scripts across threads, which produced "Event loop is closed" on the second call. I ended up giving every SDK call its own event loop and client on a dedicated worker thread. It costs milliseconds and removed the problem entirely.&lt;/p&gt;

&lt;p&gt;Lessons&lt;br&gt;
A runbook should be derived, not authored. Regenerate it at triage time from every relevant incident, and it can never be stale.&lt;br&gt;
Store failures, not just fixes. "Don't bother trying X — failed in INC-017" saved more time in testing than any ranked fix list. Static runbooks almost never record this.&lt;br&gt;
Write memories for the retriever. Exact error strings, config keys and service names, split into small focused memories. Recall quality was mostly data quality.&lt;br&gt;
Memory needs an ingestion gate. A wrong memory compounds. Review before retain, and forbid the structuring model from guessing.&lt;br&gt;
Make recall visible. A live trace of every retain, recall and reflect turned "trust the agent" into "here are the four memories it used," and made recall problems debuggable.&lt;br&gt;
The code is at github.com/25A31A05LI/hindsight-incident-responder.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>automation</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
