<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: saanvii248</title>
    <description>The latest articles on DEV Community by saanvii248 (@saanvii248).</description>
    <link>https://dev.to/saanvii248</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150805%2F18463c1c-8056-4712-82d5-9777a9d60f0d.png</url>
      <title>DEV Community: saanvii248</title>
      <link>https://dev.to/saanvii248</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saanvii248"/>
    <language>en</language>
    <item>
      <title>Title: Giving On-Call Engineers a Memory: Building On Call Memory with Hindsight</title>
      <dc:creator>saanvii248</dc:creator>
      <pubDate>Tue, 29 Sep 2026 19:17:59 +0000</pubDate>
      <link>https://dev.to/saanvii248/title-giving-on-call-engineers-a-memory-building-on-call-memory-with-hindsight-3ckp</link>
      <guid>https://dev.to/saanvii248/title-giving-on-call-engineers-a-memory-building-on-call-memory-with-hindsight-3ckp</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayr2k7k46c0qbdix4w82.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayr2k7k46c0qbdix4w82.jpeg" alt=" " width="800" height="415"&gt;&lt;/a&gt;Article &lt;br&gt;
Production incidents have a strange property: the same ones keep coming back. A database connection pool fills up in March, a different service leaks connections in August, and each time an engineer starts from zero. The knowledge exists somewhere, in a postmortem doc, a Slack thread, or the head of someone who left the company. It just isn't available at 3 a.m. when the pager goes off.&lt;br&gt;
For the Hindsight hackathon, our team built OnCall Memory, an incident response agent for a fictional fintech, Northwind Pay. It analyzes a new alert and stack trace, then recommends a root cause and fix. Its main feature is that it remembers every past incident, what fixed it, and what made things worse.&lt;br&gt;
The problem with stateless assistants&lt;br&gt;
A general-purpose LLM can read a stack trace and suggest that a connection pool is exhausted. That advice is generic. It doesn't know that your team's checkout-service leaks connections from Celery tasks, or that rebooting the RDS instance last spring caused an 85-minute thundering-herd outage. Without memory, an assistant gives the same answer on the twentieth incident as on the first.&lt;br&gt;
Architecture&lt;br&gt;
OnCall Memory is a FastAPI backend with a single-page frontend. When an engineer submits an alert, the backend runs the same analysis twice, side by side:&lt;br&gt;
Without memory: the LLM sees only the alert and logs.&lt;br&gt;
With memory: the backend first queries Hindsight for relevant past incidents and passes them to the LLM as context.&lt;br&gt;
The UI shows both answers next to each other, so the value of memory is visible in seconds. We added a Demo Mode that walks through three scenarios: an initial alert, the same root cause under different symptoms, and a case where the agent warns against a fix that failed before.&lt;br&gt;
How we use Hindsight&lt;br&gt;
Hindsight gives an agent three operations, and we used all three.&lt;br&gt;
Retain. We seeded a memory bank named northwind-oncall with 25 realistic incidents. Each includes a date, affected service, error logs, root cause, fix steps, time to resolve, and whether the fix worked. Hindsight breaks these into smaller memories, so 25 incidents became about 90 entries. Retain is also the learning loop. After an incident, the engineer marks the fix as worked or failed and describes what happened. That text is retained, and the memory counter in the UI goes up. The next similar alert can use it immediately.&lt;br&gt;
Recall. When a new alert arrives, we query the bank with the alert text and logs. Hindsight returns the most relevant memories, which we show in the UI as inspectable chunks with their source incident. Showing them matters. A judge, or a skeptical engineer, can see which past incident drove the recommendation.&lt;br&gt;
Reflect. The Systemic Insights button uses reflect across the whole bank. Instead of matching one alert, it asks what patterns keep causing incidents. In our data it found recurring database connection exhaustion, Redis memory problems from missing TTLs, Kafka consumer rebalance loops, and failures from automated infrastructure changes. It also proposed permanent fixes such as connection proxies, TTL enforcement, and pre-deploy validation.&lt;br&gt;
What memory changes&lt;br&gt;
The difference shows up over several interactions. In the first, the agent gives a reasonable diagnosis backed by similar incidents. In the second, the symptoms differ, but recall links the alert to the same underlying cause. By the third, the agent has learned from a failure and tells the engineer not to reboot the database, because that made things worse last time. The recommendation improves because the agent accumulated experience, not because the prompt got better.&lt;br&gt;
Design choices&lt;br&gt;
Visible memory. Recalled chunks are shown, not hidden, because trust in an incident tool depends on evidence.&lt;br&gt;
Failed fixes count. Remembering what didn't work is as valuable as remembering what did.&lt;br&gt;
Realistic data. Incident IDs, log lines and numbers look like real ones, which makes the demo believable.&lt;br&gt;
Where it could go&lt;br&gt;
A production version could ingest PagerDuty alerts and postmortems automatically, run one memory bank per team, and track which fixes worked over time.&lt;br&gt;
OnCall Memory shows that the gap between a generic assistant and a useful one is often memory. Hindsight made adding it a matter of three well-designed operations.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbcz81g3e6slc107dqq63.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbcz81g3e6slc107dqq63.jpeg" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;br&gt;
Code:&lt;a href="https://github.com/Divyasri03-ux/oncall-memory" rel="noopener noreferrer"&gt;https://github.com/Divyasri03-ux/oncall-memory&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>webdev</category>
      <category>ai</category>
    </item>
    <item>
      <title>OnCall Memory:Building an Incident Response Agent That Learns From Production</title>
      <dc:creator>saanvii248</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:39:28 +0000</pubDate>
      <link>https://dev.to/saanvii248/oncall-memorybuilding-an-incident-response-agent-that-learns-from-production-401m</link>
      <guid>https://dev.to/saanvii248/oncall-memorybuilding-an-incident-response-agent-that-learns-from-production-401m</guid>
      <description>&lt;p&gt;OnCall Memory: Building an Incident Response Agent That Learns From Production History&lt;/p&gt;

&lt;p&gt;Production incidents are rarely completely new.&lt;/p&gt;

&lt;p&gt;A database connection pool can become exhausted again. A deployment configuration can break a service again. A payment provider can become rate-limited again. The symptoms may change, but the underlying patterns often repeat.&lt;/p&gt;

&lt;p&gt;The problem is that a typical AI assistant starts each incident with very little knowledge of an organization's previous experiences.&lt;/p&gt;

&lt;p&gt;It can analyze the logs in front of it and provide technically reasonable suggestions, but it does not automatically know what happened during the last incident, which fix worked, or which attempted solution failed.&lt;/p&gt;

&lt;p&gt;That is the problem I wanted to address with OnCall Memory, an incident-response website designed around persistent AI memory.&lt;/p&gt;

&lt;p&gt;The Idea:&lt;/p&gt;

&lt;p&gt;OnCall Memory is built for a fictional fintech company called Northwind Pay.&lt;br&gt;
When an engineer receives a production alert, the website provides two perspectives:&lt;/p&gt;

&lt;p&gt;-&amp;gt;Without Memory a response generated from the current incident.&lt;/p&gt;

&lt;p&gt;-&amp;gt;With Hindsight Memory a response generated after retrieving relevant historical incidents from the organization's memory.&lt;/p&gt;

&lt;p&gt;This side-by-side comparison is the central idea of the website.&lt;/p&gt;

&lt;p&gt;The goal isn't simply to make an AI response longer. It is to give the model access to something it normally doesn't have: the team's previous incident experience.&lt;/p&gt;

&lt;p&gt;Building the incident history&lt;/p&gt;

&lt;p&gt;The project contains a collection of realistic Northwind Pay production incidents.&lt;/p&gt;

&lt;p&gt;The incident dataset includes situations such as PostgreSQL connection pool exhaustion, Redis eviction, expired TLS certificates, bad deployment configuration, Kafka consumer lag, OOMKilled pods, and third-party payment API rate limiting.&lt;/p&gt;

&lt;p&gt;Each incident contains information such as the symptoms, logs, root cause, resolution steps, outcome, and whether attempted fixes worked.&lt;/p&gt;

&lt;p&gt;This is important because a useful memory isn't just:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"There was a database incident."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should capture the experience surrounding that incident.&lt;/p&gt;

&lt;p&gt;What happened?&lt;/p&gt;

&lt;p&gt;Why did it happen?&lt;/p&gt;

&lt;p&gt;What was tried?&lt;/p&gt;

&lt;p&gt;What actually fixed it?&lt;/p&gt;

&lt;p&gt;Did the first solution fail?&lt;/p&gt;

&lt;p&gt;That information becomes useful context for future incidents.&lt;/p&gt;

&lt;p&gt;Hindsight as the memory layer&lt;/p&gt;

&lt;p&gt;Hindsight is at the center of the architecture.&lt;/p&gt;

&lt;p&gt;The website uses Hindsight to perform three important operations: retain, recall, and reflect.&lt;/p&gt;

&lt;p&gt;When useful incident information is available, it can be retained as memory.&lt;/p&gt;

&lt;p&gt;When a new incident arrives, the system recalls relevant historical incidents before generating the memory-enhanced diagnosis.&lt;/p&gt;

&lt;p&gt;Finally, the website can use reflection to look across the accumulated incident history and ask a broader question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What patterns keep causing our incidents and what should we fix permanently?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This moves the system beyond responding to individual alerts and toward learning from patterns across incidents.&lt;/p&gt;

&lt;p&gt;The architecture&lt;/p&gt;

&lt;p&gt;The website uses a lightweight architecture.&lt;/p&gt;

&lt;p&gt;The frontend is a single-page HTML, CSS, and JavaScript dashboard.&lt;/p&gt;

&lt;p&gt;The backend is implemented with FastAPI and handles the incident analysis workflow, memory operations, and language-model requests.&lt;/p&gt;

&lt;p&gt;Groq provides the language-model layer, while Hindsight provides persistent memory.&lt;/p&gt;

&lt;p&gt;The project also separates configuration from application code. API credentials are kept locally in environment variables rather than being hardcoded into the source code, while &lt;code&gt;.env.example&lt;/code&gt; provides the required configuration structure.&lt;/p&gt;

&lt;p&gt;Memory changes the workflow&lt;/p&gt;

&lt;p&gt;Suppose a new payment failure arrives.&lt;/p&gt;

&lt;p&gt;Without historical context, an AI assistant might suggest checking the payment provider, network connectivity, credentials, rate limits, retries, and application logs.&lt;/p&gt;

&lt;p&gt;Those are reasonable suggestions, but the engineer still has to determine which one matches the organization's previous experience.&lt;/p&gt;

&lt;p&gt;With Hindsight memory, the workflow becomes different.&lt;/p&gt;

&lt;p&gt;The system first looks for similar historical incidents.&lt;/p&gt;

&lt;p&gt;If a previous payment incident involved third-party API rate limiting, for example, that historical information can become part of the diagnosis.&lt;/p&gt;

&lt;p&gt;The engineer can also see the recalled incidents rather than receiving a completely opaque recommendation.&lt;/p&gt;

&lt;p&gt;That visibility was an important part of the design.&lt;/p&gt;

&lt;p&gt;Learning doesn't stop after the first answer&lt;/p&gt;

&lt;p&gt;OnCall Memory also includes a feedback loop.&lt;/p&gt;

&lt;p&gt;After applying a suggested fix, the engineer can indicate whether the fix worked or failed and provide the actual solution.&lt;/p&gt;

&lt;p&gt;That information can then become another memory.&lt;/p&gt;

&lt;p&gt;The intended loop is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Incident → Recall → Diagnosis → Fix → Feedback → Retain → Better future diagnosis&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This creates a system where future incidents can benefit from experiences that were not available when the original system was built.&lt;/p&gt;

&lt;p&gt;The current Hindsight memory bank contains &lt;strong&gt;139 memories&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I learned
&lt;/h3&gt;

&lt;p&gt;One of the biggest lessons from building the website was that memory quality matters as much as memory quantity.&lt;/p&gt;

&lt;p&gt;Adding more memories does not automatically create a better assistant.&lt;/p&gt;

&lt;p&gt;A useful incident memory needs meaningful context: symptoms, root cause, attempted fixes, final resolution, and outcome.&lt;/p&gt;

&lt;p&gt;I also learned that memory should be visible to the user.&lt;/p&gt;

&lt;p&gt;If an AI gives a very specific operational recommendation, an engineer should have some way to understand why that information was relevant.&lt;/p&gt;

&lt;p&gt;Displaying recalled incidents makes the memory layer much easier to inspect.&lt;/p&gt;

&lt;p&gt;Limitations&lt;/p&gt;

&lt;p&gt;OnCall Memory currently uses a fictional Northwind Pay incident history rather than real production data.&lt;/p&gt;

&lt;p&gt;That means the system demonstrates the workflow but should not be interpreted as a replacement for an organization's actual incident-management practices.&lt;/p&gt;

&lt;p&gt;The quality of the recommendations also depends on the quality of the stored memories. Poor or incomplete incident records can lead to less useful historical context.&lt;/p&gt;

&lt;p&gt;Final thoughts&lt;/p&gt;

&lt;p&gt;The central idea behind OnCall Memory is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An incident-response assistant should not forget what happened yesterday.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Language models are good at reasoning about the information they receive. Persistent memory adds another dimension: the ability to use accumulated organizational experience.&lt;/p&gt;

&lt;p&gt;By combining Groq for reasoning with Hindsight for persistent memory, OnCall Memory turns incident history into a resource that can be recalled during future incidents and analyzed for recurring patterns.&lt;/p&gt;

&lt;p&gt;The result is not an AI replacing an on-call engineer.&lt;/p&gt;

&lt;p&gt;It is an AI assistant that can increasingly say:&lt;br&gt;
"We've seen something like this before."&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2nlu1j8j11037vuj6cf.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2nlu1j8j11037vuj6cf.jpeg" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
