<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Archana Doddi</title>
    <description>The latest articles on DEV Community by Archana Doddi (@archana_doddi_00f72c49e22).</description>
    <link>https://dev.to/archana_doddi_00f72c49e22</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147386%2F04a675de-b624-47b4-ab15-5391bdca605c.png</url>
      <title>DEV Community: Archana Doddi</title>
      <link>https://dev.to/archana_doddi_00f72c49e22</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/archana_doddi_00f72c49e22"/>
    <language>en</language>
    <item>
      <title>Using Hindsight to Investigate Incidents Without Guessing</title>
      <dc:creator>Archana Doddi</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:15:08 +0000</pubDate>
      <link>https://dev.to/archana_doddi_00f72c49e22/using-hindsight-to-investigate-incidents-without-guessing-4ji8</link>
      <guid>https://dev.to/archana_doddi_00f72c49e22/using-hindsight-to-investigate-incidents-without-guessing-4ji8</guid>
      <description>&lt;p&gt;The alert looks familiar. The Payment API is timing out against its database, and I have a nagging feeling someone fixed this before. But that fix is in a closed ticket, a Slack thread, or one colleague's head, so the on-call engineer starts from scratch.&lt;/p&gt;

&lt;p&gt;That gap is why our team built RecallOps, an AI incident-response copilot. This article covers the investigation workflow: how it uses past incidents, how it treats uncertainty, and where the engineer stays in charge. I'm one of five people who built it, so when I describe a component, I'm describing the system our team built, not claiming I wrote every part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The incident: INC-017
&lt;/h2&gt;

&lt;p&gt;Here is the scenario from our demo data. INC-017 is a Payment API database timeout. The current evidence is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;98% connection utilization&lt;/li&gt;
&lt;li&gt;a 14.2% timeout rate&lt;/li&gt;
&lt;li&gt;a deployment 23 minutes earlier&lt;/li&gt;
&lt;li&gt;a connection timeout error&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Any one of these could have several explanations. The deployment might be a coincidence. High utilization might just be heavy traffic. The useful question isn't "what is the root cause?" but "have we seen this shape before?"&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjsqty46zriq4r5vlkl9g.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjsqty46zriq4r5vlkl9g.jpeg" alt="RecallOps incident workspace showing INC-017 during investigation" width="800" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screenshot 1: The incident workspace for INC-017 during investigation. It shows the starting point: the current evidence next to the recalled historical memory, so you can see what an engineer sees when the investigation begins.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why historical context matters
&lt;/h2&gt;

&lt;p&gt;Experienced engineers investigate faster because they pattern-match against incidents they've lived through. That knowledge usually stays personal. We wanted a memory layer with a clear contract, not a pile of old tickets stuffed into a prompt.&lt;/p&gt;

&lt;p&gt;Memory goes through Hindsight, which gives us two operations: &lt;strong&gt;retain&lt;/strong&gt;, which stores what we learned from an incident, and &lt;strong&gt;recall&lt;/strong&gt;, which brings back what's relevant to a new one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recalling INC-001
&lt;/h2&gt;

&lt;p&gt;When INC-017 is investigated, the backend exposes a recall endpoint. This is a simplified, representative example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@router.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/incidents/{incident_id}/recall&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Depends&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get_db&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;incident&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_incident_or_404&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;degraded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;MemoryUnavailable&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;db_fallback_recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;degraded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;matches&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;matches&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;degraded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;degraded&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For INC-017, recall returns &lt;strong&gt;INC-001&lt;/strong&gt; as the top match. INC-001 was also a Payment API database timeout. Its root cause was connection-pool exhaustion from a connection leak, and the resolution was to fix the leak and increase pool capacity. It had been resolved and retained, which is why it was available to recall.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "91% similar" means
&lt;/h2&gt;

&lt;p&gt;The match comes back at 91% similarity, and the response carries the reasons with it: same service, same service family, similar symptoms, similar database behavior, and similar timing.&lt;/p&gt;

&lt;p&gt;I care more about the reasons than the number. A bare score asks the engineer to trust it. A score with reasons lets them check whether the match makes sense. We haven't benchmarked the system, so I'm not making claims about how reliable similarity scores are in general. I'm describing what the demo shows and how it's presented.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcesi4csm3kpkwl5jkks.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcesi4csm3kpkwl5jkks.jpeg" alt="INC-001 historical memory with 91% similarity and match reasons" width="445" height="796"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screenshot 2: INC-001 as historical memory at 91% similarity. It shows the match reasons and the historical root cause displayed together, demonstrating that the score is explained rather than shown alone.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Current evidence vs. historical evidence
&lt;/h2&gt;

&lt;p&gt;The system keeps four things separate: current evidence, historical evidence, the AI recommendation, and uncertainty. For INC-017 the two evidence sets look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Current (INC-017)&lt;/th&gt;
&lt;th&gt;Historical (INC-001)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;98% connection utilization&lt;/td&gt;
&lt;td&gt;High connection utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14.2% timeout rate&lt;/td&gt;
&lt;td&gt;Same service and error family&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deployment 23 minutes ago&lt;/td&gt;
&lt;td&gt;Confirmed connection leak&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Connection timeout&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The overlap is real: high utilization, the same service, the same error family. But look at what's on each side. INC-001's leak was &lt;em&gt;confirmed&lt;/em&gt;. For INC-017, nothing has confirmed a leak yet.&lt;/p&gt;

&lt;p&gt;Keeping the sources separate means a recommendation can be traced back to specific past incidents, and each side can be reviewed and debugged on its own. Blending them into one blob would make the output harder to trust.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flc36nmhjlhzagxmfbf6m.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flc36nmhjlhzagxmfbf6m.jpeg" alt="Current vs Historical Evidence side by side" width="696" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screenshot 3: Current evidence and historical evidence side by side. It shows the separation of sources and lets you see exactly which observations overlap and which don't.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How the old root cause guides the investigation
&lt;/h2&gt;

&lt;p&gt;INC-001's connection leak becomes a lead, not a conclusion. The workspace offers structured investigation paths, and one of them fits the deployment 23 minutes before the alert: check whether the recent deployment introduced a new leak.&lt;/p&gt;

&lt;p&gt;The history narrows where to look first. It doesn't say what the answer is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendation vs. certainty
&lt;/h2&gt;

&lt;p&gt;A 91% match is worth investigating, but a similar past incident is not proof of the same root cause. Two incidents can look alike and fail for different reasons.&lt;/p&gt;

&lt;p&gt;So the UI labels INC-001 as &lt;strong&gt;evidence for investigation&lt;/strong&gt;, not confirmation of the current root cause. The system also surfaces uncertainty as its own element instead of folding it into a confident-sounding recommendation. If a copilot pretends to know, engineers either over-trust it or learn to ignore it, and neither helps at 3 a.m.&lt;/p&gt;

&lt;h2&gt;
  
  
  When memory is unavailable
&lt;/h2&gt;

&lt;p&gt;We couldn't assume a remote memory service would always be reachable. In the recall snippet above, if Hindsight fails, the endpoint catches &lt;code&gt;MemoryUnavailable&lt;/code&gt; and recalls from the system's own persisted incident data through &lt;code&gt;db_fallback_recall&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The important detail is the &lt;code&gt;degraded&lt;/code&gt; flag. The application keeps working, but the response says a fallback was used, so the UI can show a degraded state instead of silently passing off the fallback as the primary memory provider. Designing this early forced us to decide what "degraded" should look like to the person investigating.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the loop, and the engineer's role
&lt;/h2&gt;

&lt;p&gt;Once the engineer resolves INC-017, it can be retained so future incidents can recall it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resolved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Resolve the incident before retaining it&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;retained&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The check keeps unfinished investigations out of memory.&lt;/p&gt;

&lt;p&gt;RecallOps never changes production systems automatically. It provides evidence, history, recommendations, investigation paths, and uncertainty. The engineer makes the final call.&lt;/p&gt;

&lt;p&gt;After retention, a Learning area looks across retained incidents. On the demo data it identified a recurring Payment API pattern across five related incidents, with provenance showing which incidents each lesson came from.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fik9u9c3h1qagazn8el5v.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fik9u9c3h1qagazn8el5v.jpeg" alt="fig-4" width="799" height="265"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screenshot 4: The Learning page. It shows the payoff of retention: a recurring pattern across incidents, each lesson traced back to its source incidents.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What was and wasn't verified
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Verified:&lt;/strong&gt; backend startup, API health, backend tests, SQLite persistence, memory recall, retention, learning/reflection, and the browser end-to-end workflow (Launch Demo through INC-001, INC-017, recall, resolve, retain, Learning, and demo reset). The verified demo runs on SQLite with the deterministic AI fallback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not live-verified:&lt;/strong&gt; a production PostgreSQL deployment, a remote Hindsight service, and live Groq/OpenAI inference. They're configured integrations in the architecture, but we haven't tested them live, so I make no claims about them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this taught me
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Memory should explain itself.&lt;/strong&gt; Match reasons made the recall result something I could evaluate, not just accept.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A similar incident is a lead, not a verdict.&lt;/strong&gt; Separating evidence from recommendation from uncertainty keeps the tool honest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degraded modes are part of the design.&lt;/strong&gt; Deciding early how failure looks made the whole system more trustworthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The engineer stays the decision-maker.&lt;/strong&gt; The tool's value is shortening the path to the right question, not answering it for you.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Our next step is validating the remote integrations so the configured architecture gets tested too. If you're interested in agent memory, the Hindsight repository is a good place to start.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
      <category>fastapi</category>
    </item>
  </channel>
</rss>
