<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chunduru Rushi kumar</title>
    <description>The latest articles on DEV Community by Chunduru Rushi kumar (@rushikumar-06).</description>
    <link>https://dev.to/rushikumar-06</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149576%2F44e91318-6ef1-4a28-a95a-4cbef612f716.png</url>
      <title>DEV Community: Chunduru Rushi kumar</title>
      <link>https://dev.to/rushikumar-06</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rushikumar-06"/>
    <language>en</language>
    <item>
      <title>I gave my on-call agent Hindsight memory. It started anchoring.</title>
      <dc:creator>Chunduru Rushi kumar</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:10:17 +0000</pubDate>
      <link>https://dev.to/rushikumar-06/i-gave-my-on-call-agent-hindsight-memory-it-started-anchoring-4h5a</link>
      <guid>https://dev.to/rushikumar-06/i-gave-my-on-call-agent-hindsight-memory-it-started-anchoring-4h5a</guid>
      <description>&lt;p&gt;The first time I measured my on-call agent with memory against the same model without it, memory lost. Replayed&lt;br&gt;
over six months of incidents, the agent that could remember every postmortem found the root cause 50% of the&lt;br&gt;
time. The one that remembered nothing got 61%.&lt;/p&gt;

&lt;p&gt;This is the story of why that happened, and what it took to build an agent where memory makes the plan better&lt;br&gt;
without making the diagnosis worse.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;Déjà Vu is an incident-response agent for the engineer holding the pager. When an alert fires, it recalls what&lt;br&gt;
happened the last time this service broke and writes a plan: what to do, what &lt;strong&gt;not&lt;/strong&gt; to do (with the incident&lt;br&gt;
where it backfired), and who fixed it before. When the incident is over, the engineer marks each suggested step as&lt;br&gt;
&lt;em&gt;worked&lt;/em&gt;, &lt;em&gt;no effect&lt;/em&gt; or &lt;em&gt;made it worse&lt;/em&gt;, and that verdict becomes memory for the next page.&lt;/p&gt;

&lt;p&gt;The memory layer is &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight, the open-source agent memory system from Vectorize&lt;/a&gt;.&lt;br&gt;
I didn't want a vector store I'd have to babysit. I wanted something that behaves like a colleague who reads every&lt;br&gt;
postmortem: it extracts facts, links entities, consolidates repeated evidence into beliefs, and answers questions&lt;br&gt;
across all of it. If you're new to the idea, Vectorize has a good explainer on&lt;br&gt;
&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;what agent memory is and why it differs from retrieval&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Every postmortem goes in with its real timestamp, a stable document id and service tags, so later recalls can be&lt;br&gt;
scoped and cited:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retain_incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;aretain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;postmortem_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident postmortem for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;timestamp&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromisoformat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;started_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
        &lt;span class="n"&gt;document_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind:postmortem&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;service_tag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Team rules ("rollbacks must use the exact &lt;code&gt;argocd app rollback&lt;/code&gt; command") and engineer verdicts go in the same way.&lt;br&gt;
Each service also gets a Hindsight mental model: a standing question ("how does checkout-api fail, what worked, what&lt;br&gt;
made it worse?") that Hindsight re-answers after new memories consolidate. Those became living runbooks nobody&lt;br&gt;
maintains by hand. The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; covers retain, recall, reflect and&lt;br&gt;
mental models in detail.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8f080hc09q679wre3i6b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8f080hc09q679wre3i6b.png" alt="Where Hindsight sits in Déjà Vu" width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  How I measured it
&lt;/h2&gt;

&lt;p&gt;A demo where the agent says the right thing once proves nothing, so I built a replay. The history is six months of a&lt;br&gt;
(fictional) UPI and card payments gateway: 23 incidents with real-looking HikariCP, pgbouncer, Kafka and CoreDNS log&lt;br&gt;
lines, and every action someone tried along with its outcome. Six failure modes recur, usually with a &lt;em&gt;different&lt;br&gt;
trigger&lt;/em&gt; each time. Connection-pool exhaustion shows up four times, caused by a missing index, a migration lock, a&lt;br&gt;
connection leak and autoscaling. Eight incidents are one-offs, and the first occurrence of each recurring failure is new too. Those 14&lt;br&gt;
first-time incidents are the control group.&lt;/p&gt;

&lt;p&gt;The replay walks the history in date order into a fresh memory bank. Before each incident is retained, two agents&lt;br&gt;
triage it blind: the same model (&lt;code&gt;gpt-oss-120b&lt;/code&gt;) with the same instructions, one with memory and one without. A&lt;br&gt;
grader model compares both plans with the real postmortem. Because an incident is only retained after it has been&lt;br&gt;
graded, nothing is ever in memory before it happens.&lt;/p&gt;
&lt;h2&gt;
  
  
  Failure one: memory as answers
&lt;/h2&gt;

&lt;p&gt;My first prompt said, in effect, "trust memory over generic advice." The memory agent recognised every repeat&lt;br&gt;
incident (9 of 9) and halved harmful suggestions. Its diagnosis was also worse than the agent with no memory at all.&lt;/p&gt;

&lt;p&gt;The clearest case was a ledger consumer falling behind. The logs said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;WARN consumer poll timeout has expired ... longer than the configured max.poll.interval.ms
ledger-worker: fx_rates.get(currency='USD') took 4212ms (timeout 5000ms)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A four-second FX lookup inside the poll loop. The memory agent answered "max.poll.records is too high", because&lt;br&gt;
that had been the cause the last time this service lagged. Same symptom, new trigger, wrong answer. It's exactly how&lt;br&gt;
a mediocre senior engineer behaves: pattern-match first, read second.&lt;/p&gt;

&lt;p&gt;The fix was to treat memory as hypotheses. The plan must now quote the evidence in &lt;em&gt;this&lt;/em&gt; alert, mark for every&lt;br&gt;
past incident it matches whether that incident's trigger is actually present (&lt;code&gt;same_trigger&lt;/code&gt;), and say what is&lt;br&gt;
different this time. In the UI that became a distinct verdict: &lt;em&gt;"Seen this symptom before, but the trigger is new."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fivk3g0nd8dy1tk3aoi0d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fivk3g0nd8dy1tk3aoi0d.png" alt="Same symptom, new trigger" width="554" height="2050"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Failure two: the same thing, more subtly
&lt;/h2&gt;

&lt;p&gt;That helped, but the replay caught it again. A checkout alert carried a stack trace from HikariCP's leak detector,&lt;br&gt;
pointing at &lt;code&gt;DbRetryTemplate.execute&lt;/code&gt;. The agent without memory read it correctly: a new retry path wasn't returning&lt;br&gt;
connections. The agent with memory blamed pgbouncer's connection limit. pgbouncer appears nowhere in that alert. It&lt;br&gt;
came from two past pool-exhaustion incidents that did involve pgbouncer.&lt;/p&gt;

&lt;p&gt;So I changed the order of operations. Déjà Vu now reads the evidence first, with no memory at all, and memory has to&lt;br&gt;
earn the right to change that diagnosis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;FIRST_READ&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Your own first read of this alert, made from the evidence alone before looking at memory:
- evidence: {evidence}
- likely root cause: {root}
It is the working diagnosis. Change the root cause only if a past incident&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s own signature (its log line, component
or metric) appears in this alert and explains the evidence better. Either way, use memory for everything else below.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the side-by-side view this costs nothing extra: the memory lane reuses the no-memory lane's read, then checks it&lt;br&gt;
against what Hindsight recalls. Run alone, it makes its own read while the recalls are in flight.&lt;/p&gt;

&lt;h2&gt;
  
  
  What memory is actually good for
&lt;/h2&gt;

&lt;p&gt;Here is the final replay, graded blind:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;With memory&lt;/th&gt;
&lt;th&gt;Without&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Root cause right, first-time incidents (control)&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Root cause right, repeat incidents&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;td&gt;72%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan included the fix that actually worked, repeats&lt;/td&gt;
&lt;td&gt;6 / 9&lt;/td&gt;
&lt;td&gt;4 / 9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan included the fix that actually worked, all 23&lt;/td&gt;
&lt;td&gt;14 / 23&lt;/td&gt;
&lt;td&gt;9 / 23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recommended something that backfired last time&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The control row is the one I care about most. On incidents with no history, memory neither helps nor hurts, which&lt;br&gt;
is what evidence-first is for. Diagnosis on repeats is only a little better, and that's honest: both agents read the&lt;br&gt;
same logs. The real gap is operational knowledge. With memory, the plan contains the fix that worked last time and&lt;br&gt;
the rollback in the form this team uses, and it knows what not to touch.&lt;/p&gt;

&lt;p&gt;On the festive-sale alert (26 pods, &lt;code&gt;no more connections allowed (max_client_conn)&lt;/code&gt;), the agent without memory&lt;br&gt;
suggests raising the pgbouncer limit and resizing the pool. Sensible and generic. Déjà Vu says &lt;em&gt;"Déjà vu: we have&lt;br&gt;
seen this before"&lt;/em&gt;, links INC-2107, applies the fix that worked there, rolls back with the exact Argo CD command the&lt;br&gt;
team mandated, warns against restarting pgbouncer (it did nothing in INC-2014) and pages the engineer who fixed it&lt;br&gt;
last time. The answer without memory arrives in about four seconds; Déjà Vu's in about eleven.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqf31oty5t12by21mb5v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqf31oty5t12by21mb5v.png" alt="The same alert, with and without memory" width="800" height="1381"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The grader was part of the system too
&lt;/h2&gt;

&lt;p&gt;One thing surprised me. My first grader scored each plan on its own, and it gave near-identical answers different&lt;br&gt;
scores: 1.0 in one lane, 0.5 in the other, for the same root cause. Now one call grades both plans side by side,&lt;br&gt;
blind and in a fixed random order, at temperature 0, with the explicit rule that equivalent answers get equal grades.&lt;br&gt;
The gap between the lanes shrank, and I trust it more.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Memory makes anchoring the default.&lt;/strong&gt; Retrieval plus "trust your memory" produces an agent that diagnoses the
last incident instead of this one. Make it read the evidence first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run a control group.&lt;/strong&gt; If memory "helps" on incidents it has never seen, something is leaking. The first-time
incidents told me when a change was real.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure the thing memory is for.&lt;/strong&gt; Root-cause accuracy barely moved. "Did the plan contain the fix that worked?"
and "did it avoid what backfired?" moved a lot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The grader is code you have to debug.&lt;/strong&gt; Grade both answers in one blind call, or you'll optimise for noise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the loop honest.&lt;/strong&gt; Each fix here came from reading where the previous version missed, on the same 23
incidents. That's fine for learning; just don't call the result a benchmark.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The most useful thing an on-call agent can say isn't "I know what this is." It's "Don't do that. It made this&lt;br&gt;
worse on 22 April, and here's who fixed it."&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Code: &lt;a href="https://github.com/Rushikumar-06/dejavu" rel="noopener noreferrer"&gt;https://github.com/Rushikumar-06/dejavu&lt;/a&gt; · Built with the Déjà Vu team · Thanks to @Code.in&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>python</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
