<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sahasra Patwari</title>
    <description>The latest articles on DEV Community by Sahasra Patwari (@sahasra_p).</description>
    <link>https://dev.to/sahasra_p</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148292%2Fe049fe78-3149-4d04-b717-db242ffbf643.png</url>
      <title>DEV Community: Sahasra Patwari</title>
      <link>https://dev.to/sahasra_p</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sahasra_p"/>
    <language>en</language>
    <item>
      <title>Single Bank or Many? A Hindsight Design Decision I Almost Got Wrong</title>
      <dc:creator>Sahasra Patwari</dc:creator>
      <pubDate>Tue, 29 Sep 2026 01:49:15 +0000</pubDate>
      <link>https://dev.to/sahasra_p/single-bank-or-many-a-hindsight-design-decision-i-almost-got-wrong-49j6</link>
      <guid>https://dev.to/sahasra_p/single-bank-or-many-a-hindsight-design-decision-i-almost-got-wrong-49j6</guid>
      <description>&lt;p&gt;At 3 a.m., the person on call rarely needs more dashboards. They need someone who was there the last time this happened to lean over and say, "That's the connection pool. We saw it in August. Don't scale the pods, fix the pool."&lt;/p&gt;

&lt;p&gt;I built an incident response agent to be that colleague. The interesting part turned out not to be the LLM. It was one boring-sounding decision about how the agent's memory is laid out: &lt;strong&gt;one shared memory bank for the whole company, with every record tagged by service.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This post covers why I made that call, what it looks like in code, and what I'd tell someone else designing agent memory on [Hindsight]&lt;a href="https://dev.tourl"&gt;(https://github.com/vectorize-io/hindsight).&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmlwd2q6xa9sl4o6e1oj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmlwd2q6xa9sl4o6e1oj.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The RecallOps incident detail page: error signal and root cause on the left, suggested fix and Hindsight-recalled related incidents on the right.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the system does&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RecallOps is a FastAPI and React incident-response application. It stores incident records in SQLAlchemy, uses Hindsight as its operational memory layer, and uses a Groq-hosted model (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt;) to turn an incoming alert plus retrieved incident history into an evidence-backed briefing.&lt;/p&gt;

&lt;p&gt;The loop is simple. A new alert arrives, the agent recalls similar past incidents, and it answers with the relevant history cited. When an engineer marks an incident resolved, the incident's symptom, error signature, root cause, fix, and outcome get written back to memory. The next alert benefits from it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb9zfrmbmwn09tq2qv88s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb9zfrmbmwn09tq2qv88s.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Where Hindsight sits in the stack.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Most of the effort went into the middle layer, the part between "an incident happened" and "the agent remembered it usefully."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The decision: one bank, tagged by service&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Hindsight organises memory into banks. The obvious first design for a multi-service company is one bank per service: 'checkout-service ', 'payment-gateway', 'auth-service', 'inventory-service'. It feels tidy. Each team owns its history, and recall is naturally scoped.&lt;/p&gt;

&lt;p&gt;I didn't go that way because of how production incidents actually behave. The worst ones cross service boundaries. A checkout 502 might really be a payment-gateway problem. A spike of 401s on mobile might be a deployment issue that started in a different service. If each service's history sits in its own bank, the agent can only reason inside one silo at a time. The most valuable question, "have we seen something &lt;em&gt;shaped like this&lt;/em&gt; anywhere?", becomes N queries and a merge step that I'd have to write and maintain.&lt;/p&gt;

&lt;p&gt;So the whole company shares one bank, and each memory carries a tag for the service it came from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# src/memory/schema.py
&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;meridian-commerce-incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# src/memory/hindsight_tools.py
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Writes one resolved incident into memory, tagged by service. The tag
    is what lets recall_similar_incidents() later filter to just one
    service&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s history, while an untagged recall still searches everything.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;format_incident_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recall then has two modes from a single code path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall_similar_incidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;alert_description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;kwargs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bank_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;alert_description&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tags&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Leave &lt;code&gt;service&lt;/code&gt; empty and the search runs across everything. Pass &lt;code&gt;"checkout-service"&lt;/code&gt; and it's scoped. That gives me both behaviors from one bank. Hindsight's docs describe a single bank with tags as the option for applications that need reasoning across entities, while per-user banks are the more common pattern when users must be kept apart. My case is the first kind: incidents frequently cross service boundaries. You can always narrow with a tag filter later, but if you split banks up front, joining them afterward is much harder.&lt;/p&gt;

&lt;p&gt;The trade-off is real. Tags are a retrieval aid, not an authorization boundary, and Hindsight's default match mode includes untagged memories. If you have hard isolation requirements, such as different customers' data, separate banks are the right tool. The right boundary is whoever the on-call team is allowed to reason over, not a microservice name. Service history inside one company doesn't have that constraint, so I took the flexibility.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory quality is a formatting problem&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Hindsight stores free text and finds it later through semantic search. It doesn't impose a schema, which is convenient, but it means recall quality depends heavily on how consistently you write things down. Free-form postmortem prose retained as-is makes the search's job harder.&lt;/p&gt;

&lt;p&gt;So every incident goes through one template before it's retained:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;format_incident_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;symptom&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;error_signature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                           &lt;span class="n"&gt;root_cause&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolution_steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;runbook_used&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                           &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolved_by&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;INCIDENT &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; — &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Timestamp: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;timestamp&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Symptom: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;symptom&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Error signature: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;error_signature&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Root cause: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;root_cause&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Resolution steps taken: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;resolution_steps&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Runbook used: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runbook_used&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Outcome: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Resolved by: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;resolved_by&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every field is labeled, and that matters more than it looks. A new alert might resemble an old incident's symptom ("502s during peak traffic"), its error signature (&lt;code&gt;connection pool exhausted&lt;/code&gt;), or its root cause. Labeled fields give the retrieval step something to latch onto whichever way the new alert happens to be phrased. I treat this template as part of the API. Changing it is a migration, not a refactor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Letting the model decide when to remember&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The agent is a small tool-calling loop. Two functions are exposed to the model: &lt;code&gt;recall_similar_incidents&lt;/code&gt; and &lt;code&gt;log_incident&lt;/code&gt;. The system prompt is deliberately narrow about when each applies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are an incident response assistant.
When given a new alert, decide whether you need to recall similar past
incidents from memory before answering. Use the recall_similar_incidents
tool when historical context would help diagnose the issue. Use
log_incident only when explicitly told an incident has just been
resolved and needs to be recorded. Always give a clear, actionable
answer citing what you found in memory, if anything.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The write path is gated on purpose. I don't want a model deciding on its own that a half-understood outage is "resolved" and writing a wrong root cause into long-term memory. Memory that's wrong is worse than memory that's empty, because it gets cited with confidence.&lt;/p&gt;

&lt;p&gt;Two smaller choices in the loop are worth copying. Tool errors go back to the model as text instead of raising, so the agent can say "memory lookup failed, here's what I can say without it" instead of the request dying with a 500 at the worst possible time. And the loop is capped at three tool hops. If the model can't get to an answer in three, a human is better off looking directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recall is retrieval, but the briefing is synthesis&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Hindsight exposes more than &lt;code&gt;retain&lt;/code&gt; and &lt;code&gt;recall&lt;/code&gt;. &lt;code&gt;reflect&lt;/code&gt; synthesizes an answer from relevant memories in the bank, while &lt;code&gt;recall&lt;/code&gt; returns ranked memory results the agent can expose as evidence. I use both, for different jobs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;recall&lt;/code&gt; returns the top matching incident records verbatim. The agent cites these, and the engineer can click through to the source.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;reflect&lt;/code&gt; (wrapped as &lt;code&gt;get_incident_briefing&lt;/code&gt;) produces a readable summary of what memory knows about an alert.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I keep them separate because they answer different questions. "Show me the evidence" and "tell me what it adds up to" shouldn't be the same call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mental models: the part that compounds&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The feature that changed how I think about the system is Hindsight's mental models. A mental model is a saved &lt;code&gt;reflect&lt;/code&gt; answer attached to a bank that you can refresh as memory grows. I created one team-wide model, "Team-wide Production Reliability Profile," from a single source query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SOURCE_QUERY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;What recurring production failure patterns has our operations team learned from approved runbooks and resolved incidents?

For each pattern, identify the services and dependencies involved, early warning signals, verified root causes, safe mitigations, risky actions to avoid, and the first checks an on-call engineer should perform. Prioritize repeated, high-severity, cross-service incidents and clearly label uncertain evidence.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that the query asks for &lt;em&gt;risky actions to avoid&lt;/em&gt; and asks the model to &lt;em&gt;label uncertain evidence&lt;/em&gt;. Those two phrases do more for on-call usefulness than any prompt tuning I did on the agent itself. The model's output for each pattern comes back as the same shape: services and dependencies, early warning signals, verified root cause, safe mitigations, risky actions, first checks, and the matching runbook.&lt;/p&gt;

&lt;p&gt;Refreshing looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;refresh_team_mental_model&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;find_team_mental_model&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;refresh_mental_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mental_model_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;operation_id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Refresh is asynchronous, and reading the content right after triggering it can return the previous version. So the read path polls &lt;code&gt;last_refreshed_at&lt;/code&gt; until it changes. That's a small detail, and it was the one that bit me first when I built a page that showed stale content right after a refresh.&lt;/p&gt;

&lt;p&gt;I refresh this model on a schedule I control during development, but in production I'd rather not run a cron. Hindsight supports &lt;code&gt;trigger={"refresh_after_consolidation": True}&lt;/code&gt;, which re-runs the model when the bank consolidates new memories. With that trigger, the reliability profile can refresh after newly retained incidents are consolidated, so nobody has to remember to update the wiki. Because the refresh is asynchronous, the UI should show when the profile was last refreshed instead of pretending it changed instantly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it does with a real alert&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's how the pieces behave against the historical incidents in the bank. These four cover a connection-pool exhaustion on checkout, payment confirmations timing out after retry bursts got the company rate-limited by its processor, a JWT key rotation that only reached two of six pods, and repeated OOM kills during a nightly inventory sync.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Alert:&lt;/strong&gt; &lt;code&gt;checkout-service returning HTTP 502 errors, latency spiking on database calls&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The agent chooses to call &lt;code&gt;recall_similar_incidents&lt;/code&gt;, scoped to &lt;code&gt;checkout-service&lt;/code&gt;. The matching record is the earlier incident where Postgres connections were exhausted during a flash-sale burst. Its retained fields include the error signature (&lt;code&gt;connection pool exhausted&lt;/code&gt;), the root cause (pool sized for average load, not burst traffic), the fix (larger pool plus a circuit breaker to fail fast), and the runbook used. The agent's answer points the engineer at connection pool metrics first, and names the risky move: adding capacity without checking database limits.&lt;/p&gt;

&lt;p&gt;A payment-timeout alert run without a service filter shows what the shared bank is for. Recall can surface the earlier payment incident, and the mental model already contains a pattern about retries without backoff triggering rate limiting, so the suggestion is to check external API error logs for rate-limit responses first, not to retry harder.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkrt60bps1w2g5eq3c20.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkrt60bps1w2g5eq3c20.png" alt=" " width="800" height="432"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Seeding the bank and the labeled record Hindsight retains for a checkout-service incident.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'm not going to claim a percentage improvement in resolution time, because I haven't measured one. What I can say is that the answers cite specific past incidents, and an on-call engineer can verify every claim by opening the source record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lessons learned&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. If your entities need to see each other, start with one bank and tags.&lt;/strong&gt; Scope down with a filter when you need to. Splitting by entity up front makes the cross-entity questions, usually the interesting ones, expensive to answer later. If they must never see each other, use separate banks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Treat your memory format as a schema, even if the store doesn't.&lt;/strong&gt; Semantic search over inconsistent text gives inconsistent results. A labeled template gave me more recall quality than any other single change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Gate the write path.&lt;/strong&gt; Reads can be liberal and writes should be deliberate. A wrong memory is cited with confidence, so I let the model recall on its own, but only log an incident when it's told that one has been resolved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Separate "show evidence" from "synthesise."&lt;/strong&gt; &lt;code&gt;recall&lt;/code&gt; and &lt;code&gt;reflect&lt;/code&gt; do different jobs. Keeping them as different calls kept the agent's citations honest and the briefings readable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Write the mental-model query like a runbook template.&lt;/strong&gt; Asking explicitly for early warning signals, risky actions, first checks, and uncertainty labels shaped the output more than anything else I tried. Let the refresh trigger drive updates in production instead of relying on someone to press a button.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where to go from here&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you're building something similar, start with the [Hindsight documentation]&lt;a href="https://dev.tourl"&gt;(https://hindsight.vectorize.io/)&lt;/a&gt; for the retain, recall, reflect, and mental model APIs, and read the [Vectorize overview of agent memory]&lt;a href="https://dev.tourl"&gt;(https://vectorize.io/what-is-agent-memory)&lt;/a&gt; for how these pieces fit together conceptually. The Hindsight repository on GitHub is where the code lives.&lt;/p&gt;

&lt;p&gt;My advice: keep the memory layer small, make retained records structurally consistent, and choose bank boundaries based on trust and access, not your first guess at system topology. For one company's production incident history, cross-service visibility is the feature.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
