DEV Community

Cover image for How I Gave an SRE Agent Memory with Hindsight
Mandha Varshitha
Mandha Varshitha

Posted on

How I Gave an SRE Agent Memory with Hindsight

payments-api is returning HTTP 503s. The logs say active=50 idle=0 waiting=317, and the database connection pool is exhausted. Ask a language model what to do and you get sensible, generic advice: raise the pool size, or restart the service.

Those two answers are not equivalent. In the incident history I seeded for IncidentIQ, a restart on payments-api (INC-103) cleared the connections for a moment and the exhaustion came back, because long-running queries were still holding them. Raising the pool worked twice (INC-101 and INC-102). A model reading only the current alert cannot know any of that.

That is the problem I set out to solve: the agent needs to know what happened last time, whether it worked, and whether it stuck.

What I Built

IncidentIQ is an incident-triage agent with four parts: a React/TypeScript frontend, a FastAPI backend, Hindsight for persistent memory, and Groq (openai/gpt-oss-120b) for reasoning. An engineer submits a service, severity, alert and logs. The backend recalls past incidents from Hindsight, hands the relevant ones to the model, and returns ranked recommendations. The engineer marks each recommendation Worked or Failed, and that verdict is written back to Hindsight.

I did not want the model starting every incident from zero. I also did not want it inventing history. So the model reasons about the incident, and the backend owns everything that claims to be historical fact.

If you want background on the idea, these are the resources I used:

Hindsight on GitHub

Hindsight Documentation

Vectorize: What is Agent Memory?

This article is about what it took to make the memory useful.

How the System Works

React frontend
| POST /api/triage
v
FastAPI
|-- search_memory() ------> Hindsight recall
|-- filter_memories_by_service()
|-- extract_historical_facts() -> calculate_action_statistics()
|-- analyze_incident() ---> Groq (JSON-schema output)
|-- overwrite success rate + evidence IDs with backend values
v
Ranked recommendations
| engineer clicks Worked / Failed
v
POST /api/outcome -> store_memory() -> Hindsight retain

The backend has three endpoints: /api/triage, /api/predeploy and /api/outcome. Triage and pre-deploy both take a memory_enabled flag. When it is false, Hindsight is never called.

The Incident That Made Memory Useful

Before memory. With memory_enabled set to false, the recalled list is empty and build_memory_text gives the model this instead of history: no relevant memories were found, do not claim previous incidents existed, do not invent success rates or evidence IDs. The model can still diagnose from the logs alone. Pool exhaustion, undersized pool, raise it. What it cannot produce is any statement about what has been tried before. The backend also blanks each recommendation's historical_success_rate and evidence_ids, so the UI has nothing to show.

With memory. The seed data (Backend/seed_incidents.py) contains four payments-api incidents. INC-101 and INC-102 were pool exhaustion resolved by increasing the pool. INC-103 was pool exhaustion where restarting the service failed, with the root cause recorded as slow queries keeping connections occupied. INC-104 was high database latency resolved by terminating long-running queries. Now the model sees a successful fix, a failed fix and the reason it failed, and it can rank accordingly: pool increase first, long-running query investigation as the root-cause follow-up, and no confident case for a restart.

The screenshot below is a live run against my Hindsight bank. Note that the bank also held older test memories, so it recalled INC-001, INC-002 and INC-102 rather than only the current seed IDs. It shows 16 memories recalled, the pool increase ranked first, and evidence IDs attached.

Look at the numbers on that card, though: 100% historical success with a sample size of 1. I come back to that below.

How I Integrated Hindsight

The whole Hindsight surface is two functions in Backend/app/hindsight_service.py:

def store_memory(content: str):
return client.retain(
bank_id=HINDSIGHT_BANK_ID,
content=content
)

def search_memory(query: str):
return client.recall(
bank_id=HINDSIGHT_BANK_ID,
query=query
)

Everything else is about what to do with what comes back. Recall is semantic, so it returns memories about other services too. I filter them in main.py:

def memory_matches_service(
memory,
service: str,
) -> bool:

service = service.strip().lower()

if not service:
    return False

memory_text = (
    getattr(memory, "text", "")
    or ""
)

memory_text_lower = memory_text.lower()

# Service: payments-api
service_pattern = (
    rf"\bservice\s*:\s*"
    rf"{re.escape(service)}\b"
)

if re.search(
    service_pattern,
    memory_text_lower,
):
    return True
Enter fullscreen mode Exit fullscreen mode

Two more patterns follow, for "the payments-api service" and a bare service name. This is text matching, not a structured filter. It works because I control the memory format, and it is the first thing I would replace.

Historical Evidence Is Computed Outside the Model

This is the design decision I care most about. The LLM returns recommendations that include a historical_success_rate and evidence_ids, because the response schema has those fields. In /api/triage I throw the model's values away and recompute them:

for recommendation in (
    analysis.recommendations
):

    recommendation.historical_success_rate = None
    recommendation.evidence_ids = []

    if not request.memory_enabled:
        continue
Enter fullscreen mode Exit fullscreen mode

Further down, each recommendation is matched to a bucket in action_statistics and takes its historical_success_rate and evidence_ids from there. Those statistics come from extract_historical_facts and calculate_action_statistics in incident_analysis.py. The first pulls out (incident, action, outcome) facts from recalled text. The second counts each incident once per action and divides successful incidents by attempts, keeping the Hindsight memory IDs as evidence.

So a "100%" is never something the model said. It is a number the backend counted, with memory IDs a human can open.

The Outcome Loop

/api/outcome is how failures become memory. The frontend sends the incident ID, service, action and outcome, and the backend writes a structured memory:

outcome_memory = f"""
Enter fullscreen mode Exit fullscreen mode

Incident {request.incident_id} affected the
{request.service} service.

Resolution attempt:
{request.action}

Outcome:
{outcome_description}

Engineer notes:
{request.notes}

Interpretation:
The engineer marked this resolution attempt as
{outcome_description}.
"""

result = store_memory(
    outcome_memory
)
Enter fullscreen mode Exit fullscreen mode

worked is stored as successful, failed as failed, and partial as temporary. The loop is: incident, recommendation, engineer verdict, retained memory, next incident. It is engineer-in-the-loop feedback, not autonomous learning. Nothing is retained unless a person clicks a button.

Failed attempts matter most. "Restart was attempted" tells a future incident almost nothing. INC-103's record says the restart cleared connections briefly and exhaustion returned, with slow queries as the root cause. That is what changes a ranking.

Pre-Deployment Analysis

The same memory answers a different question: is this change safe to ship? /api/predeploy accepts a service and a proposed change. The backend builds a recall query from both, filters the results with the same service matcher, and asks Groq for a risk level (HIGH, MEDIUM or LOW), a 0-to-1 risk score, a rationale and safeguards. Related incident and deployment IDs are pulled from the recalled text with regexes (INC-\d+, DEPLOY-\d+), not by the model.

With "Increase connection pool: 50 → 100" on payments-api, the screenshot shows LOW risk (20%), related incidents surfaced, and safeguards such as monitoring connection counts and keeping a rollback path. It also shows 65 memories recalled before filtering and 16 after. That gap is why the filter exists.

It also shows zero related deployments. My seed data contains incidents only and no deployment records, so I don't claim deployment history the system doesn't have.

Memory and Learning Pages

The Memory explorer and Learning dashboard are not backed by Hindsight yet. They read from Frontend/src/data/mockData.ts, and the console labels them as mock or demo data. The IDs on them (INC-142, DEPLOY-204, the fix-drift chart) are sample data, not memories from the bank. The live parts are the Incidents and Pre-deploy pages, which call the backend.

An Honest Limitation

Extraction is keyword and regex based. extract_historical_facts looks for phrases like "permanently resolved" and "temporarily restored", and action normalization collapses anything mentioning "connection pool" into one bucket. Recommendations are matched to statistics by the words "restart", "connection pool" and "timeout". That is enough for this dataset and would not survive a messier one.

Small samples look authoritative. The screenshot's 100% at sample size 1 is exactly that, and the UI shows the sample size only as a small number next to a large percentage.

Outcomes recorded from the UI may not be counted per incident. Triage generates IDs like INC-5EBB28AA, while extract_incident_id only matches digits (INC-\d+), so those facts carry no reliable incident ID. Fix drift is in the response model (drift_detected) but the backend never computes it. And the only tests are two scripts, test_hindsight.py and test_groq.py, that call the services directly.

Lessons Learned

Retrieval relevance matters more than retrieval. Recall returned 65 memories and 16 were about the right service. Without filtering, the model sees other services' incidents.

Outcomes are worth more than descriptions. A stored failure with a reason changes what gets recommended. A stored incident description mostly doesn't.

Compute numbers outside the model. The model reasons well about a log. It should not be the source of a success rate.

Show the evidence and the sample size. Attach memory IDs to every claim, and don't let small counts look like certainty.

A human still closes the loop. A past fix can stop working, and only an engineer marking it Failed tells the system so.

What I Would Improve

None of this exists yet. I would store service and incident ID as structured fields instead of parsing text, so filtering stops depending on regexes. I would replace the Memory and Learning mock data with real endpoints over Hindsight. I would compute fix drift from timestamped outcomes and show sample sizes prominently. And I would write real tests around fact extraction, since that is the code most likely to break quietly.

The Takeaway

Giving an agent access to memory is a few lines of code. The two functions above are the whole Hindsight integration. The hard part is deciding which memories are relevant, which numbers can be trusted, and how that history should change a recommendation. Hindsight provides the persistent memory layer. Filtering, counting, evidence IDs and the engineer's verdict are what make it relevant and traceable.

The complete source code for IncidentIQ is available on GitHub:

IncidentIQ-SRE-Agent on GitHub

Top comments (0)