DEV Community

Sahasra Patwari
Sahasra Patwari

Posted on

Single Bank or Many? A Hindsight Design Decision I Almost Got Wrong

At 3 a.m., the person on call rarely needs more dashboards. They need someone who was there the last time this happened to lean over and say, "That's the connection pool. We saw it in August. Don't scale the pods, fix the pool."

I built an incident response agent to be that colleague. The interesting part turned out not to be the LLM. It was one boring-sounding decision about how the agent's memory is laid out: one shared memory bank for the whole company, with every record tagged by service.

This post covers why I made that call, what it looks like in code, and what I'd tell someone else designing agent memory on [Hindsight](https://github.com/vectorize-io/hindsight).


The RecallOps incident detail page: error signal and root cause on the left, suggested fix and Hindsight-recalled related incidents on the right.

What the system does

RecallOps is a FastAPI and React incident-response application. It stores incident records in SQLAlchemy, uses Hindsight as its operational memory layer, and uses a Groq-hosted model (openai/gpt-oss-120b) to turn an incoming alert plus retrieved incident history into an evidence-backed briefing.

The loop is simple. A new alert arrives, the agent recalls similar past incidents, and it answers with the relevant history cited. When an engineer marks an incident resolved, the incident's symptom, error signature, root cause, fix, and outcome get written back to memory. The next alert benefits from it.

Where Hindsight sits in the stack.

Most of the effort went into the middle layer, the part between "an incident happened" and "the agent remembered it usefully."

The decision: one bank, tagged by service

Hindsight organises memory into banks. The obvious first design for a multi-service company is one bank per service: 'checkout-service ', 'payment-gateway', 'auth-service', 'inventory-service'. It feels tidy. Each team owns its history, and recall is naturally scoped.

I didn't go that way because of how production incidents actually behave. The worst ones cross service boundaries. A checkout 502 might really be a payment-gateway problem. A spike of 401s on mobile might be a deployment issue that started in a different service. If each service's history sits in its own bank, the agent can only reason inside one silo at a time. The most valuable question, "have we seen something shaped like this anywhere?", becomes N queries and a merge step that I'd have to write and maintain.

So the whole company shares one bank, and each memory carries a tag for the service it came from:

# src/memory/schema.py
BANK_ID = "meridian-commerce-incidents"
Enter fullscreen mode Exit fullscreen mode
# src/memory/hindsight_tools.py
def log_incident(**kwargs) -> None:
    """
    Writes one resolved incident into memory, tagged by service. The tag
    is what lets recall_similar_incidents() later filter to just one
    service's history, while an untagged recall still searches everything.
    """
    content = format_incident_memory(**kwargs)
    _client.retain(
        bank_id=BANK_ID,
        content=content,
        tags=[kwargs["service"]],
    )
Enter fullscreen mode Exit fullscreen mode

Recall then has two modes from a single code path:

def recall_similar_incidents(alert_description: str, service: str = None, limit: int = 3):
    kwargs = {"bank_id": BANK_ID, "query": alert_description}
    if service:
        kwargs["tags"] = [service]
    result = _client.recall(**kwargs)
    return [r.text for r in result.results[:limit]]
Enter fullscreen mode Exit fullscreen mode

Leave service empty and the search runs across everything. Pass "checkout-service" and it's scoped. That gives me both behaviors from one bank. Hindsight's docs describe a single bank with tags as the option for applications that need reasoning across entities, while per-user banks are the more common pattern when users must be kept apart. My case is the first kind: incidents frequently cross service boundaries. You can always narrow with a tag filter later, but if you split banks up front, joining them afterward is much harder.

The trade-off is real. Tags are a retrieval aid, not an authorization boundary, and Hindsight's default match mode includes untagged memories. If you have hard isolation requirements, such as different customers' data, separate banks are the right tool. The right boundary is whoever the on-call team is allowed to reason over, not a microservice name. Service history inside one company doesn't have that constraint, so I took the flexibility.

Memory quality is a formatting problem

Hindsight stores free text and finds it later through semantic search. It doesn't impose a schema, which is convenient, but it means recall quality depends heavily on how consistently you write things down. Free-form postmortem prose retained as-is makes the search's job harder.

So every incident goes through one template before it's retained:

def format_incident_memory(incident_id, service, symptom, error_signature,
                           root_cause, resolution_steps, runbook_used,
                           outcome, resolved_by, timestamp) -> str:
    return f"""INCIDENT {incident_id} — {service}
Timestamp: {timestamp}
Symptom: {symptom}
Error signature: {error_signature}
Root cause: {root_cause}
Resolution steps taken: {resolution_steps}
Runbook used: {runbook_used}
Outcome: {outcome}
Resolved by: {resolved_by}
""".strip()
Enter fullscreen mode Exit fullscreen mode

Every field is labeled, and that matters more than it looks. A new alert might resemble an old incident's symptom ("502s during peak traffic"), its error signature (connection pool exhausted), or its root cause. Labeled fields give the retrieval step something to latch onto whichever way the new alert happens to be phrased. I treat this template as part of the API. Changing it is a migration, not a refactor.

Letting the model decide when to remember

The agent is a small tool-calling loop. Two functions are exposed to the model: recall_similar_incidents and log_incident. The system prompt is deliberately narrow about when each applies:

SYSTEM_PROMPT = """You are an incident response assistant.
When given a new alert, decide whether you need to recall similar past
incidents from memory before answering. Use the recall_similar_incidents
tool when historical context would help diagnose the issue. Use
log_incident only when explicitly told an incident has just been
resolved and needs to be recorded. Always give a clear, actionable
answer citing what you found in memory, if anything."""
Enter fullscreen mode Exit fullscreen mode

The write path is gated on purpose. I don't want a model deciding on its own that a half-understood outage is "resolved" and writing a wrong root cause into long-term memory. Memory that's wrong is worse than memory that's empty, because it gets cited with confidence.

Two smaller choices in the loop are worth copying. Tool errors go back to the model as text instead of raising, so the agent can say "memory lookup failed, here's what I can say without it" instead of the request dying with a 500 at the worst possible time. And the loop is capped at three tool hops. If the model can't get to an answer in three, a human is better off looking directly.

Recall is retrieval, but the briefing is synthesis

Hindsight exposes more than retain and recall. reflect synthesizes an answer from relevant memories in the bank, while recall returns ranked memory results the agent can expose as evidence. I use both, for different jobs:

  • recall returns the top matching incident records verbatim. The agent cites these, and the engineer can click through to the source.
  • reflect (wrapped as get_incident_briefing) produces a readable summary of what memory knows about an alert.

I keep them separate because they answer different questions. "Show me the evidence" and "tell me what it adds up to" shouldn't be the same call.

Mental models: the part that compounds

The feature that changed how I think about the system is Hindsight's mental models. A mental model is a saved reflect answer attached to a bank that you can refresh as memory grows. I created one team-wide model, "Team-wide Production Reliability Profile," from a single source query:

SOURCE_QUERY = """What recurring production failure patterns has our operations team learned from approved runbooks and resolved incidents?

For each pattern, identify the services and dependencies involved, early warning signals, verified root causes, safe mitigations, risky actions to avoid, and the first checks an on-call engineer should perform. Prioritize repeated, high-severity, cross-service incidents and clearly label uncertain evidence."""
Enter fullscreen mode Exit fullscreen mode

Notice that the query asks for risky actions to avoid and asks the model to label uncertain evidence. Those two phrases do more for on-call usefulness than any prompt tuning I did on the agent itself. The model's output for each pattern comes back as the same shape: services and dependencies, early warning signals, verified root cause, safe mitigations, risky actions, first checks, and the matching runbook.

Refreshing looks like this:

def refresh_team_mental_model():
    model = find_team_mental_model()
    result = _client.refresh_mental_model(bank_id=BANK_ID, mental_model_id=model.id)
    return result.operation_id
Enter fullscreen mode Exit fullscreen mode

Refresh is asynchronous, and reading the content right after triggering it can return the previous version. So the read path polls last_refreshed_at until it changes. That's a small detail, and it was the one that bit me first when I built a page that showed stale content right after a refresh.

I refresh this model on a schedule I control during development, but in production I'd rather not run a cron. Hindsight supports trigger={"refresh_after_consolidation": True}, which re-runs the model when the bank consolidates new memories. With that trigger, the reliability profile can refresh after newly retained incidents are consolidated, so nobody has to remember to update the wiki. Because the refresh is asynchronous, the UI should show when the profile was last refreshed instead of pretending it changed instantly.

What it does with a real alert

Here's how the pieces behave against the historical incidents in the bank. These four cover a connection-pool exhaustion on checkout, payment confirmations timing out after retry bursts got the company rate-limited by its processor, a JWT key rotation that only reached two of six pods, and repeated OOM kills during a nightly inventory sync.

Alert: checkout-service returning HTTP 502 errors, latency spiking on database calls

The agent chooses to call recall_similar_incidents, scoped to checkout-service. The matching record is the earlier incident where Postgres connections were exhausted during a flash-sale burst. Its retained fields include the error signature (connection pool exhausted), the root cause (pool sized for average load, not burst traffic), the fix (larger pool plus a circuit breaker to fail fast), and the runbook used. The agent's answer points the engineer at connection pool metrics first, and names the risky move: adding capacity without checking database limits.

A payment-timeout alert run without a service filter shows what the shared bank is for. Recall can surface the earlier payment incident, and the mental model already contains a pattern about retries without backoff triggering rate limiting, so the suggestion is to check external API error logs for rate-limit responses first, not to retry harder.


Seeding the bank and the labeled record Hindsight retains for a checkout-service incident.

I'm not going to claim a percentage improvement in resolution time, because I haven't measured one. What I can say is that the answers cite specific past incidents, and an on-call engineer can verify every claim by opening the source record.

Lessons learned

1. If your entities need to see each other, start with one bank and tags. Scope down with a filter when you need to. Splitting by entity up front makes the cross-entity questions, usually the interesting ones, expensive to answer later. If they must never see each other, use separate banks.

2. Treat your memory format as a schema, even if the store doesn't. Semantic search over inconsistent text gives inconsistent results. A labeled template gave me more recall quality than any other single change.

3. Gate the write path. Reads can be liberal and writes should be deliberate. A wrong memory is cited with confidence, so I let the model recall on its own, but only log an incident when it's told that one has been resolved.

4. Separate "show evidence" from "synthesise." recall and reflect do different jobs. Keeping them as different calls kept the agent's citations honest and the briefings readable.

5. Write the mental-model query like a runbook template. Asking explicitly for early warning signals, risky actions, first checks, and uncertainty labels shaped the output more than anything else I tried. Let the refresh trigger drive updates in production instead of relying on someone to press a button.

Where to go from here

If you're building something similar, start with the [Hindsight documentation](https://hindsight.vectorize.io/) for the retain, recall, reflect, and mental model APIs, and read the [Vectorize overview of agent memory](https://vectorize.io/what-is-agent-memory) for how these pieces fit together conceptually. The Hindsight repository on GitHub is where the code lives.

My advice: keep the memory layer small, make retained records structurally consistent, and choose bank boundaries based on trust and access, not your first guess at system topology. For one company's production incident history, cross-service visibility is the feature.

Top comments (0)