DEV Community

MehekJhawer
MehekJhawer

Posted on

My On-Call Agent Says "I Don't Know," and That's Why I Trust It.

It's 2:07 AM, Redis is getting OOM-killed every few minutes, and checkout is degrading. Six weeks ago a teammate fixed this exact failure. The postmortem exists, but it's somewhere in a wiki nobody can search at that hour. The engineer who fixed the other recurring failure left the company in June.

I built On-Call Hero to close that gap. It's an incident copilot that recalls your team's past incidents when an alert fires, tells you plainly when it has no match, and writes every new outcome back into memory, including the fixes that failed.

What it does and how it hangs together

An alert comes in with a message and a few log lines. Two things happen in parallel: a stateless LLM answers with no history (the "before"), and a memory-backed path answers using recalled incidents (the "after"). Both are shown side by side, because I wanted the difference to be visible on screen, not something I had to argue for.

alert + logs ─▶ workflow.triage()
                 ├─ handle_alert_stateless()      generic LLM
                 └─ handle_alert_with_memory()
                      ├─ recall()  similar incidents ┐
                      ├─ recall()  FAILED attempts   ├─▶ LLM structures one recommendation
                      └─ reflect() reasoning         ┘   (incident IDs validated)
                            │
                 guardrails.gate() ─▶ auto | approval | escalate | blocked
                            │
            workflow.resolve()
                 ├─ retain() outcome ─▶ wait_until_recallable()
                 └─ draft postmortem
Enter fullscreen mode Exit fullscreen mode

Memory is Hindsight, the open-source agent memory system from Vectorize. Reasoning runs on Groq-hosted models with a fallback chain. Every Hindsight call lives in one file, src/hindsight_memory.py, so the whole "how does the memory work" question can be answered by reading a single module top to bottom.

The through-line: memory of outcomes, not just documents

My first instinct was RAG over a folder of postmortems. It gets you the "this looks like INC-104" moment, and that's genuinely useful. But it's a read-only system. It knows what your team wrote down after the last bad night, and nothing about what happened during this one.

The thing I actually cared about was outcomes. A fix that worked should make the next match stronger. A fix that failed should make the agent stop recommending it. That's not a retrieval problem, it's a write-path problem, and it's where agent memory earns its keep compared with a static index. The Hindsight docs frame memory around three operations, retain, recall and reflect, and I ended up using all three for distinct jobs.

Retain: write it like a human would

Hindsight extracts entities and facts from text, so I write each incident as a plain-language postmortem rather than a JSON blob. One line matters more than it looks: if the resolver has left the company, that goes into the memory text.

def incident_to_memory_text(inc: dict[str, Any]) -> str:
    left = " (has since left the company)" if inc.get("resolver_left_company") else ""
    logs = " | ".join(inc.get("logs", []))
    return (
        f"Incident {inc['incident_id']} ({inc.get('severity', 'SEV-?')}) on {inc['service']} "
        f"[category: {inc.get('category', 'unknown')}]. "
        f"Symptom: {inc['symptom']} "
        f"Root cause: {inc['root_cause']} "
        f"Resolution: {inc['resolution']} "
        f"Resolved by {inc['resolved_by']}{left} in {inc['resolution_time_minutes']} minutes."
        + (f" Log evidence: {logs}" if logs else "")
    )
Enter fullscreen mode Exit fullscreen mode

Including the raw log lines was deliberate. Alerts arrive with logs, so recall can match on OOMKilled exitCode=137 and not only on the prose description of a symptom.

Failed fixes are first-class memories

When an incident resolves, the outcome is retained. For a success, the text records the fix and who applied it. For a failure, it says so explicitly:

content = (
    f"Incident {incident_id} ({service}) [category: {category}]. Symptom: {symptom} "
    f"Attempted remediation: {resolution} Attempted by {resolved_by}. "
    f"OUTCOME: FAILED - this did not resolve the incident; "
    f"do not recommend it again for this symptom."
)
Enter fullscreen mode Exit fullscreen mode

Then, on every alert, there's a second, dedicated recall that asks specifically for remediations recorded as failed:

def recall_failed_attempts(client, bank_id: str, alert_text: str) -> list:
    query = f"Remediation that FAILED and did not resolve this kind of incident: {alert_text}"
    try:
        res = _retry(client.recall, bank_id=bank_id, query=query, attempts=2)
    except Exception:
        return []
    return [r for r in getattr(res, "results", []) if "OUTCOME: FAILED" in r.text]
Enter fullscreen mode Exit fullscreen mode

I did this as a separate query because a single "find similar incidents" recall ranks by similarity to the symptom, and a failed attempt often reads like a successful one. The explicit OUTCOME: FAILED filter afterwards is unglamorous string matching, but it's cheap and it's checkable. It's also best-effort: if the call errors, it returns nothing and triage continues, because a secondary lookup should never block the primary one.

Don't race your own write

Retention may be processed asynchronously, so "I just wrote a memory" and "I can recall it" aren't the same moment. After every successful resolution I poll until the new incident is actually searchable:

def wait_until_recallable(client, bank_id, incident_id, probe_query,
                          timeout=15.0, interval=1.0) -> float | None:
    start = time.time()
    while time.time() - start < timeout:
        res = recall_similar_incidents(client, bank_id, probe_query)
        if any(incident_id in r.text for r in getattr(res, "results", [])):
            return round(time.time() - start, 1)
        time.sleep(interval)
    return None
Enter fullscreen mode Exit fullscreen mode

It returns how long it waited, or None on timeout, and the timeout path is tested. This is the piece that lets me say "the memory written by this run is what the next alert recalled" and have it be true rather than assumed.

Saying "I don't know"

The most important behavior in the system is one that looks like a non-feature: when nothing genuinely matches, it doesn't guess.

Two mechanisms enforce that. The prompt tells the model to set matched_incident_id to null and confidence to low instead of reaching. But prompts are requests, so there's also a check in code. The model is only allowed to cite an incident ID that actually appeared in the recalled memory or the reflection text:

known_ids = set(re.findall(r"INC-\d+", " ".join(s.raw_matches) + " " + s.reflection_text))
mid = parsed.get("matched_incident_id")
if mid and mid not in known_ids:  # only cite IDs that were actually recalled
    warnings.append(f"Model cited {mid}, which is not in recalled memory; discarded.")
    mid = None
conf = str(parsed.get("confidence", "low")).lower()
if conf not in {"high", "medium", "low"} or not mid:
    conf = "low"
Enter fullscreen mode Exit fullscreen mode

No valid ID means confidence is forced to low. Low confidence means the safety gate escalates to a human and runs nothing:

def gate(command, confidence, matched_incident_id) -> GateDecision:
    if not command or not matched_incident_id or confidence == "low":
        return GateDecision("escalate", "n/a",
                            ["No confident memory match: paging a human, running nothing."])
    risk, reasons = assess_command(command)
    if risk == "blocked":
        return GateDecision("blocked", risk, reasons + ["Destructive command blocked outright."])
    if risk == "low" and confidence == "high":
        return GateDecision("auto", risk, ["Read-only command with high confidence."])
    return GateDecision("approval", risk, reasons + ["Mutates production: needs on-call engineer approval."])
Enter fullscreen mode Exit fullscreen mode

The gate classifies each segment of a pipeline and takes the riskiest one. Read-only commands with high confidence can auto-run. Anything that mutates production needs a human to approve. Destructive patterns (flushall, rm -rf /, drop table) are blocked outright. Unknown commands are treated as mutating, so the failure mode is "ask a person," not "run it."

The interesting consequence is what happens next. When the agent escalates, a human resolves the incident, and that resolution is retained like any other. The next time the same failure appears, the agent recognizes it. The "I don't know" isn't a dead end, it's the moment the system acquires new knowledge.

What it looks like in practice

Take a Redis alert: memory at 94% and climbing on prod-cluster-3, repeated OOM kills, cache hit rate dropping.

The stateless pane does what a general-purpose model does: check logs, inspect resource usage, review recent deployments, restart the pod. None of that is wrong, and none of it helps you at 2 AM.

The memory pane recalls INC-104 from the seed history: session cache keys written without a TTL after a config change, growing until Redis hit its limit. It names who fixed it last time and how long it took (Sarah Chen, 40 minutes), and it proposes the exact remediation from that incident:

kubectl rollout undo deployment/redis-cache
redis-cli --scan --pattern 'session:*' | xargs -L 1000 redis-cli del
Enter fullscreen mode Exit fullscreen mode

Both commands mutate production, so the gate routes this to approval rather than running it. An engineer confirms, the fix runs, the outcome is retained as a new incident, and wait_until_recallable confirms it's searchable before the workflow moves on.

The team-insights view uses reflect over the whole bank instead of a single alert. It surfaces recurring failure themes, remediations that failed, and, the one that got the strongest reaction when I showed it to people, which runbooks depend on an engineer who has already left. Because the departed status is written into the memory text at retain time, that question is answerable rather than a guess.

Lessons learned

1. Call memory from code, not as an LLM tool. I don't expose Hindsight to the model as a function to call. The workflow always runs recall, then reflect, then answer, in that order. There's no tool-call step that can fail or be skipped in the middle of an incident, and the behavior is deterministic enough to test.

2. Validate what the model claims against what you retrieved. Asking the model not to hallucinate incident IDs is a request. Checking its citation against the recalled text is a guarantee. Treat the LLM's output as untrusted input to your own code.

3. Store failures deliberately. Most memory setups quietly accumulate successes. A remembered mistake is worth as much as a remembered fix, but only if you write it down in a form that recall can find and the prompt can act on.

4. Verify your writes are readable. Any system that writes and then immediately reads needs to account for asynchronous processing. Polling with a timeout and testing the timeout path is a few lines of code and removes a whole class of flaky behavior.

5. Fail closed, and build for a flaky dependency. The LLM client retries and falls back across a chain of models, treats empty output as failure, and parses JSON tolerantly. The Hindsight wrapper retries and falls back on optional arguments. There's an offline backend that runs the same code paths, and it's what the 23-test suite runs on, so the tests need neither network access nor API keys.

Where this goes next

The pieces I'd add next are the ones that connect the loop to where incidents actually happen: PagerDuty or Opsgenie webhooks for intake, a Slack front end with an Approve button for the gate, real kubectl execution behind it, and per-team memory banks with access control.

The idea underneath doesn't change. Incident response is a domain where the same failures come back in slightly different clothes, and the most valuable thing a copilot can do is recognize them, admit when it can't, and remember how the night ended. That's a memory problem first and a reasoning problem second.

The code is at github.com/mehekjhawer-blip/oncall-hero.


Top comments (0)