DEV Community

Cover image for Why I Gave OpsMind Hindsight Instead of More Prompting
Charan Jalluri
Charan Jalluri

Posted on AI-assisted

Why I Gave OpsMind Hindsight Instead of More Prompting

At 2:14 AM on a Tuesday, an alert woke our on-call engineer: HTTP 502 error rates on payment-api had jumped to 18% immediately following a routine deployment. Operating on adrenaline and minimal sleep, the engineer did what most people would do when seeing unresponsive pods—they ran kubectl rollout restart deployment/payment-api to clear the pressure, unaware that six months earlier that exact reflexive command severed in-flight database transactions, triggered retry storms, and created $12,400 in duplicate customer charges.

The original failure had been thoroughly analyzed, documented in a postmortem, and filed away in an internal wiki that nobody checked during a live outage. When we started building OpsMind, our initial impulse was the same as everyone else's in the current AI cycle: give an LLM access to our monitoring telemetry, write a comprehensive system prompt instructing it to act like a principal site reliability engineer, and let it triage. But prompt engineering doesn't fix amnesia. The model still suggested restarting the pods because, across millions of generic pre-training tokens, restarting an unhealthy microservice is statistically common advice.

To prevent our incident response system from repeating catastrophic mistakes, we had to stop treating production triage as a prompt-engineering problem and treat it as a memory problem. Here is why and how we built OpsMind with persistent memory, what happened when we wired it into our operational loop, and the concrete lessons we learned along the way.


What OpsMind Does and How It Hangs Together

OpsMind is an automated incident investigation and institutional memory engine designed for platform and SRE teams. Rather than operating as a stateless chatbot that generates speculative suggestions, OpsMind sits between active production monitoring and our engineering team's historical operational experience.

The runtime architecture consists of four tightly coupled components:

  1. The SRE Console: A dashboard built with React, Vite, and TypeScript that displays live incidents, telemetry symptoms, historical evidence matches, and root-cause hypotheses alongside explicit operational warnings.
  2. The Investigation Engine: A FastAPI service written in Python that orchestrates incident ingestion, memory queries, LLM reasoning, and postmortem retention.
  3. The Persistent Memory Layer: A managed memory bank powered by Hindsight, accessed via the official client library from the Hindsight GitHub repository. It exposes explicit RETAIN and RECALL primitives for structured engineering knowledge.
  4. The Reasoning Engine: Meta Muse Spark 1.3 (muse-spark-1.3), accessed through an OpenAI-compatible API endpoint. We run it at a low temperature (0.2) with strict JSON schema enforcement to ensure deterministic output.

The operational lifecycle runs as a closed loop:

Alert Ingestion (Symptoms + Service)
         │
         ▼
Hindsight RECALL ──► Recalls historical postmortems, anti-patterns & policies
         │
         ▼
Meta Muse Spark 1.3 ──► Grounded investigation citing past incidents & warnings
         │
         ▼
SRE Operator Triage ──► Targeted, non-destructive remediation
         │
         ▼
Hindsight RETAIN ──► Structured postmortem stored under deterministic ID
Enter fullscreen mode Exit fullscreen mode

When an alert fires—such as elevated gateway timeouts on our payment ingress—OpsMind constructs a query containing the affected service name, deployment tag, environment, and observed symptoms. Before calling the LLM, it executes a semantic recall against the Hindsight memory bank. The retrieved evidence—including prior root causes, effective remediation steps, and explicitly tagged anti-patterns—is injected into the model's prompt.

Once the engineer resolves the outage, OpsMind captures the postmortem, formats it into an immutable experience record, and retains it back into Hindsight under a deterministic identifier. The next time a similar failure pattern emerges, the system does not start from zero.


The Core Technical Story: Memory vs. Context Stuffing

When developers build LLM-assisted DevOps tools, they usually take one of two approaches: stuffing runbooks into system prompts or performing standard Retrieval-Augmented Generation (RAG) over Markdown docs in a vector database. Both approaches break down under real production pressure.

Context stuffing fails because operational knowledge is non-linear and unbounded. You cannot fit every service's historical quirks, configuration regressions, and database failover oddities into a single context window without paying steep latency penalties, suffering from attention degradation ("lost in the middle"), and bloating inference costs.

Naive RAG fails for a more dangerous reason: semantic similarity does not equal operational relevance. If you index raw postmortems and search for "payment-api 502 bad gateway," cosine similarity often pulls up paragraphs describing how someone once solved a 502 by restarting the service, while burying the subsequent incident where restarting the service caused financial data corruption. Unstructured vector search over flat text cannot distinguish between an effective fix and a disastrous anti-pattern.

We realized that what autonomous operational tools actually require is persistent agent memory. Dedicated agent memory differs fundamentally from generic document retrieval. It requires:

  • Episodic memory: Storing exact incident episodes with their initial symptoms, suspected causes, confirmed root causes, and verified resolutions.
  • Explicit negative knowledge: Remembering not just what worked, but what failed and what caused harm.
  • Deterministic identity: Storing records under stable keys so that incident updates refine existing knowledge rather than cluttering the index with duplicate, conflicting vectors.
  • Evidence-based constraints: Forcing the reasoning engine to ground its conclusions in empirical facts and explicitly acknowledge when no precedent exists.

By adopting Hindsight, we separated the reasoning capabilities of the language model from the storage of organizational experience. The model doesn't need to know the entire history of our infrastructure during pre-training; it simply needs to be a competent analytical engine evaluating historical evidence provided at inference time.


Code-Backed Implementation: How It Works Under the Hood

To make this architecture concrete, here is how the core loop is implemented across the OpsMind codebase.

1. Structuring Engineering Experience for Retention

A critical design decision was refusing to store unstructured conversational prose in Hindsight. If an SRE leaves a 500-word stream-of-consciousness Slack message, extracting reliable signal later becomes impossible. Instead, we defined a strict domain model using Pydantic in memory_types.py:

class EngineeringMemory(BaseModel):
    incident_id: Optional[str] = Field(default=None)
    memory_type: str = Field(default="incident")  # "incident" or "decision"
    service: str = Field(...)
    environment: str = Field(default="production")
    deployment_version: Optional[str] = Field(default=None)
    symptoms: List[str] = Field(default_factory=list)
    confirmed_root_cause: Optional[str] = Field(default=None)
    resolution: str = Field(...)
    failed_approaches: List[str] = Field(default_factory=list)  # anti-patterns
    warnings: List[str] = Field(default_factory=list)           # operational hazards
    lessons_learned: List[str] = Field(default_factory=list)
Enter fullscreen mode Exit fullscreen mode

Notice the explicit separation of failed_approaches and warnings. When serializing this record into an experience document for Hindsight, these fields receive dedicated markdown sections prefixed with high-priority markers (! and +).

In retain.py, we pass this structured document into Hindsight alongside deterministic document IDs and operational tags:

def retain_engineering_memory(
    memory: EngineeringMemory,
    hindsight_memory: HindsightMemory,
    bank_id: Optional[str] = None,
) -> MemoryRetainResponse:
    target_bank = bank_id or hindsight_memory.default_bank_id
    doc_id = memory.incident_id or f"MEM-{memory.service}-{memory.memory_type}"

    response = hindsight_memory.retain(
        bank_id=target_bank,
        content=memory.to_experience_document(),
        document_id=doc_id,
        context=f"Engineering experience record for {memory.service}",
        metadata=memory.extract_metadata(),
        tags=memory.extract_tags(),
    )
    return MemoryRetainResponse(success=getattr(response, "success", True), incident_id=memory.incident_id)
Enter fullscreen mode Exit fullscreen mode

By providing document_id=doc_id, subsequent revisions or postmortem edits to INC-1042 update the existing memory entry rather than fragmenting the bank with stale duplicates.

2. Synthesizing Telemetry into Focused Recall Queries

When an incident triggers an investigation, the incoming payload contains raw symptoms and metadata. Rather than passing raw alerts directly to the memory bank, recall.py constructs a structured query vector:

def construct_incident_query(service: str, symptoms: List[str], deployment_version: Optional[str] = None) -> str:
    symptoms_text = "; ".join(symptoms) if symptoms else "service degradation"
    parts = [f"Service: {service}", f"Symptoms: {symptoms_text}"]
    if deployment_version:
        parts.append(f"Deployment: {deployment_version}")
    return " | ".join(parts)
Enter fullscreen mode Exit fullscreen mode

In investigator.py, OpsMind executes the recall call with an explicit token budget (max_tokens=1500), as documented in the Hindsight documentation. This guarantees that historical context stays compact and leaves ample room for the reasoning model's output:

recall_response = self.hindsight_memory.recall(
    query=recall_query,
    bank_id=self.hindsight_memory.default_bank_id,
    max_tokens=1500,
)
if recall_response and recall_response.results:
    historical_context = recall_response.to_prompt_string()
Enter fullscreen mode Exit fullscreen mode

3. Hardening the Reasoning Engine Against Hallucination

The biggest danger in autonomous SRE tools is hallucinated certainty. If an agent invents an incident ID or fabricates a runbook command, an on-call engineer might execute an untested command during an active outage.

In prompts.py, we enforce strict grounding rules directly in the system prompt:

CRITICAL OPERATIONAL RULES:
1. Historical memories provided in the prompt are your empirical evidence.
   Cite them explicitly by identifier (e.g. INC-1042, DECISION-PAYMENT-RESTART).
   Never claim an incident occurred unless documented in HISTORICAL MEMORY.
2. If HISTORICAL MEMORY contains no relevant records, state this clearly.
   Set memory_used to false and confidence to "Low - Insufficient Historical Evidence".
   Provide cautious triage, warning that no prior experience exists.
3. NEVER fabricate or hallucinate incident IDs, postmortems, or policies.
Enter fullscreen mode Exit fullscreen mode

By coupling this prompt with a mandatory JSON output schema requiring historical_evidence, possible_root_causes, recommended_steps, and warnings, the system produces predictable, parseable responses every time.

4. Handling Real-World Client Configurations

Defensive engineering matters in operational pipelines. When loading API keys via Pydantic BaseSettings, Pydantic wraps values in SecretStr. Passing that object directly to underlying HTTP clients causes JSON serialization errors ({"get_secret_value": ...}).

In config.py, we implemented an explicit unwrapping utility to ensure plain strings across runtime environments:

def unwrap_api_key(raw_key: Any) -> Optional[str]:
    if raw_key is None:
        return None
    if hasattr(raw_key, "get_secret_value") and callable(raw_key.get_secret_value):
        raw_key = raw_key.get_secret_value()
    if isinstance(raw_key, dict):
        for candidate in ("api_key", "meta_model_api_key", "key", "token"):
            if candidate in raw_key:
                return unwrap_api_key(raw_key[candidate])
    return str(raw_key).strip().strip("'\"") if raw_key else None
Enter fullscreen mode Exit fullscreen mode

Results: Before vs. After Persistent Memory

To evaluate how persistent memory alters incident outcomes, we tested OpsMind across realistic failure scenarios. The script demo_before_after.py demonstrates the contrast between a stateless model and one grounded in Hindsight.

Scenario: The 502 Spike on Payment API

We triggered an alert for INC-DEMO-502:

  • Service: payment-api
  • Deployment: v2.10.5
  • Symptoms: Immediate surge in HTTP 502 Bad Gateway responses, gateway timeouts on checkout endpoints.
  • Context: On-call engineer alerted; considering cycling the service pods.

Phase 1: Investigation Without Persistent Memory (Empty Bank)

When OpsMind queried an empty memory bank, it received zero historical records. Adhering to our grounding rules, the agent returned:

{
  "memory_used": false,
  "confidence": "Low - Insufficient Historical Evidence",
  "historical_evidence": [],
  "incident_summary": "payment-api experiencing HTTP 502 timeouts post v2.10.5 without recorded precedent.",
  "recommended_steps": [
    "Check container CPU and memory utilization metrics in Kubernetes",
    "Inspect payment-api container logs for unhandled exceptions",
    "Consider restarting application pods or rolling back release if errors persist"
  ],
  "warnings": [
    "No historical records exist for this failure pattern. Exercise caution before destructive actions."
  ]
}
Enter fullscreen mode Exit fullscreen mode

The output was polite, well-structured, and typical of standard AI triage tools. It suggested inspecting CPU, checking logs, and—critically—restarting the application pods. In our architecture, that restart would have dropped in-flight transactions and triggered duplicate billing.

Phase 2: Investigation With Hindsight Persistent Memory

Next, we ran the identical incident payload against an active Hindsight memory bank containing historical records (INC-1042, INC-1067, INC-1091) and engineering policies (DECISION-PAYMENT-RESTART).

Hindsight recalled the relevant records within milliseconds. Meta Muse Spark 1.3 synthesized the evidence into a dramatically different response:

{
  "memory_used": true,
  "confidence": "High - Directly Supported by Historical Precedents",
  "historical_evidence": [
    {
      "incident_id": "INC-1042",
      "summary": "Database connection pool regression in v2.9.0 caused immediate 502 spike after release.",
      "effective_actions": ["Diff configuration between release manifests", "Immediate deployment rollback"],
      "dangerous_actions": ["Scaling replicas without fixing config"]
    },
    {
      "incident_id": "INC-1091",
      "summary": "Emergency restart during active traffic caused $12,400 in duplicate charges across 47 transactions.",
      "dangerous_actions": ["kubectl rollout restart deployment/payment-api without draining in-flight requests"]
    }
  ],
  "possible_root_causes": [
    "Database connection pool configuration regression in v2.10.5 manifest (similar to INC-1042)",
    "Environment variable or timeout mismatch introduced during CI/CD deploy (similar to INC-1067)"
  ],
  "recommended_steps": [
    "1. Run git diff between v2.10.4 and v2.10.5 Helm release values, specifically checking pool_max settings.",
    "2. Inspect active RDS connection counts against pool capacity metrics.",
    "3. If configuration is verified broken, execute clean rollback to v2.10.4."
  ],
  "warnings": [
    "CRITICAL: DO NOT RESTART payment-api PODS. Policy DECISION-PAYMENT-RESTART and INC-1091 confirm that restarting under load severs in-flight transactions and causes duplicate customer billing."
  ]
}
Enter fullscreen mode Exit fullscreen mode

The difference is night and day. The grounded agent immediately warned the responder against the exact reflex that previously caused data corruption. Instead of vague suggestions to "check CPU," it told the engineer to run a git diff on connection pool parameters between the two specific release manifests.


Lessons Learned

Building OpsMind fundamentally changed how we think about using LLMs in platform operations. If you are designing intelligent systems for operational domains, keep these lessons in mind:

1. Anti-patterns are twice as valuable as happy-path documentation

Most engineering wikis document how systems are supposed to work. In an outage, you need to know how they break and what makes them worse. Structuring memory to capture failed_approaches and warnings gave our agent its most critical capability: preventing engineers from repeating known mistakes under pressure.

2. Memory is not document search

Traditional vector search treats all text as homogeneous blobs. Operational memory requires semantic structure: symptoms must match symptoms, deployment versions must filter releases, and root causes must link to verified remedies. Using Hindsight as an explicit agent memory service allowed us to retain structured experience records rather than throwing messy text files into a vector bucket.

3. Hallucination guardrails must be structural, not conversational

Telling an LLM "please don't make things up" in a prompt is a suggestion, not a guarantee. You must enforce structural guardrails: require the model to explicitly cite recalled incident IDs, provide a low-confidence escape hatch when the memory bank is empty, and enforce strict JSON schemas so downstream tools can verify claims before showing them to an operator.

4. Deterministic identity prevents memory rot

If every incident resolution creates a new, unindexed vector document, your memory bank quickly suffers from conflicting and outdated information. By assigning deterministic identifiers to documents (INC-1042, DECISION-PAYMENT-RESTART) during the RETAIN step, updates overwrite and refine existing knowledge. Clean memory hygiene is as important for AI agents as it is for databases.


Conclusion

The future of SRE tooling isn't about writing more complex prompts or waiting for models with multi-million-token context windows. When production is on fire, speed and empirical accuracy are what matter.

By combining the reasoning power of Meta Muse Spark 1.3 with the persistent memory primitives of Hindsight, OpsMind turns ephemeral incident triage into a compounding institutional asset. The team no longer loses sleep over the same outage twice—because our tooling finally remembers what we learned the hard way.

Top comments (0)