DEV Community

Cover image for How Hindsight Made Incident Advice Specific
M SAI CHARAN REDDY
M SAI CHARAN REDDY

Posted on

How Hindsight Made Incident Advice Specific

Most incident assistants can generate a checklist. The harder problem is getting them to remember what actually fixed the last outage without pretending that history has already proved today's root cause.

I built Incident Memory Agent to explore that problem. It is a Streamlit application for engineering and DevOps teams that combines persistent incident memory from Hindsight with reasoning from Groq.

The result is not an automatic remediation system. It is an investigation assistant that recalls relevant operational experience, shows the evidence it retrieved, and keeps uncertainty visible.

What the application does

The workflow is intentionally narrow:

  1. An engineer describes the currently known incident facts.
  2. Hindsight recalls semantically related historical incidents.
  3. Groq produces a baseline investigation using only the current facts.
  4. Groq produces a second investigation using the same facts plus the recalled memories.
  5. The interface displays both answers and the memory candidates side by side.
  6. After the real cause is confirmed, an engineer records the resolution so it can help during a future incident.

I describe this as Recall, Reason, and Learn. Hindsight handles persistent memory, Groq handles language-model reasoning, and the engineer remains responsible for verification.

flowchart LR
    A[Current incident] --> B[Hindsight recall]
    A --> C[Baseline reasoning]
    B --> D[Memory-aware reasoning]
    C --> E[Side-by-side comparison]
    D --> E
    E --> F[Engineer verifies evidence]
    F --> G[Record confirmed resolution]
    G --> H[Hindsight retain]

This separation matters. Memory should improve the investigation path, but it should not silently become the answer.

The failure mode I wanted to avoid

Suppose the checkout service starts returning HTTP 500 errors immediately after a deployment. A previous checkout incident had the same symptom and was caused by a missing PAYMENT_API_URL variable.

A careless agent might respond: "The environment variable is missing again."

That is a plausible guess, not a confirmed fact. The current failure could also come from a database migration, dependency outage, invalid secret, code regression, or deployment problem.

The useful response is more precise:

  • Confirm that HTTP 500 errors began after the deployment.
  • Mark the current root cause as unknown.
  • Present the previous missing variable as historical evidence.
  • Recommend checking environment configuration early.
  • Require logs, metrics, configuration, and deployment data before concluding.

This evidence-versus-proof boundary became the central design decision in the project.

Connecting Hindsight

The application reads its service configuration from environment variables. Secrets remain outside the repository.

load_dotenv()

hindsight_url = os.getenv("HINDSIGHT_API_URL")
hindsight_key = os.getenv("HINDSIGHT_API_KEY")
bank_id = os.getenv("HINDSIGHT_BANK_ID")

hindsight = Hindsight(
    base_url=hindsight_url,
    api_key=hindsight_key,
)
Enter fullscreen mode Exit fullscreen mode

For every investigation, the current incident description becomes the recall query:

memory_response = hindsight.recall(
    bank_id=bank_id,
    query=current_incident,
)

recalled_memories = [
    memory.text for memory in memory_response.results
]
Enter fullscreen mode Exit fullscreen mode

I used one Hindsight bank named incident-memory-agent. It contains structured incident narratives describing the service, symptom, impact, cause, resolution, prevention step, and date. I also seeded eight realistic synthetic incidents so the retrieval behavior can be tested across different services.

The value of semantic recall became obvious quickly. The current description does not need to exactly match the wording of the stored incident. A notification failure can still retrieve a past SMTP credential incident because the operational meaning is similar.

The Hindsight documentation was useful for understanding the retain and recall workflow. The broader Vectorize explanation of agent memory also helped clarify why memory should be treated as a separate system rather than an ever-growing prompt.

Comparing the same incident twice

The most useful interface decision was generating two answers.

The baseline request receives only the current incident. The memory-aware request receives the same incident plus the recalled memories. Everything else stays as similar as possible.

response = groq.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[
        {
            "role": "system",
            "content": (
                "Use historical memories as evidence, not as proof. "
                "Never invent facts, logs, causes or completed actions."
            ),
        },
        {"role": "user", "content": prompt},
    ],
    temperature=0.2,
    max_completion_tokens=1000,
)
Enter fullscreen mode Exit fullscreen mode

The prompt forces the model to separate its response into confirmed current facts, unknown information, relevant historical evidence, investigation steps, and an uncertainty warning.

That structure is more important than making the answer sound confident. During an outage, a confidently invented log line is worse than a short list of clearly marked unknowns.

Screenshot: New Investigation showing the baseline and Hindsight responses side by side.

A concrete learning-loop test

I tested the complete loop with an invoice-service incident.

First, I recorded this confirmed resolution:

  • Symptom: PDF invoices stopped being generated.
  • Impact: Customers could not download invoices for 25 minutes.
  • Root cause: INVOICE_TEMPLATE_PATH was missing.
  • Resolution: The variable was restored and the service restarted.
  • Prevention: Required-variable validation was added.

The Record Resolution page retained this structured information in Hindsight. The write action is protected by an admin password so public visitors cannot casually add data to the shared bank.

Next, I submitted a new current incident:

The invoice service stopped generating PDF invoices after today's deployment. The root cause has not yet been confirmed.

The baseline response suggested general checks such as reviewing logs and deployment changes. The Hindsight response retrieved the earlier invoice incident and prioritized validating INVOICE_TEMPLATE_PATH.

Crucially, it still described that variable as a historical clue. It did not claim that the present root cause had been confirmed.

That single interaction demonstrates the behavior I wanted: the system learned from a confirmed resolution, recalled it later, and used it to improve the investigation without replacing verification.

Screenshot: Retrieved Memory Candidates showing the stored invoice-service resolution.

What I learned

1. Memory must be visible

If users cannot see what the agent recalled, they cannot judge whether the recommendation is grounded or irrelevant. Showing the retrieved candidates makes the reasoning easier to inspect.

2. Current facts and historical evidence need separate blocks

Combining them into one prompt paragraph encourages the model to blur the boundary. Explicit labels produced clearer responses and made unsupported claims easier to notice.

3. Only confirmed outcomes should become durable memory

Early incident hypotheses are often wrong. Saving them as truth would make future investigations worse. The learning workflow therefore asks for the confirmed cause, resolution, and prevention step after the incident is understood.

4. A before-and-after comparison explains memory better than a feature list

Saying that an agent has persistent memory is abstract. Showing a generic baseline beside a historically informed response makes the difference concrete.

5. Retrieval quality depends on memory quality

Duplicated, vague, or incomplete incident records create noisy recall results. Stable document identifiers, structured incident summaries, and careful confirmation are important as the bank grows.

Current limitations

The application does not yet connect directly to Kubernetes, Grafana, Datadog, cloud logs, or ticketing systems. Engineers currently enter the incident facts manually. The seeded incidents are synthetic, and the admin password is a lightweight write-protection mechanism rather than full user authentication.

Those limitations keep the system focused on the memory workflow. A production version should add identity-based access, memory provenance, duplicate control, automated tests, observability, and connectors for live operational evidence.

The safety rule should remain unchanged: memory can decide what to investigate first, but only current evidence can confirm what is happening now.

Try the project

You can explore the live Incident Memory Agent, review the source code on GitHub, or watch the complete video walkthrough.

Building this changed how I think about incident assistants. The model was never the missing piece. The missing piece was a memory layer with a disciplined boundary between experience and evidence. Hindsight made the investigation more specific; the uncertainty rules kept it honest.


Top comments (1)

Collapse
 
reddysaicharan985 profile image
M SAI CHARAN REDDY •

hi