DEV Community

Cover image for Building an Incident Dashboard Around Hindsight Memory
GADDI KOPULA VENKATESH
GADDI KOPULA VENKATESH

Posted on

Building an Incident Dashboard Around Hindsight Memory

Building an Incident Dashboard Around Hindsight Memory
An incident-response agent can have good reasoning and still be difficult to use.

For OpsMind, the frontend was therefore treated as more than a place to display an AI-generated answer. The dashboard needed to make the incident state, evidence, historical memory, diagnosis, remediation, and learning lifecycle visible to an engineer.

The interface follows the same principle as the backend:

Current evidence first. Historical memory second. Action only after review.

From Backend State to Operator View
The OpsMind backend exposes endpoints for:

  • Health
  • Incident listing
  • Individual incident details
  • Incident analysis
  • Incident resolution The frontend consumes these APIs and turns the resulting state into an incident-response workflow.

The main dashboard provides an incident selector and displays information such as:

  • Severity
  • Service
  • Metrics
  • Health signals
  • Evidence timeline
  • Historical context
  • Hindsight memory
  • AI diagnosis
  • Confidence
  • Recommended actions
  • Runbook
  • Resolution state OpsMind and Hindsight memory architecture visual

Figure 1 — The dashboard reflects the same incident-response and memory lifecycle as the backend.

The goal is to allow an engineer to understand the incident without jumping between multiple screens.

Making Current Evidence Visible
The first information shown after selecting an incident is its current operational state.

For example, INC-008 displays:

  • Payment API
  • Critical severity
  • 6.1-second latency
  • 26% HTTP 500 error rate
  • 97% database connection utilization
  • The dashboard also shows the related log signals.

This gives the engineer immediate visibility into what is happening now.

The UI should not make an engineer open the historical memory section before seeing the current evidence.

That mirrors the reasoning architecture.

Making Memory Explicit

One of the most important frontend decisions was to make historical memory visible rather than hiding it inside the AI prompt.

The dashboard can show a memory match and identify the historical incidents retrieved by Hindsight.

For INC-008, historical context included incidents such as:

  • INC-007
  • INC-006
  • INC-001 After INC-008 was resolved and retained, another investigation could retrieve INC-008 as historical context.

This makes the memory loop observable.

AI incident and memory context visual

Figure 2 — The incident dashboard exposes diagnosis, evidence, confidence, and historical memory together.

An engineer can therefore see not only what the AI concluded, but also the context that influenced the conclusion.

Why Memory Should Not Be Hidden

If historical context is completely invisible, an engineer may have difficulty understanding why an agent recommended a particular action.

Showing the historical incident references provides a basic explanation of where additional context came from.

It also makes the system easier to debug.

If an irrelevant incident appears in memory, an engineer can identify that problem rather than simply seeing an unexplained AI recommendation.

This is especially useful when working with retrieval systems.

Retrieval quality becomes part of the application's observable behavior.

The Confidence and Evidence Sections
OpsMind also displays a confidence value and evidence signals.

The confidence field gives a concise indication of how strongly the agent's reasoning supports the diagnosis.

The evidence section shows the concrete signals associated with the incident.

Root Cause

Database connection pool exhaustion.

Evidence

Database connection utilization, connection acquisition delays, latency, and HTTP 500 errors.

This makes the diagnosis easier to inspect.

The interface does not need to expose every internal model token or reasoning detail.

Instead, it provides structured information relevant to an operator's decision.

Designing the Approval Experience

After diagnosis, the dashboard presents a human approval gate.

The interface makes the distinction between recommendation and execution visible.

The engineer can review the diagnosis and runbook before approving remediation.

The runbook is marked as simulation-only.

This is important because the dashboard should communicate the system's operational boundaries clearly.

The user should never be left wondering whether clicking the action button will change a real production service.

Showing the Learning Event

Once the simulated remediation succeeds, the interface changes state.

The incident becomes:

RESOLVED

and the dashboard shows that the outcome was retained as organizational memory.

Hindsight memory lifecycle visual

Figure 3 — The dashboard shows that the resolved incident has become persistent organizational memory.

This visual state is important because it communicates that resolution is not the end of the workflow.

The incident has moved into the memory lifecycle.

Demonstrating Future Recall
The strongest frontend demonstration occurs when a different incident is analyzed after the memory has been retained.

When INC-007 is analyzed, the dashboard can show INC-008 among its historical context.

AI incident investigation visual

Figure 4 — The dashboard makes the cross-incident memory relationship visible.

This creates a clear visual narrative:

  • ## INC-008 was resolved.
  • ## INC-008 was retained.
  • ## INC-007 was investigated later.
  • ## INC-008 appeared as historical context.

That is much easier to understand when the interface makes the state transitions visible.

Keeping the Frontend Focused

One lesson from building the dashboard was that an AI interface can become cluttered very quickly.

There are many possible pieces of information:

  • Logs
  • Metrics
  • Memory
  • Diagnosis
  • Reasoning
  • Recommendations
  • Runbooks
  • Resolution
  • Learning Displaying everything with equal visual importance can make the interface harder to use.

The useful hierarchy is:

Current incident → Evidence → Diagnosis → Historical context → Recommended action → Approval → Outcome

That sequence follows the engineer's decision process.

Frontend and Backend Boundaries
The frontend does not perform the incident reasoning itself.

It calls the backend API.

The backend coordinates:

  • Incident retrieval
  • Evidence retrieval
  • Hindsight recall
  • AI diagnosis
  • Runbook selection
  • Resolution
  • Memory retention This separation keeps the frontend focused on presentation and interaction.

It also makes it possible to change the AI or memory implementation without redesigning the entire interface.

What We Learned

1. Observability should include the AI workflow

If memory influences an AI decision, users should have some visibility into that context.

2. UI state should match backend state

The dashboard should clearly distinguish between analyzed, awaiting approval, resolved, and memory-retained states.

3. Current evidence needs visual priority

Historical memory is useful, but the latest telemetry should remain prominent.

4. The learning loop should be visible

Showing the transition from resolution to retained memory makes the value of persistent memory much easier to understand.

5. Good AI UX is about decision support

The interface should help an engineer inspect and decide, not simply display generated text.

Conclusion

Building the OpsMind dashboard made one thing clear: an AI SRE agent is not only a backend problem.

The frontend determines whether an engineer can understand what the agent knows, why it reached a conclusion, what historical context influenced it, and what will happen if remediation is approved.

The dashboard therefore mirrors the architecture of the agent itself.

  • Current evidence is visible first.
  • Historical memory is explicit.
  • Recommendations are inspectable.
  • Execution requires approval.
  • Successful outcomes become visible as retained organizational memory.

That makes the memory loop something an engineer can actually see and reason about rather than an invisible mechanism behind an AI response.

Top comments (3)

Collapse
 
octyn profile image
OCTYN •

the visible-memory point has a second half: what happens when the engineer spots the irrelevant incident. if they can flag "this shouldn't have been retrieved" right there, the dashboard is cleaning memory, not just explaining it. if they can't, that wrong incident shows up in every future diagnosis and trust erodes quietly. memory that can't be corrected becomes confident wrongness.