DEV Community

Vishal Vishwanadhula
Vishal Vishwanadhula

Posted on

DeployLens: Building a Production Incident Investigator That Actually Remembers

Production incidents rarely happen at a convenient time.

It might be 2 AM. An API suddenly starts returning 5xx errors, latency is increasing, alerts are firing, and an SRE has to answer one critical question:

  1. What changed, and have we seen this problem before? Most incident investigation workflows require engineers to jump between deployment histories, configuration changes, monitoring dashboards, Git commits, logs, Slack conversations, and old postmortems.

The problem isn't only finding information.

The problem is remembering what actually worked.

That is the idea behind DeployLens — a memory-powered production incident investigation system designed for DevOps, SRE, and platform engineering teams.

GitHub: https://github.com/yelemvenkatarachika/DeployLens

The Problem With Traditional Incident Investigation

Imagine a checkout service suddenly begins returning errors.

An engineer may need to investigate:

  • Recent deployments
  • Configuration changes
  • Alerts and metrics
  • Previous incidents
  • Historical postmortems
  • Previously attempted fixes

A traditional AI assistant can summarize the information available to it, but it usually doesn't maintain a reliable organizational memory of previous incidents and their outcomes.

For example, suppose restarting checkout-api was attempted several times during previous incidents.

A generic assistant might still suggest:

Restart the checkout service and monitor the error rate.

But an experienced SRE might know:

We already tried that three times.
It only helped temporarily.
The real issue was Redis connection pool exhaustion.

That difference is what DeployLens focuses on.

What Is DeployLens?

DeployLens is a production change intelligence and incident investigation agent with persistent operational memory.

Its central question is:

What changed before production broke, have we seen this failure before, and what actually worked last time?

Instead of treating every incident as a brand-new problem, DeployLens attempts to connect today's incident with the organization's historical operational experience.

The core architecture combines:

Next.js
↓
FastAPI
↓
Incident Investigation Engine
↓
Hindsight Long-Term Memory
↓
LLM Reasoning

The project uses:

  • Next.js + TypeScript for the frontend
  • FastAPI + Python for the backend
  • SQLite for relational application state
  • Hindsight by Vectorize for persistent operational memory
  • Groq LLM for reasoning
  • Tailwind CSS for the UI
  • Recharts for visualizations
  • Docker Compose for containerized deployment

The Architecture

The system is divided into several layers.

  1. Frontend

The frontend is built with Next.js and TypeScript.

It provides interfaces for:

Incident investigation
Historical memory search
Deployment timelines
Memory exploration
Pattern discovery
Before/after comparisons
Incident resolution

Some of the key UI components include:

InvestigationView
HaveWeSeenThisView
FailedFixCard
BeforeAfterComparison
ResolutionModal
TimelineView
MemoryCard

The goal is not just to display AI-generated text, but to expose the evidence and historical context that contributed to the investigation.

  1. FastAPI Backend

The backend acts as the orchestration layer.

It connects the frontend with:

Incident data
Deployment information
Configuration changes
Hindsight memory
The LLM
Investigation logic

The backend is organized into areas such as:

backend/
└── app/
├── api/
├── agents/
├── memory/
├── models/
├── schemas/
└── core/

FastAPI provides REST endpoints while SQLAlchemy handles the relational application state.

Pydantic is used for schema validation, helping ensure that agent outputs follow the expected structure.

The Most Important Component: Hindsight Memory

This is where DeployLens differs from a conventional chatbot.

DeployLens uses Hindsight by Vectorize as a persistent memory layer.

The idea is simple:

Don't just remember documents. Remember operational experiences.

The project stores different types of memories, including:

Deployment Memory

Example:

payment-service v2.7.4 was deployed.
The deployment changed Redis connection handling
and introduced PAYMENT_REDIS_POOL_SIZE=20.
Incident Memory
INC-1042 affected checkout-api.
Latency increased from 420ms to 2.8 seconds.
HTTP 5xx errors reached 17%.
Investigation Memory
Restarting checkout-api temporarily reduced latency,
but the problem returned within eight minutes.
Resolution Memory
The incident was resolved by increasing the Redis
connection pool size from 20 to 50.
Failure Memory
Restarting checkout-api repeatedly provided only
temporary recovery and did not permanently resolve
Redis pool exhaustion.
Learning Memory
For checkout incidents involving Redis timeout warnings
after payment deployments, investigate connection pool
configuration before restarting the service.

This distinction is extremely important.

A system that remembers only what happened is useful.

A system that remembers what happened, what was tried, what failed, and what ultimately worked becomes much more valuable during future incidents.

How an Investigation Works

Let's walk through a representative DeployLens investigation.

Suppose a production incident occurs:

INC-2051
Checkout Error Rate Spike

The system sees a significant increase in checkout errors.

Instead of immediately generating a generic troubleshooting response, DeployLens starts correlating evidence.

Step 1: Identify the incident

The active incident becomes the starting point.

Incident
↓
Checkout API
↓
Error-rate increase
Step 2: Examine recent changes

The system checks recent deployment and configuration information.

For example:

23 minutes before incident:
payment-service v2.8.1 deployed

The deployment introduced:

PAYMENT_REDIS_POOL_SIZE=20

Now there is a potentially relevant change.

Step 3: Search historical memory

DeployLens queries the Hindsight memory bank.

Instead of asking:

"What is Redis pool exhaustion?"

it asks a much more useful operational question:

"Have we experienced a similar incident before,
and how did we resolve it?"

A historical incident such as INC-1042 can then become relevant.

The project demo uses a representative 91% similarity match for this scenario.

Step 4: Compare previous actions

The memory system can surface previous remediation attempts.

For example:

Restart checkout-api
↓
Temporary improvement
↓
Problem returned

This gives the current investigation important context.

Step 5: Retrieve the verified resolution

The historical incident indicates that changing the Redis connection pool was the effective fix.

Instead of blindly suggesting another restart, the investigation can present the historical resolution:

PAYMENT_REDIS_POOL_SIZE
20 → 50

The engineer still makes the operational decision, but DeployLens makes the relevant organizational history much easier to access.

Why "Failed Fix Memory" Matters

One of the most interesting ideas in DeployLens is the concept of remembering failed or temporary fixes.

Traditional knowledge systems often capture successful solutions:

Problem → Solution

But incident response is messier than that.

Real investigations also look like:

Problem
↓
Attempt #1 → Failed
↓
Attempt #2 → Temporary recovery
↓
Attempt #3 → Failed
↓
Root cause discovered
↓
Verified resolution

DeployLens explicitly models this history.

That means the system can surface information such as:

⚠ Historical Fix Warning

Restarting checkout-api has previously provided
temporary relief but did not permanently resolve
the underlying Redis issue.

This is valuable because knowing what not to repeat can be as important as knowing what worked.

The Incident Learning Loop

Another important design decision is that DeployLens does not stop when an incident is investigated.

Once an incident is resolved, the outcome can be written back into memory.

The process becomes:

Incident
↓
Investigation
↓
Historical Recall
↓
Engineer Action
↓
Verified Resolution
↓
Store Incident Outcome
↓
New Organizational Memory

This creates a feedback loop.

Every resolved incident has the potential to improve future investigations.

In other words:

The system is designed to become more useful as operational history accumulates.

That is fundamentally different from a stateless question-and-answer workflow.

Before vs After: Stateless AI vs Memory-Powered AI

DeployLens also includes a dedicated comparison view.

Without persistent memory

A generic AI system might say:

Check recent deployments.
Inspect Redis configuration.
Restart the service.
Check application logs.
Monitor latency.

These are reasonable troubleshooting suggestions, but they are generic.

With persistent operational memory

DeployLens can connect:

Current incident
↓
Recent deployment
↓
Configuration change
↓
Similar historical incident
↓
Previous failed fix
↓
Verified resolution

The result is not simply an AI-generated answer.

It is contextual reasoning informed by organizational experience.

Memory Explorer

DeployLens also includes a dedicated /memory interface.

This allows users to inspect the operational memory accumulated by the system.

The interface is intended to provide visibility into:

Stored memories
Memory metadata
Search results
Memory categories
Memory growth

This is important for transparency.

When AI systems make decisions, engineers should be able to understand what information the system is working from.

Discovered Operational Patterns

DeployLens also includes a /patterns view.

The goal is to identify recurring operational relationships across incidents.

For example:

Pattern:
Redis timeout warnings frequently appear after
payment-service configuration changes.

Or:

Anti-pattern:
Service restarts repeatedly provide temporary relief
for connection-pool exhaustion incidents.

These observations can eventually become operational knowledge that helps engineers investigate incidents more systematically.

Project Structure

The repository is organized to keep the frontend, backend, memory layer, and documentation separated:

DeployLens/
│
├── frontend/
│ ├── app/
│ ├── components/
│ ├── lib/
│ └── types/
│
├── backend/
│ ├── app/
│ │ ├── api/
│ │ ├── agents/
│ │ ├── memory/
│ │ ├── models/
│ │ ├── schemas/
│ │ └── core/
│ │
│ ├── data/
│ └── tests/
│
├── scripts/
├── docs/
├── docker-compose.yml
└── README.md

This separation makes it easier to evolve individual parts of the system independently.

Top comments (0)