Production incidents rarely happen at a convenient time.
It might be 2 AM. An API suddenly starts returning 5xx errors, latency is increasing, alerts are firing, and an SRE has to answer one critical question:
- What changed, and have we seen this problem before? Most incident investigation workflows require engineers to jump between deployment histories, configuration changes, monitoring dashboards, Git commits, logs, Slack conversations, and old postmortems.
The problem isn't only finding information.
The problem is remembering what actually worked.
That is the idea behind DeployLens — a memory-powered production incident investigation system designed for DevOps, SRE, and platform engineering teams.
GitHub: https://github.com/yelemvenkatarachika/DeployLens
The Problem With Traditional Incident Investigation
Imagine a checkout service suddenly begins returning errors.
An engineer may need to investigate:
- Recent deployments
- Configuration changes
- Alerts and metrics
- Previous incidents
- Historical postmortems
- Previously attempted fixes
A traditional AI assistant can summarize the information available to it, but it usually doesn't maintain a reliable organizational memory of previous incidents and their outcomes.
For example, suppose restarting checkout-api was attempted several times during previous incidents.
A generic assistant might still suggest:
Restart the checkout service and monitor the error rate.
But an experienced SRE might know:
We already tried that three times.
It only helped temporarily.
The real issue was Redis connection pool exhaustion.
That difference is what DeployLens focuses on.
What Is DeployLens?
DeployLens is a production change intelligence and incident investigation agent with persistent operational memory.
Its central question is:
What changed before production broke, have we seen this failure before, and what actually worked last time?
Instead of treating every incident as a brand-new problem, DeployLens attempts to connect today's incident with the organization's historical operational experience.
The core architecture combines:
Next.js
↓
FastAPI
↓
Incident Investigation Engine
↓
Hindsight Long-Term Memory
↓
LLM Reasoning
The project uses:
- Next.js + TypeScript for the frontend
- FastAPI + Python for the backend
- SQLite for relational application state
- Hindsight by Vectorize for persistent operational memory
- Groq LLM for reasoning
- Tailwind CSS for the UI
- Recharts for visualizations
- Docker Compose for containerized deployment
The Architecture
The system is divided into several layers.
- Frontend
The frontend is built with Next.js and TypeScript.
It provides interfaces for:
Incident investigation
Historical memory search
Deployment timelines
Memory exploration
Pattern discovery
Before/after comparisons
Incident resolution
Some of the key UI components include:
InvestigationView
HaveWeSeenThisView
FailedFixCard
BeforeAfterComparison
ResolutionModal
TimelineView
MemoryCard
The goal is not just to display AI-generated text, but to expose the evidence and historical context that contributed to the investigation.
- FastAPI Backend
The backend acts as the orchestration layer.
It connects the frontend with:
Incident data
Deployment information
Configuration changes
Hindsight memory
The LLM
Investigation logic
The backend is organized into areas such as:
backend/
└── app/
├── api/
├── agents/
├── memory/
├── models/
├── schemas/
└── core/
FastAPI provides REST endpoints while SQLAlchemy handles the relational application state.
Pydantic is used for schema validation, helping ensure that agent outputs follow the expected structure.
The Most Important Component: Hindsight Memory
This is where DeployLens differs from a conventional chatbot.
DeployLens uses Hindsight by Vectorize as a persistent memory layer.
The idea is simple:
Don't just remember documents. Remember operational experiences.
The project stores different types of memories, including:
Deployment Memory
Example:
payment-service v2.7.4 was deployed.
The deployment changed Redis connection handling
and introduced PAYMENT_REDIS_POOL_SIZE=20.
Incident Memory
INC-1042 affected checkout-api.
Latency increased from 420ms to 2.8 seconds.
HTTP 5xx errors reached 17%.
Investigation Memory
Restarting checkout-api temporarily reduced latency,
but the problem returned within eight minutes.
Resolution Memory
The incident was resolved by increasing the Redis
connection pool size from 20 to 50.
Failure Memory
Restarting checkout-api repeatedly provided only
temporary recovery and did not permanently resolve
Redis pool exhaustion.
Learning Memory
For checkout incidents involving Redis timeout warnings
after payment deployments, investigate connection pool
configuration before restarting the service.
This distinction is extremely important.
A system that remembers only what happened is useful.
A system that remembers what happened, what was tried, what failed, and what ultimately worked becomes much more valuable during future incidents.
How an Investigation Works
Let's walk through a representative DeployLens investigation.
Suppose a production incident occurs:
INC-2051
Checkout Error Rate Spike
The system sees a significant increase in checkout errors.
Instead of immediately generating a generic troubleshooting response, DeployLens starts correlating evidence.
Step 1: Identify the incident
The active incident becomes the starting point.
Incident
↓
Checkout API
↓
Error-rate increase
Step 2: Examine recent changes
The system checks recent deployment and configuration information.
For example:
23 minutes before incident:
payment-service v2.8.1 deployed
The deployment introduced:
PAYMENT_REDIS_POOL_SIZE=20
Now there is a potentially relevant change.
Step 3: Search historical memory
DeployLens queries the Hindsight memory bank.
Instead of asking:
"What is Redis pool exhaustion?"
it asks a much more useful operational question:
"Have we experienced a similar incident before,
and how did we resolve it?"
A historical incident such as INC-1042 can then become relevant.
The project demo uses a representative 91% similarity match for this scenario.
Step 4: Compare previous actions
The memory system can surface previous remediation attempts.
For example:
Restart checkout-api
↓
Temporary improvement
↓
Problem returned
This gives the current investigation important context.
Step 5: Retrieve the verified resolution
The historical incident indicates that changing the Redis connection pool was the effective fix.
Instead of blindly suggesting another restart, the investigation can present the historical resolution:
PAYMENT_REDIS_POOL_SIZE
20 → 50
The engineer still makes the operational decision, but DeployLens makes the relevant organizational history much easier to access.
Why "Failed Fix Memory" Matters
One of the most interesting ideas in DeployLens is the concept of remembering failed or temporary fixes.
Traditional knowledge systems often capture successful solutions:
Problem → Solution
But incident response is messier than that.
Real investigations also look like:
Problem
↓
Attempt #1 → Failed
↓
Attempt #2 → Temporary recovery
↓
Attempt #3 → Failed
↓
Root cause discovered
↓
Verified resolution
DeployLens explicitly models this history.
That means the system can surface information such as:
⚠ Historical Fix Warning
Restarting checkout-api has previously provided
temporary relief but did not permanently resolve
the underlying Redis issue.
This is valuable because knowing what not to repeat can be as important as knowing what worked.
The Incident Learning Loop
Another important design decision is that DeployLens does not stop when an incident is investigated.
Once an incident is resolved, the outcome can be written back into memory.
The process becomes:
Incident
↓
Investigation
↓
Historical Recall
↓
Engineer Action
↓
Verified Resolution
↓
Store Incident Outcome
↓
New Organizational Memory
This creates a feedback loop.
Every resolved incident has the potential to improve future investigations.
In other words:
The system is designed to become more useful as operational history accumulates.
That is fundamentally different from a stateless question-and-answer workflow.
Before vs After: Stateless AI vs Memory-Powered AI
DeployLens also includes a dedicated comparison view.
Without persistent memory
A generic AI system might say:
Check recent deployments.
Inspect Redis configuration.
Restart the service.
Check application logs.
Monitor latency.
These are reasonable troubleshooting suggestions, but they are generic.
With persistent operational memory
DeployLens can connect:
Current incident
↓
Recent deployment
↓
Configuration change
↓
Similar historical incident
↓
Previous failed fix
↓
Verified resolution
The result is not simply an AI-generated answer.
It is contextual reasoning informed by organizational experience.
Memory Explorer
DeployLens also includes a dedicated /memory interface.
This allows users to inspect the operational memory accumulated by the system.
The interface is intended to provide visibility into:
Stored memories
Memory metadata
Search results
Memory categories
Memory growth
This is important for transparency.
When AI systems make decisions, engineers should be able to understand what information the system is working from.
Discovered Operational Patterns
DeployLens also includes a /patterns view.
The goal is to identify recurring operational relationships across incidents.
For example:
Pattern:
Redis timeout warnings frequently appear after
payment-service configuration changes.
Or:
Anti-pattern:
Service restarts repeatedly provide temporary relief
for connection-pool exhaustion incidents.
These observations can eventually become operational knowledge that helps engineers investigate incidents more systematically.
Project Structure
The repository is organized to keep the frontend, backend, memory layer, and documentation separated:
DeployLens/
│
├── frontend/
│ ├── app/
│ ├── components/
│ ├── lib/
│ └── types/
│
├── backend/
│ ├── app/
│ │ ├── api/
│ │ ├── agents/
│ │ ├── memory/
│ │ ├── models/
│ │ ├── schemas/
│ │ └── core/
│ │
│ ├── data/
│ └── tests/
│
├── scripts/
├── docs/
├── docker-compose.yml
└── README.md
This separation makes it easier to evolve individual parts of the system independently.
Top comments (0)