<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vishal Vishwanadhula</title>
    <description>The latest articles on DEV Community by Vishal Vishwanadhula (@vishal_vishwanadhula_8194).</description>
    <link>https://dev.to/vishal_vishwanadhula_8194</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150156%2Fab5e6199-fb17-4df7-a5e6-07182edca516.jpg</url>
      <title>DEV Community: Vishal Vishwanadhula</title>
      <link>https://dev.to/vishal_vishwanadhula_8194</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vishal_vishwanadhula_8194"/>
    <language>en</language>
    <item>
      <title>DeployLens: Building a Production Incident Investigator That Actually Remembers</title>
      <dc:creator>Vishal Vishwanadhula</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:54:21 +0000</pubDate>
      <link>https://dev.to/vishal_vishwanadhula_8194/deploylens-building-a-production-incident-investigator-that-actually-remembers-37a4</link>
      <guid>https://dev.to/vishal_vishwanadhula_8194/deploylens-building-a-production-incident-investigator-that-actually-remembers-37a4</guid>
      <description>&lt;p&gt;Production incidents rarely happen at a convenient time.&lt;/p&gt;

&lt;p&gt;It might be 2 AM. An API suddenly starts returning 5xx errors, latency is increasing, alerts are firing, and an SRE has to answer one critical question:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What changed, and have we seen this problem before?
Most incident investigation workflows require engineers to jump between deployment histories, configuration changes, monitoring dashboards, Git commits, logs, Slack conversations, and old postmortems.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The problem isn't only finding information.&lt;/p&gt;

&lt;p&gt;The problem is remembering what actually worked.&lt;/p&gt;

&lt;p&gt;That is the idea behind DeployLens — a memory-powered production incident investigation system designed for DevOps, SRE, and platform engineering teams.&lt;/p&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/yelemvenkatarachika/DeployLens" rel="noopener noreferrer"&gt;https://github.com/yelemvenkatarachika/DeployLens&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Traditional Incident Investigation
&lt;/h2&gt;

&lt;p&gt;Imagine a checkout service suddenly begins returning errors.&lt;/p&gt;

&lt;p&gt;An engineer may need to investigate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recent deployments&lt;/li&gt;
&lt;li&gt;Configuration changes&lt;/li&gt;
&lt;li&gt;Alerts and metrics&lt;/li&gt;
&lt;li&gt;Previous incidents&lt;/li&gt;
&lt;li&gt;Historical postmortems&lt;/li&gt;
&lt;li&gt;Previously attempted fixes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A traditional AI assistant can summarize the information available to it, but it usually doesn't maintain a reliable organizational memory of previous incidents and their outcomes.&lt;/p&gt;

&lt;p&gt;For example, suppose restarting checkout-api was attempted several times during previous incidents.&lt;/p&gt;

&lt;p&gt;A generic assistant might still suggest:&lt;/p&gt;

&lt;p&gt;Restart the checkout service and monitor the error rate.&lt;/p&gt;

&lt;p&gt;But an experienced SRE might know:&lt;/p&gt;

&lt;p&gt;We already tried that three times.&lt;br&gt;
It only helped temporarily.&lt;br&gt;
The real issue was Redis connection pool exhaustion.&lt;/p&gt;

&lt;p&gt;That difference is what DeployLens focuses on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is DeployLens?
&lt;/h2&gt;

&lt;p&gt;DeployLens is a production change intelligence and incident investigation agent with persistent operational memory.&lt;/p&gt;

&lt;p&gt;Its central question is:&lt;/p&gt;

&lt;p&gt;What changed before production broke, have we seen this failure before, and what actually worked last time?&lt;/p&gt;

&lt;p&gt;Instead of treating every incident as a brand-new problem, DeployLens attempts to connect today's incident with the organization's historical operational experience.&lt;/p&gt;

&lt;p&gt;The core architecture combines:&lt;/p&gt;

&lt;p&gt;Next.js&lt;br&gt;
     ↓&lt;br&gt;
FastAPI&lt;br&gt;
     ↓&lt;br&gt;
Incident Investigation Engine&lt;br&gt;
     ↓&lt;br&gt;
Hindsight Long-Term Memory&lt;br&gt;
     ↓&lt;br&gt;
LLM Reasoning&lt;/p&gt;

&lt;p&gt;The project uses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Next.js + TypeScript for the frontend&lt;/li&gt;
&lt;li&gt;FastAPI + Python for the backend&lt;/li&gt;
&lt;li&gt;SQLite for relational application state&lt;/li&gt;
&lt;li&gt;Hindsight by Vectorize for persistent operational memory&lt;/li&gt;
&lt;li&gt;Groq LLM for reasoning&lt;/li&gt;
&lt;li&gt;Tailwind CSS for the UI&lt;/li&gt;
&lt;li&gt;Recharts for visualizations&lt;/li&gt;
&lt;li&gt;Docker Compose for containerized deployment&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;The system is divided into several layers.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Frontend&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The frontend is built with Next.js and TypeScript.&lt;/p&gt;

&lt;p&gt;It provides interfaces for:&lt;/p&gt;

&lt;p&gt;Incident investigation&lt;br&gt;
Historical memory search&lt;br&gt;
Deployment timelines&lt;br&gt;
Memory exploration&lt;br&gt;
Pattern discovery&lt;br&gt;
Before/after comparisons&lt;br&gt;
Incident resolution&lt;/p&gt;

&lt;p&gt;Some of the key UI components include:&lt;/p&gt;

&lt;p&gt;InvestigationView&lt;br&gt;
HaveWeSeenThisView&lt;br&gt;
FailedFixCard&lt;br&gt;
BeforeAfterComparison&lt;br&gt;
ResolutionModal&lt;br&gt;
TimelineView&lt;br&gt;
MemoryCard&lt;/p&gt;

&lt;p&gt;The goal is not just to display AI-generated text, but to expose the evidence and historical context that contributed to the investigation.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;FastAPI Backend&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The backend acts as the orchestration layer.&lt;/p&gt;

&lt;p&gt;It connects the frontend with:&lt;/p&gt;

&lt;p&gt;Incident data&lt;br&gt;
Deployment information&lt;br&gt;
Configuration changes&lt;br&gt;
Hindsight memory&lt;br&gt;
The LLM&lt;br&gt;
Investigation logic&lt;/p&gt;

&lt;p&gt;The backend is organized into areas such as:&lt;/p&gt;

&lt;p&gt;backend/&lt;br&gt;
└── app/&lt;br&gt;
    ├── api/&lt;br&gt;
    ├── agents/&lt;br&gt;
    ├── memory/&lt;br&gt;
    ├── models/&lt;br&gt;
    ├── schemas/&lt;br&gt;
    └── core/&lt;/p&gt;

&lt;p&gt;FastAPI provides REST endpoints while SQLAlchemy handles the relational application state.&lt;/p&gt;

&lt;p&gt;Pydantic is used for schema validation, helping ensure that agent outputs follow the expected structure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Most Important Component: Hindsight Memory
&lt;/h2&gt;

&lt;p&gt;This is where DeployLens differs from a conventional chatbot.&lt;/p&gt;

&lt;p&gt;DeployLens uses Hindsight by Vectorize as a persistent memory layer.&lt;/p&gt;

&lt;p&gt;The idea is simple:&lt;/p&gt;

&lt;p&gt;Don't just remember documents. Remember operational experiences.&lt;/p&gt;

&lt;p&gt;The project stores different types of memories, including:&lt;/p&gt;

&lt;p&gt;Deployment Memory&lt;/p&gt;

&lt;p&gt;Example:&lt;/p&gt;

&lt;p&gt;payment-service v2.7.4 was deployed.&lt;br&gt;
The deployment changed Redis connection handling&lt;br&gt;
and introduced PAYMENT_REDIS_POOL_SIZE=20.&lt;br&gt;
Incident Memory&lt;br&gt;
INC-1042 affected checkout-api.&lt;br&gt;
Latency increased from 420ms to 2.8 seconds.&lt;br&gt;
HTTP 5xx errors reached 17%.&lt;br&gt;
Investigation Memory&lt;br&gt;
Restarting checkout-api temporarily reduced latency,&lt;br&gt;
but the problem returned within eight minutes.&lt;br&gt;
Resolution Memory&lt;br&gt;
The incident was resolved by increasing the Redis&lt;br&gt;
connection pool size from 20 to 50.&lt;br&gt;
Failure Memory&lt;br&gt;
Restarting checkout-api repeatedly provided only&lt;br&gt;
temporary recovery and did not permanently resolve&lt;br&gt;
Redis pool exhaustion.&lt;br&gt;
Learning Memory&lt;br&gt;
For checkout incidents involving Redis timeout warnings&lt;br&gt;
after payment deployments, investigate connection pool&lt;br&gt;
configuration before restarting the service.&lt;/p&gt;

&lt;p&gt;This distinction is extremely important.&lt;/p&gt;

&lt;p&gt;A system that remembers only what happened is useful.&lt;/p&gt;

&lt;p&gt;A system that remembers what happened, what was tried, what failed, and what ultimately worked becomes much more valuable during future incidents.&lt;/p&gt;

&lt;h2&gt;
  
  
  How an Investigation Works
&lt;/h2&gt;

&lt;p&gt;Let's walk through a representative DeployLens investigation.&lt;/p&gt;

&lt;p&gt;Suppose a production incident occurs:&lt;/p&gt;

&lt;p&gt;INC-2051&lt;br&gt;
Checkout Error Rate Spike&lt;/p&gt;

&lt;p&gt;The system sees a significant increase in checkout errors.&lt;/p&gt;

&lt;p&gt;Instead of immediately generating a generic troubleshooting response, DeployLens starts correlating evidence.&lt;/p&gt;

&lt;p&gt;Step 1: Identify the incident&lt;/p&gt;

&lt;p&gt;The active incident becomes the starting point.&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Checkout API&lt;br&gt;
   ↓&lt;br&gt;
Error-rate increase&lt;br&gt;
Step 2: Examine recent changes&lt;/p&gt;

&lt;p&gt;The system checks recent deployment and configuration information.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;23 minutes before incident:&lt;br&gt;
payment-service v2.8.1 deployed&lt;/p&gt;

&lt;p&gt;The deployment introduced:&lt;/p&gt;

&lt;p&gt;PAYMENT_REDIS_POOL_SIZE=20&lt;/p&gt;

&lt;p&gt;Now there is a potentially relevant change.&lt;/p&gt;

&lt;p&gt;Step 3: Search historical memory&lt;/p&gt;

&lt;p&gt;DeployLens queries the Hindsight memory bank.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;p&gt;"What is Redis pool exhaustion?"&lt;/p&gt;

&lt;p&gt;it asks a much more useful operational question:&lt;/p&gt;

&lt;p&gt;"Have we experienced a similar incident before,&lt;br&gt;
and how did we resolve it?"&lt;/p&gt;

&lt;p&gt;A historical incident such as INC-1042 can then become relevant.&lt;/p&gt;

&lt;p&gt;The project demo uses a representative 91% similarity match for this scenario.&lt;/p&gt;

&lt;p&gt;Step 4: Compare previous actions&lt;/p&gt;

&lt;p&gt;The memory system can surface previous remediation attempts.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Restart checkout-api&lt;br&gt;
       ↓&lt;br&gt;
Temporary improvement&lt;br&gt;
       ↓&lt;br&gt;
Problem returned&lt;/p&gt;

&lt;p&gt;This gives the current investigation important context.&lt;/p&gt;

&lt;p&gt;Step 5: Retrieve the verified resolution&lt;/p&gt;

&lt;p&gt;The historical incident indicates that changing the Redis connection pool was the effective fix.&lt;/p&gt;

&lt;p&gt;Instead of blindly suggesting another restart, the investigation can present the historical resolution:&lt;/p&gt;

&lt;p&gt;PAYMENT_REDIS_POOL_SIZE&lt;br&gt;
20 → 50&lt;/p&gt;

&lt;p&gt;The engineer still makes the operational decision, but DeployLens makes the relevant organizational history much easier to access.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "Failed Fix Memory" Matters
&lt;/h2&gt;

&lt;p&gt;One of the most interesting ideas in DeployLens is the concept of remembering failed or temporary fixes.&lt;/p&gt;

&lt;p&gt;Traditional knowledge systems often capture successful solutions:&lt;/p&gt;

&lt;p&gt;Problem → Solution&lt;/p&gt;

&lt;p&gt;But incident response is messier than that.&lt;/p&gt;

&lt;p&gt;Real investigations also look like:&lt;/p&gt;

&lt;p&gt;Problem&lt;br&gt;
   ↓&lt;br&gt;
Attempt #1 → Failed&lt;br&gt;
   ↓&lt;br&gt;
Attempt #2 → Temporary recovery&lt;br&gt;
   ↓&lt;br&gt;
Attempt #3 → Failed&lt;br&gt;
   ↓&lt;br&gt;
Root cause discovered&lt;br&gt;
   ↓&lt;br&gt;
Verified resolution&lt;/p&gt;

&lt;p&gt;DeployLens explicitly models this history.&lt;/p&gt;

&lt;p&gt;That means the system can surface information such as:&lt;/p&gt;

&lt;p&gt;⚠ Historical Fix Warning&lt;/p&gt;

&lt;p&gt;Restarting checkout-api has previously provided&lt;br&gt;
temporary relief but did not permanently resolve&lt;br&gt;
the underlying Redis issue.&lt;/p&gt;

&lt;p&gt;This is valuable because knowing what not to repeat can be as important as knowing what worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Incident Learning Loop
&lt;/h2&gt;

&lt;p&gt;Another important design decision is that DeployLens does not stop when an incident is investigated.&lt;/p&gt;

&lt;p&gt;Once an incident is resolved, the outcome can be written back into memory.&lt;/p&gt;

&lt;p&gt;The process becomes:&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Investigation&lt;br&gt;
   ↓&lt;br&gt;
Historical Recall&lt;br&gt;
   ↓&lt;br&gt;
Engineer Action&lt;br&gt;
   ↓&lt;br&gt;
Verified Resolution&lt;br&gt;
   ↓&lt;br&gt;
Store Incident Outcome&lt;br&gt;
   ↓&lt;br&gt;
New Organizational Memory&lt;/p&gt;

&lt;p&gt;This creates a feedback loop.&lt;/p&gt;

&lt;p&gt;Every resolved incident has the potential to improve future investigations.&lt;/p&gt;

&lt;p&gt;In other words:&lt;/p&gt;

&lt;p&gt;The system is designed to become more useful as operational history accumulates.&lt;/p&gt;

&lt;p&gt;That is fundamentally different from a stateless question-and-answer workflow.&lt;/p&gt;

&lt;p&gt;Before vs After: Stateless AI vs Memory-Powered AI&lt;/p&gt;

&lt;p&gt;DeployLens also includes a dedicated comparison view.&lt;/p&gt;

&lt;p&gt;Without persistent memory&lt;/p&gt;

&lt;p&gt;A generic AI system might say:&lt;/p&gt;

&lt;p&gt;Check recent deployments.&lt;br&gt;
Inspect Redis configuration.&lt;br&gt;
Restart the service.&lt;br&gt;
Check application logs.&lt;br&gt;
Monitor latency.&lt;/p&gt;

&lt;p&gt;These are reasonable troubleshooting suggestions, but they are generic.&lt;/p&gt;

&lt;p&gt;With persistent operational memory&lt;/p&gt;

&lt;p&gt;DeployLens can connect:&lt;/p&gt;

&lt;p&gt;Current incident&lt;br&gt;
       ↓&lt;br&gt;
Recent deployment&lt;br&gt;
       ↓&lt;br&gt;
Configuration change&lt;br&gt;
       ↓&lt;br&gt;
Similar historical incident&lt;br&gt;
       ↓&lt;br&gt;
Previous failed fix&lt;br&gt;
       ↓&lt;br&gt;
Verified resolution&lt;/p&gt;

&lt;p&gt;The result is not simply an AI-generated answer.&lt;/p&gt;

&lt;p&gt;It is contextual reasoning informed by organizational experience.&lt;/p&gt;

&lt;p&gt;Memory Explorer&lt;/p&gt;

&lt;p&gt;DeployLens also includes a dedicated /memory interface.&lt;/p&gt;

&lt;p&gt;This allows users to inspect the operational memory accumulated by the system.&lt;/p&gt;

&lt;p&gt;The interface is intended to provide visibility into:&lt;/p&gt;

&lt;p&gt;Stored memories&lt;br&gt;
Memory metadata&lt;br&gt;
Search results&lt;br&gt;
Memory categories&lt;br&gt;
Memory growth&lt;/p&gt;

&lt;p&gt;This is important for transparency.&lt;/p&gt;

&lt;p&gt;When AI systems make decisions, engineers should be able to understand what information the system is working from.&lt;/p&gt;

&lt;p&gt;Discovered Operational Patterns&lt;/p&gt;

&lt;p&gt;DeployLens also includes a /patterns view.&lt;/p&gt;

&lt;p&gt;The goal is to identify recurring operational relationships across incidents.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Pattern:&lt;br&gt;
Redis timeout warnings frequently appear after&lt;br&gt;
payment-service configuration changes.&lt;/p&gt;

&lt;p&gt;Or:&lt;/p&gt;

&lt;p&gt;Anti-pattern:&lt;br&gt;
Service restarts repeatedly provide temporary relief&lt;br&gt;
for connection-pool exhaustion incidents.&lt;/p&gt;

&lt;p&gt;These observations can eventually become operational knowledge that helps engineers investigate incidents more systematically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Structure
&lt;/h2&gt;

&lt;p&gt;The repository is organized to keep the frontend, backend, memory layer, and documentation separated:&lt;/p&gt;

&lt;p&gt;DeployLens/&lt;br&gt;
│&lt;br&gt;
├── frontend/&lt;br&gt;
│   ├── app/&lt;br&gt;
│   ├── components/&lt;br&gt;
│   ├── lib/&lt;br&gt;
│   └── types/&lt;br&gt;
│&lt;br&gt;
├── backend/&lt;br&gt;
│   ├── app/&lt;br&gt;
│   │   ├── api/&lt;br&gt;
│   │   ├── agents/&lt;br&gt;
│   │   ├── memory/&lt;br&gt;
│   │   ├── models/&lt;br&gt;
│   │   ├── schemas/&lt;br&gt;
│   │   └── core/&lt;br&gt;
│   │&lt;br&gt;
│   ├── data/&lt;br&gt;
│   └── tests/&lt;br&gt;
│&lt;br&gt;
├── scripts/&lt;br&gt;
├── docs/&lt;br&gt;
├── docker-compose.yml&lt;br&gt;
└── README.md&lt;/p&gt;

&lt;p&gt;This separation makes it easier to evolve individual parts of the system independently.&lt;/p&gt;

</description>
      <category>automation</category>
      <category>devops</category>
      <category>github</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
