DEV Community

Cover image for IncidentPilot — An Evidence-First AI Agent for Production Incidents
aarthirs
aarthirs

Posted on

IncidentPilot — An Evidence-First AI Agent for Production Incidents

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

Production incidents rarely come with a single clean answer.

A spike in API latency might be caused by database saturation, connection pool exhaustion, a bad deployment, or a failing workload. The difficult part isn't generating another AI answer — it's finding the right evidence and knowing when the available guidance conflicts.

That's why I built IncidentPilot.

What is IncidentPilot?

IncidentPilot is an AI-powered incident investigation agent that helps engineers investigate production incidents using structured operational knowledge.

Instead of relying only on the model's general knowledge, IncidentPilot retrieves relevant information from a Sanity Knowledge Base through Sanity Context MCP.

It can connect:

  • Runbooks
  • Historical incidents
  • Service documentation
  • Alert definitions
  • Architecture documentation
  • Incident investigation policies

and uses those sources to build an evidence-backed investigation plan.

The interesting part

I intentionally created conflicting operational guidance.

For example, one older runbook says:

Restart the affected pods.

while the newer runbook says to:

Collect termination and container evidence before restarting.

When I ask:

"Should I restart the pods?"

IncidentPilot doesn't silently choose one answer.

It surfaces both sources, shows their dates, explains the conflict, and recommends following the newer investigation procedure.

That was an important design goal:

AI should expose conflicting evidence instead of hiding it.

Example investigation

I can ask:

"API latency increased to 4 seconds. 3 pods are restarting and database CPU is 92%. What should I investigate first?"

IncidentPilot identifies:

  1. Database saturation
  2. Pod restart causes
  3. Recent deployments/configuration changes
  4. Similar historical incidents

It also provides the source behind each recommendation.

Architecture

React UI
↓
Express backend
↓
AI agent
↓
Sanity Context MCP
↓
Sanity Knowledge Base
↓
Runbooks / Incidents / Alerts / Documentation
↓
Evidence-backed investigation

Why Sanity Context?

The main reason I used Sanity Context is that the agent needs access to structured, curated operational knowledge rather than relying on model memory.

The Knowledge Base contains the source material, while Context MCP provides the read-only interface for the agent.

This makes the answer traceable back to the underlying operational knowledge.

Knowledge Base

The IncidentPilot Knowledge Base contains:

  • Payment API service documentation
  • Database CPU alert
  • API latency alert
  • Pod restart alert
  • Database CPU runbook
  • Current pod restart runbook
  • Legacy pod restart runbook
  • Historical incident INC-248
  • Payment architecture documentation
  • Incident investigation policy

Challenge

This project was built for the DEV Community × Sanity Context challenge.

The project focuses specifically on the challenge requirements:

  • Meaningful use of Sanity Context
  • Structured operational content
  • Knowledge Base retrieval
  • Agent-based investigation
  • Source/provenance visibility
  • Conflicting guidance detection
  • Usable investigation UI

Tech Stack

  • React
  • Vite
  • Node.js
  • Express
  • OpenAI
  • Sanity Context MCP
  • Sanity Knowledge Base
  • Markdown knowledge sources

Hithub link : IncidentPilot

Try it

Ask IncidentPilot:

"What should I investigate first?"

or:

"Have we seen an incident like this before?"

or my favorite:

"Should I restart the pods?"

The last question demonstrates why grounding the agent in structured knowledge matters.

Top comments (0)